Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking

Reinforcement Learning from Human Feedback remains vulnerable to reward hacking, where models exploit imperfections in learned reward models while violating human intent. Adversarial Reward Auditing (ARA) reframes reward hacking as a competitive game: a Hacker policy discovers vulnerabilities, an Auditor detects exploitation from latent representations, and Auditor-Guided RLHF gates reward signals to penalize detected hacking. Across sycophancy, verbosity, and code-gaming scenarios, ARA improves the alignment-utility tradeoff and generalizes detection and mitigation across domains.

Authors

Mohammad Beigi

Ming Jin

Junshan Zhang

Qifan Wang

Lifu Huang

Published

October 6, 2026