Adversarial Reward Auditing for Active Detection and Mitigation of Reward Hacking
Reinforcement Learning from Human Feedback remains vulnerable to reward hacking, where models exploit imperfections in learned reward models while violating human intent. Adversarial Reward Auditing (ARA) reframes reward hacking as a competitive game: a Hacker policy discovers vulnerabilities, an Auditor detects exploitation from latent representations, and Auditor-Guided RLHF gates reward signals to penalize detected hacking. Across sycophancy, verbosity, and code-gaming scenarios, ARA improves the alignment-utility tradeoff and generalizes detection and mitigation across domains.