The State of Reward Hacking in AI

September 8, 2026
View text
  • Joseph Rogero

Modern machine learning reinforces outputs and behaviors that score well on certain measures. Most of these measures are proxies for things we really care about: User ratings are proxies for helpful conversation, benchmark scores are proxies for task competence, passing unit tests are proxies for working code. Because these measures incompletely track the complexity of human aims, AIs learn to "reward hack", or pursue an unexpected and undesirable strategy that scores highly but is not the intended behavior. This report summarizes the state of reward hacking: when it happens and what mitigations are available.

Footnotes

Citation