Reward hacking

A system finding an unintended way to maximise its reward or objective without doing what its designers meant 1. Answered by recording the objective and testing for unintended strategies, not only for the intended task.

Developed in
ch. 11, By learning paradigm
Chapters
ch. 11, AI Defined
Source
1 numbered reference, listed below

Where it is used

2 chapters of the Body of Knowledge use the term. Each link opens the first section that does.

Sources

  1. [1] Concrete Problems in AI Safety (reward hacking among five practical problems; arXiv 1606.06565). Amodei et al.. 2016-06-21. https://arxiv.org/abs/1606.06565 (verified: primary)

Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .

Cite this term

García Aibar, J. (2026). Reward hacking. In AI Governance Engineering: The Thesis & Body of Knowledge (v0.5.0), Glossary. https://doi.org/10.5281/zenodo.22956197. https://aigovernanceengineer.com/glossary/reward-hacking. CC BY 4.0

BibTeX

@misc{aige2026rewardhacking,
  author  = {Jorge García Aibar},
  title   = {{Reward hacking}},
  note    = {Glossary, AI Governance Engineering: The Thesis \& Body of Knowledge, version 0.5.0},
  year    = {2026},
  doi     = {10.5281/zenodo.22956197},
  url     = {https://aigovernanceengineer.com/glossary/reward-hacking}
}