Learning to maximise a reward signal through trial and feedback 1. Its characteristic failure is reward hacking, so the reward is recorded as the system's objective and evals look for unintended strategies.
- Developed in
- ch. 11, By learning paradigm
- Chapters
- ch. 11, AI Defined
- Source
- 1 numbered reference, listed below
Where it is used
The term is not used under this name in running prose; the sections listed under "Developed in" treat it.
Related terms
Sources
- [1] ISO/IEC 22989:2022, Artificial intelligence concepts and terminology (referenced by identifier only; autonomy and heteronomy; clause 5.11 machine learning approaches: supervised, unsupervised, semi-supervised, reinforcement). ISO/IEC. 2022-07. https://www.iso.org/standard/74296.html (verified: secondary)
Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .