---
title: "Reinforcement learning"
description: "Learning to maximise a reward signal through trial and feedback."
canonical: https://aigovernanceengineer.com/glossary/reinforcement-learning
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22956197
version: "0.5.0"
updated: 2026-09-24
---

# Reinforcement learning

Learning to maximise a reward signal through trial and feedback [1]. Its characteristic failure is reward hacking, so the reward is recorded as the system's objective and evals look for unintended strategies.

- Developed in: [ch. 11, By learning paradigm](https://aigovernanceengineer.com/bok/ai-defined#by-learning-paradigm)
- Chapters: [ch. 11, AI Defined](https://aigovernanceengineer.com/bok/ai-defined)
- In the glossary chapter: https://aigovernanceengineer.com/bok/glossary#t-reinforcement-learning

## Sources

[1] ISO/IEC 22989:2022, Artificial intelligence concepts and terminology (referenced by identifier only; autonomy and heteronomy; clause 5.11 machine learning approaches: supervised, unsupervised, semi-supervised, reinforcement). ISO/IEC. 2022-07. https://www.iso.org/standard/74296.html (verified: secondary)
