Reinforcement learning from human feedback (RLHF)

A way to align a pre-trained model: supervised fine-tuning on human demonstrations, then reinforcement learning against a reward model trained on human rankings of outputs 1. The raters' instructions and the reward model become governed artefacts, because they shape what the model refuses and prefers.

Developed in
ch. 11, By learning paradigm
Chapters
ch. 11, AI Defined
Contrast with
Fine-tuning
Source
1 numbered reference, listed below

Where it is used

The term is not used under this name in running prose; the sections listed under "Developed in" treat it.

Sources

  1. [1] "Training language models to follow instructions with human feedback" (Ouyang et al.; supervised fine-tuning on demonstrations, then reinforcement learning from human feedback on ranked outputs; arXiv 2203.02155). arXiv. 2022-03-04. https://arxiv.org/abs/2203.02155 (verified: primary)

Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .

Cite this term

García Aibar, J. (2026). Reinforcement learning from human feedback (RLHF). In AI Governance Engineering: The Thesis & Body of Knowledge (v0.5.0), Glossary. https://doi.org/10.5281/zenodo.22956197. https://aigovernanceengineer.com/glossary/reinforcement-learning-from-human-feedback-rlhf. CC BY 4.0

BibTeX

@misc{aige2026reinforcementlearningfromhumanfeedbackrlhf,
  author  = {Jorge García Aibar},
  title   = {{Reinforcement learning from human feedback (RLHF)}},
  note    = {Glossary, AI Governance Engineering: The Thesis \& Body of Knowledge, version 0.5.0},
  year    = {2026},
  doi     = {10.5281/zenodo.22956197},
  url     = {https://aigovernanceengineer.com/glossary/reinforcement-learning-from-human-feedback-rlhf}
}