---
title: "Reinforcement learning from human feedback (RLHF)"
description: "A way to align a pre-trained model: supervised fine-tuning on human demonstrations, then reinforcement learning against a reward model trained on human rankings of outputs."
canonical: https://aigovernanceengineer.com/glossary/reinforcement-learning-from-human-feedback-rlhf
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22956197
version: "0.5.0"
updated: 2026-09-24
---

# Reinforcement learning from human feedback (RLHF)

A way to align a pre-trained model: supervised fine-tuning on human demonstrations, then reinforcement learning against a reward model trained on human rankings of outputs [1]. The raters' instructions and the reward model become governed artefacts, because they shape what the model refuses and prefers.

- Developed in: [ch. 11, By learning paradigm](https://aigovernanceengineer.com/bok/ai-defined#by-learning-paradigm)
- Chapters: [ch. 11, AI Defined](https://aigovernanceengineer.com/bok/ai-defined)
- Contrast with: [Fine-tuning](https://aigovernanceengineer.com/glossary/fine-tuning)
- In the glossary chapter: https://aigovernanceengineer.com/bok/glossary#t-reinforcement-learning-from-human-feedback-rlhf

## Sources

[1] "Training language models to follow instructions with human feedback" (Ouyang et al.; supervised fine-tuning on demonstrations, then reinforcement learning from human feedback on ranked outputs; arXiv 2203.02155). arXiv. 2022-03-04. https://arxiv.org/abs/2203.02155 (verified: primary)
