A way to align a pre-trained model: supervised fine-tuning on human demonstrations, then reinforcement learning against a reward model trained on human rankings of outputs 1. The raters' instructions and the reward model become governed artefacts, because they shape what the model refuses and prefers.
- Developed in
- ch. 11, By learning paradigm
- Chapters
- ch. 11, AI Defined
- Contrast with
- Fine-tuning
- Source
- 1 numbered reference, listed below
Where it is used
The term is not used under this name in running prose; the sections listed under "Developed in" treat it.
Related terms
Sources
- [1] "Training language models to follow instructions with human feedback" (Ouyang et al.; supervised fine-tuning on demonstrations, then reinforcement learning from human feedback on ranked outputs; arXiv 2203.02155). arXiv. 2022-03-04. https://arxiv.org/abs/2203.02155 (verified: primary)
Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .