Evals

Automated tests of a model's or agent's behaviour (capability, safety and adversarial), run as controls, not as one-off research.

Developed in
ch. 04, Layer 03: Evals & Red Teaming as Evidence
Chapters
ch. 04, The Stack
Source
Defined by this Body of Knowledge: the section it is developed in is the source

Where it is used

20 chapters of the Body of Knowledge use the term. Each link opens the first section that does.

Patterns that use this term

10 pattern pages use the term, most mentions first.

Sources

No external source: the term is coined or used in a specific sense by this Body of Knowledge, and the section linked under "Developed in" is its source.

Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .

Cite this term

García Aibar, J. (2026). Evals. In AI Governance Engineering: The Thesis & Body of Knowledge (v0.5.0), Glossary. https://doi.org/10.5281/zenodo.22956197. https://aigovernanceengineer.com/glossary/evals. CC BY 4.0

BibTeX

@misc{aige2026evals,
  author  = {Jorge García Aibar},
  title   = {{Evals}},
  note    = {Glossary, AI Governance Engineering: The Thesis \& Body of Knowledge, version 0.5.0},
  year    = {2026},
  doi     = {10.5281/zenodo.22956197},
  url     = {https://aigovernanceengineer.com/glossary/evals}
}