---
title: "Evals"
description: "Automated tests of a model's or agent's behaviour (capability, safety and adversarial), run as controls, not as one-off research."
canonical: https://aigovernanceengineer.com/glossary/evals
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22956197
version: "0.5.0"
updated: 2026-09-24
---

# Evals

Automated tests of a model's or agent's behaviour (capability, safety and adversarial), run as controls, not as one-off research.

- Developed in: [ch. 04, Layer 03: Evals & Red Teaming as Evidence](https://aigovernanceengineer.com/bok/the-stack#layer-03-evals--red-teaming-as-evidence)
- Chapters: [ch. 04, The Stack](https://aigovernanceengineer.com/bok/the-stack)
- In the glossary chapter: https://aigovernanceengineer.com/bok/glossary#t-evals

## Sources

Defined by this Body of Knowledge: the section it is developed in is the source.
