Test-set contamination

The presence of evaluation items in a model's training data, which inflates its scores; it can be demonstrated even for black-box language models 1. Mitigated with private held-out sets, rotated items and dated test items.

Developed in
ch. 14, Statistical validity of evals
Chapters
ch. 14, Development
Source
1 numbered reference, listed below

Where it is used

The term is not used under this name in running prose; the sections listed under "Developed in" treat it.

Sources

  1. [1] Proving Test Set Contamination in Black Box Language Models (Oren et al.; arXiv 2310.17623). arXiv. 2023-10-26. https://arxiv.org/abs/2310.17623 (verified: primary)

Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .

Cite this term

García Aibar, J. (2026). Test-set contamination. In AI Governance Engineering: The Thesis & Body of Knowledge (v0.5.0), Glossary. https://doi.org/10.5281/zenodo.22956197. https://aigovernanceengineer.com/glossary/test-set-contamination. CC BY 4.0

BibTeX

@misc{aige2026testsetcontamination,
  author  = {Jorge García Aibar},
  title   = {{Test-set contamination}},
  note    = {Glossary, AI Governance Engineering: The Thesis \& Body of Knowledge, version 0.5.0},
  year    = {2026},
  doi     = {10.5281/zenodo.22956197},
  url     = {https://aigovernanceengineer.com/glossary/test-set-contamination}
}