The presence of evaluation items in a model's training data, which inflates its scores; it can be demonstrated even for black-box language models 1. Mitigated with private held-out sets, rotated items and dated test items.
- Developed in
- ch. 14, Statistical validity of evals
- Chapters
- ch. 14, Development
- Source
- 1 numbered reference, listed below
Where it is used
The term is not used under this name in running prose; the sections listed under "Developed in" treat it.
Related terms
Sources
- [1] Proving Test Set Contamination in Black Box Language Models (Oren et al.; arXiv 2310.17623). arXiv. 2023-10-26. https://arxiv.org/abs/2310.17623 (verified: primary)
Definitions of legal terms paraphrase the cited text, which governs. Dated statements are as of .