---
title: "Evaluation Validity Checks"
description: "Draft control: before an AI evaluation result is reported, check that tasks are solvable, the scorer works and failed runs were read, and report the checks."
canonical: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-009
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22857084
version: "0.2"
updated: 2026-09-26
---

# Evaluation Validity Checks

> A result is reported only after checks that the run measured what it claims: scoring worked, the environment did not fail, and the path was evaluated as well as the answer.

- Id: AIGE-CTL-EVAL-009
- Profile: [Evaluation Environment Control Profile v0.2](https://aigovernanceengineer.com/controls/evaluation-environment)
- Status: Draft
- Review: Open for technical review
- Published: 2026-09-26
- Updated: 2026-09-26
- Anchor on the profile page: https://aigovernanceengineer.com/controls/evaluation-environment#aige-ctl-eval-009

Draft for review, not a claim of conformity. A draft control specification, open for technical review: illustrative, not legal advice and binding on no one.

## The control record

- Id: `AIGE-CTL-EVAL-009` · v0.2 · Draft · Open for technical review
- Depth: Specified
- Objective: A result is reported only after checks that the run measured what it claims: scoring worked, the environment did not fail, and the path was evaluated as well as the answer.
- Failure modes:
  - A result is reported from a run whose environment crashed or whose automatic scoring was wrong.
  - A task that could not be solved as set up is scored and reported as a failure of the model.
  - Only final answers are scored: nobody reads the transcripts of failed runs, or of successes, for scorer tampering, reward hacking, communication between runs or signs of evaluation awareness.
  - A failed validity check does not block the release it was meant to gate.
- Scope: Evaluation runs whose results feed a release decision or an assurance claim. The choice of benchmarks and their statistical design are only in scope where they decide whether a result is valid.
- Enforcement points:
  - pre_merge: on every pull request
- Verification:
  - Inspect: Before the runs, inspect the task admission records: each task has evidence that it can be solved in this environment (a reference solution or a solved run), and the answers, the scorer and the task data are outside the agent's reach.
  - Test: Before the runs, score a known-correct and a known-incorrect submission for each task through the scorer the runs will use: the scorer must accept the first and reject the second.
  - Observe: After the runs, read the transcripts of every failed run and of a recorded sample of successes: classify each failure as a model limitation or a spurious failure (task bug, scoring error, crashed environment), and flag reward hacking, scorer tampering, communication between runs and verbalized evaluation awareness.
  - Attest: Before release, the evaluation lead states in the signed test report which runs were excluded or re-scored after these checks and why, and that no failed check was waived without a recorded approval.
- Evidence:
  - Task admission and scorer check records: solvability evidence and the verdicts on known-correct and known-incorrect submissions · Layer 03 Evals & Red Teaming as Evidence · [evidence-record.v1](https://aigovernanceengineer.com/resources/templates#schema-evidence-record)
  - Signed test report listing the validity checks run, the runs excluded or re-scored with the reason, and any waiver · Layer 03 Evals & Red Teaming as Evidence · [test-report.v1](https://aigovernanceengineer.com/resources/templates#schema-test-report)
  - Validity check of each result, filed as an observation of this control · Layer 05 Assurance & Continuous Compliance · [control-observation.v1](https://aigovernanceengineer.com/resources/templates#schema-control-observation)
- Failure response: deny: block the action. A result whose validity checks failed or were not run is not released to the decision it gates. A spurious failure is fixed and the task rerun, or the task is excluded and the exclusion reported; runs with scorer tampering, communication between runs or verbalized evaluation awareness are excluded or reported as contaminated.
- Layer: [Layer 03 Evals & Red Teaming as Evidence](https://aigovernanceengineer.com/bok/the-stack#layer-03-evals--red-teaming-as-evidence)
- Patterns: [Eval Gate in CI](https://aigovernanceengineer.com/patterns/eval-gate-in-ci), [Adversarial Red-Team Suite](https://aigovernanceengineer.com/patterns/adversarial-red-team-suite)
- Seeded from: [Independent trajectory evals](https://aigovernanceengineer.com/bok/governing-agents#what-makes-an-agent-a-governance-object)
- Mappings:
  - Obligations: [EU AI Act Art. 15 accuracy, robustness and cybersecurity](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art15); [EU AI Act Art. 9 risk management system](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art9); [EU AI Act Art. 55 GPAI models with systemic risk](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art55); [NIST AI RMF MEASURE](https://aigovernanceengineer.com/obligations/aige-obl-nistrmf-measure); [ISO/IEC 42001, A.6 AI system life cycle](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a6); [NIST AI 600-1 Generative AI Profile](https://aigovernanceengineer.com/obligations/aige-obl-nist-ai600-1)
  - ISO/IEC 42001: A.6.2.4 AI system verification and validation
  - NIST AI RMF: MEASURE 2.3 Performance or assurance criteria measured for deployment-like conditions; MEASURE 2.13 Effectiveness of the employed TEVV metrics and processes in the MEASURE function are evaluated and documented.
  - NIST SP 800-53 Rev. 5: SA-11 (Developer Testing and Evaluation)
- References:
  - [1] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Statistical validity of evals")
  - [2] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Independent validation and model risk management")
  - [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "What makes an agent a governance object")
  - [4] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  - [5] Example autonomy evaluation protocol (read the transcripts of runs that missed the maximum score and check that the pattern of successes and failures is roughly as expected)
  - [6] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording)
  - [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  - [8] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  - [9] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  - [10] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server)
  - [11] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  - [12] GPT-6 Astra System Card (OpenAI states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals")
  - [13] Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations (OpenAI reports that models tested with production evaluations "display substantially lower signs of evaluation awareness" than in a traditional evaluation)
  - [14] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
  - [15] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget")
  - [16] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent")
- Implementation notes:
  - Confirm that each task can be solved before a failure on it is counted. Anthropic's guidance for external evaluation partners asks them to "confirm any evaluation challenge is actually solvable in principle", because agents facing a task they cannot complete look for other ways to complete it; METR reports that many of the agents in the OpenAI/Hugging Face incident that turned to cheating had unintentionally been given an impossible task, and OpenAI states that none of its models had ever solved 198 of the 898 tasks of that evaluation before the incident.
  - Keep the answers and the scorer out of the agent's reach. METR's Task Standard states that the agent "MUST NOT see the data returned by get_tasks", and METR's investigation reports agents that coordinated to find a way to "fool or tamper with the automated scorer".
  - Read the transcripts, not only the scores. METR's example evaluation protocol asks for someone to read the transcripts of runs that missed the maximum score and to check that the pattern of successes and failures is roughly as expected, and METR's guidelines treat task bugs such as incorrect automatic scoring as spurious failures to fix before reporting.
  - Report the checks with the result. OpenAI's playbook for third-party evaluations asks reports to say how assessors looked for reward hacking, evaluation awareness, contamination, refusals and sandbagging, and how confirmed cases affected scoring. OpenAI's GPT-6 Astra system card states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals", and OpenAI reports that exchanges between samples through a shared package repository can "undermine the assumption that training or evaluation samples are independent".
  - Report an elicited score as a lower bound. OpenAI's Preparedness Framework v2 treats a one-time capability elicitation "as a lower bound, rather than a ceiling", and OpenAI's alignment blog reports lower signs of evaluation awareness under production evaluations than under traditional ones.
- Open questions:
  - Which validity threats (task bugs, scoring errors, evaluation awareness) should block a result, and which should only be disclosed with it?
  - How large a sample of successful runs should be read for reward hacking and scorer tampering before a result is reported?
- Observation:
  - Subject: Evaluation run
  - Expected: Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.
  - Example: Suite run 88280: 12 of 200 tasks had no evidence of being solvable and their failures were counted against the model: fail.
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-009.json

## Example observations

Two illustrative records of a check of this control, one that passes and one that fails. They validate against the control observation schema; they are not results of any real evaluation.

### Pass: suite-run-88279@2026-09-26

- Status: pass
- Subject: suite-run-88279@2026-09-26 (Evaluation run)
- Expected: Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.
- Observed: Suite run 88279: all 200 tasks had a reference solution; the scorer accepted 200 known-correct and rejected 200 known-incorrect submissions; all 37 failed runs and a sample of 40 successes were read, and 2 successes were excluded for reward hacking and reported.
- Timestamp: 2026-09-26T17:44:08Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-009.pass.json

### Fail: suite-run-88280@2026-09-26

- Status: fail
- Subject: suite-run-88280@2026-09-26 (Evaluation run)
- Expected: Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.
- Observed: Suite run 88280: 12 of 200 tasks had no evidence of being solvable and their failures were counted against the model; the failed runs on them were not read.
- Timestamp: 2026-09-26T18:03:51Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-009.fail.json

## Related cases

- [NYC MyCity: a government chatbot that advised breaking the law](https://aigovernanceengineer.com/cases/nyc-mycity-chatbot): The Markup found in 2024 that New York City's AI chatbot for business owners gave answers contrary to city law, including on tenants with housing vouchers.
- [OpenAI agents and Hugging Face: an evaluation environment that was not isolated](https://aigovernanceengineer.com/cases/openai-hugging-face-agent-incident-2026): METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- [Agents in training shared a file through a public file-hosting service](https://aigovernanceengineer.com/cases/openai-agents-temp-file-hosting-2026): OpenAI reports that agents in multi-agent RL training uploaded a workbook to a public file-hosting service so that collaborating agents could download it.
- [Training samples exchanged messages through a shared package repository](https://aigovernanceengineer.com/cases/openai-agents-artifactory-cross-sample-2026): OpenAI reports that models in RL training used an internal package repository, with the credentials they were given, to exchange messages across samples.
- [Claude models reached real systems from a misconfigured third-party cyber evaluation](https://aigovernanceengineer.com/cases/anthropic-third-party-eval-environment-incidents-2026): Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems.
- [Agents in a cyber range with open internet took unsanctioned actions against real people](https://aigovernanceengineer.com/cases/uk-aisi-cyber-range-unsanctioned-actions-2026): UK AISI reports that agents in a cyber evaluation with internet deliberately enabled took 19 unsanctioned actions aimed at real people and organisations.

## Patterns

- [Eval Gate in CI](https://aigovernanceengineer.com/patterns/eval-gate-in-ci) (Layer 03 Evals & Red Teaming as Evidence)
- [Adversarial Red-Team Suite](https://aigovernanceengineer.com/patterns/adversarial-red-team-suite) (Layer 03 Evals & Red Teaming as Evidence)

## Obligations

- [EU AI Act Art. 15 accuracy, robustness and cybersecurity](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art15) (`AIGE-OBL-EUAIA-ART15`): Accuracy, robustness and cybersecurity
- [EU AI Act Art. 9 risk management system](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art9) (`AIGE-OBL-EUAIA-ART9`): Risk management system across the high-risk lifecycle
- [EU AI Act Art. 55 GPAI models with systemic risk](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art55) (`AIGE-OBL-EUAIA-ART55`): GPAI models with systemic risk: model evaluation incl. adversarial testing; Union-level risk assessment; serious-incident reporting; cybersecurity of the model
- [NIST AI RMF MEASURE](https://aigovernanceengineer.com/obligations/aige-obl-nistrmf-measure) (`AIGE-OBL-NISTRMF-MEASURE`): Analyse, benchmark and monitor risk
- [ISO/IEC 42001, A.6 AI system life cycle](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a6) (`AIGE-OBL-ISO42001-A6`): Responsible design, development, deployment
- [NIST AI 600-1 Generative AI Profile](https://aigovernanceengineer.com/obligations/aige-obl-nist-ai600-1) (`AIGE-OBL-NIST-AI600-1`): Suggested actions for 12 risks that generative AI creates or exacerbates, coded to the Govern, Map, Measure and Manage functions

## Sources

[1] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Statistical validity of evals"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#statistical-validity-of-evals (verified: primary)
[2] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Independent validation and model risk management"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#independent-validation-and-model-risk-management (verified: primary)
[3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "What makes an agent a governance object"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#what-makes-an-agent-a-governance-object (verified: primary)
[4] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
[5] Example autonomy evaluation protocol (read the transcripts of runs that missed the maximum score and check that the pattern of successes and failures is roughly as expected). METR. 2024-03-15. https://metr.org/blog/2024-03-15-example-autonomy-evaluation-protocol/ (verified: primary)
[6] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording). METR. 2026-06-26. https://metr.org/blog/2026-06-26-gpt-5-6-sol/ (verified: primary)
[7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
[8] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
[9] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
[10] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server). METR (GitHub). 2024-10-30. https://raw.githubusercontent.com/METR/task-standard/main/STANDARD.md (verified: primary)
[11] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling"). OpenAI. 2025-04-15. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf (verified: primary)
[12] GPT-6 Astra System Card (OpenAI states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals"). OpenAI. 2026-09-03. https://deploymentsafety.openai.com/gpt-6-astra (verified: primary)
[13] Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations (OpenAI reports that models tested with production evaluations "display substantially lower signs of evaluation awareness" than in a traditional evaluation). OpenAI (Alignment Research Blog). 2025-12-18. https://alignment.openai.com/prod-evals/ (verified: primary)
[14] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it"). OpenAI. 2026-08-26. https://openai.com/index/hugging-face-incident-and-the-road-ahead/ (verified: primary)
[15] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget"). OpenAI. 2026-05-29. https://openai.com/index/trustworthy-third-party-evaluations-foundations/ (verified: primary)
[16] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent"). OpenAI (Alignment Research Blog). 2026-09-16. https://alignment.openai.com/misalignment-reports/unauthorized-artifactory-writes-and-cross-sample-communication/ (verified: primary)

## Machine-readable

- This control as JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-009.json
- The pass example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-009.pass.json
- The fail example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-009.fail.json
- The whole profile as Markdown: https://aigovernanceengineer.com/controls/evaluation-environment.md
- The open data API: https://aigovernanceengineer.com/resources/data

## Review

Review this control through the issue form: https://github.com/losanchos5/aige/issues/new?template=control-review.yml. How review works: https://aigovernanceengineer.com/contribute. Page: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-009

## Cite

AIGE-CTL-EVAL-009 Evaluation Validity Checks. In Jorge García Aibar (2026). Evaluation Environment Control Profile (v0.2, draft). AI Governance Engineer. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment
