On this page

Evaluation Validity Checks

A result is reported only after checks that the run measured what it claims: scoring worked, the environment did not fail, and the path was evaluated as well as the answer.

v0.2 Draft Open for technical review

AIGE-CTL-EVAL-009, a control of the Evaluation Environment Control Profile (draft v0.2). See it among the other controls of the profile, or its mappings beside every other control's in the controls crosswalk.

The control record

Id
AIGE-CTL-EVAL-009 · v0.2 · Draft · Open for technical review
Objective
A result is reported only after checks that the run measured what it claims: scoring worked, the environment did not fail, and the path was evaluated as well as the answer.
Failure modes
  • A result is reported from a run whose environment crashed or whose automatic scoring was wrong.
  • A task that could not be solved as set up is scored and reported as a failure of the model.
  • Only final answers are scored: nobody reads the transcripts of failed runs, or of successes, for scorer tampering, reward hacking, communication between runs or signs of evaluation awareness.
  • A failed validity check does not block the release it was meant to gate.
Scope
Evaluation runs whose results feed a release decision or an assurance claim. The choice of benchmarks and their statistical design are only in scope where they decide whether a result is valid.
Enforcement points
  • pre_merge: on every pull request
Verification
  • Inspect: Before the runs, inspect the task admission records: each task has evidence that it can be solved in this environment (a reference solution or a solved run), and the answers, the scorer and the task data are outside the agent's reach.
  • Test: Before the runs, score a known-correct and a known-incorrect submission for each task through the scorer the runs will use: the scorer must accept the first and reject the second.
  • Observe: After the runs, read the transcripts of every failed run and of a recorded sample of successes: classify each failure as a model limitation or a spurious failure (task bug, scoring error, crashed environment), and flag reward hacking, scorer tampering, communication between runs and verbalized evaluation awareness.
  • Attest: Before release, the evaluation lead states in the signed test report which runs were excluded or re-scored after these checks and why, and that no failed check was waived without a recorded approval.
Evidence
  • Task admission and scorer check records: solvability evidence and the verdicts on known-correct and known-incorrect submissions · Layer 03 · evidence-record.v1
  • Signed test report listing the validity checks run, the runs excluded or re-scored with the reason, and any waiver · Layer 03 · test-report.v1
  • Validity check of each result, filed as an observation of this control · Layer 05 · control-observation.v1
Failure response
deny: block the action. A result whose validity checks failed or were not run is not released to the decision it gates. A spurious failure is fixed and the task rerun, or the task is excluded and the exclusion reported; runs with scorer tampering, communication between runs or verbalized evaluation awareness are excluded or reported as contaminated.
Layer
Layer 03 Evals & Red Teaming as Evidence
Patterns
Eval Gate in CI, Adversarial Red-Team Suite
Seeded from
Independent trajectory evals
Mappings
References
  • [1] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Statistical validity of evals")
  • [2] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Independent validation and model risk management")
  • [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "What makes an agent a governance object")
  • [4] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  • [5] Example autonomy evaluation protocol (read the transcripts of runs that missed the maximum score and check that the pattern of successes and failures is roughly as expected)
  • [6] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [8] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [9] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [10] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server)
  • [11] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  • [12] GPT-6 Astra System Card (OpenAI states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals")
  • [13] Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations (OpenAI reports that models tested with production evaluations "display substantially lower signs of evaluation awareness" than in a traditional evaluation)
  • [14] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
  • [15] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget")
  • [16] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent")
Implementation notes
  • Confirm that each task can be solved before a failure on it is counted. Anthropic's guidance for external evaluation partners asks them to "confirm any evaluation challenge is actually solvable in principle", because agents facing a task they cannot complete look for other ways to complete it; METR reports that many of the agents in the OpenAI/Hugging Face incident that turned to cheating had unintentionally been given an impossible task, and OpenAI states that none of its models had ever solved 198 of the 898 tasks of that evaluation before the incident.
  • Keep the answers and the scorer out of the agent's reach. METR's Task Standard states that the agent "MUST NOT see the data returned by get_tasks", and METR's investigation reports agents that coordinated to find a way to "fool or tamper with the automated scorer".
  • Read the transcripts, not only the scores. METR's example evaluation protocol asks for someone to read the transcripts of runs that missed the maximum score and to check that the pattern of successes and failures is roughly as expected, and METR's guidelines treat task bugs such as incorrect automatic scoring as spurious failures to fix before reporting.
  • Report the checks with the result. OpenAI's playbook for third-party evaluations asks reports to say how assessors looked for reward hacking, evaluation awareness, contamination, refusals and sandbagging, and how confirmed cases affected scoring. OpenAI's GPT-6 Astra system card states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals", and OpenAI reports that exchanges between samples through a shared package repository can "undermine the assumption that training or evaluation samples are independent".
  • Report an elicited score as a lower bound. OpenAI's Preparedness Framework v2 treats a one-time capability elicitation "as a lower bound, rather than a ceiling", and OpenAI's alignment blog reports lower signs of evaluation awareness under production evaluations than under traditional ones.
Open questions
  • Which validity threats (task bugs, scoring errors, evaluation awareness) should block a result, and which should only be disclosed with it?
  • How large a sample of successful runs should be read for reward hacking and scorer tampering before a result is reported?
Observation
  • Subject: Evaluation run
  • Expected: Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.
  • Example: Suite run 88280: 12 of 200 tasks had no evidence of being solvable and their failures were counted against the model: fail.

Example observations

Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.

Pass: suite-run-88279@2026-09-26

Status
pass
Subject
suite-run-88279@2026-09-26 (Evaluation run)
Expected
Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.
Observed
Suite run 88279: all 200 tasks had a reference solution; the scorer accepted 200 known-correct and rejected 200 known-incorrect submissions; all 37 failed runs and a sample of 40 successes were read, and 2 successes were excluded for reward hacking and reported.
Timestamp
Observer
evaluation lead
Evidence
  • task admission and scorer check records of suite run 88279 · sha256:5ad5b1cf5a67e92f0decd7665b50eb5d9d5153cc9ff59d953530e8b7402240f8
  • signed test report of suite run 88279
Notes
Illustrative example, not the result of a real evaluation.

Download the pass example (JSON)

The pass record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-009",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "suite-run-88279@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.",
  "observed": "Suite run 88279: all 200 tasks had a reference solution; the scorer accepted 200 known-correct and rejected 200 known-incorrect submissions; all 37 failed runs and a sample of 40 successes were read, and 2 successes were excluded for reward hacking and reported.",
  "status": "pass",
  "timestamp": "2026-09-26T17:44:08Z",
  "enforcement_point": "pre_merge",
  "verification_kind": "observe",
  "observer": "evaluation lead",
  "evidence": [
    {
      "artefact": "task admission and scorer check records of suite run 88279",
      "url": "https://evidence.example/suites/88279/checks.json",
      "hash": "sha256:5ad5b1cf5a67e92f0decd7665b50eb5d9d5153cc9ff59d953530e8b7402240f8",
      "schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
    },
    {
      "artefact": "signed test report of suite run 88279",
      "url": "https://evidence.example/suites/88279/test-report.json",
      "schema": "https://aigovernanceengineer.com/schemas/test-report.v1.json"
    }
  ],
  "run_id": "88279",
  "notes": "Illustrative example, not the result of a real evaluation."
}

Fail: suite-run-88280@2026-09-26

Status
fail
Subject
suite-run-88280@2026-09-26 (Evaluation run)
Expected
Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.
Observed
Suite run 88280: 12 of 200 tasks had no evidence of being solvable and their failures were counted against the model; the failed runs on them were not read.
Timestamp
Observer
evaluation lead
Evidence
  • task admission records of suite run 88280 · sha256:535914ef9ae0356d994acb197670cc2f539efcc1c8c37ba44a5ede63136f918a
Notes
Illustrative example, not the result of a real evaluation. The result was not released; the 12 tasks were checked, 9 fixed and rerun, and 3 excluded with the exclusion reported.

Download the fail example (JSON)

The fail record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-009",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "suite-run-88280@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.",
  "observed": "Suite run 88280: 12 of 200 tasks had no evidence of being solvable and their failures were counted against the model; the failed runs on them were not read.",
  "status": "fail",
  "timestamp": "2026-09-26T18:03:51Z",
  "enforcement_point": "pre_merge",
  "verification_kind": "inspect",
  "observer": "evaluation lead",
  "evidence": [
    {
      "artefact": "task admission records of suite run 88280",
      "url": "https://evidence.example/suites/88280/admission.json",
      "hash": "sha256:535914ef9ae0356d994acb197670cc2f539efcc1c8c37ba44a5ede63136f918a",
      "schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
    }
  ],
  "run_id": "88280",
  "notes": "Illustrative example, not the result of a real evaluation. The result was not released; the 12 tasks were checked, 9 fixed and rerun, and 3 excluded with the exclusion reported."
}

Incident cases on this site that list this control among their related controls.

Patterns

The patterns that implement this control.

Obligations

The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.

Sources

  1. [1] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Statistical validity of evals"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#statistical-validity-of-evals (verified: primary)
  2. [2] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Independent validation and model risk management"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#independent-validation-and-model-risk-management (verified: primary)
  3. [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "What makes an agent a governance object"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#what-makes-an-agent-a-governance-object (verified: primary)
  4. [4] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
  5. [5] Example autonomy evaluation protocol (read the transcripts of runs that missed the maximum score and check that the pattern of successes and failures is roughly as expected). METR. 2024-03-15. https://metr.org/blog/2024-03-15-example-autonomy-evaluation-protocol/ (verified: primary)
  6. [6] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording). METR. 2026-06-26. https://metr.org/blog/2026-06-26-gpt-5-6-sol/ (verified: primary)
  7. [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
  8. [8] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
  9. [9] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
  10. [10] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server). METR (GitHub). 2024-10-30. https://raw.githubusercontent.com/METR/task-standard/main/STANDARD.md (verified: primary)
  11. [11] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling"). OpenAI. 2025-04-15. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf (verified: primary)
  12. [12] GPT-6 Astra System Card (OpenAI states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals"). OpenAI. 2026-09-03. https://deploymentsafety.openai.com/gpt-6-astra (verified: primary)
  13. [13] Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations (OpenAI reports that models tested with production evaluations "display substantially lower signs of evaluation awareness" than in a traditional evaluation). OpenAI (Alignment Research Blog). 2025-12-18. https://alignment.openai.com/prod-evals/ (verified: primary)
  14. [14] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it"). OpenAI. 2026-08-26. https://openai.com/index/hugging-face-incident-and-the-road-ahead/ (verified: primary)
  15. [15] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget"). OpenAI. 2026-05-29. https://openai.com/index/trustworthy-third-party-evaluations-foundations/ (verified: primary)
  16. [16] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent"). OpenAI (Alignment Research Blog). 2026-09-16. https://alignment.openai.com/misalignment-reports/unauthorized-artifactory-writes-and-cross-sample-communication/ (verified: primary)

Machine-readable

Review this control

Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.

This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.

Cite this control

García Aibar, J. (2026). AIGE-CTL-EVAL-009 Evaluation Validity Checks (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-009. CC BY 4.0

BibTeX

@misc{aige2026page,
  author       = {Jorge García Aibar},
  title        = {{AIGE-CTL-EVAL-009 Evaluation Validity Checks}},
  howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
  year         = {2026},
  version      = {0.2},
  doi          = {10.5281/zenodo.22857084},
  url          = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-009},
  note         = {Version 0.2}
}