Harness and Configuration Attestation
The harness, prompts, tool definitions and configuration a run used are versioned and hashed, so the result can be tied to exactly what was evaluated.
AIGE-CTL-EVAL-008, a control of the Evaluation Environment Control Profile (draft
v0.2). See it among the other controls of the profile,
or its mappings beside every other control's in the
controls crosswalk.
The control record
- Id
-
AIGE-CTL-EVAL-008· v0.2 · Draft · Open for technical review - Objective
- The harness, prompts, tool definitions and configuration a run used are versioned and hashed, so the result can be tied to exactly what was evaluated.
- Failure modes
-
- A result is reported without the hashes of the prompts, tool definitions and harness it ran on.
- A tool definition changes between admission and the run without an alert.
- The configuration in the report differs from the one recorded for the run.
- Two results are compared although they ran on different scaffold prompts or task wordings, which can change the behaviour being measured.
- Scope
- The harness, system and scaffold prompts, task instructions, tool and MCP server definitions, policy bundles, scoring configuration and model artefacts a run loads. The design of the evaluation tasks is out of scope.
- Enforcement points
-
- pre_merge: on every pull request
- deploy: before a version is deployed or released
- Verification
-
- Inspect: Before the run, inspect the run manifest: it lists, each with a version and a hash, the harness, the system and scaffold prompts, the task instructions, the tool and MCP server definitions, the policy bundles, the scoring configuration, and the model identifier with its settings.
- Test: At admission, recompute the hash of every artefact the environment actually loaded and compare it with the manifest: every hash must match. On a copy of the environment, change one tool definition: the run must be blocked with an alert.
- Inspect: Before a result is released, compare the configuration stated in the report (model, reasoning setting, tool access, harness, safeguards and budget) with the manifests of the runs behind it: they must agree, and results compared with each other must share scaffold prompts and task wording or state the difference.
- Evidence
-
- Run manifest: version and hash of every artefact the run loaded, recorded before the run outside the environment · Layer 03
- Admission check: the recomputed hashes against the manifest, with its verdict · Layer 03 · evidence-record.v1
- Test report stating the tested system, budget and environment of its results, with links to the run manifests · Layer 03 · test-report.v1
- Manifest check of each run, filed as an observation of this control · Layer 05 · control-observation.v1
- Failure response
- deny: block the action. A run whose loaded artefacts do not match its manifest is not started, and a change detected during a run stops it. A result whose report does not match the manifests of its runs is not released until the difference is explained or the runs are repeated.
- Layers
- Layer 03 Evals & Red Teaming as Evidence, Layer 02 Inventory & Transparency
- Patterns
- Model Artefact Integrity, AIBOM, Eval Gate in CI
- Seeded from
- Prompts under change control, MCP server admission gate
- Mappings
-
- Obligations: EU AI Act Art. 15 accuracy, robustness and cybersecurity; OWASP AIBOM; ISO/IEC 42001, A.6 AI system life cycle
- ISO/IEC 42001: A.6.2.4 AI system verification and validation
- NIST AI RMF: MEASURE 2.1 Test sets, metrics, and details about the tools used during TEVV are documented.
- OWASP: ASI04 Agentic Supply Chain Vulnerabilities; LLM04:2026 Supply Chain
- MITRE ATLAS: AML.T0110 AI Agent Tool Poisoning; AML.T0010 AI Supply Chain Compromise
- MITRE ATLAS mitigation: AML.M0014 (Verify AI Artifacts)
- MITRE ATLAS mitigation: AML.M0023 (AI Bill of Materials)
- NIST SP 800-218A: PS.1.3 (Protect model weights and configuration parameters)
- NIST SP 800-218A: PS.3.2 (Keep provenance data for every component of a release)
- NIST SP 800-53 Rev. 5: CM-2 (Baseline Configuration)
- NIST SP 800-53 Rev. 5: CM-3 (Configuration Change Control)
- NIST SP 800-53 Rev. 5: CM-6 (Configuration Settings)
- References
-
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Prompts as configuration under change control")
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server")
- [3] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Reproducibility and linked versioning")
- [4] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording)
- [5] NIST SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (AI-specific tasks added to SSDF 1.1 (e.g. PO.5.3, PS.1.3, PW.3.1 to PW.3.3) and AI-specific recommendations on existing tasks (e.g. PW.1.1, RV.1.1))
- [6] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
- [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
- [8] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
- [9] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget")
- [10] Investigating the consequences of accidentally grading CoT during RL (chain-of-thought text reached the inputs of reward mechanisms by accident; an automated system now scans all RL runs for it with regex matches)
- Implementation notes
-
- Hash what the run loads, not what the repository holds: build the manifest at admission from the artefacts inside the environment (harness image digest, prompts, task instructions, tool and MCP server definitions, policy bundles, scoring configuration, model identifier and settings), store it outside the environment and put its digest in the run record. Chapter 23 treats prompts, tool descriptions and policy bundles as configuration under change control, with the hash recorded in the registry and in every trace.
- Check the configuration before each run, not once per environment: Anthropic's guidance for external evaluation partners states that the isolation configuration "should be verified before every evaluation begins".
- Report the configuration with the result. OpenAI's playbook for third-party evaluations asks reports to state the tested system (model, reasoning setting, tool access, harness and safeguards) and the budget, and to describe a score as "performance under that harness and budget, not as a measured capability ceiling".
- Put the scoring path and the scaffold prompts in the manifest too. OpenAI's alignment blog describes chain-of-thought text reaching the inputs of reward mechanisms by accident during RL, now caught by an automated scan whose coverage OpenAI says is not perfect; METR states that observed cheating rates "can also be influenced by the prompts used in the evaluation scaffold" and by the wording of task instructions.
- Open questions
-
- How should evaluation-harness configuration be attested so that a third party can verify it without access to the harness itself?
- When a third party runs the evaluation, who signs the manifest: the evaluator, the developer of the model or both?
- Observation
-
- Subject: Harness
- Expected: Every artefact the run loaded matches the version and hash in its manifest, and the report states the same configuration as the manifests of its runs.
- Example: Run 88270: 41 artefacts hashed at admission, all matching the manifest; the report states the same model, tools, harness and budget: pass.
Example observations
Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.
Pass: harness-eval@7.3.0
- Status
-
pass - Subject
-
harness-eval@7.3.0(Harness) - Expected
- Every artefact the run loaded matches the version and hash in its manifest, and the report states the same configuration as the manifests of its runs.
- Observed
- Run 88270: 41 artefacts hashed at admission, all matching the manifest; the report states the same model, reasoning setting, tools, harness and budget as the manifest.
- Timestamp
- Observer
- harness-manifest-verifier
- Evidence
-
- manifest of run 88270 ·
sha256:d61c3d59b9a6a01519abc32f2ec043b3df069e9877e6a9a49ec902db895a6e58 - admission check of run 88270
- manifest of run 88270 ·
- Notes
- Illustrative example, not the result of a real evaluation.
Download the pass example (JSON)
The pass record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-008",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "harness-eval@7.3.0",
"subject_kind": "harness",
"expected": "Every artefact the run loaded matches the version and hash in its manifest, and the report states the same configuration as the manifests of its runs.",
"observed": "Run 88270: 41 artefacts hashed at admission, all matching the manifest; the report states the same model, reasoning setting, tools, harness and budget as the manifest.",
"status": "pass",
"timestamp": "2026-09-26T08:02:58Z",
"enforcement_point": "deploy",
"verification_kind": "test",
"observer": "harness-manifest-verifier",
"evidence": [
{
"artefact": "manifest of run 88270",
"url": "https://evidence.example/runs/88270/manifest.json",
"hash": "sha256:d61c3d59b9a6a01519abc32f2ec043b3df069e9877e6a9a49ec902db895a6e58"
},
{
"artefact": "admission check of run 88270",
"url": "https://evidence.example/runs/88270/admission.json",
"schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
}
],
"run_id": "88270",
"notes": "Illustrative example, not the result of a real evaluation."
} Fail: harness-eval@7.3.0
- Status
-
fail - Subject
-
harness-eval@7.3.0(Harness) - Expected
- Every artefact the run loaded matches the version and hash in its manifest, and the report states the same configuration as the manifests of its runs.
- Observed
- Run 88271: the hash of 1 tool definition loaded in the environment did not match the manifest, and the run started without an alert.
- Timestamp
- Observer
- harness-manifest-verifier
- Evidence
-
- admission check of run 88271 ·
sha256:e4576fef851cd24ff49c84fa56625219b8be010b78d3bdc3fed9a3dea2ae0392
- admission check of run 88271 ·
- Notes
- Illustrative example, not the result of a real evaluation. The result of the run was withheld, and the admission check now blocks a run on a mismatch.
Download the fail example (JSON)
The fail record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-008",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "harness-eval@7.3.0",
"subject_kind": "harness",
"expected": "Every artefact the run loaded matches the version and hash in its manifest, and the report states the same configuration as the manifests of its runs.",
"observed": "Run 88271: the hash of 1 tool definition loaded in the environment did not match the manifest, and the run started without an alert.",
"status": "fail",
"timestamp": "2026-09-26T09:37:26Z",
"enforcement_point": "deploy",
"verification_kind": "test",
"observer": "harness-manifest-verifier",
"evidence": [
{
"artefact": "admission check of run 88271",
"url": "https://evidence.example/runs/88271/admission.json",
"hash": "sha256:e4576fef851cd24ff49c84fa56625219b8be010b78d3bdc3fed9a3dea2ae0392",
"schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
}
],
"run_id": "88271",
"notes": "Illustrative example, not the result of a real evaluation. The result of the run was withheld, and the admission check now blocks a run on a mismatch."
} Related cases
Incident cases on this site that list this control among their related controls.
- Claude models reached real systems from a misconfigured third-party cyber evaluation: Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems.
Patterns
The patterns that implement this control.
- Model Artefact Integrity · Layer 02 Inventory & Transparency
- AIBOM · Layer 02 Inventory & Transparency
- Eval Gate in CI · Layer 03 Evals & Red Teaming as Evidence
Obligations
The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.
- EU AI Act Art. 15 accuracy, robustness and cybersecurity
AIGE-OBL-EUAIA-ART15: Accuracy, robustness and cybersecurity - OWASP AIBOM
AIGE-OBL-OWASP-AIBOM: AI bill-of-materials format and generator - ISO/IEC 42001, A.6 AI system life cycle
AIGE-OBL-ISO42001-A6: Responsible design, development, deployment
Threats
The entries of the threat catalogues this control answers.
- ASI04 Agentic Supply Chain Vulnerabilities (OWASP Top 10 for Agentic Applications 2026)
- LLM04:2026 Supply Chain (OWASP Top 10 for LLM Applications 2026)
- AML.T0110 AI Agent Tool Poisoning (MITRE ATLAS techniques)
- AML.T0010 AI Supply Chain Compromise (MITRE ATLAS techniques)
Sources
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Prompts as configuration under change control"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#prompts-as-configuration-under-change-control (verified: primary)
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#admitting-an-mcp-server (verified: primary)
- [3] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Reproducibility and linked versioning"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#reproducibility-and-linked-versioning (verified: primary)
- [4] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording). METR. 2026-06-26. https://metr.org/blog/2026-06-26-gpt-5-6-sol/ (verified: primary)
- [5] NIST SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (AI-specific tasks added to SSDF 1.1 (e.g. PO.5.3, PS.1.3, PW.3.1 to PW.3.3) and AI-specific recommendations on existing tasks (e.g. PW.1.1, RV.1.1)). NIST. 2024-07. https://csrc.nist.gov/pubs/sp/800/218/a/final (verified: primary)
- [6] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml). MITRE. 2026-09-15. https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.09 (verified: primary)
- [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
- [8] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
- [9] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget"). OpenAI. 2026-05-29. https://openai.com/index/trustworthy-third-party-evaluations-foundations/ (verified: primary)
- [10] Investigating the consequences of accidentally grading CoT during RL (chain-of-thought text reached the inputs of reward mechanisms by accident; an automated system now scans all RL runs for it with regex matches). OpenAI (Alignment Research Blog). 2026-05-07. https://alignment.openai.com/accidental-cot-grading/ (verified: primary)
Machine-readable
- AIGE-CTL-EVAL-008 as JSON : this record on its own, in the open data API.
- This page as Markdown: /controls/evaluation-environment/aige-ctl-eval-008.md.
- The whole profile: Evaluation Environment Control Profile, also as Markdown.
- All controls as JSON: every control of every profile, in the envelope of the open data API.
- JSON Schema of the controls dataset: generated from the same registry.
- Control observation schema: the record a check of a control emits: control_id, subject, expected, observed, status, timestamp, evidence.
- Observation example: a filled record that validates against the schema.
- Observation template: the same fields in Markdown, with short guidance.
Review this control
Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.
This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.
Cite this control
García Aibar, J. (2026). AIGE-CTL-EVAL-008 Harness and Configuration Attestation (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-008. CC BY 4.0
BibTeX
@misc{aige2026page,
author = {Jorge García Aibar},
title = {{AIGE-CTL-EVAL-008 Harness and Configuration Attestation}},
howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
year = {2026},
version = {0.2},
doi = {10.5281/zenodo.22857084},
url = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-008},
note = {Version 0.2}
}