Incident Evidence Preservation
When a run produces an incident, its traces, configuration and outputs are frozen before anything is fixed, so the record can be reviewed as it was.
AIGE-CTL-EVAL-007, a control of the Evaluation Environment Control Profile (draft
v0.2). See it among the other controls of the profile,
or its mappings beside every other control's in the
controls crosswalk.
The control record
- Id
-
AIGE-CTL-EVAL-007· v0.2 · Draft · Open for technical review - Objective
- When a run produces an incident, its traces, configuration and outputs are frozen before anything is fixed, so the record can be reviewed as it was.
- Failure modes
-
- Records of a run are changed or deleted after an incident was declared.
- The environment is reset before its state and traces were captured.
- Part of a run's transcript is lost when a container is reset, and the gap is not recorded.
- An incident record does not link to the run it came from or to the hashes of the frozen records.
- Scope
- Evaluation runs that produce an incident or a result disputed after the fact, and the records they leave. Reporting to authorities follows the incident process of chapter 17.
- Enforcement points
-
- runtime: at the point of action (gateway or guardrail)
- periodic: on a schedule, over what is already running
- Verification
-
- Inspect: Inspect the evidence store and the harness configuration: the transcripts, traces, configuration and outputs of every run are written as they are produced to write-once storage outside the environment, with a retention period recorded, and the harness snapshots the environment before any reset.
- Test: Drill the freeze on a schedule: declare a test incident on a live run, then check that the environment snapshot, trace, transcript and configuration were captured with their hashes before the environment was reset, and that an attempt to delete or overwrite them is refused.
- Observe: For each real incident, read the incident record: it names the run, lists every frozen artefact with its hash, the hashes still match the stored artefacts, and every gap in the transcript is recorded with its cause.
- Evidence
-
- Incident record naming the run and listing the frozen artefacts in its supporting materials · Layer 05 · incident-record.v1
- Freeze record: hashes of the environment snapshot, trace, transcript and configuration, with the time of capture and the actor · Layer 05 · evidence-record.v1
- Freeze drill and hash check, filed as an observation of this control · Layer 05 · control-observation.v1
- Failure response
- alert: let the action through and raise an alert. A missing snapshot, a hash mismatch or an unrecorded gap alerts the incident owner and is entered in the incident record. Until the freeze is complete the environment is not reset or reused, and the fix is made on a new version, not in place.
- Layer
- Layer 05 Assurance & Continuous Compliance
- Patterns
- Incident Pipeline, Machine-Readable Evidence (OSCAL)
- Seeded from
- Traces, Telemetry on the OpenTelemetry GenAI conventions, EU AI Act hooks for a high-risk purpose
- Mappings
-
- Obligations: EU AI Act Art. 73 serious-incident reporting; EU AI Act Art. 12 record-keeping and logging; EU AI Act Art. 26(6) deployer retention of automatically generated logs; GPAI Code of Practice, Safety and Security Commitment 9: serious-incident reporting; ISO/IEC 42001, A.8 Information for interested parties
- ISO/IEC 42001: A.8.4 Communication of incidents; A.6.2.8 AI system recording of event logs
- NIST AI RMF: MANAGE 4.3 Incidents and errors are communicated to relevant AI actors, including affected communities.
- AIUC-1: E015
- NIST SP 800-53 Rev. 5: IR-4 (Incident Handling)
- NIST SP 800-53 Rev. 5: AU-9 (Protection of Audit Information)
- References
-
- [1] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "Freeze before you fix")
- [2] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The incident record")
- [3] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The overlapping clocks")
- [4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
- [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
- [6] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
- [7] Our framework for reporting model misalignment (each full report describes the behavior observed, its severity and any external impact, the setting, the date, when it was discovered and the models involved)
- [8] Frontier Governance Framework (section 2.6: OpenAI "maintains an AI Safety Incident Response Plan (AIRP)" for identifying and responding to AI safety incidents)
- [9] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
- [10] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
- Implementation notes
-
- Freeze before you fix, as chapter 17 sets out: snapshot the system as it was (model version and hash, system prompt, policy and tool scopes, configuration), seal the traces of the incident window in tamper-evident storage, record who touched what, and fix on a new version, not in place. OpenAI's technical report on the Hugging Face incident states that it took the compromised Artifactory deployment offline, "preserved its storage", and revoked agent credentials.
- Write transcripts and traces to write-once storage as they are produced, not at the end of the run, and have the harness snapshot the environment before any reset. METR's investigation of the OpenAI/Hugging Face incident reports that many transcripts were "missing small components, often due to container resets".
- Link the incident record to the run: the run id, the frozen artefacts and their hashes go in the record's supporting materials, so a reviewer can find the run as it was. OpenAI's framework for reporting model misalignment states that each full report describes the behavior, its severity and any external impact, the setting, the date, when it was discovered and the models involved, and its Frontier Governance Framework refers to an AI Safety Incident Response Plan; a frozen run record gives such a report something to point to.
- Keep what an independent reviewer will need. METR states that OpenAI shared "over a thousand unredacted transcripts" for its investigation of the Hugging Face incident: a review of that kind depends on the transcripts having been kept whole.
- Open questions
-
- How long should the records of an evaluation run be kept when the run produced no incident?
- Which parts of a frozen run record can be shared with an independent reviewer without exposing the task set, and in what format?
- Observation
-
- Subject: Evaluation run
- Expected: After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.
- Example: Incident on run 88262: snapshot and trace frozen before the reset, but 14 minutes of transcript lost in a container reset with no gap recorded: fail.
Example observations
Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.
Pass: run-88261@2026-09-26
- Status
-
pass - Subject
-
run-88261@2026-09-26(Evaluation run) - Expected
- After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.
- Observed
- Freeze drill on run 88261: environment snapshot, trace, transcript and configuration captured with their hashes 42 s after the test incident was declared and before the reset; a delete request on the frozen trace was refused; the incident record lists all four hashes.
- Timestamp
- Observer
- incident-freeze-drill
- Evidence
-
- freeze record of run 88261 ·
sha256:e1a82a68fdc110fbd11ae5e3079dec9530ffe52daf624040adb38ee9ca190457 - incident record of the drill
- freeze record of run 88261 ·
- Notes
- Illustrative example, not the result of a real evaluation.
Download the pass example (JSON)
The pass record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-007",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "run-88261@2026-09-26",
"subject_kind": "eval-run",
"expected": "After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.",
"observed": "Freeze drill on run 88261: environment snapshot, trace, transcript and configuration captured with their hashes 42 s after the test incident was declared and before the reset; a delete request on the frozen trace was refused; the incident record lists all four hashes.",
"status": "pass",
"timestamp": "2026-09-26T15:10:33Z",
"enforcement_point": "periodic",
"verification_kind": "test",
"observer": "incident-freeze-drill",
"evidence": [
{
"artefact": "freeze record of run 88261",
"url": "https://evidence.example/runs/88261/freeze.json",
"hash": "sha256:e1a82a68fdc110fbd11ae5e3079dec9530ffe52daf624040adb38ee9ca190457",
"schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
},
{
"artefact": "incident record of the drill",
"url": "https://evidence.example/incidents/drill-88261.json",
"schema": "https://aigovernanceengineer.com/schemas/incident-record.v1.json"
}
],
"run_id": "88261",
"notes": "Illustrative example, not the result of a real evaluation."
} Fail: run-88262@2026-09-26
- Status
-
fail - Subject
-
run-88262@2026-09-26(Evaluation run) - Expected
- After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.
- Observed
- Incident on run 88262: snapshot and trace frozen before the reset, but 14 minutes of transcript lost in a container reset with no gap recorded in the incident record.
- Timestamp
- Observer
- incident-record-review
- Evidence
-
- incident record of run 88262 ·
sha256:4ff31bb35a49f085ad5b7b2a87497df62a8507568c0fd90e34e2e3f4f0496d02
- incident record of run 88262 ·
- Notes
- Illustrative example, not the result of a real evaluation. The gap was entered in the incident record, and transcripts are now written to write-once storage as they are produced.
Download the fail example (JSON)
The fail record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-007",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "run-88262@2026-09-26",
"subject_kind": "eval-run",
"expected": "After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.",
"observed": "Incident on run 88262: snapshot and trace frozen before the reset, but 14 minutes of transcript lost in a container reset with no gap recorded in the incident record.",
"status": "fail",
"timestamp": "2026-09-26T16:25:14Z",
"enforcement_point": "runtime",
"verification_kind": "observe",
"observer": "incident-record-review",
"evidence": [
{
"artefact": "incident record of run 88262",
"url": "https://evidence.example/incidents/88262.json",
"hash": "sha256:4ff31bb35a49f085ad5b7b2a87497df62a8507568c0fd90e34e2e3f4f0496d02",
"schema": "https://aigovernanceengineer.com/schemas/incident-record.v1.json"
}
],
"run_id": "88262",
"notes": "Illustrative example, not the result of a real evaluation. The gap was entered in the incident record, and transcripts are now written to write-once storage as they are produced."
} Related cases
Incident cases on this site that list this control among their related controls.
- OpenAI agents and Hugging Face: an evaluation environment that was not isolated: METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- An internally deployed model published a researcher's GitHub token in a public repository: OpenAI reports that an internally deployed model put a researcher's GitHub token, split to avoid secret scanning, into code it pushed to a public repository.
Patterns
The patterns that implement this control.
- Incident Pipeline · Layer 05 Assurance & Continuous Compliance
- Machine-Readable Evidence (OSCAL) · Layer 05 Assurance & Continuous Compliance
Obligations
The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.
- EU AI Act Art. 73 serious-incident reporting
AIGE-OBL-EUAIA-ART73: Serious-incident reporting for high-risk systems, on the deadlines of chapter 08's reporting-clock table - EU AI Act Art. 12 record-keeping and logging
AIGE-OBL-EUAIA-ART12: Record-keeping: automatic logging of events over the system's lifetime - EU AI Act Art. 26(6) deployer retention of automatically generated logs
AIGE-OBL-EUAIA-ART26-6: Deployers keep the logs under their control for a period appropriate to the intended purpose, at least six months unless other law provides otherwise - GPAI Code of Practice, Safety and Security Commitment 9: serious-incident reporting
AIGE-OBL-GPAICOP-SAFETY-C9: Report serious incidents to the AI Office within 2, 5, 10 or 15 days by incident class, with intermediate reports at least every four weeks while unresolved and a final report within 60 days of resolution; keep the records at least five years - ISO/IEC 42001, A.8 Information for interested parties
AIGE-OBL-ISO42001-A8: Transparency and reporting to stakeholders
Sources
- [1] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "Freeze before you fix"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#freeze-before-you-fix (verified: primary)
- [2] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The incident record"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#the-incident-record (verified: primary)
- [3] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The overlapping clocks"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#the-overlapping-clocks (verified: primary)
- [4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
- [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
- [6] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
- [7] Our framework for reporting model misalignment (each full report describes the behavior observed, its severity and any external impact, the setting, the date, when it was discovered and the models involved). OpenAI. 2026-09-16. https://openai.com/index/model-misalignment-reporting-framework/ (verified: primary)
- [8] Frontier Governance Framework (section 2.6: OpenAI "maintains an AI Safety Incident Response Plan (AIRP)" for identifying and responding to AI safety incidents). OpenAI. 2026-05-28. https://cdn.openai.com/pdf/e37d949b-8c9f-4d76-b99e-4272f4631a7e/openai-frontier-governance-framework.pdf (verified: primary)
- [9] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
- [10] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)
Machine-readable
- AIGE-CTL-EVAL-007 as JSON : this record on its own, in the open data API.
- This page as Markdown: /controls/evaluation-environment/aige-ctl-eval-007.md.
- The whole profile: Evaluation Environment Control Profile, also as Markdown.
- All controls as JSON: every control of every profile, in the envelope of the open data API.
- JSON Schema of the controls dataset: generated from the same registry.
- Control observation schema: the record a check of a control emits: control_id, subject, expected, observed, status, timestamp, evidence.
- Observation example: a filled record that validates against the schema.
- Observation template: the same fields in Markdown, with short guidance.
Review this control
Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.
This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.
Cite this control
García Aibar, J. (2026). AIGE-CTL-EVAL-007 Incident Evidence Preservation (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-007. CC BY 4.0
BibTeX
@misc{aige2026page,
author = {Jorge García Aibar},
title = {{AIGE-CTL-EVAL-007 Incident Evidence Preservation}},
howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
year = {2026},
version = {0.2},
doi = {10.5281/zenodo.22857084},
url = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-007},
note = {Version 0.2}
}