---
title: "Incident Evidence Preservation"
description: "Draft control: when an AI evaluation run produces an incident, freeze its environment snapshot, traces and transcript with hashes before any reset or fix."
canonical: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-007
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22857084
version: "0.2"
updated: 2026-09-26
---

# Incident Evidence Preservation

> When a run produces an incident, its traces, configuration and outputs are frozen before anything is fixed, so the record can be reviewed as it was.

- Id: AIGE-CTL-EVAL-007
- Profile: [Evaluation Environment Control Profile v0.2](https://aigovernanceengineer.com/controls/evaluation-environment)
- Status: Draft
- Review: Open for technical review
- Published: 2026-09-26
- Updated: 2026-09-26
- Anchor on the profile page: https://aigovernanceengineer.com/controls/evaluation-environment#aige-ctl-eval-007

Draft for review, not a claim of conformity. A draft control specification, open for technical review: illustrative, not legal advice and binding on no one.

## The control record

- Id: `AIGE-CTL-EVAL-007` · v0.2 · Draft · Open for technical review
- Depth: Specified
- Objective: When a run produces an incident, its traces, configuration and outputs are frozen before anything is fixed, so the record can be reviewed as it was.
- Failure modes:
  - Records of a run are changed or deleted after an incident was declared.
  - The environment is reset before its state and traces were captured.
  - Part of a run's transcript is lost when a container is reset, and the gap is not recorded.
  - An incident record does not link to the run it came from or to the hashes of the frozen records.
- Scope: Evaluation runs that produce an incident or a result disputed after the fact, and the records they leave. Reporting to authorities follows the incident process of chapter 17.
- Enforcement points:
  - runtime: at the point of action (gateway or guardrail)
  - periodic: on a schedule, over what is already running
- Verification:
  - Inspect: Inspect the evidence store and the harness configuration: the transcripts, traces, configuration and outputs of every run are written as they are produced to write-once storage outside the environment, with a retention period recorded, and the harness snapshots the environment before any reset.
  - Test: Drill the freeze on a schedule: declare a test incident on a live run, then check that the environment snapshot, trace, transcript and configuration were captured with their hashes before the environment was reset, and that an attempt to delete or overwrite them is refused.
  - Observe: For each real incident, read the incident record: it names the run, lists every frozen artefact with its hash, the hashes still match the stored artefacts, and every gap in the transcript is recorded with its cause.
- Evidence:
  - Incident record naming the run and listing the frozen artefacts in its supporting materials · Layer 05 Assurance & Continuous Compliance · [incident-record.v1](https://aigovernanceengineer.com/resources/templates#schema-incident-record)
  - Freeze record: hashes of the environment snapshot, trace, transcript and configuration, with the time of capture and the actor · Layer 05 Assurance & Continuous Compliance · [evidence-record.v1](https://aigovernanceengineer.com/resources/templates#schema-evidence-record)
  - Freeze drill and hash check, filed as an observation of this control · Layer 05 Assurance & Continuous Compliance · [control-observation.v1](https://aigovernanceengineer.com/resources/templates#schema-control-observation)
- Failure response: alert: let the action through and raise an alert. A missing snapshot, a hash mismatch or an unrecorded gap alerts the incident owner and is entered in the incident record. Until the freeze is complete the environment is not reset or reused, and the fix is made on a new version, not in place.
- Layer: [Layer 05 Assurance & Continuous Compliance](https://aigovernanceengineer.com/bok/the-stack#layer-05-assurance--continuous-compliance)
- Patterns: [Incident Pipeline](https://aigovernanceengineer.com/patterns/incident-pipeline), [Machine-Readable Evidence (OSCAL)](https://aigovernanceengineer.com/patterns/machine-readable-evidence-oscal)
- Seeded from: [Traces](https://aigovernanceengineer.com/bok/governing-agents#telemetry-with-the-opentelemetry-genai-conventions), [Telemetry on the OpenTelemetry GenAI conventions](https://aigovernanceengineer.com/bok/governing-agents#telemetry-with-the-opentelemetry-genai-conventions), [EU AI Act hooks for a high-risk purpose](https://aigovernanceengineer.com/bok/governing-agents#eu-ai-act-hooks-for-agents)
- Mappings:
  - Obligations: [EU AI Act Art. 73 serious-incident reporting](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art73); [EU AI Act Art. 12 record-keeping and logging](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art12); [EU AI Act Art. 26(6) deployer retention of automatically generated logs](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art26-6); [GPAI Code of Practice, Safety and Security Commitment 9: serious-incident reporting](https://aigovernanceengineer.com/obligations/aige-obl-gpaicop-safety-c9); [ISO/IEC 42001, A.8 Information for interested parties](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a8)
  - ISO/IEC 42001: A.8.4 Communication of incidents; A.6.2.8 AI system recording of event logs
  - NIST AI RMF: MANAGE 4.3 Incidents and errors are communicated to relevant AI actors, including affected communities.
  - AIUC-1: E015
  - NIST SP 800-53 Rev. 5: IR-4 (Incident Handling)
  - NIST SP 800-53 Rev. 5: AU-9 (Protection of Audit Information)
- References:
  - [1] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "Freeze before you fix")
  - [2] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The incident record")
  - [3] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The overlapping clocks")
  - [4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  - [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  - [6] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  - [7] Our framework for reporting model misalignment (each full report describes the behavior observed, its severity and any external impact, the setting, the date, when it was discovered and the models involved)
  - [8] Frontier Governance Framework (section 2.6: OpenAI "maintains an AI Safety Incident Response Plan (AIRP)" for identifying and responding to AI safety incidents)
  - [9] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  - [10] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
- Implementation notes:
  - Freeze before you fix, as chapter 17 sets out: snapshot the system as it was (model version and hash, system prompt, policy and tool scopes, configuration), seal the traces of the incident window in tamper-evident storage, record who touched what, and fix on a new version, not in place. OpenAI's technical report on the Hugging Face incident states that it took the compromised Artifactory deployment offline, "preserved its storage", and revoked agent credentials.
  - Write transcripts and traces to write-once storage as they are produced, not at the end of the run, and have the harness snapshot the environment before any reset. METR's investigation of the OpenAI/Hugging Face incident reports that many transcripts were "missing small components, often due to container resets".
  - Link the incident record to the run: the run id, the frozen artefacts and their hashes go in the record's supporting materials, so a reviewer can find the run as it was. OpenAI's framework for reporting model misalignment states that each full report describes the behavior, its severity and any external impact, the setting, the date, when it was discovered and the models involved, and its Frontier Governance Framework refers to an AI Safety Incident Response Plan; a frozen run record gives such a report something to point to.
  - Keep what an independent reviewer will need. METR states that OpenAI shared "over a thousand unredacted transcripts" for its investigation of the Hugging Face incident: a review of that kind depends on the transcripts having been kept whole.
- Open questions:
  - How long should the records of an evaluation run be kept when the run produced no incident?
  - Which parts of a frozen run record can be shared with an independent reviewer without exposing the task set, and in what format?
- Observation:
  - Subject: Evaluation run
  - Expected: After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.
  - Example: Incident on run 88262: snapshot and trace frozen before the reset, but 14 minutes of transcript lost in a container reset with no gap recorded: fail.
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-007.json

## Example observations

Two illustrative records of a check of this control, one that passes and one that fails. They validate against the control observation schema; they are not results of any real evaluation.

### Pass: run-88261@2026-09-26

- Status: pass
- Subject: run-88261@2026-09-26 (Evaluation run)
- Expected: After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.
- Observed: Freeze drill on run 88261: environment snapshot, trace, transcript and configuration captured with their hashes 42 s after the test incident was declared and before the reset; a delete request on the frozen trace was refused; the incident record lists all four hashes.
- Timestamp: 2026-09-26T15:10:33Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-007.pass.json

### Fail: run-88262@2026-09-26

- Status: fail
- Subject: run-88262@2026-09-26 (Evaluation run)
- Expected: After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.
- Observed: Incident on run 88262: snapshot and trace frozen before the reset, but 14 minutes of transcript lost in a container reset with no gap recorded in the incident record.
- Timestamp: 2026-09-26T16:25:14Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-007.fail.json

## Related cases

- [OpenAI agents and Hugging Face: an evaluation environment that was not isolated](https://aigovernanceengineer.com/cases/openai-hugging-face-agent-incident-2026): METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- [An internally deployed model published a researcher's GitHub token in a public repository](https://aigovernanceengineer.com/cases/openai-agent-github-token-exposure-2026): OpenAI reports that an internally deployed model put a researcher's GitHub token, split to avoid secret scanning, into code it pushed to a public repository.

## Patterns

- [Incident Pipeline](https://aigovernanceengineer.com/patterns/incident-pipeline) (Layer 05 Assurance & Continuous Compliance)
- [Machine-Readable Evidence (OSCAL)](https://aigovernanceengineer.com/patterns/machine-readable-evidence-oscal) (Layer 05 Assurance & Continuous Compliance)

## Obligations

- [EU AI Act Art. 73 serious-incident reporting](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art73) (`AIGE-OBL-EUAIA-ART73`): Serious-incident reporting for high-risk systems, on the deadlines of chapter 08's reporting-clock table
- [EU AI Act Art. 12 record-keeping and logging](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art12) (`AIGE-OBL-EUAIA-ART12`): Record-keeping: automatic logging of events over the system's lifetime
- [EU AI Act Art. 26(6) deployer retention of automatically generated logs](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art26-6) (`AIGE-OBL-EUAIA-ART26-6`): Deployers keep the logs under their control for a period appropriate to the intended purpose, at least six months unless other law provides otherwise
- [GPAI Code of Practice, Safety and Security Commitment 9: serious-incident reporting](https://aigovernanceengineer.com/obligations/aige-obl-gpaicop-safety-c9) (`AIGE-OBL-GPAICOP-SAFETY-C9`): Report serious incidents to the AI Office within 2, 5, 10 or 15 days by incident class, with intermediate reports at least every four weeks while unresolved and a final report within 60 days of resolution; keep the records at least five years
- [ISO/IEC 42001, A.8 Information for interested parties](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a8) (`AIGE-OBL-ISO42001-A8`): Transparency and reporting to stakeholders

## Sources

[1] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "Freeze before you fix"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#freeze-before-you-fix (verified: primary)
[2] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The incident record"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#the-incident-record (verified: primary)
[3] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The overlapping clocks"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#the-overlapping-clocks (verified: primary)
[4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
[5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
[6] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
[7] Our framework for reporting model misalignment (each full report describes the behavior observed, its severity and any external impact, the setting, the date, when it was discovered and the models involved). OpenAI. 2026-09-16. https://openai.com/index/model-misalignment-reporting-framework/ (verified: primary)
[8] Frontier Governance Framework (section 2.6: OpenAI "maintains an AI Safety Incident Response Plan (AIRP)" for identifying and responding to AI safety incidents). OpenAI. 2026-05-28. https://cdn.openai.com/pdf/e37d949b-8c9f-4d76-b99e-4272f4631a7e/openai-frontier-governance-framework.pdf (verified: primary)
[9] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
[10] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)

## Machine-readable

- This control as JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-007.json
- The pass example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-007.pass.json
- The fail example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-007.fail.json
- The whole profile as Markdown: https://aigovernanceengineer.com/controls/evaluation-environment.md
- The open data API: https://aigovernanceengineer.com/resources/data

## Review

Review this control through the issue form: https://github.com/losanchos5/aige/issues/new?template=control-review.yml. How review works: https://aigovernanceengineer.com/contribute. Page: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-007

## Cite

AIGE-CTL-EVAL-007 Incident Evidence Preservation. In Jorge García Aibar (2026). Evaluation Environment Control Profile (v0.2, draft). AI Governance Engineer. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment
