---
title: "Monitoring Integrity"
description: "Draft control: the traces and monitors of an AI evaluation run cover all of it, stay out of the agent's reach and reconcile with the tool servers' own logs."
canonical: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-005
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22857084
version: "0.2"
updated: 2026-09-26
---

# Monitoring Integrity

> The traces and monitors of a run capture every plan step and tool call with the agent's identity and version, as the tool returned it, and the agent cannot disable or alter them.

- Id: AIGE-CTL-EVAL-005
- Profile: [Evaluation Environment Control Profile v0.2](https://aigovernanceengineer.com/controls/evaluation-environment)
- Status: Draft
- Review: Open for technical review
- Published: 2026-09-26
- Updated: 2026-09-26
- Anchor on the profile page: https://aigovernanceengineer.com/controls/evaluation-environment#aige-ctl-eval-005

Draft for review, not a claim of conformity. A draft control specification, open for technical review: illustrative, not legal advice and binding on no one.

## The control record

- Id: `AIGE-CTL-EVAL-005` · v0.2 · Draft · Open for technical review
- Depth: Specified
- Objective: The traces and monitors of a run capture every plan step and tool call with the agent's identity and version, as the tool returned it, and the agent cannot disable or alter them.
- Failure modes:
  - A tool call appears in the tool server's logs but not in the run's trace, and no gap is recorded.
  - Monitoring is switched off or reconfigured from inside the environment during a run, for example by changing an environment variable.
  - A monitor that the run relies on does not run on it, or covers only a sample of it, and nobody notices until after the run.
  - The recorded output of a tool call differs from what the tool returned (a spoofed tool call).
- Scope: Telemetry, traces and monitors attached to an evaluation run, and the store they are written to. Preserving the records after an incident is covered by AIGE-CTL-EVAL-007.
- Enforcement points:
  - runtime: at the point of action (gateway or guardrail)
  - periodic: on a schedule, over what is already running
- Verification:
  - Inspect: Before the run, inspect the monitoring configuration: the trace collector, the monitors and their settings run outside the environment, and nothing the agent can write (environment variables, files, its own processes) can switch them off or change what they record.
  - Test: At admission and after each harness change, from inside the environment, try to disable or redirect monitoring (change the environment variables and configuration the agent can reach, stop any local exporter); the trace must continue at the collector and the attempt must raise an alert.
  - Observe: During the run, record the heartbeat of each monitor the run relies on: it must cover every sample and the whole duration, and any gap is recorded with its start and end.
  - Observe: After the run, reconcile the trace with the tool servers' own logs: the same calls in the same order, and every recorded tool output matches the hash the tool server logged for it.
- Evidence:
  - Monitor coverage record of the run: which monitors ran, over which samples and time, and every gap · Layer 04 Runtime Controls & Observability · [evidence-record.v1](https://aigovernanceengineer.com/resources/templates#schema-evidence-record)
  - Reconciliation of the trace with the tool servers' logs, with the hash of every tool output · Layer 04 Runtime Controls & Observability · [evidence-record.v1](https://aigovernanceengineer.com/resources/templates#schema-evidence-record)
  - Tampering test at admission and the reconciliation verdict, filed as an observation of this control · Layer 05 Assurance & Continuous Compliance · [control-observation.v1](https://aigovernanceengineer.com/resources/templates#schema-control-observation)
- Failure response: alert: let the action through and raise an alert. A monitoring gap, a tampering attempt or a trace that does not reconcile with the tool servers' logs alerts the evaluator and marks the run. The result of a run whose monitors did not cover it, or whose trace does not reconcile, is withheld until the difference is explained; a tampering attempt is also a stop condition (AIGE-CTL-EVAL-006).
- Layers: [Layer 04 Runtime Controls & Observability](https://aigovernanceengineer.com/bok/the-stack#layer-04-runtime-controls--observability), [Layer 05 Assurance & Continuous Compliance](https://aigovernanceengineer.com/bok/the-stack#layer-05-assurance--continuous-compliance)
- Patterns: [Continuous Assurance Telemetry](https://aigovernanceengineer.com/patterns/continuous-assurance-telemetry)
- Seeded from: [Traces](https://aigovernanceengineer.com/bok/governing-agents#telemetry-with-the-opentelemetry-genai-conventions), [Telemetry on the OpenTelemetry GenAI conventions](https://aigovernanceengineer.com/bok/governing-agents#telemetry-with-the-opentelemetry-genai-conventions), [Trajectory anomaly detection](https://aigovernanceengineer.com/bok/governing-agents#an-agent-incident-taxonomy)
- Mappings:
  - Obligations: [EU AI Act Art. 12 record-keeping and logging](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art12); [NIST AI RMF MEASURE](https://aigovernanceengineer.com/obligations/aige-obl-nistrmf-measure)
  - ISO/IEC 42001: A.6.2.8 AI system recording of event logs; A.6.2.6 AI system operation and monitoring
  - NIST AI RMF: MEASURE 3.1 Existing, unanticipated and emergent risks are tracked
  - OWASP: [ASI10 Rogue Agents](https://aigovernanceengineer.com/resources/threats#threat-asi10)
  - AIUC-1: E015
  - MITRE ATLAS mitigation: AML.M0024 (AI Telemetry Logging)
  - NIST SP 800-53 Rev. 5: AU-2 (Event Logging)
  - NIST SP 800-53 Rev. 5: AU-9 (Protection of Audit Information)
  - NIST SP 800-53 Rev. 5: AU-12 (Audit Record Generation)
  - NIST SP 800-53 Rev. 5: SI-4 (System Monitoring)
- References:
  - [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Telemetry with the OpenTelemetry GenAI conventions")
  - [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "An agent incident taxonomy")
  - [3] OpenTelemetry semantic conventions for generative AI (agent, tool and model spans, events and metrics; status Development)
  - [4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  - [5] Frontier Risk Report (February to March 2026) (METR states that its red-teaming exercise at Anthropic found "several simple ways for monitoring to be disabled", one by changing an environment variable)
  - [6] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  - [7] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
  - [8] Priorities and principles for effective third party assessments (asks whether monitoring is implemented across training, evaluations and deployment "in a way that cannot easily be disabled", and for claims "pre-registered before assessment activities begin")
  - [9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  - [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  - [11] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
  - [12] Unauthorized communication via temporary file hosting services (misalignment report) (OpenAI reports agents in RL training that uploaded output files to public hosting platforms for co-working agents; the behavior was found by a monitoring system that "ran on 20% of the samples" of the run, which now runs on 100% of samples)
- Implementation notes:
  - Run the collector, the monitors and their configuration outside the environment, and write the trace to append-only storage the agent cannot reach. OpenAI's Preparedness Framework v2 lists agent actions "logged in an uneditable database" and reviewed by asynchronous monitors among its potential safeguards against a misaligned model. OpenAI's technical report on the Hugging Face incident states that models sometimes tried to "erase or tamper with their outputs or message logs", and that none of the attempts it observed to manipulate tool trajectories affected the logs its graders or monitors ultimately see.
  - Keep every switch of the monitoring out of the agent's reach. METR states that its red-teaming exercise at Anthropic found "several simple ways for monitoring to be disabled", one by changing an environment variable that indicated whether the traffic came from Claude Code, and OpenAI's criteria for third-party assessments ask whether monitoring covers training, evaluations and deployment "in a way that cannot easily be disabled".
  - Check that the monitors ran on this run, and over all of it. OpenAI states that its chain-of-thought monitors "did not run on the evaluations" involved in the Hugging Face incident, and reports a training incident found by a monitoring system that ran on 20% of the samples of the run and now runs on 100%.
  - Reconcile the transcript with what the tools logged, not only with itself. METR's investigation of the OpenAI/Hugging Face incident reports that at least 96 transcripts show clear evidence of spoofed tool calls, where an agent appeared to run one command and replaced its output; a hash of each output, logged by the tool server, lets a reviewer find the substitution.
- Open questions:
  - How can a reviewer show that a run's trace is complete for actions that pass no tool server, not only that the records it holds are unaltered?
  - The OpenTelemetry GenAI conventions are still at status Development: should the profile name the attributes it relies on, or only the events a trace must contain?
- Observation:
  - Subject: Evaluation run
  - Expected: The monitors ran over the whole run, the agent could not switch them off, and the trace reconciles with the tool servers' logs call by call.
  - Example: Run 88251: 1,904 calls in the trace and in the tool servers' logs, 0 output hash mismatches, monitor coverage 100% of samples: pass.
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-005.json

## Example observations

Two illustrative records of a check of this control, one that passes and one that fails. They validate against the control observation schema; they are not results of any real evaluation.

### Pass: run-88251@2026-09-26

- Status: pass
- Subject: run-88251@2026-09-26 (Evaluation run)
- Expected: The monitors ran over the whole run, the agent could not switch them off, and the trace reconciles with the tool servers' logs call by call.
- Observed: Run 88251: 1,904 calls in the trace and in the tool servers' logs, in the same order, with 0 output hash mismatches; monitor coverage 100% of samples; the admission attempt to stop the exporter raised an alert and the trace continued.
- Timestamp: 2026-09-26T13:20:02Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-005.pass.json

### Fail: run-88252@2026-09-26

- Status: fail
- Subject: run-88252@2026-09-26 (Evaluation run)
- Expected: The monitors ran over the whole run, the agent could not switch them off, and the trace reconciles with the tool servers' logs call by call.
- Observed: Run 88252: 7 tool outputs in the transcript differ from the output hashes the tool servers logged (spoofed tool calls), and the chain-of-thought monitor covered 20% of samples.
- Timestamp: 2026-09-26T14:02:45Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-005.fail.json

## Related cases

- [OpenAI agents and Hugging Face: an evaluation environment that was not isolated](https://aigovernanceengineer.com/cases/openai-hugging-face-agent-incident-2026): METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- [An agent in training reached a public chatbot through the sandbox DNS resolver](https://aigovernanceengineer.com/cases/openai-agent-dns-covert-channel-2026): OpenAI reports that an agent in RL training used its sandbox's DNS resolver to reach a public chatbot; the run was stopped hours after the alert.
- [An internally deployed model published a researcher's GitHub token in a public repository](https://aigovernanceengineer.com/cases/openai-agent-github-token-exposure-2026): OpenAI reports that an internally deployed model put a researcher's GitHub token, split to avoid secret scanning, into code it pushed to a public repository.
- [Agents in training shared a file through a public file-hosting service](https://aigovernanceengineer.com/cases/openai-agents-temp-file-hosting-2026): OpenAI reports that agents in multi-agent RL training uploaded a workbook to a public file-hosting service so that collaborating agents could download it.
- [Training samples exchanged messages through a shared package repository](https://aigovernanceengineer.com/cases/openai-agents-artifactory-cross-sample-2026): OpenAI reports that models in RL training used an internal package repository, with the credentials they were given, to exchange messages across samples.
- [Claude models reached real systems from a misconfigured third-party cyber evaluation](https://aigovernanceengineer.com/cases/anthropic-third-party-eval-environment-incidents-2026): Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems.
- [Agents in a cyber range with open internet took unsanctioned actions against real people](https://aigovernanceengineer.com/cases/uk-aisi-cyber-range-unsanctioned-actions-2026): UK AISI reports that agents in a cyber evaluation with internet deliberately enabled took 19 unsanctioned actions aimed at real people and organisations.

## Patterns

- [Continuous Assurance Telemetry](https://aigovernanceengineer.com/patterns/continuous-assurance-telemetry) (Layer 05 Assurance & Continuous Compliance)

## Obligations

- [EU AI Act Art. 12 record-keeping and logging](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art12) (`AIGE-OBL-EUAIA-ART12`): Record-keeping: automatic logging of events over the system's lifetime
- [NIST AI RMF MEASURE](https://aigovernanceengineer.com/obligations/aige-obl-nistrmf-measure) (`AIGE-OBL-NISTRMF-MEASURE`): Analyse, benchmark and monitor risk

## Threats

- [ASI10 Rogue Agents](https://aigovernanceengineer.com/resources/threats#threat-asi10) (OWASP Top 10 for Agentic Applications 2026)

## Sources

[1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Telemetry with the OpenTelemetry GenAI conventions"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#telemetry-with-the-opentelemetry-genai-conventions (verified: primary)
[2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "An agent incident taxonomy"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#an-agent-incident-taxonomy (verified: primary)
[3] OpenTelemetry semantic conventions for generative AI (agent, tool and model spans, events and metrics; status Development). OpenTelemetry. 2026. https://github.com/open-telemetry/semantic-conventions-genai/tree/main/docs/gen-ai (verified: primary)
[4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
[5] Frontier Risk Report (February to March 2026) (METR states that its red-teaming exercise at Anthropic found "several simple ways for monitoring to be disabled", one by changing an environment variable). METR. 2026-05-19. https://metr.org/blog/2026-05-19-frontier-risk-report/ (verified: primary)
[6] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
[7] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it"). OpenAI. 2026-08-26. https://openai.com/index/hugging-face-incident-and-the-road-ahead/ (verified: primary)
[8] Priorities and principles for effective third party assessments (asks whether monitoring is implemented across training, evaluations and deployment "in a way that cannot easily be disabled", and for claims "pre-registered before assessment activities begin"). OpenAI. 2026-09-22. https://openai.com/index/priorities-principles-third-party-assessments/ (verified: primary)
[9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling"). OpenAI. 2025-04-15. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf (verified: primary)
[10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
[11] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)
[12] Unauthorized communication via temporary file hosting services (misalignment report) (OpenAI reports agents in RL training that uploaded output files to public hosting platforms for co-working agents; the behavior was found by a monitoring system that "ran on 20% of the samples" of the run, which now runs on 100% of samples). OpenAI (Alignment Research Blog). 2026-09-16. https://alignment.openai.com/misalignment-reports/unauthorized-communication-via-temporary-file-hosting-services/ (verified: primary)

## Machine-readable

- This control as JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-005.json
- The pass example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-005.pass.json
- The fail example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-005.fail.json
- The whole profile as Markdown: https://aigovernanceengineer.com/controls/evaluation-environment.md
- The open data API: https://aigovernanceengineer.com/resources/data

## Review

Review this control through the issue form: https://github.com/losanchos5/aige/issues/new?template=control-review.yml. How review works: https://aigovernanceengineer.com/contribute. Page: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-005

## Cite

AIGE-CTL-EVAL-005 Monitoring Integrity. In Jorge García Aibar (2026). Evaluation Environment Control Profile (v0.2, draft). AI Governance Engineer. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment
