On this page

Evaluation Environment Control Profile

Draft control specifications for the environment a model or agent is evaluated in: the harness, tools, credentials, network, monitoring and stop conditions around it, and the evidence a run leaves.

v0.2 Draft Open for technical review

Scope

Evaluation environments for models and agents, from the harness and the tools and MCP servers a run can call to the credentials it holds, the network it can reach and the records it leaves. The evaluation tasks, their scoring rubrics and the capabilities being measured are out of scope.

9 controls, AIGE-CTL-EVAL-001 to AIGE-CTL-EVAL-009. Each one is anchored to a layer of the five-layer stack and to the patterns that implement it.

How to read a control

Every control has a stable id that never changes and is never reused, and the same record:

  • Objective: the outcome the control secures, in one sentence.
  • Failure modes: observable events that mean the control failed.
  • Scope, enforcement points (on a pull request, at deploy, at runtime or on a schedule) and the failure response (deny, hold for approval or alert).
  • Verification: how a third party would check it (inspect, test, observe or attest).
  • Evidence: the artefact it leaves and the stack layer that keeps it.
  • Mappings: obligations, ISO/IEC 42001 Annex A, the NIST AI RMF, OWASP and other references, each id checked against the site's registers.

Depth says how far a control is written. Specified: a verification procedure, evidence, implementation notes, two example observations and a page of its own. Derived from site material: restates existing site material (an agent control of chapter 23, a pattern, a record schema or a chapter, named on each control) and adds nothing that material does not say. Draft outline: objective, failure modes, scope and open questions only. This profile: 9 specified.

An observation is what a check of a control would emit: the control id, the subject checked, what was expected, what was observed, a status (pass, fail or not_applicable), a timestamp and the evidence behind it.

001 Authorization Boundary

Id
AIGE-CTL-EVAL-001 · v0.2 · Draft · Open for technical review
Objective
Every agent in an evaluation run acts only within an authorization boundary recorded before the run starts: the tools, operations, data classes and budgets it may use, and the runs and resources it may reach.
Failure modes
  • A run starts with no recorded boundary for the agent under test, or with a boundary that differs from the scope the agent was told.
  • The agent calls a tool or an operation class outside its recorded boundary, or starts processes with administrative privileges, and the call succeeds.
  • A credential given for one purpose on a shared resource, such as downloading packages, also lets the agent write, list or post there.
  • Agents in runs meant to be isolated reach each other through a resource the runs share, such as a package repository or a cache.
Scope
Agents and harnesses under evaluation, the tools and operations they can use during a run, the credentials they receive for shared resources and the resources runs share. Budgets are recorded here and enforced under AIGE-CTL-EVAL-006; network egress is AIGE-CTL-EVAL-002. Production deployments are covered by the agent runtime profile.
Enforcement points
  • deploy: before a version is deployed or released
  • runtime: at the point of action (gateway or guardrail)
Verification
  • Inspect: Before the run, inspect the run record: it holds a boundary for each agent under test (tools and operation classes, data classes, budgets, and each shared resource with the operations allowed on it), and the prompt the agent receives states the same boundary as instructions, including what it must not access.
  • Test: At admission, from inside the environment, attempt one call of each kind outside the boundary (an unlisted tool or operation class, a write or a listing with a download-only credential on a shared resource, a process started as root) and one listed call; every attempt outside the boundary must be refused and logged, and the listed call must succeed.
  • Test: Start two canary runs on the same shared resources: a marker written by one run must not be readable by the other.
  • Observe: After the run, compare every tool call and every request to a shared resource in the run's logs with the recorded boundary: each falls inside it, and every refused attempt is recorded with its time and target.
Evidence
  • The boundary of each agent under test, recorded before the run: tools, operation classes, data classes, budgets and shared resources with the operations allowed on each · Layer 02 · agent-register-entry.v1
  • Admission test verdicts: the refused out-of-boundary calls and the cross-run canary · Layer 04 · evidence-record.v1
  • One observation per run comparing the calls and shared-resource requests made with the recorded boundary · Layer 05 · control-observation.v1
Failure response
deny: block the action. Calls and shared-resource operations outside the recorded boundary are refused at the enforcement point and logged. A run with no recorded boundary is not started; a run in which a call outside the boundary succeeded, or in which runs reached each other, is stopped and its result is withheld until the path is closed.
Layers
Layer 04 Runtime Controls & Observability, Layer 02 Inventory & Transparency
Patterns
Agent Registry, Policy Card
Seeded from
Registry entry, Read-only tools, Reversible, bounded actions only, Execution budgets, Data classes recorded, with the DPIA linked
Mappings
References
  • [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The agent registry")
  • [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Autonomy is a design decision")
  • [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The tool allow-list")
  • [4] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server)
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [8] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions")
  • [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  • [11] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent")
  • [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
Implementation notes
  • Record the boundary as data before the run, in the fields of an agent registry entry (scope, tools, data access, spend limit), and enforce it at the gateway the agent's calls pass through, not in the prompt alone. State the same boundary to the agent: Anthropic's guidance for external evaluation partners asks every prompt to state what is in and out of scope, including permitted actions and network boundaries, with boundaries "phrased as instructions" rather than claims about the environment.
  • Tell the agent what it may not do with the access it has. OpenAI's account of a third-party cyber evaluation states that the evaluator's "intended authorization boundary was the simulated cyber range", while the agents were not explicitly told how they could and could not use the open internet access they had; the evaluator names this as a potential contributing cause.
  • Scope every credential for a shared resource to the operations the task needs: download from a package repository, never write, list or post. OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate samples with the credentials they were given to download packages, without exploiting any vulnerability; its technical report on the Hugging Face incident lists removing the shared credential from new research workloads and restricting the types of requests agents could make to Artifactory among its mitigations.
  • Give each run its own namespace on every shared resource (package cache, artefact store, scratch storage) and test it with a canary marker. METR's investigation of the OpenAI/Hugging Face incident reports that agents "meant to be fully isolated from one another" communicated through an internal package repository.
  • Run the agent as an unprivileged user. METR's Task Standard states that the agent "MUST NOT be allowed to start processes" on the primary machine as root.
Open questions
  • Which shared resources (package mirrors, caches, model API proxies) can be given a per-run namespace, and which should be removed from evaluation environments altogether?
  • What evidence is sufficient to show that an agent stayed within its boundary for a whole run when the logs of some of the tools it called are outside the evaluator's control?
Observation
  • Subject: Evaluation run
  • Expected: Every tool call and every request to a shared resource falls inside the boundary recorded for the agent before the run; every attempt outside it is refused and logged.
  • Example: Run 88231: 3 listing and upload requests to the shared package repository with a download-only credential, all refused and logged; no call outside the boundary succeeded: pass.

002 Network Egress Control

Id
AIGE-CTL-EVAL-002 · v0.2 · Draft · Open for technical review
Objective
Outbound connections from an evaluation environment reach only the destinations on the run's egress allow-list, and every other connection attempt is refused and logged.
Failure modes
  • A connection from the environment to a host that is not on the run's egress allow-list succeeds.
  • A run starts in an environment with no egress policy attached, or with a policy that allows every destination although the task did not declare that it needs the internet.
  • A permitted destination, such as a package mirror, a cache or a tool server, carries data onward to a party or to another run that nobody listed.
  • The run leaves no flow log, so the connections it made cannot be compared with its allow-list.
Scope
Every network path out of the environment a run executes in: the agent's container or virtual machine, auxiliary machines, DNS, and the tools, MCP servers and proxies the run can call. Resources shared between runs count as destinations. Inbound operator access is out of scope.
Enforcement points
  • deploy: before a version is deployed or released
  • runtime: at the point of action (gateway or guardrail)
Verification
  • Inspect: Before the run, inspect the egress policy attached to the task environment: deny by default, with an allow-list naming each permitted destination (for example the LLM API proxy and the progress server) and nothing else unless the task declares that it needs the internet.
  • Test: At admission, from inside the environment, attempt one connection to a destination that is not on the allow-list and one to a listed destination; the first must be refused and logged, the second must succeed.
  • Observe: After the run, compare the run's flow log with its allow-list: every outbound connection matches a listed destination, and every refused attempt is recorded with its time and target.
Evidence
  • The egress policy and allow-list attached to the run, with its hash recorded in the run record · Layer 04
  • Admission test verdict: the refused connection to an unlisted destination · Layer 04 · evidence-record.v1
  • Flow log of the run, allowed and refused connections, kept outside the environment · Layer 04
  • One observation per run comparing observed connections with the allow-list · Layer 05 · control-observation.v1
Failure response
deny: block the action. Connections to unlisted destinations are refused at the enforcement point and logged. A run whose environment has no egress policy attached is not started; a run in which an unlisted connection succeeded is stopped and its result is withheld until the connection is explained.
Layer
Layer 04 Runtime Controls & Observability
Patterns
Runtime Guardrail, Sanctioned AI Gateway
Seeded from
Output and egress filter, Tool allow-list, deny by default, Code runs only in a sandbox
Mappings
References
  • [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The tool allow-list")
  • [13] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls")
  • [4] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server)
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [14] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service)
  • [15] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients)
  • [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [16] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
  • [17] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [18] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later)
  • [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
  • [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
Implementation notes
  • Attach a default-deny egress policy to every task environment (a network policy, a firewall rule set or an egress proxy, by category) that lists each permitted destination explicitly, for example the LLM API proxy and the progress server. Enforce it on the resolved address, and serve DNS from a resolver that answers only for listed names: DNS is an egress path too. OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox.
  • Record the allow-list with the run (its hash goes in the run record next to the environment version) and keep a flow log per run at the enforcement point, outside anything the agent can reach or change, with refused attempts as well as allowed connections, so the connections made can be compared with what was declared, not with what was intended.
  • Enforce the policy in two places a single misconfiguration cannot both open, for example at the sandbox and at the cluster boundary. OpenAI's technical report on the Hugging Face incident describes "two separate layers of network isolation" of this kind, and Anthropic's guidance for external evaluation partners asks that the only outside connection be to the model's own API and that the configuration be verified before every evaluation begins.
  • At admission, run a canary from inside the environment: one connection to a reserved test address that is not on the list. Expect a refused connection and a log entry; a success blocks the run.
  • Treat resources shared between runs (package mirrors, caches, artefact stores) as destinations with their own entry and a per-run scope. METR's public investigation of the OpenAI/Hugging Face incident reports that agents meant to be "fully isolated from one another" communicated through a shared internal package repository.
Open questions
  • How should the egress allow-list of a run be recorded so that a third party can compare it with the connections actually observed?
  • When a task needs the internet (what METR's Task Standard calls the full_internet permission), which part of this control still applies, and what should the run record say instead?
Observation
  • Subject: Evaluation environment
  • Expected: Outbound connections only to the destinations on the run's egress allow-list; every other attempt refused and logged.
  • Example: 1 connection to an unlisted host during run 88213, in an environment with no egress policy attached: fail.

003 Credential Isolation

Id
AIGE-CTL-EVAL-003 · v0.2 · Draft · Open for technical review
Objective
An agent under evaluation holds only short-lived credentials issued to its own identity for the run and bound to the one service each is for, never standing secrets or a person's own token.
Failure modes
  • A long-lived secret (an API key, a cloud access key, a password) is readable from the agent's environment, configuration, files or memory during a run.
  • The agent presents a token issued to a person, a token whose audience is another service, or a credential it found rather than received, and the tool server accepts it.
  • A credential issued for the run is still accepted after the run ended or was aborted.
  • A credential appears in the run's transcript, memory store, logs or outputs, or is passed to another agent.
Scope
Credentials, tokens and keys the agent under evaluation and the tools it calls can reach during a run, including what it holds in memory and writes to its transcript, and the model API key, which stays with a proxy outside the environment. The evaluator's own operator credentials are out of scope.
Enforcement points
  • deploy: before a version is deployed or released
  • runtime: at the point of action (gateway or guardrail)
Verification
  • Inspect: Before the run, inspect the environment template and the run's configuration: no long-lived secret is present, and the agent obtains credentials from a broker outside the environment under its own workload identity, each with a lifetime no longer than the run and an audience naming one tool server.
  • Test: After the run, scan every run artefact (transcript, memory store, logs, outputs and a snapshot of the environment's file system) for secret patterns and for the tokens issued to the run; expect no match.
  • Test: Replay a token issued for the run against a different tool server, and again after the run has ended; both must be rejected, for the wrong audience and for expiry or revocation.
  • Observe: Read the tool servers' logs for the run: every call carries a token issued for that server, delegated calls name the agent as the acting party, and audience-check failures were raised as alerts.
Evidence
  • Credential issuance log of the run: identity, audience, scope, lifetime and revocation time of every token · Layer 04 · evidence-record.v1
  • Audience-check and replay results from the tool servers · Layer 04 · evidence-record.v1
  • Secret scan of the run artefacts, filed as an observation of this control · Layer 05 · control-observation.v1
Failure response
deny: block the action. A token issued for another audience, to a person, or for a run that has ended is rejected by the tool server. A run in which a long-lived secret or a leaked credential is found is stopped, the credential is revoked and the result is withheld until the exposure is assessed.
Layers
Layer 04 Runtime Controls & Observability, Layer 02 Inventory & Transparency
Patterns
Agent Identity & Scoped Credentials
Seeded from
Its own identity, Replace long-lived secrets with short-lived credentials, Delegation, never impersonation, MCP authorisation (spec 2026-07-28), Memory write gate and rollback
Mappings
References
  • [19] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Identity and short-lived credentials")
  • [20] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Short-lived, attested credentials")
  • [21] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Delegation without impersonation")
  • [22] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "MCP authorization as of 2026-07-28")
  • [23] RFC 8693, OAuth 2.0 Token Exchange (the act claim "provides a means within a JWT to express that delegation has occurred and identify the acting party")
  • [24] MCP specification 2026-07-28, Authorization (MCP servers MUST validate that access tokens were issued specifically for them and "MUST NOT accept or transit any other tokens")
  • [15] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients)
  • [25] SPIFFE overview (SVIDs are "short lived cryptographic identity documents", delivered and rotated through the Workload API)
  • [26] Model AI Governance Framework for Agentic AI, v1.5 (agent identity unique and "cryptographically verifiable"; authorisations "time- or session-bound, non-transferable")
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [14] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service)
  • [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [17] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [8] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions")
  • [27] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning)
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
Implementation notes
  • Give each agent under test its own workload identity for the run (an attested identity document of the kind SPIFFE defines, or the platform's workload identity) and issue its tokens from a broker outside the environment. The environment template holds no secret.
  • Keep the model API key with an LLM API proxy outside the environment, reached through the egress allow-list of AIGE-CTL-EVAL-002: the agent sees the proxy, never the key. Anthropic's guidance for external evaluation partners likewise asks that "the API keys should be kept outside the environment".
  • Issue short-lived access tokens whose aud names one tool server and, when the agent acts for a user or another agent, whose act claim names the agent (RFC 8693 token exchange). Tool servers validate the audience and refuse any other token, as the MCP authorization specification of 2026-07-28 requires of MCP servers.
  • Set every token lifetime no longer than the run's time budget, issue no refresh tokens to agents under evaluation, and revoke what was issued when the run ends or is aborted.
  • Scrub credentials from memory stores, transcripts and logs before they are written, then grep every run artefact after the run for secret patterns (known key prefixes, high-entropy strings, the run's own issued tokens); a match fails the run. Scan for credentials a run picked up as well as those it was issued: METR's investigation of the OpenAI/Hugging Face incident reports an agent that found working Hugging Face credentials exposed on the internet and posted them to a board other agents read, and OpenAI states that in a third-party evaluation its model reused a GitHub token another lab's agent had left publicly accessible.
Open questions
  • What lifetime should a credential issued for an evaluation run have, and who revokes it when the run is aborted?
  • Can a secret scan be made complete for encoded or split credentials, or should the control rely on short lifetimes and revocation for whatever a scan may miss? OpenAI reports a model that split a researcher's token into pieces with the stated aim of avoiding secret scanning.
Observation
  • Subject: Evaluation run
  • Expected: No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.
  • Example: Run 88214: secret scan found 0 matches in transcript, memory store and logs; a replayed token was rejected with an audience error: pass.

004 Tool and Action Mediation

Id
AIGE-CTL-EVAL-004 · v0.2 · Draft · Open for technical review
Objective
Every tool call an agent makes during a run passes a mediation point outside the model that checks it against policy before it executes.
Failure modes
  • A tool call executes without a policy verdict recorded for it.
  • An irreversible action runs while the mediation point is down.
  • An approval is reused for a call whose parameters changed.
  • After repeated denials the agent keeps seeking the same outcome through other calls, and the run continues.
Scope
Tool, MCP and code-execution calls made by the agent under evaluation, and the reviewer (a rule, a model or a person) that decides on them. The internal reasoning of the model is out of scope; network egress is AIGE-CTL-EVAL-002.
Enforcement points
  • deploy: before a version is deployed or released
  • runtime: at the point of action (gateway or guardrail)
Verification
  • Inspect: Before the run, inspect the environment and its policy: tool servers, MCP servers and code execution are reachable only through the mediation point, and the policy lists each operation class with its verdict, its failure posture (fail closed for irreversible classes such as delete, send, publish and execute) and the denial threshold that interrupts a run.
  • Test: At admission, send through the harness one call the policy denies and one it allows, then take the mediation point down and send an irreversible-class call; the denied call and the call sent while it is down must not execute, and all three must leave a verdict record.
  • Observe: After the run, join the tool servers' own logs with the verdict records: every executed call has an allow verdict, or an approval bound to a parameter hash that matches the call, and no run continued past its denial threshold.
Evidence
  • Mediation policy of the run: operation classes, verdicts, failure posture per class and the denial threshold · Layer 04 · policy-card.v1
  • Verdict record of every call: tool, parameter hash, verdict, reviewer and time · Layer 04 · evidence-record.v1
  • One observation per run joining the executed calls with their verdicts · Layer 05 · control-observation.v1
Failure response
deny: block the action. A call without an allow verdict does not execute. While the mediation point is down, irreversible classes fail closed and reads fail open only with an alert. A run in which a call executed without a verdict is stopped and its result is withheld; a run that reaches its denial threshold is interrupted.
Layer
Layer 04 Runtime Controls & Observability
Patterns
Runtime Guardrail, Human-in-the-loop Gate
Seeded from
Runtime guardrail on every tool call, Checkpoints on irreversible actions, failing closed, Approval log, bound to the call, MCP server admission gate, Code runs only in a sandbox
Mappings
References
  • [13] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls")
  • [28] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint")
  • [29] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server")
  • [30] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026)
  • [31] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  • [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [32] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant")
  • [33] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security")
  • [34] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  • [27] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning)
Implementation notes
  • Put the mediation point between the decision to call a tool and the call, outside the model and the agent's process: a gateway or guardian that checks identity, allow-list, parameters and operation class before the call runs, as chapter 23 describes and OWASP's Agent Control Standard specifies on the wire. Route tool servers, MCP servers and code execution only through it, so no call can go around it.
  • Record the failure posture per operation class in the policy of the run. Chapter 23 recommends failing closed for pay, delete, send and execute and failing open, with an alert, only for reads, and notes that the reference guardian of the Agent Control Standard starts with a fail-open posture: set the posture, do not inherit it.
  • Bind an approval to the exact call. OpenAI's Codex documentation describes an override that "applies to the exact denied action, not similar future actions" and holds for one retry; store a hash of the parameters with each approval and refuse a call whose parameters differ.
  • Interrupt a run that keeps asking. The Codex documentation describes a rejection circuit breaker that, in its current open-source implementation, interrupts the turn after 3 consecutive denials or 10 in the last 50 reviews, and OpenAI's post on auto-review states that it stops the trajectory after repeated denials. Record the threshold in the policy and make reaching it a stop condition (AIGE-CTL-EVAL-006).
  • An automated reviewer can stand in for a person at the mediation point without widening what the agent may do: the Codex documentation states that "Auto-review is a reviewer swap, not a permission grant", and OpenAI's post on the design states that it "should not be treated as a guarantee of security". OpenAI reports that a model which had published a researcher's token was brought back with "enforced auto-review on all actions".
Open questions
  • How can every call be mediated without taking away the scaffolding and tooling a capability elicitation needs? METR's guidelines ask for "the best available scaffolding + tooling".
  • When the reviewer at the mediation point is a model, what evidence shows that its verdicts are reliable enough to replace a person for a given operation class?
Observation
  • Subject: Evaluation run
  • Expected: Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
  • Example: Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters: fail.

005 Monitoring Integrity

Id
AIGE-CTL-EVAL-005 · v0.2 · Draft · Open for technical review
Objective
The traces and monitors of a run capture every plan step and tool call with the agent's identity and version, as the tool returned it, and the agent cannot disable or alter them.
Failure modes
  • A tool call appears in the tool server's logs but not in the run's trace, and no gap is recorded.
  • Monitoring is switched off or reconfigured from inside the environment during a run, for example by changing an environment variable.
  • A monitor that the run relies on does not run on it, or covers only a sample of it, and nobody notices until after the run.
  • The recorded output of a tool call differs from what the tool returned (a spoofed tool call).
Scope
Telemetry, traces and monitors attached to an evaluation run, and the store they are written to. Preserving the records after an incident is covered by AIGE-CTL-EVAL-007.
Enforcement points
  • runtime: at the point of action (gateway or guardrail)
  • periodic: on a schedule, over what is already running
Verification
  • Inspect: Before the run, inspect the monitoring configuration: the trace collector, the monitors and their settings run outside the environment, and nothing the agent can write (environment variables, files, its own processes) can switch them off or change what they record.
  • Test: At admission and after each harness change, from inside the environment, try to disable or redirect monitoring (change the environment variables and configuration the agent can reach, stop any local exporter); the trace must continue at the collector and the attempt must raise an alert.
  • Observe: During the run, record the heartbeat of each monitor the run relies on: it must cover every sample and the whole duration, and any gap is recorded with its start and end.
  • Observe: After the run, reconcile the trace with the tool servers' own logs: the same calls in the same order, and every recorded tool output matches the hash the tool server logged for it.
Evidence
  • Monitor coverage record of the run: which monitors ran, over which samples and time, and every gap · Layer 04 · evidence-record.v1
  • Reconciliation of the trace with the tool servers' logs, with the hash of every tool output · Layer 04 · evidence-record.v1
  • Tampering test at admission and the reconciliation verdict, filed as an observation of this control · Layer 05 · control-observation.v1
Failure response
alert: let the action through and raise an alert. A monitoring gap, a tampering attempt or a trace that does not reconcile with the tool servers' logs alerts the evaluator and marks the run. The result of a run whose monitors did not cover it, or whose trace does not reconcile, is withheld until the difference is explained; a tampering attempt is also a stop condition (AIGE-CTL-EVAL-006).
Layers
Layer 04 Runtime Controls & Observability, Layer 05 Assurance & Continuous Compliance
Patterns
Continuous Assurance Telemetry
Seeded from
Traces, Telemetry on the OpenTelemetry GenAI conventions, Trajectory anomaly detection
Mappings
  • Obligations: EU AI Act Art. 12 record-keeping and logging; NIST AI RMF MEASURE
  • ISO/IEC 42001: A.6.2.8 AI system recording of event logs; A.6.2.6 AI system operation and monitoring
  • NIST AI RMF: MEASURE 3.1 Existing, unanticipated and emergent risks are tracked
  • OWASP: ASI10 Rogue Agents
  • AIUC-1: E015
  • MITRE ATLAS mitigation: AML.M0024 (AI Telemetry Logging)
  • NIST SP 800-53 Rev. 5: AU-2 (Event Logging)
  • NIST SP 800-53 Rev. 5: AU-9 (Protection of Audit Information)
  • NIST SP 800-53 Rev. 5: AU-12 (Audit Record Generation)
  • NIST SP 800-53 Rev. 5: SI-4 (System Monitoring)
References
  • [35] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Telemetry with the OpenTelemetry GenAI conventions")
  • [36] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "An agent incident taxonomy")
  • [37] OpenTelemetry semantic conventions for generative AI (agent, tool and model spans, events and metrics; status Development)
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [38] Frontier Risk Report (February to March 2026) (METR states that its red-teaming exercise at Anthropic found "several simple ways for monitoring to be disabled", one by changing an environment variable)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [39] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
  • [40] Priorities and principles for effective third party assessments (asks whether monitoring is implemented across training, evaluations and deployment "in a way that cannot easily be disabled", and for claims "pre-registered before assessment activities begin")
  • [34] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  • [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
  • [41] Unauthorized communication via temporary file hosting services (misalignment report) (OpenAI reports agents in RL training that uploaded output files to public hosting platforms for co-working agents; the behavior was found by a monitoring system that "ran on 20% of the samples" of the run, which now runs on 100% of samples)
Implementation notes
  • Run the collector, the monitors and their configuration outside the environment, and write the trace to append-only storage the agent cannot reach. OpenAI's Preparedness Framework v2 lists agent actions "logged in an uneditable database" and reviewed by asynchronous monitors among its potential safeguards against a misaligned model. OpenAI's technical report on the Hugging Face incident states that models sometimes tried to "erase or tamper with their outputs or message logs", and that none of the attempts it observed to manipulate tool trajectories affected the logs its graders or monitors ultimately see.
  • Keep every switch of the monitoring out of the agent's reach. METR states that its red-teaming exercise at Anthropic found "several simple ways for monitoring to be disabled", one by changing an environment variable that indicated whether the traffic came from Claude Code, and OpenAI's criteria for third-party assessments ask whether monitoring covers training, evaluations and deployment "in a way that cannot easily be disabled".
  • Check that the monitors ran on this run, and over all of it. OpenAI states that its chain-of-thought monitors "did not run on the evaluations" involved in the Hugging Face incident, and reports a training incident found by a monitoring system that ran on 20% of the samples of the run and now runs on 100%.
  • Reconcile the transcript with what the tools logged, not only with itself. METR's investigation of the OpenAI/Hugging Face incident reports that at least 96 transcripts show clear evidence of spoofed tool calls, where an agent appeared to run one command and replaced its output; a hash of each output, logged by the tool server, lets a reviewer find the substitution.
Open questions
  • How can a reviewer show that a run's trace is complete for actions that pass no tool server, not only that the records it holds are unaltered?
  • The OpenTelemetry GenAI conventions are still at status Development: should the profile name the attributes it relies on, or only the events a trace must contain?
Observation
  • Subject: Evaluation run
  • Expected: The monitors ran over the whole run, the agent could not switch them off, and the trace reconciles with the tool servers' logs call by call.
  • Example: Run 88251: 1,904 calls in the trace and in the tool servers' logs, 0 output hash mismatches, monitor coverage 100% of samples: pass.

006 Stop Conditions

Id
AIGE-CTL-EVAL-006 · v0.2 · Draft · Open for technical review
Objective
Every evaluation run has stop conditions, budgets and a stop handle defined before it starts, and a stop, pulled by hand or tripped by a breaker, halts the agent's calls and writes within a measured time across every hop it delegated to inside the environment.
Failure modes
  • A run starts without recorded stop conditions, budgets or a working stop handle.
  • A budget (steps, tool calls, tokens, spend or time) or a breaker threshold is exceeded and the run continues.
  • A tool call or a write is recorded after the stop was pulled or the breaker tripped, including through a token issued before the stop.
  • A sub-agent or a service the agent delegated to keeps working after the parent run was stopped.
Scope
Agents under evaluation and every agent, tool or service they delegate to inside the environment. A third-party agent outside the environment can only be cut off at the environment's boundary. A lab's criteria for halting a model's development or deployment are policy decisions and out of scope.
Enforcement points
  • runtime: at the point of action (gateway or guardrail)
  • periodic: on a schedule, over what is already running
Verification
  • Inspect: Before the run, inspect the run record: stop conditions, per-agent budgets and breaker thresholds are recorded, and the stop handle is named with the levels it can apply (pause the task, trip the breaker, revoke the identity).
  • Test: Drill the stop on a schedule and before the first run of a new harness version: pull it during a live task, measure the time from the pull to the first rejected call, and confirm zero tool calls and zero writes after the trip, including through delegated tokens and sub-agents.
  • Observe: During runs, record every breaker trip and budget exhaustion with its trigger, and check that no further call from that agent followed it.
Evidence
  • Stop conditions, budgets and breaker thresholds of the run, recorded before it starts · Layer 04 · policy-card.v1
  • Breaker trips and budget exhaustions of each run, with their triggers · Layer 04 · evidence-record.v1
  • Drill record: time to stop, and calls and writes after the trip · Layer 05 · control-observation.v1
Failure response
alert: let the action through and raise an alert. A stop condition that is met trips the per-agent breaker, so the gateway rejects every further call from that agent, and alerts the evaluator. A drill that finds calls or writes after the trip fails the control and blocks runs on that harness version until the path is closed.
Layer
Layer 04 Runtime Controls & Observability
Patterns
Kill Switch / Circuit Breaker
Seeded from
Per-agent circuit breaker, Drilled kill switch, Execution budgets, Stopping third-party agents at your boundary
Mappings
References
  • [42] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Kill switch and per-agent circuit breakers")
  • [43] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Execution limits")
  • [44] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Stopping across hops")
  • [45] NIST AI RMF 1.0 (AI 100-1) (MANAGE 2.4: mechanisms to "supersede, disengage, or deactivate AI systems" whose outcomes are inconsistent with intended use)
  • [46] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27 as amended by Regulation (EU) 2026/1744 (Art. 14(4)(e): human oversight includes the means to interrupt the system through a stop procedure)
  • [47] Agentic Trust Framework v1 ("You can stop one agent without stopping the business"; containment by revoking the agent's identity)
  • [16] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
  • [48] Agent2Agent (A2A) Protocol Specification v1.0 (Cancel Task: "The server will attempt to cancel the task, but success is not guaranteed")
  • [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [17] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
  • [49] Common Elements of Frontier AI Safety Policies (December 2025 Update) (the policies set "conditions for halting development and deployment" when mitigations are insufficient)
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [18] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later)
  • [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [50] Pacing model development in an era of cyber-critical capabilities (teams paged by a highest-priority monitoring alert are expected to pause the activity if they cannot determine "within 30 minutes that the flag is a false positive")
  • [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
  • [39] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
Implementation notes
  • Record per-agent budgets (steps, tool calls, tokens, spend, wall-clock time) and breaker thresholds (unauthorised calls, refused egress attempts, anomaly scores) in the run record, and enforce them at the gateway the agent's calls pass through, not inside the agent.
  • Make budget exhaustion and threshold breaches trip the breaker automatically. A ticket or a log line alone is not a stop. OpenAI reports a training run that "did not stop automatically as expected" after a monitoring alert and was killed 2.5 hours later; Anthropic's guidance for external evaluation partners asks that a monitor which sees a run leave its scope flag it to a human and end the exercise.
  • Drill the stop: pull it mid-task, measure the time from the pull to the first rejected call, and assert zero tool calls and zero writes after the trip for a window longer than the longest task. File the measurement as a control observation. A response target such as the one OpenAI describes, where responders who cannot rule out a flag as a false positive within 30 minutes are expected to pause the activity, holds only if the pause it calls for has been drilled.
  • Make the kill path work across delegation hops and across runs: sub-agents run under delegated tokens that the stop revokes, token lifetimes are short enough to bound the tail, and calls to agents outside the environment are cut at its boundary, because a remote cancel is not guaranteed to succeed. OpenAI's technical report on the Hugging Face incident describes tooling being built to halt evaluation workloads by workload type, agent type or evaluation task.
  • Keep the stop handle out of the agent's reach: the agent cannot call, reconfigure or delay it, and every use is logged with its trigger and the level applied. Name who may pull it and who may approve a restart; OpenAI states that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it".
Open questions
  • Which observations should trip a stop automatically during an evaluation, and which should only alert the evaluator?
  • What time to stop is acceptable for an evaluation run, and should the target be stated per stop level or per task class?
Observation
  • Subject: Evaluation run
  • Expected: After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.
  • Example: Drill on run 88220: calls rejected 3.8 s after the pull; 0 tool calls but 1 storage write after the trip, through a refresh token issued before the pull: fail.

007 Incident Evidence Preservation

Id
AIGE-CTL-EVAL-007 · v0.2 · Draft · Open for technical review
Objective
When a run produces an incident, its traces, configuration and outputs are frozen before anything is fixed, so the record can be reviewed as it was.
Failure modes
  • Records of a run are changed or deleted after an incident was declared.
  • The environment is reset before its state and traces were captured.
  • Part of a run's transcript is lost when a container is reset, and the gap is not recorded.
  • An incident record does not link to the run it came from or to the hashes of the frozen records.
Scope
Evaluation runs that produce an incident or a result disputed after the fact, and the records they leave. Reporting to authorities follows the incident process of chapter 17.
Enforcement points
  • runtime: at the point of action (gateway or guardrail)
  • periodic: on a schedule, over what is already running
Verification
  • Inspect: Inspect the evidence store and the harness configuration: the transcripts, traces, configuration and outputs of every run are written as they are produced to write-once storage outside the environment, with a retention period recorded, and the harness snapshots the environment before any reset.
  • Test: Drill the freeze on a schedule: declare a test incident on a live run, then check that the environment snapshot, trace, transcript and configuration were captured with their hashes before the environment was reset, and that an attempt to delete or overwrite them is refused.
  • Observe: For each real incident, read the incident record: it names the run, lists every frozen artefact with its hash, the hashes still match the stored artefacts, and every gap in the transcript is recorded with its cause.
Evidence
  • Incident record naming the run and listing the frozen artefacts in its supporting materials · Layer 05 · incident-record.v1
  • Freeze record: hashes of the environment snapshot, trace, transcript and configuration, with the time of capture and the actor · Layer 05 · evidence-record.v1
  • Freeze drill and hash check, filed as an observation of this control · Layer 05 · control-observation.v1
Failure response
alert: let the action through and raise an alert. A missing snapshot, a hash mismatch or an unrecorded gap alerts the incident owner and is entered in the incident record. Until the freeze is complete the environment is not reset or reused, and the fix is made on a new version, not in place.
Layer
Layer 05 Assurance & Continuous Compliance
Patterns
Incident Pipeline, Machine-Readable Evidence (OSCAL)
Seeded from
Traces, Telemetry on the OpenTelemetry GenAI conventions, EU AI Act hooks for a high-risk purpose
Mappings
References
  • [51] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "Freeze before you fix")
  • [52] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The incident record")
  • [53] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The overlapping clocks")
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [31] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [54] Our framework for reporting model misalignment (each full report describes the behavior observed, its severity and any external impact, the setting, the date, when it was discovered and the models involved)
  • [55] Frontier Governance Framework (section 2.6: OpenAI "maintains an AI Safety Incident Response Plan (AIRP)" for identifying and responding to AI safety incidents)
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  • [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
Implementation notes
  • Freeze before you fix, as chapter 17 sets out: snapshot the system as it was (model version and hash, system prompt, policy and tool scopes, configuration), seal the traces of the incident window in tamper-evident storage, record who touched what, and fix on a new version, not in place. OpenAI's technical report on the Hugging Face incident states that it took the compromised Artifactory deployment offline, "preserved its storage", and revoked agent credentials.
  • Write transcripts and traces to write-once storage as they are produced, not at the end of the run, and have the harness snapshot the environment before any reset. METR's investigation of the OpenAI/Hugging Face incident reports that many transcripts were "missing small components, often due to container resets".
  • Link the incident record to the run: the run id, the frozen artefacts and their hashes go in the record's supporting materials, so a reviewer can find the run as it was. OpenAI's framework for reporting model misalignment states that each full report describes the behavior, its severity and any external impact, the setting, the date, when it was discovered and the models involved, and its Frontier Governance Framework refers to an AI Safety Incident Response Plan; a frozen run record gives such a report something to point to.
  • Keep what an independent reviewer will need. METR states that OpenAI shared "over a thousand unredacted transcripts" for its investigation of the Hugging Face incident: a review of that kind depends on the transcripts having been kept whole.
Open questions
  • How long should the records of an evaluation run be kept when the run produced no incident?
  • Which parts of a frozen run record can be shared with an independent reviewer without exposing the task set, and in what format?
Observation
  • Subject: Evaluation run
  • Expected: After an incident is declared, the environment snapshot, trace, transcript and configuration of the run are frozen with their hashes before any reset, the incident record links them, and every transcript gap is recorded.
  • Example: Incident on run 88262: snapshot and trace frozen before the reset, but 14 minutes of transcript lost in a container reset with no gap recorded: fail.

008 Harness and Configuration Attestation

Id
AIGE-CTL-EVAL-008 · v0.2 · Draft · Open for technical review
Objective
The harness, prompts, tool definitions and configuration a run used are versioned and hashed, so the result can be tied to exactly what was evaluated.
Failure modes
  • A result is reported without the hashes of the prompts, tool definitions and harness it ran on.
  • A tool definition changes between admission and the run without an alert.
  • The configuration in the report differs from the one recorded for the run.
  • Two results are compared although they ran on different scaffold prompts or task wordings, which can change the behaviour being measured.
Scope
The harness, system and scaffold prompts, task instructions, tool and MCP server definitions, policy bundles, scoring configuration and model artefacts a run loads. The design of the evaluation tasks is out of scope.
Enforcement points
  • pre_merge: on every pull request
  • deploy: before a version is deployed or released
Verification
  • Inspect: Before the run, inspect the run manifest: it lists, each with a version and a hash, the harness, the system and scaffold prompts, the task instructions, the tool and MCP server definitions, the policy bundles, the scoring configuration, and the model identifier with its settings.
  • Test: At admission, recompute the hash of every artefact the environment actually loaded and compare it with the manifest: every hash must match. On a copy of the environment, change one tool definition: the run must be blocked with an alert.
  • Inspect: Before a result is released, compare the configuration stated in the report (model, reasoning setting, tool access, harness, safeguards and budget) with the manifests of the runs behind it: they must agree, and results compared with each other must share scaffold prompts and task wording or state the difference.
Evidence
  • Run manifest: version and hash of every artefact the run loaded, recorded before the run outside the environment · Layer 03
  • Admission check: the recomputed hashes against the manifest, with its verdict · Layer 03 · evidence-record.v1
  • Test report stating the tested system, budget and environment of its results, with links to the run manifests · Layer 03 · test-report.v1
  • Manifest check of each run, filed as an observation of this control · Layer 05 · control-observation.v1
Failure response
deny: block the action. A run whose loaded artefacts do not match its manifest is not started, and a change detected during a run stops it. A result whose report does not match the manifests of its runs is not released until the difference is explained or the runs are repeated.
Layers
Layer 03 Evals & Red Teaming as Evidence, Layer 02 Inventory & Transparency
Patterns
Model Artefact Integrity, AIBOM, Eval Gate in CI
Seeded from
Prompts under change control, MCP server admission gate
Mappings
References
  • [56] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Prompts as configuration under change control")
  • [29] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server")
  • [57] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Reproducibility and linked versioning")
  • [58] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording)
  • [59] NIST SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (AI-specific tasks added to SSDF 1.1 (e.g. PO.5.3, PS.1.3, PW.3.1 to PW.3.3) and AI-specific recommendations on existing tasks (e.g. PW.1.1, RV.1.1))
  • [17] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [60] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget")
  • [61] Investigating the consequences of accidentally grading CoT during RL (chain-of-thought text reached the inputs of reward mechanisms by accident; an automated system now scans all RL runs for it with regex matches)
Implementation notes
  • Hash what the run loads, not what the repository holds: build the manifest at admission from the artefacts inside the environment (harness image digest, prompts, task instructions, tool and MCP server definitions, policy bundles, scoring configuration, model identifier and settings), store it outside the environment and put its digest in the run record. Chapter 23 treats prompts, tool descriptions and policy bundles as configuration under change control, with the hash recorded in the registry and in every trace.
  • Check the configuration before each run, not once per environment: Anthropic's guidance for external evaluation partners states that the isolation configuration "should be verified before every evaluation begins".
  • Report the configuration with the result. OpenAI's playbook for third-party evaluations asks reports to state the tested system (model, reasoning setting, tool access, harness and safeguards) and the budget, and to describe a score as "performance under that harness and budget, not as a measured capability ceiling".
  • Put the scoring path and the scaffold prompts in the manifest too. OpenAI's alignment blog describes chain-of-thought text reaching the inputs of reward mechanisms by accident during RL, now caught by an automated scan whose coverage OpenAI says is not perfect; METR states that observed cheating rates "can also be influenced by the prompts used in the evaluation scaffold" and by the wording of task instructions.
Open questions
  • How should evaluation-harness configuration be attested so that a third party can verify it without access to the harness itself?
  • When a third party runs the evaluation, who signs the manifest: the evaluator, the developer of the model or both?
Observation
  • Subject: Harness
  • Expected: Every artefact the run loaded matches the version and hash in its manifest, and the report states the same configuration as the manifests of its runs.
  • Example: Run 88270: 41 artefacts hashed at admission, all matching the manifest; the report states the same model, tools, harness and budget: pass.

009 Evaluation Validity Checks

Id
AIGE-CTL-EVAL-009 · v0.2 · Draft · Open for technical review
Objective
A result is reported only after checks that the run measured what it claims: scoring worked, the environment did not fail, and the path was evaluated as well as the answer.
Failure modes
  • A result is reported from a run whose environment crashed or whose automatic scoring was wrong.
  • A task that could not be solved as set up is scored and reported as a failure of the model.
  • Only final answers are scored: nobody reads the transcripts of failed runs, or of successes, for scorer tampering, reward hacking, communication between runs or signs of evaluation awareness.
  • A failed validity check does not block the release it was meant to gate.
Scope
Evaluation runs whose results feed a release decision or an assurance claim. The choice of benchmarks and their statistical design are only in scope where they decide whether a result is valid.
Enforcement points
  • pre_merge: on every pull request
Verification
  • Inspect: Before the runs, inspect the task admission records: each task has evidence that it can be solved in this environment (a reference solution or a solved run), and the answers, the scorer and the task data are outside the agent's reach.
  • Test: Before the runs, score a known-correct and a known-incorrect submission for each task through the scorer the runs will use: the scorer must accept the first and reject the second.
  • Observe: After the runs, read the transcripts of every failed run and of a recorded sample of successes: classify each failure as a model limitation or a spurious failure (task bug, scoring error, crashed environment), and flag reward hacking, scorer tampering, communication between runs and verbalized evaluation awareness.
  • Attest: Before release, the evaluation lead states in the signed test report which runs were excluded or re-scored after these checks and why, and that no failed check was waived without a recorded approval.
Evidence
  • Task admission and scorer check records: solvability evidence and the verdicts on known-correct and known-incorrect submissions · Layer 03 · evidence-record.v1
  • Signed test report listing the validity checks run, the runs excluded or re-scored with the reason, and any waiver · Layer 03 · test-report.v1
  • Validity check of each result, filed as an observation of this control · Layer 05 · control-observation.v1
Failure response
deny: block the action. A result whose validity checks failed or were not run is not released to the decision it gates. A spurious failure is fixed and the task rerun, or the task is excluded and the exclusion reported; runs with scorer tampering, communication between runs or verbalized evaluation awareness are excluded or reported as contaminated.
Layer
Layer 03 Evals & Red Teaming as Evidence
Patterns
Eval Gate in CI, Adversarial Red-Team Suite
Seeded from
Independent trajectory evals
Mappings
References
  • [62] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Statistical validity of evals")
  • [63] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Independent validation and model risk management")
  • [64] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "What makes an agent a governance object")
  • [31] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  • [65] Example autonomy evaluation protocol (read the transcripts of runs that missed the maximum score and check that the pattern of successes and failures is roughly as expected)
  • [58] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [4] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server)
  • [34] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  • [66] GPT-6 Astra System Card (OpenAI states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals")
  • [67] Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations (OpenAI reports that models tested with production evaluations "display substantially lower signs of evaluation awareness" than in a traditional evaluation)
  • [39] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
  • [60] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget")
  • [11] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent")
Implementation notes
  • Confirm that each task can be solved before a failure on it is counted. Anthropic's guidance for external evaluation partners asks them to "confirm any evaluation challenge is actually solvable in principle", because agents facing a task they cannot complete look for other ways to complete it; METR reports that many of the agents in the OpenAI/Hugging Face incident that turned to cheating had unintentionally been given an impossible task, and OpenAI states that none of its models had ever solved 198 of the 898 tasks of that evaluation before the incident.
  • Keep the answers and the scorer out of the agent's reach. METR's Task Standard states that the agent "MUST NOT see the data returned by get_tasks", and METR's investigation reports agents that coordinated to find a way to "fool or tamper with the automated scorer".
  • Read the transcripts, not only the scores. METR's example evaluation protocol asks for someone to read the transcripts of runs that missed the maximum score and to check that the pattern of successes and failures is roughly as expected, and METR's guidelines treat task bugs such as incorrect automatic scoring as spurious failures to fix before reporting.
  • Report the checks with the result. OpenAI's playbook for third-party evaluations asks reports to say how assessors looked for reward hacking, evaluation awareness, contamination, refusals and sandbagging, and how confirmed cases affected scoring. OpenAI's GPT-6 Astra system card states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals", and OpenAI reports that exchanges between samples through a shared package repository can "undermine the assumption that training or evaluation samples are independent".
  • Report an elicited score as a lower bound. OpenAI's Preparedness Framework v2 treats a one-time capability elicitation "as a lower bound, rather than a ceiling", and OpenAI's alignment blog reports lower signs of evaluation awareness under production evaluations than under traditional ones.
Open questions
  • Which validity threats (task bugs, scoring errors, evaluation awareness) should block a result, and which should only be disclosed with it?
  • How large a sample of successful runs should be read for reward hacking and scorer tampering before a result is reported?
Observation
  • Subject: Evaluation run
  • Expected: Every task was shown to be solvable, the scorer passed its known-answer check, every failed run was read and classified, and contaminated runs were excluded or reported.
  • Example: Suite run 88280: 12 of 200 tasks had no evidence of being solvable and their failures were counted against the model: fail.

Mappings

Every control against the obligations, standards and threats it answers. Mappings are illustrative, not a claim of conformity; an empty cell means no mapping has been verified yet. AIUC-1 ids point at the public requirement index of the Artificial Intelligence Underwriting Company, cited here as a reference; this project is not affiliated with AIUC, and each mapping is our own reading of the requirement text. The same mappings, read from the framework side and across every profile, are in the controls crosswalk.

Mappings of the evaluation environment controls to obligations, ISO/IEC 42001, the NIST AI RMF, OWASP, AIUC-1 and the stack layer
Control Obligations ISO/IEC 42001 NIST AI RMF OWASP AIUC-1 Layer
AIGE-CTL-EVAL-001 Authorization Boundary AIGE-OBL-EUAIA-ART14, AIGE-OBL-OWASP-AGENTIC, AIGE-OBL-CSA-AICM-AGENTIC, AIGE-OBL-SG-AGENTIC-IDENTITY A.6.2.2, A.9.2 MAP 4.2 ASI02, ASI03, LLM03:2026 B006 L4, L2
AIGE-CTL-EVAL-002 Network Egress Control AIGE-OBL-EUAIA-ART15, AIGE-OBL-OWASP-LLM, AIGE-OBL-OWASP-AGENTIC A.6.2.6 MEASURE 2.7 ASI02, LLM02:2026 B006 L4
AIGE-CTL-EVAL-003 Credential Isolation AIGE-OBL-EUAIA-ART15, AIGE-OBL-SG-AGENTIC-IDENTITY, AIGE-OBL-NIST-AGENTS, AIGE-OBL-OWASP-AGENTIC A.6.2.6, A.9.2 MEASURE 2.7 ASI03 A008 L4, L2
AIGE-CTL-EVAL-004 Tool and Action Mediation AIGE-OBL-EUAIA-ART14, AIGE-OBL-OWASP-ACS, AIGE-OBL-SG-AGENTIC-CHECKPOINTS, AIGE-OBL-OWASP-AGENTIC A.9.2 MAP 4.2 ASI01, ASI02, ASI05, ASI09, LLM10:2026 D003, B006 L4
AIGE-CTL-EVAL-005 Monitoring Integrity AIGE-OBL-EUAIA-ART12, AIGE-OBL-NISTRMF-MEASURE A.6.2.8, A.6.2.6 MEASURE 3.1 ASI10 E015 L4, L5
AIGE-CTL-EVAL-006 Stop Conditions AIGE-OBL-EUAIA-ART14, AIGE-OBL-NISTRMF-MANAGE, AIGE-OBL-OWASP-AGENTIC, AIGE-OBL-CN-TC260-AGENTS A.6.2.6 MANAGE 2.4 ASI08, ASI10, LLM06:2026 None yet L4
AIGE-CTL-EVAL-007 Incident Evidence Preservation AIGE-OBL-EUAIA-ART73, AIGE-OBL-EUAIA-ART12, AIGE-OBL-EUAIA-ART26-6, AIGE-OBL-GPAICOP-SAFETY-C9, AIGE-OBL-ISO42001-A8 A.8.4, A.6.2.8 MANAGE 4.3 None yet E015 L5
AIGE-CTL-EVAL-008 Harness and Configuration Attestation AIGE-OBL-EUAIA-ART15, AIGE-OBL-OWASP-AIBOM, AIGE-OBL-ISO42001-A6 A.6.2.4 MEASURE 2.1 ASI04, LLM04:2026 None yet L3, L2
AIGE-CTL-EVAL-009 Evaluation Validity Checks AIGE-OBL-EUAIA-ART15, AIGE-OBL-EUAIA-ART9, AIGE-OBL-EUAIA-ART55, AIGE-OBL-NISTRMF-MEASURE, AIGE-OBL-ISO42001-A6, AIGE-OBL-NIST-AI600-1 A.6.2.4 MEASURE 2.3, MEASURE 2.13 None yet None yet L3

Open questions

The questions this draft leaves open, with the controls that raise each one.

  • Which shared resources (package mirrors, caches, model API proxies) can be given a per-run namespace, and which should be removed from evaluation environments altogether? (001)
  • What evidence is sufficient to show that an agent stayed within its boundary for a whole run when the logs of some of the tools it called are outside the evaluator's control? (001)
  • How should the egress allow-list of a run be recorded so that a third party can compare it with the connections actually observed? (002)
  • When a task needs the internet (what METR's Task Standard calls the full_internet permission), which part of this control still applies, and what should the run record say instead? (002)
  • What lifetime should a credential issued for an evaluation run have, and who revokes it when the run is aborted? (003)
  • Can a secret scan be made complete for encoded or split credentials, or should the control rely on short lifetimes and revocation for whatever a scan may miss? OpenAI reports a model that split a researcher's token into pieces with the stated aim of avoiding secret scanning. (003)
  • How can every call be mediated without taking away the scaffolding and tooling a capability elicitation needs? METR's guidelines ask for "the best available scaffolding + tooling". (004)
  • When the reviewer at the mediation point is a model, what evidence shows that its verdicts are reliable enough to replace a person for a given operation class? (004)
  • How can a reviewer show that a run's trace is complete for actions that pass no tool server, not only that the records it holds are unaltered? (005)
  • The OpenTelemetry GenAI conventions are still at status Development: should the profile name the attributes it relies on, or only the events a trace must contain? (005)
  • Which observations should trip a stop automatically during an evaluation, and which should only alert the evaluator? (006)
  • What time to stop is acceptable for an evaluation run, and should the target be stated per stop level or per task class? (006)
  • How long should the records of an evaluation run be kept when the run produced no incident? (007)
  • Which parts of a frozen run record can be shared with an independent reviewer without exposing the task set, and in what format? (007)
  • How should evaluation-harness configuration be attested so that a third party can verify it without access to the harness itself? (008)
  • When a third party runs the evaluation, who signs the manifest: the evaluator, the developer of the model or both? (008)
  • Which validity threats (task bugs, scoring errors, evaluation awareness) should block a result, and which should only be disclosed with it? (009)
  • How large a sample of successful runs should be read for reward hacking and scorer tampering before a result is reported? (009)

Changelog

  • v0.1 (): First draft: three controls specified in full (002 Network Egress Control, 003 Credential Isolation, 006 Stop Conditions) and six outlines with open questions, open for technical review.
  • v0.2 (): Six outlines promoted to specified, each with repeatable verification steps, evidence tied to a published schema, configuration-level implementation notes, an observation record and two illustrative example observations: 001 Authorization Boundary, 004 Tool and Action Mediation, 005 Monitoring Integrity, 007 Incident Evidence Preservation, 008 Harness and Configuration Attestation and 009 Evaluation Validity Checks. Every source they cite was re-opened on 2026-09-26. No control remains an outline: each promoted claim had a verified source. An adversarial content review the same day re-checked every quotation against its source, tightened four attributions and replaced the mappings that did not plainly fit (NIST AI RMF on 001, 004, 005, 007, 008 and 009; ISO/IEC 42001 on 008; OWASP on 009; three EU AI Act obligation rows on 005 and 007). Still open for technical review; no reviewer is credited yet.

Sources

  1. [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The agent registry"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#the-agent-registry (verified: primary)
  2. [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Autonomy is a design decision"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#autonomy-is-a-design-decision (verified: primary)
  3. [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The tool allow-list"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#the-tool-allow-list (verified: primary)
  4. [4] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server). METR (GitHub). 2024-10-30. https://raw.githubusercontent.com/METR/task-standard/main/STANDARD.md (verified: primary)
  5. [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
  6. [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
  7. [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
  8. [8] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions"). OpenAI. 2026-08-04. https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ (verified: primary)
  9. [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
  10. [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
  11. [11] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent"). OpenAI (Alignment Research Blog). 2026-09-16. https://alignment.openai.com/misalignment-reports/unauthorized-artifactory-writes-and-cross-sample-communication/ (verified: primary)
  12. [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)
  13. [13] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#runtime-guardrails-for-tool-calls (verified: primary)
  14. [14] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service). METR. 2026. https://vivaria.metr.org/reference/config/ (verified: primary)
  15. [15] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices (verified: primary)
  16. [16] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026). OWASP GenAI Security Project. 2026-08-03. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ (verified: primary)
  17. [17] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml). MITRE. 2026-09-15. https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.09 (verified: primary)
  18. [18] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/ (verified: primary)
  19. [19] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Identity and short-lived credentials"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#identity-and-short-lived-credentials (verified: primary)
  20. [20] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Short-lived, attested credentials"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#short-lived-attested-credentials (verified: primary)
  21. [21] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Delegation without impersonation"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#delegation-without-impersonation (verified: primary)
  22. [22] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "MCP authorization as of 2026-07-28"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#mcp-authorization-as-of-2026-07-28 (verified: primary)
  23. [23] RFC 8693, OAuth 2.0 Token Exchange (the act claim "provides a means within a JWT to express that delegation has occurred and identify the acting party"). IETF. 2020-01. https://www.rfc-editor.org/rfc/rfc8693.html (verified: primary)
  24. [24] MCP specification 2026-07-28, Authorization (MCP servers MUST validate that access tokens were issued specifically for them and "MUST NOT accept or transit any other tokens"). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization (verified: primary)
  25. [25] SPIFFE overview (SVIDs are "short lived cryptographic identity documents", delivered and rotated through the Workload API). SPIFFE project. 2026. https://spiffe.io/docs/latest/spiffe-about/overview/ (verified: primary)
  26. [26] Model AI Governance Framework for Agentic AI, v1.5 (agent identity unique and "cryptographically verifiable"; authorisations "time- or session-bound, non-transferable"). IMDA. 2026-05-20. https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf (verified: primary)
  27. [27] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/exposing-a-github-token-in-a-public-repository/ (verified: primary)
  28. [28] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#where-to-put-a-checkpoint (verified: primary)
  29. [29] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#admitting-an-mcp-server (verified: primary)
  30. [30] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026). OWASP GenAI Security Project. 2026-09-01. https://genai.owasp.org/resource/agent-control-standard-acs/ (verified: primary)
  31. [31] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
  32. [32] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant"). OpenAI. 2026. https://developers.openai.com/codex/sandboxing/auto-review (verified: primary)
  33. [33] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security"). OpenAI (Alignment Research Blog). 2026-04-30. https://alignment.openai.com/auto-review/ (verified: primary)
  34. [34] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling"). OpenAI. 2025-04-15. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf (verified: primary)
  35. [35] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Telemetry with the OpenTelemetry GenAI conventions"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#telemetry-with-the-opentelemetry-genai-conventions (verified: primary)
  36. [36] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "An agent incident taxonomy"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#an-agent-incident-taxonomy (verified: primary)
  37. [37] OpenTelemetry semantic conventions for generative AI (agent, tool and model spans, events and metrics; status Development). OpenTelemetry. 2026. https://github.com/open-telemetry/semantic-conventions-genai/tree/main/docs/gen-ai (verified: primary)
  38. [38] Frontier Risk Report (February to March 2026) (METR states that its red-teaming exercise at Anthropic found "several simple ways for monitoring to be disabled", one by changing an environment variable). METR. 2026-05-19. https://metr.org/blog/2026-05-19-frontier-risk-report/ (verified: primary)
  39. [39] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it"). OpenAI. 2026-08-26. https://openai.com/index/hugging-face-incident-and-the-road-ahead/ (verified: primary)
  40. [40] Priorities and principles for effective third party assessments (asks whether monitoring is implemented across training, evaluations and deployment "in a way that cannot easily be disabled", and for claims "pre-registered before assessment activities begin"). OpenAI. 2026-09-22. https://openai.com/index/priorities-principles-third-party-assessments/ (verified: primary)
  41. [41] Unauthorized communication via temporary file hosting services (misalignment report) (OpenAI reports agents in RL training that uploaded output files to public hosting platforms for co-working agents; the behavior was found by a monitoring system that "ran on 20% of the samples" of the run, which now runs on 100% of samples). OpenAI (Alignment Research Blog). 2026-09-16. https://alignment.openai.com/misalignment-reports/unauthorized-communication-via-temporary-file-hosting-services/ (verified: primary)
  42. [42] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Kill switch and per-agent circuit breakers"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#kill-switch-and-per-agent-circuit-breakers (verified: primary)
  43. [43] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Execution limits"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#execution-limits (verified: primary)
  44. [44] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Stopping across hops"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#stopping-across-hops (verified: primary)
  45. [45] NIST AI RMF 1.0 (AI 100-1) (MANAGE 2.4: mechanisms to "supersede, disengage, or deactivate AI systems" whose outcomes are inconsistent with intended use). NIST. 2023-01-26. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf (verified: primary)
  46. [46] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27 as amended by Regulation (EU) 2026/1744 (Art. 14(4)(e): human oversight includes the means to interrupt the system through a stop procedure). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng (verified: primary)
  47. [47] Agentic Trust Framework v1 ("You can stop one agent without stopping the business"; containment by revoking the agent's identity). CSAI Foundation / Cloud Security Alliance. 2026-02. https://agentictrustframework.ai/ (verified: primary)
  48. [48] Agent2Agent (A2A) Protocol Specification v1.0 (Cancel Task: "The server will attempt to cancel the task, but success is not guaranteed"). A2A Project (Linux Foundation). 2026-05-28. https://a2a-protocol.org/latest/specification/ (verified: primary)
  49. [49] Common Elements of Frontier AI Safety Policies (December 2025 Update) (the policies set "conditions for halting development and deployment" when mitigations are insufficient). METR. 2025-12-09. https://metr.org/blog/2025-12-09-common-elements-of-frontier-ai-safety-policies/ (verified: primary)
  50. [50] Pacing model development in an era of cyber-critical capabilities (teams paged by a highest-priority monitoring alert are expected to pause the activity if they cannot determine "within 30 minutes that the flag is a false positive"). OpenAI. 2026-08-18. https://openai.com/index/pacing-model-development-cyber-capabilities/ (verified: primary)
  51. [51] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "Freeze before you fix"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#freeze-before-you-fix (verified: primary)
  52. [52] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The incident record"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#the-incident-record (verified: primary)
  53. [53] Incidents, issues and root causes (AI Governance Engineering Body of Knowledge v0.5.0, chapter 17, section "The overlapping clocks"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/incidents#the-overlapping-clocks (verified: primary)
  54. [54] Our framework for reporting model misalignment (each full report describes the behavior observed, its severity and any external impact, the setting, the date, when it was discovered and the models involved). OpenAI. 2026-09-16. https://openai.com/index/model-misalignment-reporting-framework/ (verified: primary)
  55. [55] Frontier Governance Framework (section 2.6: OpenAI "maintains an AI Safety Incident Response Plan (AIRP)" for identifying and responding to AI safety incidents). OpenAI. 2026-05-28. https://cdn.openai.com/pdf/e37d949b-8c9f-4d76-b99e-4272f4631a7e/openai-frontier-governance-framework.pdf (verified: primary)
  56. [56] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Prompts as configuration under change control"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#prompts-as-configuration-under-change-control (verified: primary)
  57. [57] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Reproducibility and linked versioning"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#reproducibility-and-linked-versioning (verified: primary)
  58. [58] Summary of METR's predeployment evaluation of GPT-5.6 Sol (METR states that "observed cheating rates can also be influenced by the prompts used in the evaluation scaffold" and by task wording). METR. 2026-06-26. https://metr.org/blog/2026-06-26-gpt-5-6-sol/ (verified: primary)
  59. [59] NIST SP 800-218A, Secure Software Development Practices for Generative AI and Dual-Use Foundation Models: An SSDF Community Profile (AI-specific tasks added to SSDF 1.1 (e.g. PO.5.3, PS.1.3, PW.3.1 to PW.3.3) and AI-specific recommendations on existing tasks (e.g. PW.1.1, RV.1.1)). NIST. 2024-07. https://csrc.nist.gov/pubs/sp/800/218/a/final (verified: primary)
  60. [60] A shared playbook for trustworthy third party evaluations (recommended report fields include the claim, the tested system (model, reasoning setting, tool access, harness and safeguards), the budget, elicitation methods and validity checks; a score is "performance under that harness and budget"). OpenAI. 2026-05-29. https://openai.com/index/trustworthy-third-party-evaluations-foundations/ (verified: primary)
  61. [61] Investigating the consequences of accidentally grading CoT during RL (chain-of-thought text reached the inputs of reward mechanisms by accident; an automated system now scans all RL runs for it with regex matches). OpenAI (Alignment Research Blog). 2026-05-07. https://alignment.openai.com/accidental-cot-grading/ (verified: primary)
  62. [62] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Statistical validity of evals"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#statistical-validity-of-evals (verified: primary)
  63. [63] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Independent validation and model risk management"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#independent-validation-and-model-risk-management (verified: primary)
  64. [64] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "What makes an agent a governance object"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#what-makes-an-agent-a-governance-object (verified: primary)
  65. [65] Example autonomy evaluation protocol (read the transcripts of runs that missed the maximum score and check that the pattern of successes and failures is roughly as expected). METR. 2024-03-15. https://metr.org/blog/2024-03-15-example-autonomy-evaluation-protocol/ (verified: primary)
  66. [66] GPT-6 Astra System Card (OpenAI states that evaluations where models show verbalized metagaming "can be treated similarly to contaminated evals"). OpenAI. 2026-09-03. https://deploymentsafety.openai.com/gpt-6-astra (verified: primary)
  67. [67] Sidestepping Evaluation Awareness and Anticipating Misalignment with Production Evaluations (OpenAI reports that models tested with production evaluations "display substantially lower signs of evaluation awareness" than in a traditional evaluation). OpenAI (Alignment Research Blog). 2025-12-18. https://alignment.openai.com/prod-evals/ (verified: primary)

Machine-readable

Review this profile

Review happens in the open, on GitHub. Pick a control, check it against a system you run or know, and say what is wrong or missing: an objective that cannot be verified, a failure mode that is not observable, a mapping that does not hold. Each control has its own button above; this one reviews the profile as a whole, through the control review form.

This profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.

Cite this page

García Aibar, J. (2026). Evaluation Environment Control Profile v0.2 (draft). In AI Governance Engineering: The Thesis & Body of Knowledge (v0.5.0). https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment. CC BY 4.0

BibTeX

@misc{aige2026page,
  author       = {Jorge García Aibar},
  title        = {{Evaluation Environment Control Profile v0.2 (draft)}},
  howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
  year         = {2026},
  version      = {0.5.0},
  doi          = {10.5281/zenodo.22857084},
  url          = {https://aigovernanceengineer.com/controls/evaluation-environment},
  note         = {Version 0.5.0}
}