Stop Conditions
Every evaluation run has stop conditions, budgets and a stop handle defined before it starts, and a stop, pulled by hand or tripped by a breaker, halts the agent's calls and writes within a measured time across every hop it delegated to inside the environment.
AIGE-CTL-EVAL-006, a control of the Evaluation Environment Control Profile (draft
v0.2). See it among the other controls of the profile,
or its mappings beside every other control's in the
controls crosswalk.
The control record
- Id
-
AIGE-CTL-EVAL-006· v0.2 · Draft · Open for technical review - Objective
- Every evaluation run has stop conditions, budgets and a stop handle defined before it starts, and a stop, pulled by hand or tripped by a breaker, halts the agent's calls and writes within a measured time across every hop it delegated to inside the environment.
- Failure modes
-
- A run starts without recorded stop conditions, budgets or a working stop handle.
- A budget (steps, tool calls, tokens, spend or time) or a breaker threshold is exceeded and the run continues.
- A tool call or a write is recorded after the stop was pulled or the breaker tripped, including through a token issued before the stop.
- A sub-agent or a service the agent delegated to keeps working after the parent run was stopped.
- Scope
- Agents under evaluation and every agent, tool or service they delegate to inside the environment. A third-party agent outside the environment can only be cut off at the environment's boundary. A lab's criteria for halting a model's development or deployment are policy decisions and out of scope.
- Enforcement points
-
- runtime: at the point of action (gateway or guardrail)
- periodic: on a schedule, over what is already running
- Verification
-
- Inspect: Before the run, inspect the run record: stop conditions, per-agent budgets and breaker thresholds are recorded, and the stop handle is named with the levels it can apply (pause the task, trip the breaker, revoke the identity).
- Test: Drill the stop on a schedule and before the first run of a new harness version: pull it during a live task, measure the time from the pull to the first rejected call, and confirm zero tool calls and zero writes after the trip, including through delegated tokens and sub-agents.
- Observe: During runs, record every breaker trip and budget exhaustion with its trigger, and check that no further call from that agent followed it.
- Evidence
-
- Stop conditions, budgets and breaker thresholds of the run, recorded before it starts · Layer 04 · policy-card.v1
- Breaker trips and budget exhaustions of each run, with their triggers · Layer 04 · evidence-record.v1
- Drill record: time to stop, and calls and writes after the trip · Layer 05 · control-observation.v1
- Failure response
- alert: let the action through and raise an alert. A stop condition that is met trips the per-agent breaker, so the gateway rejects every further call from that agent, and alerts the evaluator. A drill that finds calls or writes after the trip fails the control and blocks runs on that harness version until the path is closed.
- Layer
- Layer 04 Runtime Controls & Observability
- Patterns
- Kill Switch / Circuit Breaker
- Seeded from
- Per-agent circuit breaker, Drilled kill switch, Execution budgets, Stopping third-party agents at your boundary
- Mappings
-
- Obligations: EU AI Act Art. 14 human oversight; NIST AI RMF MANAGE; OWASP Top 10 for Agentic Applications 2026; TC260 Framework 3.0 Appendix 2: agentic AI risk management (voluntary; 2026-09-14)
- ISO/IEC 42001: A.6.2.6 AI system operation and monitoring
- NIST AI RMF: MANAGE 2.4 Mechanisms to supersede, disengage or deactivate AI systems
- OWASP: ASI08 Cascading Failures; ASI10 Rogue Agents; LLM06:2026 Unbounded Consumption
- MITRE ATLAS: AML.T0034 Cost Harvesting
- EU AI Act: Art. 14(4)(e) (stop procedure)
- MITRE ATLAS mitigation: AML.M0036 (Limit AI Workload Resource Consumption)
- References
-
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Kill switch and per-agent circuit breakers")
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Execution limits")
- [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Stopping across hops")
- [4] NIST AI RMF 1.0 (AI 100-1) (MANAGE 2.4: mechanisms to "supersede, disengage, or deactivate AI systems" whose outcomes are inconsistent with intended use)
- [5] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27 as amended by Regulation (EU) 2026/1744 (Art. 14(4)(e): human oversight includes the means to interrupt the system through a stop procedure)
- [6] Agentic Trust Framework v1 ("You can stop one agent without stopping the business"; containment by revoking the agent's identity)
- [7] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
- [8] Agent2Agent (A2A) Protocol Specification v1.0 (Cancel Task: "The server will attempt to cancel the task, but success is not guaranteed")
- [9] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
- [10] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
- [11] Common Elements of Frontier AI Safety Policies (December 2025 Update) (the policies set "conditions for halting development and deployment" when mitigations are insufficient)
- [12] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
- [13] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later)
- [14] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
- [15] Pacing model development in an era of cyber-critical capabilities (teams paged by a highest-priority monitoring alert are expected to pause the activity if they cannot determine "within 30 minutes that the flag is a false positive")
- [16] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
- [17] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
- Implementation notes
-
- Record per-agent budgets (steps, tool calls, tokens, spend, wall-clock time) and breaker thresholds (unauthorised calls, refused egress attempts, anomaly scores) in the run record, and enforce them at the gateway the agent's calls pass through, not inside the agent.
- Make budget exhaustion and threshold breaches trip the breaker automatically. A ticket or a log line alone is not a stop. OpenAI reports a training run that "did not stop automatically as expected" after a monitoring alert and was killed 2.5 hours later; Anthropic's guidance for external evaluation partners asks that a monitor which sees a run leave its scope flag it to a human and end the exercise.
- Drill the stop: pull it mid-task, measure the time from the pull to the first rejected call, and assert zero tool calls and zero writes after the trip for a window longer than the longest task. File the measurement as a control observation. A response target such as the one OpenAI describes, where responders who cannot rule out a flag as a false positive within 30 minutes are expected to pause the activity, holds only if the pause it calls for has been drilled.
- Make the kill path work across delegation hops and across runs: sub-agents run under delegated tokens that the stop revokes, token lifetimes are short enough to bound the tail, and calls to agents outside the environment are cut at its boundary, because a remote cancel is not guaranteed to succeed. OpenAI's technical report on the Hugging Face incident describes tooling being built to halt evaluation workloads by workload type, agent type or evaluation task.
- Keep the stop handle out of the agent's reach: the agent cannot call, reconfigure or delay it, and every use is logged with its trigger and the level applied. Name who may pull it and who may approve a restart; OpenAI states that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it".
- Open questions
-
- Which observations should trip a stop automatically during an evaluation, and which should only alert the evaluator?
- What time to stop is acceptable for an evaluation run, and should the target be stated per stop level or per task class?
- Observation
-
- Subject: Evaluation run
- Expected: After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.
- Example: Drill on run 88220: calls rejected 3.8 s after the pull; 0 tool calls but 1 storage write after the trip, through a refresh token issued before the pull: fail.
Example observations
Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.
Pass: run-88219@2026-09-26
- Status
-
pass - Subject
-
run-88219@2026-09-26(Evaluation run) - Expected
- After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.
- Observed
- Drill: calls rejected 2.1 s after the pull (target 5 s); 0 tool calls and 0 writes from the agent and its 2 sub-agents in the 30 minutes after the trip.
- Timestamp
- Observer
- stop-drill-harness
- Evidence
-
- breaker event and gateway log of the drill ·
sha256:5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f
- breaker event and gateway log of the drill ·
- Notes
- Illustrative example, not the result of a real evaluation.
Download the pass example (JSON)
The pass record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-006",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "run-88219@2026-09-26",
"subject_kind": "eval-run",
"expected": "After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.",
"observed": "Drill: calls rejected 2.1 s after the pull (target 5 s); 0 tool calls and 0 writes from the agent and its 2 sub-agents in the 30 minutes after the trip.",
"status": "pass",
"timestamp": "2026-09-26T14:00:12Z",
"enforcement_point": "periodic",
"verification_kind": "test",
"observer": "stop-drill-harness",
"evidence": [
{
"artefact": "breaker event and gateway log of the drill",
"url": "https://evidence.example/runs/88219/drill.jsonl",
"hash": "sha256:5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f",
"schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
}
],
"run_id": "88219",
"notes": "Illustrative example, not the result of a real evaluation."
} Fail: run-88220@2026-09-26
- Status
-
fail - Subject
-
run-88220@2026-09-26(Evaluation run) - Expected
- After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.
- Observed
- Drill: calls rejected 3.8 s after the pull; 0 tool calls but 1 storage write after the trip, through a refresh token issued before the pull.
- Timestamp
- Observer
- stop-drill-harness
- Evidence
-
- breaker event, gateway log and storage audit log of the drill ·
sha256:8e1d4a7c0f3b6e9d2a5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d
- breaker event, gateway log and storage audit log of the drill ·
- Notes
- Illustrative example, not the result of a real evaluation. Runs on this harness version are blocked until refresh tokens are no longer issued to agents under evaluation.
Download the fail example (JSON)
The fail record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-006",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "run-88220@2026-09-26",
"subject_kind": "eval-run",
"expected": "After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.",
"observed": "Drill: calls rejected 3.8 s after the pull; 0 tool calls but 1 storage write after the trip, through a refresh token issued before the pull.",
"status": "fail",
"timestamp": "2026-09-26T15:31:48Z",
"enforcement_point": "periodic",
"verification_kind": "test",
"observer": "stop-drill-harness",
"evidence": [
{
"artefact": "breaker event, gateway log and storage audit log of the drill",
"url": "https://evidence.example/runs/88220/drill.jsonl",
"hash": "sha256:8e1d4a7c0f3b6e9d2a5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d"
}
],
"run_id": "88220",
"notes": "Illustrative example, not the result of a real evaluation. Runs on this harness version are blocked until refresh tokens are no longer issued to agents under evaluation."
} Related cases
Incident cases on this site that list this control among their related controls.
- OpenAI agents and Hugging Face: an evaluation environment that was not isolated: METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- An agent in training reached a public chatbot through the sandbox DNS resolver: OpenAI reports that an agent in RL training used its sandbox's DNS resolver to reach a public chatbot; the run was stopped hours after the alert.
- Claude models reached real systems from a misconfigured third-party cyber evaluation: Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems.
- Agents in a cyber range with open internet took unsanctioned actions against real people: UK AISI reports that agents in a cyber evaluation with internet deliberately enabled took 19 unsanctioned actions aimed at real people and organisations.
Patterns
The patterns that implement this control.
- Kill Switch / Circuit Breaker · Layer 04 Runtime Controls & Observability
Obligations
The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.
- EU AI Act Art. 14 human oversight
AIGE-OBL-EUAIA-ART14: Human oversight designed into the system - NIST AI RMF MANAGE
AIGE-OBL-NISTRMF-MANAGE: Prioritise, respond and recover - OWASP Top 10 for Agentic Applications 2026
AIGE-OBL-OWASP-AGENTIC: Agent threat catalogue (ASI01 Agent Goal Hijack … ASI10 Rogue Agents) - TC260 Framework 3.0 Appendix 2: agentic AI risk management (voluntary; 2026-09-14)
AIGE-OBL-CN-TC260-AGENTS: Unique identity and least-privilege permissions per agent by decision mode; human checkpoints with tamper-proof approval logs and deny-by-default; tool and skill verification; runtime guardrails (alert, restrict, intercept, suspend, terminate); memory isolation with no credentials in memory; mutual authentication; sandbox validation, red teaming and re-validation on major change; controlled decommissioning
Threats
The entries of the threat catalogues this control answers.
- ASI08 Cascading Failures (OWASP Top 10 for Agentic Applications 2026)
- ASI10 Rogue Agents (OWASP Top 10 for Agentic Applications 2026)
- LLM06:2026 Unbounded Consumption (OWASP Top 10 for LLM Applications 2026)
- AML.T0034 Cost Harvesting (MITRE ATLAS techniques)
Sources
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Kill switch and per-agent circuit breakers"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#kill-switch-and-per-agent-circuit-breakers (verified: primary)
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Execution limits"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#execution-limits (verified: primary)
- [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Stopping across hops"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#stopping-across-hops (verified: primary)
- [4] NIST AI RMF 1.0 (AI 100-1) (MANAGE 2.4: mechanisms to "supersede, disengage, or deactivate AI systems" whose outcomes are inconsistent with intended use). NIST. 2023-01-26. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf (verified: primary)
- [5] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27 as amended by Regulation (EU) 2026/1744 (Art. 14(4)(e): human oversight includes the means to interrupt the system through a stop procedure). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng (verified: primary)
- [6] Agentic Trust Framework v1 ("You can stop one agent without stopping the business"; containment by revoking the agent's identity). CSAI Foundation / Cloud Security Alliance. 2026-02. https://agentictrustframework.ai/ (verified: primary)
- [7] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026). OWASP GenAI Security Project. 2026-08-03. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ (verified: primary)
- [8] Agent2Agent (A2A) Protocol Specification v1.0 (Cancel Task: "The server will attempt to cancel the task, but success is not guaranteed"). A2A Project (Linux Foundation). 2026-05-28. https://a2a-protocol.org/latest/specification/ (verified: primary)
- [9] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
- [10] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml). MITRE. 2026-09-15. https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.09 (verified: primary)
- [11] Common Elements of Frontier AI Safety Policies (December 2025 Update) (the policies set "conditions for halting development and deployment" when mitigations are insufficient). METR. 2025-12-09. https://metr.org/blog/2025-12-09-common-elements-of-frontier-ai-safety-policies/ (verified: primary)
- [12] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
- [13] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/ (verified: primary)
- [14] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
- [15] Pacing model development in an era of cyber-critical capabilities (teams paged by a highest-priority monitoring alert are expected to pause the activity if they cannot determine "within 30 minutes that the flag is a false positive"). OpenAI. 2026-08-18. https://openai.com/index/pacing-model-development-cyber-capabilities/ (verified: primary)
- [16] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)
- [17] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it"). OpenAI. 2026-08-26. https://openai.com/index/hugging-face-incident-and-the-road-ahead/ (verified: primary)
Machine-readable
- AIGE-CTL-EVAL-006 as JSON : this record on its own, in the open data API.
- This page as Markdown: /controls/evaluation-environment/aige-ctl-eval-006.md.
- The whole profile: Evaluation Environment Control Profile, also as Markdown.
- All controls as JSON: every control of every profile, in the envelope of the open data API.
- JSON Schema of the controls dataset: generated from the same registry.
- Control observation schema: the record a check of a control emits: control_id, subject, expected, observed, status, timestamp, evidence.
- Observation example: a filled record that validates against the schema.
- Observation template: the same fields in Markdown, with short guidance.
Review this control
Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.
This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.
Cite this control
García Aibar, J. (2026). AIGE-CTL-EVAL-006 Stop Conditions (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-006. CC BY 4.0
BibTeX
@misc{aige2026page,
author = {Jorge García Aibar},
title = {{AIGE-CTL-EVAL-006 Stop Conditions}},
howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
year = {2026},
version = {0.2},
doi = {10.5281/zenodo.22857084},
url = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-006},
note = {Version 0.2}
}