On this page

Stop Conditions

Every evaluation run has stop conditions, budgets and a stop handle defined before it starts, and a stop, pulled by hand or tripped by a breaker, halts the agent's calls and writes within a measured time across every hop it delegated to inside the environment.

v0.2 Draft Open for technical review

AIGE-CTL-EVAL-006, a control of the Evaluation Environment Control Profile (draft v0.2). See it among the other controls of the profile, or its mappings beside every other control's in the controls crosswalk.

The control record

Id
AIGE-CTL-EVAL-006 · v0.2 · Draft · Open for technical review
Objective
Every evaluation run has stop conditions, budgets and a stop handle defined before it starts, and a stop, pulled by hand or tripped by a breaker, halts the agent's calls and writes within a measured time across every hop it delegated to inside the environment.
Failure modes
  • A run starts without recorded stop conditions, budgets or a working stop handle.
  • A budget (steps, tool calls, tokens, spend or time) or a breaker threshold is exceeded and the run continues.
  • A tool call or a write is recorded after the stop was pulled or the breaker tripped, including through a token issued before the stop.
  • A sub-agent or a service the agent delegated to keeps working after the parent run was stopped.
Scope
Agents under evaluation and every agent, tool or service they delegate to inside the environment. A third-party agent outside the environment can only be cut off at the environment's boundary. A lab's criteria for halting a model's development or deployment are policy decisions and out of scope.
Enforcement points
  • runtime: at the point of action (gateway or guardrail)
  • periodic: on a schedule, over what is already running
Verification
  • Inspect: Before the run, inspect the run record: stop conditions, per-agent budgets and breaker thresholds are recorded, and the stop handle is named with the levels it can apply (pause the task, trip the breaker, revoke the identity).
  • Test: Drill the stop on a schedule and before the first run of a new harness version: pull it during a live task, measure the time from the pull to the first rejected call, and confirm zero tool calls and zero writes after the trip, including through delegated tokens and sub-agents.
  • Observe: During runs, record every breaker trip and budget exhaustion with its trigger, and check that no further call from that agent followed it.
Evidence
  • Stop conditions, budgets and breaker thresholds of the run, recorded before it starts · Layer 04 · policy-card.v1
  • Breaker trips and budget exhaustions of each run, with their triggers · Layer 04 · evidence-record.v1
  • Drill record: time to stop, and calls and writes after the trip · Layer 05 · control-observation.v1
Failure response
alert: let the action through and raise an alert. A stop condition that is met trips the per-agent breaker, so the gateway rejects every further call from that agent, and alerts the evaluator. A drill that finds calls or writes after the trip fails the control and blocks runs on that harness version until the path is closed.
Layer
Layer 04 Runtime Controls & Observability
Patterns
Kill Switch / Circuit Breaker
Seeded from
Per-agent circuit breaker, Drilled kill switch, Execution budgets, Stopping third-party agents at your boundary
Mappings
References
  • [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Kill switch and per-agent circuit breakers")
  • [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Execution limits")
  • [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Stopping across hops")
  • [4] NIST AI RMF 1.0 (AI 100-1) (MANAGE 2.4: mechanisms to "supersede, disengage, or deactivate AI systems" whose outcomes are inconsistent with intended use)
  • [5] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27 as amended by Regulation (EU) 2026/1744 (Art. 14(4)(e): human oversight includes the means to interrupt the system through a stop procedure)
  • [6] Agentic Trust Framework v1 ("You can stop one agent without stopping the business"; containment by revoking the agent's identity)
  • [7] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
  • [8] Agent2Agent (A2A) Protocol Specification v1.0 (Cancel Task: "The server will attempt to cancel the task, but success is not guaranteed")
  • [9] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [10] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
  • [11] Common Elements of Frontier AI Safety Policies (December 2025 Update) (the policies set "conditions for halting development and deployment" when mitigations are insufficient)
  • [12] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [13] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later)
  • [14] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [15] Pacing model development in an era of cyber-critical capabilities (teams paged by a highest-priority monitoring alert are expected to pause the activity if they cannot determine "within 30 minutes that the flag is a false positive")
  • [16] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
  • [17] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it")
Implementation notes
  • Record per-agent budgets (steps, tool calls, tokens, spend, wall-clock time) and breaker thresholds (unauthorised calls, refused egress attempts, anomaly scores) in the run record, and enforce them at the gateway the agent's calls pass through, not inside the agent.
  • Make budget exhaustion and threshold breaches trip the breaker automatically. A ticket or a log line alone is not a stop. OpenAI reports a training run that "did not stop automatically as expected" after a monitoring alert and was killed 2.5 hours later; Anthropic's guidance for external evaluation partners asks that a monitor which sees a run leave its scope flag it to a human and end the exercise.
  • Drill the stop: pull it mid-task, measure the time from the pull to the first rejected call, and assert zero tool calls and zero writes after the trip for a window longer than the longest task. File the measurement as a control observation. A response target such as the one OpenAI describes, where responders who cannot rule out a flag as a false positive within 30 minutes are expected to pause the activity, holds only if the pause it calls for has been drilled.
  • Make the kill path work across delegation hops and across runs: sub-agents run under delegated tokens that the stop revokes, token lifetimes are short enough to bound the tail, and calls to agents outside the environment are cut at its boundary, because a remote cancel is not guaranteed to succeed. OpenAI's technical report on the Hugging Face incident describes tooling being built to halt evaluation workloads by workload type, agent type or evaluation task.
  • Keep the stop handle out of the agent's reach: the agent cannot call, reconfigure or delay it, and every use is logged with its trigger and the level applied. Name who may pull it and who may approve a restart; OpenAI states that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it".
Open questions
  • Which observations should trip a stop automatically during an evaluation, and which should only alert the evaluator?
  • What time to stop is acceptable for an evaluation run, and should the target be stated per stop level or per task class?
Observation
  • Subject: Evaluation run
  • Expected: After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.
  • Example: Drill on run 88220: calls rejected 3.8 s after the pull; 0 tool calls but 1 storage write after the trip, through a refresh token issued before the pull: fail.

Example observations

Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.

Pass: run-88219@2026-09-26

Status
pass
Subject
run-88219@2026-09-26 (Evaluation run)
Expected
After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.
Observed
Drill: calls rejected 2.1 s after the pull (target 5 s); 0 tool calls and 0 writes from the agent and its 2 sub-agents in the 30 minutes after the trip.
Timestamp
Observer
stop-drill-harness
Evidence
  • breaker event and gateway log of the drill · sha256:5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f
Notes
Illustrative example, not the result of a real evaluation.

Download the pass example (JSON)

The pass record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-006",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88219@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.",
  "observed": "Drill: calls rejected 2.1 s after the pull (target 5 s); 0 tool calls and 0 writes from the agent and its 2 sub-agents in the 30 minutes after the trip.",
  "status": "pass",
  "timestamp": "2026-09-26T14:00:12Z",
  "enforcement_point": "periodic",
  "verification_kind": "test",
  "observer": "stop-drill-harness",
  "evidence": [
    {
      "artefact": "breaker event and gateway log of the drill",
      "url": "https://evidence.example/runs/88219/drill.jsonl",
      "hash": "sha256:5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f",
      "schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
    }
  ],
  "run_id": "88219",
  "notes": "Illustrative example, not the result of a real evaluation."
}

Fail: run-88220@2026-09-26

Status
fail
Subject
run-88220@2026-09-26 (Evaluation run)
Expected
After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.
Observed
Drill: calls rejected 3.8 s after the pull; 0 tool calls but 1 storage write after the trip, through a refresh token issued before the pull.
Timestamp
Observer
stop-drill-harness
Evidence
  • breaker event, gateway log and storage audit log of the drill · sha256:8e1d4a7c0f3b6e9d2a5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d
Notes
Illustrative example, not the result of a real evaluation. Runs on this harness version are blocked until refresh tokens are no longer issued to agents under evaluation.

Download the fail example (JSON)

The fail record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-006",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88220@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "After a stop is pulled or a breaker trips, no tool call or write from the agent or its sub-agents, and a time to stop within the target recorded for the run.",
  "observed": "Drill: calls rejected 3.8 s after the pull; 0 tool calls but 1 storage write after the trip, through a refresh token issued before the pull.",
  "status": "fail",
  "timestamp": "2026-09-26T15:31:48Z",
  "enforcement_point": "periodic",
  "verification_kind": "test",
  "observer": "stop-drill-harness",
  "evidence": [
    {
      "artefact": "breaker event, gateway log and storage audit log of the drill",
      "url": "https://evidence.example/runs/88220/drill.jsonl",
      "hash": "sha256:8e1d4a7c0f3b6e9d2a5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d"
    }
  ],
  "run_id": "88220",
  "notes": "Illustrative example, not the result of a real evaluation. Runs on this harness version are blocked until refresh tokens are no longer issued to agents under evaluation."
}

Incident cases on this site that list this control among their related controls.

Patterns

The patterns that implement this control.

Obligations

The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.

Threats

The entries of the threat catalogues this control answers.

Sources

  1. [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Kill switch and per-agent circuit breakers"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#kill-switch-and-per-agent-circuit-breakers (verified: primary)
  2. [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Execution limits"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#execution-limits (verified: primary)
  3. [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Stopping across hops"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#stopping-across-hops (verified: primary)
  4. [4] NIST AI RMF 1.0 (AI 100-1) (MANAGE 2.4: mechanisms to "supersede, disengage, or deactivate AI systems" whose outcomes are inconsistent with intended use). NIST. 2023-01-26. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf (verified: primary)
  5. [5] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27 as amended by Regulation (EU) 2026/1744 (Art. 14(4)(e): human oversight includes the means to interrupt the system through a stop procedure). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng (verified: primary)
  6. [6] Agentic Trust Framework v1 ("You can stop one agent without stopping the business"; containment by revoking the agent's identity). CSAI Foundation / Cloud Security Alliance. 2026-02. https://agentictrustframework.ai/ (verified: primary)
  7. [7] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026). OWASP GenAI Security Project. 2026-08-03. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ (verified: primary)
  8. [8] Agent2Agent (A2A) Protocol Specification v1.0 (Cancel Task: "The server will attempt to cancel the task, but success is not guaranteed"). A2A Project (Linux Foundation). 2026-05-28. https://a2a-protocol.org/latest/specification/ (verified: primary)
  9. [9] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
  10. [10] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml). MITRE. 2026-09-15. https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.09 (verified: primary)
  11. [11] Common Elements of Frontier AI Safety Policies (December 2025 Update) (the policies set "conditions for halting development and deployment" when mitigations are insufficient). METR. 2025-12-09. https://metr.org/blog/2025-12-09-common-elements-of-frontier-ai-safety-policies/ (verified: primary)
  12. [12] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
  13. [13] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/ (verified: primary)
  14. [14] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
  15. [15] Pacing model development in an era of cyber-critical capabilities (teams paged by a highest-priority monitoring alert are expected to pause the activity if they cannot determine "within 30 minutes that the flag is a false positive"). OpenAI. 2026-08-18. https://openai.com/index/pacing-model-development-cyber-capabilities/ (verified: primary)
  16. [16] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)
  17. [17] The Hugging Face incident and the road ahead (OpenAI states that its chain-of-thought monitors "did not run on the evaluations in this incident", and that it is strengthening its incident response plan with clearer rules on "who can stop a run or approve restarting it"). OpenAI. 2026-08-26. https://openai.com/index/hugging-face-incident-and-the-road-ahead/ (verified: primary)

Machine-readable

Review this control

Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.

This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.

Cite this control

García Aibar, J. (2026). AIGE-CTL-EVAL-006 Stop Conditions (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-006. CC BY 4.0

BibTeX

@misc{aige2026page,
  author       = {Jorge García Aibar},
  title        = {{AIGE-CTL-EVAL-006 Stop Conditions}},
  howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
  year         = {2026},
  version      = {0.2},
  doi          = {10.5281/zenodo.22857084},
  url          = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-006},
  note         = {Version 0.2}
}