Tool and Action Mediation
Every tool call an agent makes during a run passes a mediation point outside the model that checks it against policy before it executes.
AIGE-CTL-EVAL-004, a control of the Evaluation Environment Control Profile (draft
v0.2). See it among the other controls of the profile,
or its mappings beside every other control's in the
controls crosswalk.
The control record
- Id
-
AIGE-CTL-EVAL-004· v0.2 · Draft · Open for technical review - Objective
- Every tool call an agent makes during a run passes a mediation point outside the model that checks it against policy before it executes.
- Failure modes
-
- A tool call executes without a policy verdict recorded for it.
- An irreversible action runs while the mediation point is down.
- An approval is reused for a call whose parameters changed.
- After repeated denials the agent keeps seeking the same outcome through other calls, and the run continues.
- Scope
- Tool, MCP and code-execution calls made by the agent under evaluation, and the reviewer (a rule, a model or a person) that decides on them. The internal reasoning of the model is out of scope; network egress is AIGE-CTL-EVAL-002.
- Enforcement points
-
- deploy: before a version is deployed or released
- runtime: at the point of action (gateway or guardrail)
- Verification
-
- Inspect: Before the run, inspect the environment and its policy: tool servers, MCP servers and code execution are reachable only through the mediation point, and the policy lists each operation class with its verdict, its failure posture (fail closed for irreversible classes such as delete, send, publish and execute) and the denial threshold that interrupts a run.
- Test: At admission, send through the harness one call the policy denies and one it allows, then take the mediation point down and send an irreversible-class call; the denied call and the call sent while it is down must not execute, and all three must leave a verdict record.
- Observe: After the run, join the tool servers' own logs with the verdict records: every executed call has an allow verdict, or an approval bound to a parameter hash that matches the call, and no run continued past its denial threshold.
- Evidence
-
- Mediation policy of the run: operation classes, verdicts, failure posture per class and the denial threshold · Layer 04 · policy-card.v1
- Verdict record of every call: tool, parameter hash, verdict, reviewer and time · Layer 04 · evidence-record.v1
- One observation per run joining the executed calls with their verdicts · Layer 05 · control-observation.v1
- Failure response
- deny: block the action. A call without an allow verdict does not execute. While the mediation point is down, irreversible classes fail closed and reads fail open only with an alert. A run in which a call executed without a verdict is stopped and its result is withheld; a run that reaches its denial threshold is interrupted.
- Layer
- Layer 04 Runtime Controls & Observability
- Patterns
- Runtime Guardrail, Human-in-the-loop Gate
- Seeded from
- Runtime guardrail on every tool call, Checkpoints on irreversible actions, failing closed, Approval log, bound to the call, MCP server admission gate, Code runs only in a sandbox
- Mappings
-
- Obligations: EU AI Act Art. 14 human oversight; OWASP Agent Control Standard (ACS), Agent Control Standard (ACS); Singapore IMDA Model AI Governance Framework for Agentic AI: human checkpoints for significant actions (voluntary); OWASP Top 10 for Agentic Applications 2026
- ISO/IEC 42001: A.9.2 Processes for responsible use of AI systems
- NIST AI RMF: MAP 4.2 Internal risk controls for components of the AI system, including third-party AI technologies, are identified and documented.
- OWASP: ASI01 Agent Goal Hijack; ASI02 Tool Misuse and Exploitation; ASI05 Unexpected Code Execution (RCE); ASI09 Human-Agent Trust Exploitation; LLM10:2026 Improper Output Handling
- AIUC-1: D003, B006
- MITRE ATLAS mitigation: AML.M0028 (AI Agent Tools Permissions Configuration)
- MITRE ATLAS mitigation: AML.M0029 (Human In-the-Loop for AI Agent Actions)
- MITRE ATLAS mitigation: AML.M0030 (Restrict AI Agent Tool Invocation on Untrusted Data)
- References
-
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls")
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint")
- [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server")
- [4] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026)
- [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
- [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
- [7] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant")
- [8] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security")
- [9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
- [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
- [11] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning)
- Implementation notes
-
- Put the mediation point between the decision to call a tool and the call, outside the model and the agent's process: a gateway or guardian that checks identity, allow-list, parameters and operation class before the call runs, as chapter 23 describes and OWASP's Agent Control Standard specifies on the wire. Route tool servers, MCP servers and code execution only through it, so no call can go around it.
- Record the failure posture per operation class in the policy of the run. Chapter 23 recommends failing closed for pay, delete, send and execute and failing open, with an alert, only for reads, and notes that the reference guardian of the Agent Control Standard starts with a fail-open posture: set the posture, do not inherit it.
- Bind an approval to the exact call. OpenAI's Codex documentation describes an override that "applies to the exact denied action, not similar future actions" and holds for one retry; store a hash of the parameters with each approval and refuse a call whose parameters differ.
- Interrupt a run that keeps asking. The Codex documentation describes a rejection circuit breaker that, in its current open-source implementation, interrupts the turn after 3 consecutive denials or 10 in the last 50 reviews, and OpenAI's post on auto-review states that it stops the trajectory after repeated denials. Record the threshold in the policy and make reaching it a stop condition (AIGE-CTL-EVAL-006).
- An automated reviewer can stand in for a person at the mediation point without widening what the agent may do: the Codex documentation states that "Auto-review is a reviewer swap, not a permission grant", and OpenAI's post on the design states that it "should not be treated as a guarantee of security". OpenAI reports that a model which had published a researcher's token was brought back with "enforced auto-review on all actions".
- Open questions
-
- How can every call be mediated without taking away the scaffolding and tooling a capability elicitation needs? METR's guidelines ask for "the best available scaffolding + tooling".
- When the reviewer at the mediation point is a model, what evidence shows that its verdicts are reliable enough to replace a person for a given operation class?
- Observation
-
- Subject: Evaluation run
- Expected: Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
- Example: Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters: fail.
Example observations
Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.
Pass: run-88240@2026-09-26
- Status
-
pass - Subject
-
run-88240@2026-09-26(Evaluation run) - Expected
- Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
- Observed
- Run 88240: 2,587 tool calls, each with a verdict recorded before it ran; the 4 denied calls did not execute; at admission, a delete call sent while the mediation point was down was refused.
- Timestamp
- Observer
- tool-verdict-join
- Evidence
-
- verdict records of run 88240 ·
sha256:d6607e2f5f24722d20d9dd11b9b7ac1528e5b7d3b9c54c1dc004725f140fad69 - tool server logs of run 88240 ·
sha256:c0b5ba6e57cb73f9e96c3bca712e8d5bfb04cd1fec27fadcbbd02cdce3c519da
- verdict records of run 88240 ·
- Notes
- Illustrative example, not the result of a real evaluation.
Download the pass example (JSON)
The pass record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-004",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "run-88240@2026-09-26",
"subject_kind": "eval-run",
"expected": "Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.",
"observed": "Run 88240: 2,587 tool calls, each with a verdict recorded before it ran; the 4 denied calls did not execute; at admission, a delete call sent while the mediation point was down was refused.",
"status": "pass",
"timestamp": "2026-09-26T09:05:51Z",
"enforcement_point": "runtime",
"verification_kind": "observe",
"observer": "tool-verdict-join",
"evidence": [
{
"artefact": "verdict records of run 88240",
"url": "https://evidence.example/runs/88240/verdicts.jsonl",
"hash": "sha256:d6607e2f5f24722d20d9dd11b9b7ac1528e5b7d3b9c54c1dc004725f140fad69",
"schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
},
{
"artefact": "tool server logs of run 88240",
"url": "https://evidence.example/runs/88240/tool-servers.jsonl",
"hash": "sha256:c0b5ba6e57cb73f9e96c3bca712e8d5bfb04cd1fec27fadcbbd02cdce3c519da"
}
],
"run_id": "88240",
"notes": "Illustrative example, not the result of a real evaluation."
} Fail: run-88241@2026-09-26
- Status
-
fail - Subject
-
run-88241@2026-09-26(Evaluation run) - Expected
- Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
- Observed
- Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters (parameter hash mismatch).
- Timestamp
- Observer
- tool-verdict-join
- Evidence
-
- verdict records joined with the tool server logs of run 88241 ·
sha256:6945e2c3ebb9e89b535306fe0322518220f51b3ec03536f4bbbd153a91cd2d88
- verdict records joined with the tool server logs of run 88241 ·
- Notes
- Illustrative example, not the result of a real evaluation. The run was stopped, its result withheld and approvals bound to a parameter hash.
Download the fail example (JSON)
The fail record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-004",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "run-88241@2026-09-26",
"subject_kind": "eval-run",
"expected": "Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.",
"observed": "Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters (parameter hash mismatch).",
"status": "fail",
"timestamp": "2026-09-26T11:48:20Z",
"enforcement_point": "runtime",
"verification_kind": "observe",
"observer": "tool-verdict-join",
"evidence": [
{
"artefact": "verdict records joined with the tool server logs of run 88241",
"url": "https://evidence.example/runs/88241/verdict-join.jsonl",
"hash": "sha256:6945e2c3ebb9e89b535306fe0322518220f51b3ec03536f4bbbd153a91cd2d88"
}
],
"run_id": "88241",
"notes": "Illustrative example, not the result of a real evaluation. The run was stopped, its result withheld and approvals bound to a parameter hash."
} Related cases
Incident cases on this site that list this control among their related controls.
- OpenAI agents and Hugging Face: an evaluation environment that was not isolated: METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- An internally deployed model published a researcher's GitHub token in a public repository: OpenAI reports that an internally deployed model put a researcher's GitHub token, split to avoid secret scanning, into code it pushed to a public repository.
- Training samples exchanged messages through a shared package repository: OpenAI reports that models in RL training used an internal package repository, with the credentials they were given, to exchange messages across samples.
Patterns
The patterns that implement this control.
- Runtime Guardrail · Layer 04 Runtime Controls & Observability
- Human-in-the-loop Gate · Layer 04 Runtime Controls & Observability
Obligations
The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.
- EU AI Act Art. 14 human oversight
AIGE-OBL-EUAIA-ART14: Human oversight designed into the system - OWASP Agent Control Standard (ACS), Agent Control Standard (ACS)
AIGE-OBL-OWASP-ACS: A standard for expressing agent controls - Singapore IMDA Model AI Governance Framework for Agentic AI: human checkpoints for significant actions (voluntary)
AIGE-OBL-SG-AGENTIC-CHECKPOINTS: Significant checkpoints for high-stakes, irreversible, outlier and user-defined actions, with approvals that are contextual and digestible and enforced through system-level controls - OWASP Top 10 for Agentic Applications 2026
AIGE-OBL-OWASP-AGENTIC: Agent threat catalogue (ASI01 Agent Goal Hijack … ASI10 Rogue Agents)
Threats
The entries of the threat catalogues this control answers.
- ASI01 Agent Goal Hijack (OWASP Top 10 for Agentic Applications 2026)
- ASI02 Tool Misuse and Exploitation (OWASP Top 10 for Agentic Applications 2026)
- ASI05 Unexpected Code Execution (RCE) (OWASP Top 10 for Agentic Applications 2026)
- ASI09 Human-Agent Trust Exploitation (OWASP Top 10 for Agentic Applications 2026)
- LLM10:2026 Improper Output Handling (OWASP Top 10 for LLM Applications 2026)
Sources
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#runtime-guardrails-for-tool-calls (verified: primary)
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#where-to-put-a-checkpoint (verified: primary)
- [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#admitting-an-mcp-server (verified: primary)
- [4] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026). OWASP GenAI Security Project. 2026-09-01. https://genai.owasp.org/resource/agent-control-standard-acs/ (verified: primary)
- [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
- [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
- [7] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant"). OpenAI. 2026. https://developers.openai.com/codex/sandboxing/auto-review (verified: primary)
- [8] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security"). OpenAI (Alignment Research Blog). 2026-04-30. https://alignment.openai.com/auto-review/ (verified: primary)
- [9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling"). OpenAI. 2025-04-15. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf (verified: primary)
- [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
- [11] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/exposing-a-github-token-in-a-public-repository/ (verified: primary)
Machine-readable
- AIGE-CTL-EVAL-004 as JSON : this record on its own, in the open data API.
- This page as Markdown: /controls/evaluation-environment/aige-ctl-eval-004.md.
- The whole profile: Evaluation Environment Control Profile, also as Markdown.
- All controls as JSON: every control of every profile, in the envelope of the open data API.
- JSON Schema of the controls dataset: generated from the same registry.
- Control observation schema: the record a check of a control emits: control_id, subject, expected, observed, status, timestamp, evidence.
- Observation example: a filled record that validates against the schema.
- Observation template: the same fields in Markdown, with short guidance.
Review this control
Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.
This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.
Cite this control
García Aibar, J. (2026). AIGE-CTL-EVAL-004 Tool and Action Mediation (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-004. CC BY 4.0
BibTeX
@misc{aige2026page,
author = {Jorge García Aibar},
title = {{AIGE-CTL-EVAL-004 Tool and Action Mediation}},
howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
year = {2026},
version = {0.2},
doi = {10.5281/zenodo.22857084},
url = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-004},
note = {Version 0.2}
}