---
title: "Tool and Action Mediation"
description: "Draft control: every tool call an agent makes in an evaluation run passes a mediation point that records a verdict and fails closed for irreversible actions."
canonical: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-004
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22857084
version: "0.2"
updated: 2026-09-26
---

# Tool and Action Mediation

> Every tool call an agent makes during a run passes a mediation point outside the model that checks it against policy before it executes.

- Id: AIGE-CTL-EVAL-004
- Profile: [Evaluation Environment Control Profile v0.2](https://aigovernanceengineer.com/controls/evaluation-environment)
- Status: Draft
- Review: Open for technical review
- Published: 2026-09-26
- Updated: 2026-09-26
- Anchor on the profile page: https://aigovernanceengineer.com/controls/evaluation-environment#aige-ctl-eval-004

Draft for review, not a claim of conformity. A draft control specification, open for technical review: illustrative, not legal advice and binding on no one.

## The control record

- Id: `AIGE-CTL-EVAL-004` · v0.2 · Draft · Open for technical review
- Depth: Specified
- Objective: Every tool call an agent makes during a run passes a mediation point outside the model that checks it against policy before it executes.
- Failure modes:
  - A tool call executes without a policy verdict recorded for it.
  - An irreversible action runs while the mediation point is down.
  - An approval is reused for a call whose parameters changed.
  - After repeated denials the agent keeps seeking the same outcome through other calls, and the run continues.
- Scope: Tool, MCP and code-execution calls made by the agent under evaluation, and the reviewer (a rule, a model or a person) that decides on them. The internal reasoning of the model is out of scope; network egress is AIGE-CTL-EVAL-002.
- Enforcement points:
  - deploy: before a version is deployed or released
  - runtime: at the point of action (gateway or guardrail)
- Verification:
  - Inspect: Before the run, inspect the environment and its policy: tool servers, MCP servers and code execution are reachable only through the mediation point, and the policy lists each operation class with its verdict, its failure posture (fail closed for irreversible classes such as delete, send, publish and execute) and the denial threshold that interrupts a run.
  - Test: At admission, send through the harness one call the policy denies and one it allows, then take the mediation point down and send an irreversible-class call; the denied call and the call sent while it is down must not execute, and all three must leave a verdict record.
  - Observe: After the run, join the tool servers' own logs with the verdict records: every executed call has an allow verdict, or an approval bound to a parameter hash that matches the call, and no run continued past its denial threshold.
- Evidence:
  - Mediation policy of the run: operation classes, verdicts, failure posture per class and the denial threshold · Layer 04 Runtime Controls & Observability · [policy-card.v1](https://aigovernanceengineer.com/resources/templates#schema-policy-card)
  - Verdict record of every call: tool, parameter hash, verdict, reviewer and time · Layer 04 Runtime Controls & Observability · [evidence-record.v1](https://aigovernanceengineer.com/resources/templates#schema-evidence-record)
  - One observation per run joining the executed calls with their verdicts · Layer 05 Assurance & Continuous Compliance · [control-observation.v1](https://aigovernanceengineer.com/resources/templates#schema-control-observation)
- Failure response: deny: block the action. A call without an allow verdict does not execute. While the mediation point is down, irreversible classes fail closed and reads fail open only with an alert. A run in which a call executed without a verdict is stopped and its result is withheld; a run that reaches its denial threshold is interrupted.
- Layer: [Layer 04 Runtime Controls & Observability](https://aigovernanceengineer.com/bok/the-stack#layer-04-runtime-controls--observability)
- Patterns: [Runtime Guardrail](https://aigovernanceengineer.com/patterns/runtime-guardrail), [Human-in-the-loop Gate](https://aigovernanceengineer.com/patterns/human-in-the-loop-gate)
- Seeded from: [Runtime guardrail on every tool call](https://aigovernanceengineer.com/bok/governing-agents#runtime-guardrails-for-tool-calls), [Checkpoints on irreversible actions, failing closed](https://aigovernanceengineer.com/bok/governing-agents#where-to-put-a-checkpoint), [Approval log, bound to the call](https://aigovernanceengineer.com/bok/governing-agents#what-a-good-approval-looks-like), [MCP server admission gate](https://aigovernanceengineer.com/bok/governing-agents#admitting-an-mcp-server), [Code runs only in a sandbox](https://aigovernanceengineer.com/bok/governing-agents#runtime-guardrails-for-tool-calls)
- Mappings:
  - Obligations: [EU AI Act Art. 14 human oversight](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art14); [OWASP Agent Control Standard (ACS), Agent Control Standard (ACS)](https://aigovernanceengineer.com/obligations/aige-obl-owasp-acs); [Singapore IMDA Model AI Governance Framework for Agentic AI: human checkpoints for significant actions (voluntary)](https://aigovernanceengineer.com/obligations/aige-obl-sg-agentic-checkpoints); [OWASP Top 10 for Agentic Applications 2026](https://aigovernanceengineer.com/obligations/aige-obl-owasp-agentic)
  - ISO/IEC 42001: A.9.2 Processes for responsible use of AI systems
  - NIST AI RMF: MAP 4.2 Internal risk controls for components of the AI system, including third-party AI technologies, are identified and documented.
  - OWASP: [ASI01 Agent Goal Hijack](https://aigovernanceengineer.com/resources/threats#threat-asi01); [ASI02 Tool Misuse and Exploitation](https://aigovernanceengineer.com/resources/threats#threat-asi02); [ASI05 Unexpected Code Execution (RCE)](https://aigovernanceengineer.com/resources/threats#threat-asi05); [ASI09 Human-Agent Trust Exploitation](https://aigovernanceengineer.com/resources/threats#threat-asi09); [LLM10:2026 Improper Output Handling](https://aigovernanceengineer.com/resources/threats#threat-llm10-2026)
  - AIUC-1: D003, B006
  - MITRE ATLAS mitigation: AML.M0028 (AI Agent Tools Permissions Configuration)
  - MITRE ATLAS mitigation: AML.M0029 (Human In-the-Loop for AI Agent Actions)
  - MITRE ATLAS mitigation: AML.M0030 (Restrict AI Agent Tool Invocation on Untrusted Data)
- References:
  - [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls")
  - [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint")
  - [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server")
  - [4] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026)
  - [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  - [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  - [7] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant")
  - [8] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security")
  - [9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  - [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  - [11] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning)
- Implementation notes:
  - Put the mediation point between the decision to call a tool and the call, outside the model and the agent's process: a gateway or guardian that checks identity, allow-list, parameters and operation class before the call runs, as chapter 23 describes and OWASP's Agent Control Standard specifies on the wire. Route tool servers, MCP servers and code execution only through it, so no call can go around it.
  - Record the failure posture per operation class in the policy of the run. Chapter 23 recommends failing closed for pay, delete, send and execute and failing open, with an alert, only for reads, and notes that the reference guardian of the Agent Control Standard starts with a fail-open posture: set the posture, do not inherit it.
  - Bind an approval to the exact call. OpenAI's Codex documentation describes an override that "applies to the exact denied action, not similar future actions" and holds for one retry; store a hash of the parameters with each approval and refuse a call whose parameters differ.
  - Interrupt a run that keeps asking. The Codex documentation describes a rejection circuit breaker that, in its current open-source implementation, interrupts the turn after 3 consecutive denials or 10 in the last 50 reviews, and OpenAI's post on auto-review states that it stops the trajectory after repeated denials. Record the threshold in the policy and make reaching it a stop condition (AIGE-CTL-EVAL-006).
  - An automated reviewer can stand in for a person at the mediation point without widening what the agent may do: the Codex documentation states that "Auto-review is a reviewer swap, not a permission grant", and OpenAI's post on the design states that it "should not be treated as a guarantee of security". OpenAI reports that a model which had published a researcher's token was brought back with "enforced auto-review on all actions".
- Open questions:
  - How can every call be mediated without taking away the scaffolding and tooling a capability elicitation needs? METR's guidelines ask for "the best available scaffolding + tooling".
  - When the reviewer at the mediation point is a model, what evidence shows that its verdicts are reliable enough to replace a person for a given operation class?
- Observation:
  - Subject: Evaluation run
  - Expected: Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
  - Example: Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters: fail.
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-004.json

## Example observations

Two illustrative records of a check of this control, one that passes and one that fails. They validate against the control observation schema; they are not results of any real evaluation.

### Pass: run-88240@2026-09-26

- Status: pass
- Subject: run-88240@2026-09-26 (Evaluation run)
- Expected: Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
- Observed: Run 88240: 2,587 tool calls, each with a verdict recorded before it ran; the 4 denied calls did not execute; at admission, a delete call sent while the mediation point was down was refused.
- Timestamp: 2026-09-26T09:05:51Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-004.pass.json

### Fail: run-88241@2026-09-26

- Status: fail
- Subject: run-88241@2026-09-26 (Evaluation run)
- Expected: Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
- Observed: Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters (parameter hash mismatch).
- Timestamp: 2026-09-26T11:48:20Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-004.fail.json

## Related cases

- [OpenAI agents and Hugging Face: an evaluation environment that was not isolated](https://aigovernanceengineer.com/cases/openai-hugging-face-agent-incident-2026): METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- [An internally deployed model published a researcher's GitHub token in a public repository](https://aigovernanceengineer.com/cases/openai-agent-github-token-exposure-2026): OpenAI reports that an internally deployed model put a researcher's GitHub token, split to avoid secret scanning, into code it pushed to a public repository.
- [Training samples exchanged messages through a shared package repository](https://aigovernanceengineer.com/cases/openai-agents-artifactory-cross-sample-2026): OpenAI reports that models in RL training used an internal package repository, with the credentials they were given, to exchange messages across samples.

## Patterns

- [Runtime Guardrail](https://aigovernanceengineer.com/patterns/runtime-guardrail) (Layer 04 Runtime Controls & Observability)
- [Human-in-the-loop Gate](https://aigovernanceengineer.com/patterns/human-in-the-loop-gate) (Layer 04 Runtime Controls & Observability)

## Obligations

- [EU AI Act Art. 14 human oversight](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art14) (`AIGE-OBL-EUAIA-ART14`): Human oversight designed into the system
- [OWASP Agent Control Standard (ACS), Agent Control Standard (ACS)](https://aigovernanceengineer.com/obligations/aige-obl-owasp-acs) (`AIGE-OBL-OWASP-ACS`): A standard for expressing agent controls
- [Singapore IMDA Model AI Governance Framework for Agentic AI: human checkpoints for significant actions (voluntary)](https://aigovernanceengineer.com/obligations/aige-obl-sg-agentic-checkpoints) (`AIGE-OBL-SG-AGENTIC-CHECKPOINTS`): Significant checkpoints for high-stakes, irreversible, outlier and user-defined actions, with approvals that are contextual and digestible and enforced through system-level controls
- [OWASP Top 10 for Agentic Applications 2026](https://aigovernanceengineer.com/obligations/aige-obl-owasp-agentic) (`AIGE-OBL-OWASP-AGENTIC`): Agent threat catalogue (ASI01 Agent Goal Hijack … ASI10 Rogue Agents)

## Threats

- [ASI01 Agent Goal Hijack](https://aigovernanceengineer.com/resources/threats#threat-asi01) (OWASP Top 10 for Agentic Applications 2026)
- [ASI02 Tool Misuse and Exploitation](https://aigovernanceengineer.com/resources/threats#threat-asi02) (OWASP Top 10 for Agentic Applications 2026)
- [ASI05 Unexpected Code Execution (RCE)](https://aigovernanceengineer.com/resources/threats#threat-asi05) (OWASP Top 10 for Agentic Applications 2026)
- [ASI09 Human-Agent Trust Exploitation](https://aigovernanceengineer.com/resources/threats#threat-asi09) (OWASP Top 10 for Agentic Applications 2026)
- [LLM10:2026 Improper Output Handling](https://aigovernanceengineer.com/resources/threats#threat-llm10-2026) (OWASP Top 10 for LLM Applications 2026)

## Sources

[1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#runtime-guardrails-for-tool-calls (verified: primary)
[2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#where-to-put-a-checkpoint (verified: primary)
[3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#admitting-an-mcp-server (verified: primary)
[4] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026). OWASP GenAI Security Project. 2026-09-01. https://genai.owasp.org/resource/agent-control-standard-acs/ (verified: primary)
[5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
[6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
[7] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant"). OpenAI. 2026. https://developers.openai.com/codex/sandboxing/auto-review (verified: primary)
[8] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security"). OpenAI (Alignment Research Blog). 2026-04-30. https://alignment.openai.com/auto-review/ (verified: primary)
[9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling"). OpenAI. 2025-04-15. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf (verified: primary)
[10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
[11] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/exposing-a-github-token-in-a-public-repository/ (verified: primary)

## Machine-readable

- This control as JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-004.json
- The pass example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-004.pass.json
- The fail example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-004.fail.json
- The whole profile as Markdown: https://aigovernanceengineer.com/controls/evaluation-environment.md
- The open data API: https://aigovernanceengineer.com/resources/data

## Review

Review this control through the issue form: https://github.com/losanchos5/aige/issues/new?template=control-review.yml. How review works: https://aigovernanceengineer.com/contribute. Page: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-004

## Cite

AIGE-CTL-EVAL-004 Tool and Action Mediation. In Jorge García Aibar (2026). Evaluation Environment Control Profile (v0.2, draft). AI Governance Engineer. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment
