On this page

Tool and Action Mediation

Every tool call an agent makes during a run passes a mediation point outside the model that checks it against policy before it executes.

v0.2 Draft Open for technical review

AIGE-CTL-EVAL-004, a control of the Evaluation Environment Control Profile (draft v0.2). See it among the other controls of the profile, or its mappings beside every other control's in the controls crosswalk.

The control record

Id
AIGE-CTL-EVAL-004 · v0.2 · Draft · Open for technical review
Objective
Every tool call an agent makes during a run passes a mediation point outside the model that checks it against policy before it executes.
Failure modes
  • A tool call executes without a policy verdict recorded for it.
  • An irreversible action runs while the mediation point is down.
  • An approval is reused for a call whose parameters changed.
  • After repeated denials the agent keeps seeking the same outcome through other calls, and the run continues.
Scope
Tool, MCP and code-execution calls made by the agent under evaluation, and the reviewer (a rule, a model or a person) that decides on them. The internal reasoning of the model is out of scope; network egress is AIGE-CTL-EVAL-002.
Enforcement points
  • deploy: before a version is deployed or released
  • runtime: at the point of action (gateway or guardrail)
Verification
  • Inspect: Before the run, inspect the environment and its policy: tool servers, MCP servers and code execution are reachable only through the mediation point, and the policy lists each operation class with its verdict, its failure posture (fail closed for irreversible classes such as delete, send, publish and execute) and the denial threshold that interrupts a run.
  • Test: At admission, send through the harness one call the policy denies and one it allows, then take the mediation point down and send an irreversible-class call; the denied call and the call sent while it is down must not execute, and all three must leave a verdict record.
  • Observe: After the run, join the tool servers' own logs with the verdict records: every executed call has an allow verdict, or an approval bound to a parameter hash that matches the call, and no run continued past its denial threshold.
Evidence
  • Mediation policy of the run: operation classes, verdicts, failure posture per class and the denial threshold · Layer 04 · policy-card.v1
  • Verdict record of every call: tool, parameter hash, verdict, reviewer and time · Layer 04 · evidence-record.v1
  • One observation per run joining the executed calls with their verdicts · Layer 05 · control-observation.v1
Failure response
deny: block the action. A call without an allow verdict does not execute. While the mediation point is down, irreversible classes fail closed and reads fail open only with an alert. A run in which a call executed without a verdict is stopped and its result is withheld; a run that reaches its denial threshold is interrupted.
Layer
Layer 04 Runtime Controls & Observability
Patterns
Runtime Guardrail, Human-in-the-loop Gate
Seeded from
Runtime guardrail on every tool call, Checkpoints on irreversible actions, failing closed, Approval log, bound to the call, MCP server admission gate, Code runs only in a sandbox
Mappings
References
  • [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls")
  • [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint")
  • [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server")
  • [4] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026)
  • [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling")
  • [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [7] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant")
  • [8] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security")
  • [9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling")
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  • [11] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning)
Implementation notes
  • Put the mediation point between the decision to call a tool and the call, outside the model and the agent's process: a gateway or guardian that checks identity, allow-list, parameters and operation class before the call runs, as chapter 23 describes and OWASP's Agent Control Standard specifies on the wire. Route tool servers, MCP servers and code execution only through it, so no call can go around it.
  • Record the failure posture per operation class in the policy of the run. Chapter 23 recommends failing closed for pay, delete, send and execute and failing open, with an alert, only for reads, and notes that the reference guardian of the Agent Control Standard starts with a fail-open posture: set the posture, do not inherit it.
  • Bind an approval to the exact call. OpenAI's Codex documentation describes an override that "applies to the exact denied action, not similar future actions" and holds for one retry; store a hash of the parameters with each approval and refuse a call whose parameters differ.
  • Interrupt a run that keeps asking. The Codex documentation describes a rejection circuit breaker that, in its current open-source implementation, interrupts the turn after 3 consecutive denials or 10 in the last 50 reviews, and OpenAI's post on auto-review states that it stops the trajectory after repeated denials. Record the threshold in the policy and make reaching it a stop condition (AIGE-CTL-EVAL-006).
  • An automated reviewer can stand in for a person at the mediation point without widening what the agent may do: the Codex documentation states that "Auto-review is a reviewer swap, not a permission grant", and OpenAI's post on the design states that it "should not be treated as a guarantee of security". OpenAI reports that a model which had published a researcher's token was brought back with "enforced auto-review on all actions".
Open questions
  • How can every call be mediated without taking away the scaffolding and tooling a capability elicitation needs? METR's guidelines ask for "the best available scaffolding + tooling".
  • When the reviewer at the mediation point is a model, what evidence shows that its verdicts are reliable enough to replace a person for a given operation class?
Observation
  • Subject: Evaluation run
  • Expected: Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
  • Example: Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters: fail.

Example observations

Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.

Pass: run-88240@2026-09-26

Status
pass
Subject
run-88240@2026-09-26 (Evaluation run)
Expected
Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
Observed
Run 88240: 2,587 tool calls, each with a verdict recorded before it ran; the 4 denied calls did not execute; at admission, a delete call sent while the mediation point was down was refused.
Timestamp
Observer
tool-verdict-join
Evidence
  • verdict records of run 88240 · sha256:d6607e2f5f24722d20d9dd11b9b7ac1528e5b7d3b9c54c1dc004725f140fad69
  • tool server logs of run 88240 · sha256:c0b5ba6e57cb73f9e96c3bca712e8d5bfb04cd1fec27fadcbbd02cdce3c519da
Notes
Illustrative example, not the result of a real evaluation.

Download the pass example (JSON)

The pass record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-004",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88240@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.",
  "observed": "Run 88240: 2,587 tool calls, each with a verdict recorded before it ran; the 4 denied calls did not execute; at admission, a delete call sent while the mediation point was down was refused.",
  "status": "pass",
  "timestamp": "2026-09-26T09:05:51Z",
  "enforcement_point": "runtime",
  "verification_kind": "observe",
  "observer": "tool-verdict-join",
  "evidence": [
    {
      "artefact": "verdict records of run 88240",
      "url": "https://evidence.example/runs/88240/verdicts.jsonl",
      "hash": "sha256:d6607e2f5f24722d20d9dd11b9b7ac1528e5b7d3b9c54c1dc004725f140fad69",
      "schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
    },
    {
      "artefact": "tool server logs of run 88240",
      "url": "https://evidence.example/runs/88240/tool-servers.jsonl",
      "hash": "sha256:c0b5ba6e57cb73f9e96c3bca712e8d5bfb04cd1fec27fadcbbd02cdce3c519da"
    }
  ],
  "run_id": "88240",
  "notes": "Illustrative example, not the result of a real evaluation."
}

Fail: run-88241@2026-09-26

Status
fail
Subject
run-88241@2026-09-26 (Evaluation run)
Expected
Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.
Observed
Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters (parameter hash mismatch).
Timestamp
Observer
tool-verdict-join
Evidence
  • verdict records joined with the tool server logs of run 88241 · sha256:6945e2c3ebb9e89b535306fe0322518220f51b3ec03536f4bbbd153a91cd2d88
Notes
Illustrative example, not the result of a real evaluation. The run was stopped, its result withheld and approvals bound to a parameter hash.

Download the fail example (JSON)

The fail record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-004",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88241@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "Every executed tool call has an allow verdict, or an approval bound to its exact parameters, recorded before it ran; irreversible calls fail closed while the mediation point is down.",
  "observed": "Run 88241: 3,112 tool calls, each with a verdict; 1 publish call executed on an approval issued for different parameters (parameter hash mismatch).",
  "status": "fail",
  "timestamp": "2026-09-26T11:48:20Z",
  "enforcement_point": "runtime",
  "verification_kind": "observe",
  "observer": "tool-verdict-join",
  "evidence": [
    {
      "artefact": "verdict records joined with the tool server logs of run 88241",
      "url": "https://evidence.example/runs/88241/verdict-join.jsonl",
      "hash": "sha256:6945e2c3ebb9e89b535306fe0322518220f51b3ec03536f4bbbd153a91cd2d88"
    }
  ],
  "run_id": "88241",
  "notes": "Illustrative example, not the result of a real evaluation. The run was stopped, its result withheld and approvals bound to a parameter hash."
}

Incident cases on this site that list this control among their related controls.

Patterns

The patterns that implement this control.

Obligations

The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.

Threats

The entries of the threat catalogues this control answers.

Sources

  1. [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#runtime-guardrails-for-tool-calls (verified: primary)
  2. [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Where to put a checkpoint"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#where-to-put-a-checkpoint (verified: primary)
  3. [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Admitting an MCP server"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#admitting-an-mcp-server (verified: primary)
  4. [4] Agent Control Standard (ACS) (wire specification for a guardian agent that decides on an agent action before it runs; donated to OWASP, announced 1 Sep 2026). OWASP GenAI Security Project. 2026-09-01. https://genai.owasp.org/resource/agent-control-standard-acs/ (verified: primary)
  5. [5] Guidelines for capability elicitation (task bugs such as "The automatic scoring is incorrect" or a crashed environment are spurious failures to fix before reporting; models get "the best available scaffolding + tooling"). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
  6. [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
  7. [7] Codex auto-review (undated developer documentation, read 2026-09-26: "Auto-review is a reviewer swap, not a permission grant"). OpenAI. 2026. https://developers.openai.com/codex/sandboxing/auto-review (verified: primary)
  8. [8] Auto-review of agent actions without synchronous human oversight (a separate agent approves or denies actions that cross the sandbox boundary; OpenAI states that auto-review "should not be treated as a guarantee of security"). OpenAI (Alignment Research Blog). 2026-04-30. https://alignment.openai.com/auto-review/ (verified: primary)
  9. [9] Preparedness Framework, Version 2 (Table 5 lists potential safeguards against a misaligned model, among them limiting internet and tool access, credentials and filesystem access, with agent actions "logged in an uneditable database"; a one-time capability elicitation is treated "as a lower bound, rather than a ceiling"). OpenAI. 2025-04-15. https://cdn.openai.com/pdf/18a02b5d-6b67-4cec-ab64-68cdfbddebcd/preparedness-framework-v2.pdf (verified: primary)
  10. [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
  11. [11] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/exposing-a-github-token-in-a-public-repository/ (verified: primary)

Machine-readable

Review this control

Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.

This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.

Cite this control

García Aibar, J. (2026). AIGE-CTL-EVAL-004 Tool and Action Mediation (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-004. CC BY 4.0

BibTeX

@misc{aige2026page,
  author       = {Jorge García Aibar},
  title        = {{AIGE-CTL-EVAL-004 Tool and Action Mediation}},
  howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
  year         = {2026},
  version      = {0.2},
  doi          = {10.5281/zenodo.22857084},
  url          = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-004},
  note         = {Version 0.2}
}