Governing AI agents: the agent, not only the model.
An agent acts under delegated authority: it calls tools, keeps memory and hands work to other agents. This page is the way in: the control plane that bounds an agent, the patterns that build it and the threats it has to hold against, each linked to its section of chapter 23.
Four properties make an agent a different object.
- Delegated authority It acts on someone’s behalf. The question is whose authority it used, and whether that was narrower than theirs.
- Tools Its output is an effect in a system of record, not text a person reads first.
- Memory What it keeps persists across sessions, users and tasks, and one poisoned write outlives the session that made it.
- Autonomy Steps happen with no person between them, so the design has to say where a person decides and how anyone stops it.
OWASP’s agentic list names the first principle least agency: avoid autonomy the task does not need. The cheapest agent control is the agent you did not build.
The agent control plane, component by component.
Ten components sit around every agent in production. Each one leaves evidence, and each links to its section in chapter 23.
Text description
The control plane in the order a tool call meets it. The agent registry (Layer 02 Inventory & Transparency) holds every agent with an owner, a purpose, an autonomy level, its tools, pinned versions, stop handles and an expiry: no registry entry, no credential. The identity issuer (Layer 04 Runtime Controls & Observability) gives the agent a short-lived, attested workload credential, with delegation that names the agent, never impersonation of the user. The agent proposes a tool call. The tool gateway denies by default, with pinned tool definitions, scopes, rates and egress per tool, and admitted MCP servers only. The runtime guardrail checks every call before it runs (identity against a live registry entry, the allow-list and definition hash, parameters within policy, instruction provenance, the output and egress filter, execution budgets) and fails closed for pay, delete, send and execute, open with an alert only for reads. A human checkpoint approves where the stakes or irreversibility demand it, showing the raw call. The per-agent circuit breaker has six stop levels: pause a task, narrow the scope, trip the breaker, revoke the identity, stop a class and degrade; exhausted budgets trip it, and a tripped breaker makes the gateway reject every call from the agent. Past your boundary sit the tool, the MCP server or a remote agent: you cannot stop someone else's agent, only stop calling it and revoke what you issued to it. Every step leaves telemetry that carries identity, tool calls, verdicts and approvals, kept as evidence (Layer 05 Assurance & Continuous Compliance). Memory controls, delegation across hops and prompt change control complete the plane in chapter 23.
- Agent registry Every agent with an owner, a purpose, an autonomy level, its tools, pinned versions, stop handles and an expiry.
- Identity issuer Short-lived, attested workload credentials; delegation that names the agent, never impersonation of the user.
- Tool gateway and allow-list Deny by default: pinned tool definitions, scopes, rates and egress per tool, and admitted MCP servers only.
- Runtime guardrail Every tool call checked before it runs, with a failure posture decided per operation class and execution budgets.
- Human checkpoint Approval where the stakes or irreversibility demand it, showing the raw call and bound to its parameters.
- Circuit breaker and kill switch Stop levels from one task to a fleet segment, drilled on a schedule and verified to hold.
- Memory controls Write gates, isolation per user and task, retention as code and rollback to a known-good state.
- Delegation across hops The originating principal and every actor stay on the record, and scope narrows at each hop.
- Prompt and configuration change control System prompts and tool descriptions versioned, reviewed, gated by evals and rolled back by hash.
- Telemetry and evidence Traces that carry identity, tool calls, verdicts and approvals, kept as evidence and fed to incident response.
Chapter 23 sets out 31 agent controls across these components. The agent control profile tool picks the ones one agent needs; the agent runtime control profile states each as a draft reference control with the evidence it must leave, open for technical review.
Six patterns that build it.
- 01 Agent Registry The runtime-aware inventory every other control attaches to. Layer 02 · Inventory & Transparency
- 02 Agent Identity & Scoped Credentials One identity per agent, with an owner, a bounded scope and an expiry. Layer 04 · Runtime Controls & Observability
- 03 Human-in-the-loop Gate A designed checkpoint where the consequence justifies the latency. Layer 04 · Runtime Controls & Observability
- 04 Runtime Guardrail Input, output and tool-call mediation at the enforcement point. Layer 04 · Runtime Controls & Observability
- 05 Kill Switch / Circuit Breaker A tested way to stop one agent without breaking the fleet. Layer 04 · Runtime Controls & Observability
- 06 Shadow-AI Discovery Finds the agents that never registered, so the registry stays true. Layer 02 · Inventory & Transparency
Ten agentic threats, and the control for each.
The OWASP Top 10 for Agentic Applications 2026, each threat mapped to the control that contains it and the pattern that implements the control.
| Threat | Control | Pattern |
|---|---|---|
| ASI01 Agent Goal Hijack | Instruction provenance; checkpoints before writes; trajectory evals | Runtime Guardrail |
| ASI02 Tool Misuse and Exploitation | Tool allow-list; per-tool rate, egress and budgets | Runtime Guardrail |
| ASI03 Identity and Privilege Abuse | Workload identity; short-lived delegated tokens; audience checks | Agent Identity & Scoped Credentials |
| ASI04 Agentic Supply Chain Vulnerabilities | MCP server admission; pinned tool definitions | AIBOM |
| ASI05 Unexpected Code Execution (RCE) | Sandboxed execution; deny by default | Runtime Guardrail |
| ASI06 Memory & Context Poisoning | Memory write gate; isolation; rollback | Runtime Guardrail |
| ASI07 Insecure Inter-Agent Communication | Mutual authentication; signed Agent Cards; peer allow-list | Agent Identity & Scoped Credentials |
| ASI08 Cascading Failures | Depth and fan-out limits; per-agent breakers | Kill Switch / Circuit Breaker |
| ASI09 Human-Agent Trust Exploitation | Approvals that show the raw call; oversight metrics | Human-in-the-loop Gate |
| ASI10 Rogue Agents | Registry with expiry; discovery; drilled kill switch | Agent Registry |
Where the frameworks and the law meet it.
- The agent-control table (chapter 08) TC260 Framework 3.0 Appendix 2, the OWASP agentic list and the NIST agent initiative, control by control.
- NIST AI Agent Standards Initiative (chapter 08) Agent identity, authentication, authorisation and agent security.
- OWASP GenAI Security Project (chapter 08) The agentic threat list, the 2026 LLM list and the Agent Control Standard.
- CSA AICM and STAR for AI (chapter 08) The AI Controls Matrix and its extensions towards agents.
- Frameworks written for agents (chapter 23) NIST, CSA, Singapore’s IMDA, TC260 and OWASP, with their status.
- EU AI Act hooks for agents (chapter 23) Articles 12, 14, 15, 25, 26 and 50 and the GPAI duties, with the agent artefact for each.
- Deployer duties, Article 26 (chapter 18) Oversight by competent people, monitoring, suspension and log retention.
- Designing human oversight (chapter 04) Automation bias, oversight that degrades, and how to measure it.
Tool categories for the control plane.
The categories that fill the runtime layer for agents, with the examples the Body of Knowledge names. The category is the substance; the brands are illustrative, not an endorsement.
| Category | Examples (illustrative) | Layer |
|---|---|---|
| Agent-discovery tools | Zenity | L2 |
| Guardrail frameworks | NVIDIA NeMo Guardrails · Meta LlamaFirewall · Lakera · Guardrails AI · Llama Guard | L4 |
| Observability | Langfuse · Arize Phoenix · OpenTelemetry (GenAI semantic conventions) | L4 |
| MCP / tool-call security | MCP Inspector · Snyk Agent Scan · mcp-context-protector | L4 |
| Kill switch / circuit breaker | Unleash · Envoy · Resilience4j | L4 |
| Agent workload identity | SPIFFE/SPIRE · Microsoft Entra Agent ID · Okta Agent SSO | L4 |
When agents fail.
- An agent acting beyond its mandate Harms atlas: operational disruption when an agent holds standing credentials broader than its task.
- Prompt injection Harms atlas: security compromise through instructions hidden in content the agent reads.
- Over-reliance Harms atlas: institutions acting on generated output without independent checks.
- The agent incident taxonomy (chapter 23) Eleven failure classes with the detection signal, the first containment and the threat IDs.
- AI-specific failure modes (chapter 17) Tool misuse, cascading failures and rogue behaviour inside the wider incident process.
Read the whole chapter.
Chapter 23 sets out autonomy levels, identity and MCP authorisation, checkpoints, memory, delegation chains and the EU AI Act hooks, with every source.