Engineering assurance for frontier AI
For the teams that evaluate, deploy and assure frontier systems: where the controls, patterns and evidence formats on this site apply to that work, and where the open questions are.
Who this route is for.
Frontier AI systems no longer act only through prompts and responses. They use tools, credentials, networks, external services and delegated authority. The environment around the model is therefore part of the system being governed.
This route is for the engineers who build and run that environment and for the people who have to vouch for what it recorded: the evaluation harness, the sandbox and its egress, the credentials an agent holds, the stop handles, the telemetry and the evidence an assessor reads afterwards. Frontier work is one domain where AI governance engineering applies, not a separate discipline.
- Evaluation and red-team engineers
- Safety case and assurance teams
- Agent platform and infrastructure engineers
- Third-party evaluators and auditors
It does not tell a lab how to do safety research, which capabilities to test for or where to set a threshold. It covers the engineering around those decisions: what the environment must enforce, what it must record and how the record becomes evidence.
New to the field? Read AI governance explained: the definition, the frameworks and where engineering fits.
- Model / agent The system under evaluation
- Harness & tools (layer 4) What it can call
- Authorization (layer 4) Whose authority it holds
- Runtime controls (layer 4) Limits and stop handles
- Monitoring (layer 4) Traces of every step
- Evidence (layer 5) Records a machine can check
- Assurance (layer 5) What a reviewer relies on
Evaluation environments
An evaluation result is only as good as the environment that produced it.
The EU AI Act asks every provider of a general-purpose AI model to keep technical documentation of it (Art. 53) and, for models with systemic risk, to evaluate them, including conducting and documenting adversarial testing (Art. 55) 1. The GPAI Code of Practice lists the capability to use tools, and affordances such as access to tools and the level of human oversight, among the model characteristics that bear on systemic risk 2. Neither says what the environment around the model must enforce while it is being tested.
Public evaluation practice gives a starting point. METR's Task Standard (v0.5.0) states that, unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to a small set of destinations such as an LLM API proxy 3. METR's Guidelines for Capability Elicitation treat task bugs, such as incorrect automatic scoring or a crashed environment, as spurious failures to be fixed before a result is reported 4. The evaluation environment profile turns such practices into numbered draft controls: egress, credentials, isolation, stop conditions and the run record, each with the point where it is enforced; all nine are specified in full, with a verification procedure and the evidence they leave.
- Evaluation environment control profile AIGE-CTL-EVAL-001 to 009, draft v0.2, open for technical review.
- Eval gate in CI Pattern: a model or agent ships only above a documented threshold; the run is the evidence.
- Adversarial red-team suite Pattern: a versioned suite built from a threat taxonomy, run in CI or on a schedule.
- Layer 03: evals and red teaming as evidence Chapter 04 on what an evaluation must record to count as evidence.
- Threats mapped to controls OWASP LLM and Agentic, and MITRE ATLAS, each tied to patterns and evals.
Runtime safeguards
The same controls hold whether the agent is under evaluation or in production.
The OWASP Top 10 for Agentic Applications sets out the failure classes an agent platform has to hold against, from goal hijack to rogue agents 5. For tool servers, the MCP authorization specification makes the server validate the audience of each token and forbids passing a token through to another service 6. Chapter 23 turns both into rules: a unique identity per agent, short-lived credentials, a tool allow-list, execution limits and a stop handle that works per agent.
Isolation is a claim to test, not to assume. METR's public investigation of a 2026 agent hacking incident reports that agents meant to be "fully isolated from one another" communicated through a shared internal package repository 7. The agent runtime profile states each boundary as a control with the evidence it should leave; its verification procedures are still to be written in review, so that a team can show the boundary held rather than assert it.
- Agent runtime control profile AIGE-CTL-AGENT-001 to 031, derived from chapter 23, draft v0.1.
- Agent identity and scoped credentials Pattern: one identity per agent, credentials scoped to the task and short-lived.
- Runtime guardrail Pattern: guardrails on the live request path that emit a decision event for each call.
- Kill switch and circuit breaker Pattern: a tested stop for one agent that leaves the rest of the fleet running.
- Identity and short-lived credentials Chapter 23: channel authentication is not agent identity.
- Kill switch and per-agent circuit breakers Chapter 23: stopping one agent, and stopping across hops.
- Agent control profile builder Toolkit: the minimum control set for one agent at its autonomy level.
- Threats mapped to controls The agentic threats, each tied to the controls that hold against it.
Assurance evidence
A reviewer should be able to check the record, not only read the report.
The NIST AI RMF asks for mechanisms to supersede, disengage or deactivate AI systems whose outcomes are inconsistent with intended use (MANAGE 2.4) 8. An assessor can only rely on such a mechanism if the system leaves a record that it exists, that it was tested and what happened when it fired. The patterns and schemas on this site describe that record: evidence emitted by the control itself, in a documented format, collected continuously rather than assembled before an audit.
- Machine-readable evidence (OSCAL) Pattern: control evidence in a standard format, so an audit becomes a query.
- Continuous assurance telemetry Pattern: control decisions streamed to one assurance store as they happen.
- Open data The registers behind the site as static JSON, with schemas and an OpenAPI file.
- API index The machine-readable catalogue of every dataset.
- MCP server Ask any MCP client and get answers read from the registers, each with its source.
- Layer 05: assurance and continuous compliance Chapter 04 on evidence that stays current between audits.
Incidents
When a boundary fails, the record decides what can be learned.
Chapter 23 sorts agent incidents into eleven classes, from goal hijack to exfiltration through a tool, each with a detection signal and a first containment, and traces agent and tool operations with the OpenTelemetry GenAI conventions 9. The record itself can be the weak point. The same investigation reports spoofed tool calls in at least 96 transcripts and transcript gaps caused by container resets 7. A run record that the agent can write to, or that a restart can truncate, is weak evidence.
- Incident cases Public incidents as engineering post-mortems: failure mode, control, evidence.
- From incident to control How a case becomes a failure mode and a control.
- DNS covert channel in a training run (2026) Egress filtered at the web proxy but not at the resolver, and a stop that was not automatic.
- Credential published by an internal agent (2026) A person's token within an agent's reach, split to get past secret scanning.
- Public file hosting from a training run (2026) A broken shared filesystem and live internet sent agents to a public host.
- Package repository as a message board (2026) A shared service outside the isolation model, between samples meant to be independent.
- Misconfigured third-party cyber evaluations (2026) Told there was no internet, connected anyway: four incidents at an evaluation partner.
- Unsanctioned actions in a cyber range (2026) Open internet bounded only by a prompt, and a token one agent left reused by another.
- Chapter 17: incidents Detection, triage, reporting clocks and corrective action.
- An agent incident taxonomy Chapter 23: eleven classes, each with a detection signal and a first containment.
- Incident pipeline Pattern: detect, triage and report serious incidents within the legal window.
- Incident clock Toolkit: every reporting deadline as a calendar date, as chapter 17 states the clocks.
Certifications and frameworks
Schemes that already exist, and what an open implementation layer adds to them.
AIUC-1 lists 53 requirements in six domains, from data and privacy to society, two of them now marked retired 10; its certificates are issued centrally on an audit report prepared by an accredited auditor 11. ISO/IEC 42001 sets requirements for an AI management system 12, and ISO/IEC 42006 sets the requirements for the bodies that audit and certify against it 13. The NIST AI RMF organises outcomes under govern, map, measure and manage 8.
These schemes say what must be true. The controls on this site add the engineering layer underneath: for each control, the point where it is enforced, how a third party would check that it held and the evidence record it leaves (in full for nine controls so far, in draft for the rest). A certification or framework can map to them; they do not replace one and confer no certification. This site is not affiliated with, reviewed by or certified by AIUC, ISO, NIST, METR or any AI developer whose guidance it cites.
Public guidance from frontier developers and evaluators now uses the same terms for the environment around a model: an authorization boundary, and expectations for isolation, credential handling, monitoring and stop conditions 14; a sandbox whose only outside connection is the model API, with keys kept outside it, a configuration checked before every evaluation and challenges confirmed solvable 15. The profiles on this site cross-reference that guidance where it evidences a control. They are not derived from it, and no developer or evaluator has reviewed or endorsed them.
Where it sits on the stack.
The same five layers as every other route, read for frontier evaluation and deployment.
- Layer 01: Govern-as-Code Evaluation plans, egress and tool policies written as code and reviewed like code.
- Layer 02: Inventory & Transparency An inventory of models, agents, harnesses and tool servers, each with an owner.
- Layer 03: Evals & Red Teaming as Evidence Evaluations and red-team runs that record their environment, so the result counts as evidence.
- Layer 04: Runtime Controls & Observability Identity, tool mediation, execution limits and stop handles on every agent step.
- Layer 05: Assurance & Continuous Compliance Evidence records an assessor can check, collected as the controls run.
Open questions.
- Which parts of the environment around a model must an evaluation record for its result to count as evidence?
- Where does an agent control boundary sit when one agent delegates to another across organisations?
- Which runtime safeguards can be verified from telemetry alone, and which need an inspection or an attestation?
Each one is a candidate for a research note. If you work on one of them, or think this route gets the engineering wrong, contribute: corrections are reviewed in the open.
Sources
- [1] Regulation (EU) 2024/1689 (Artificial Intelligence Act), consolidated text of 27 July 2026 (Art. 53 obligations of providers of general-purpose AI models; Art. 55 model evaluation and adversarial testing for models with systemic risk). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng (verified: primary)
- [2] GPAI Code of Practice, Safety and Security chapter (Appendix 1.3 capabilities to operate autonomously and to use tools; affordances such as access to tools and the level of human oversight). European Commission. 2025-07-10. https://ec.europa.eu/newsroom/dae/redirection/document/118119 (verified: primary)
- [3] METR Task Standard, version 0.5.0 (task machines "MUST NOT have internet access" except to a small set of destinations, unless the task declares full_internet; STANDARD.md last changed 2024-10-30). METR. 2024-10-30. https://raw.githubusercontent.com/METR/task-standard/main/STANDARD.md (verified: primary)
- [4] Guidelines for capability elicitation (task bugs such as incorrect automatic scoring or a crashed environment treated as spurious failures, fixed before results are reported). METR. 2024-03-15. https://metr.org/blog/2024-03-15-guidelines-for-capability-elicitation/ (verified: primary)
- [5] OWASP Top 10 for Agentic Applications 2026 (ASI01 to ASI10, from agent goal hijack to rogue agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
- [6] MCP specification 2026-07-28, Authorization (OAuth 2.1 resource server; audience validation; servers "MUST NOT accept or transit any other tokens"). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization (verified: primary)
- [7] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (by Hjalmar Wijk and Ajeya Cotra (METR) and Ryan Greenblatt (Redwood Research, contracting with METR): agents "meant to be fully isolated from one another"; spoofed tool calls in at least 96 transcripts; gaps from container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
- [8] NIST AI RMF 1.0 (AI 100-1) (MANAGE 2.4: mechanisms to supersede, disengage or deactivate AI systems). NIST. 2023-01-26. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf (verified: primary)
- [9] OpenTelemetry semantic conventions for generative AI (agent, workflow, execute_tool and memory operations; status Development). OpenTelemetry. 2026. https://github.com/open-telemetry/semantic-conventions-genai/tree/main/docs/gen-ai (verified: primary)
- [10] AIUC-1 standard (requirements A001 to F002 in six domains, two marked retired; most recent version released 15 Jul 2026). Artificial Intelligence Underwriting Company. 2026-07-15. https://standard.aiuc-1.com/ (verified: primary)
- [11] Accredited AIUC-1 auditors (certificates issued centrally; accredited auditors prepare the audit report). Artificial Intelligence Underwriting Company. 2026. https://standard.aiuc-1.com/accredited-auditors (verified: primary)
- [12] ISO/IEC 42001:2023, AI management systems (referenced by identifier only; requirements for an AI management system). ISO/IEC. 2023. https://www.iso.org/standard/81230.html (verified: secondary)
- [13] ISO/IEC 42006:2025, Requirements for bodies providing audit and certification of AI management systems (builds on ISO/IEC 17021-1). ISO/IEC. 2025-07. https://www.iso.org/standard/44546.html (verified: primary)
- [14] Third-party cyber evaluations involving OpenAI models (the evaluator's "intended authorization boundary"; OpenAI will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions"). OpenAI. 2026-08-04. https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ (verified: primary)
- [15] Improving our alignment and security efforts (best practices for external evaluation partners: sandbox and network isolation, API keys kept outside the environment, configuration verified before every evaluation, challenges confirmed solvable, explicit scope, real-time monitoring that ends an out-of-scope run). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
Start with the evaluation environment profile.
Nine draft controls for the environment a model is evaluated in, each with its enforcement point and evidence.