AI governance for engineers: controls that leave evidence.
For the people who write the pipeline, the service or the agent: where a governance rule becomes code, what each stage should record, and which obligation that record answers.
Who this is for.
You own the build, the deploy and the pager. Governance tends to reach you as tickets and questionnaires. This route turns it into things you already know how to ship: a policy file the pipeline reads, an eval that can fail the build, a registry the deploy writes to, a guardrail at the enforcement point and a record each of them emits. Read it in order; each step says what it gives you.
- ML and data engineers
- Platform and MLOps engineers
- Application and agent developers
- Security engineers
- Site reliability engineers
On the learning path, start at Eval gate in CI, Observability with OpenTelemetry or Policy-as-code with OPA.
New to the field? Read AI governance explained: the definition, the frameworks and where engineering fits.
Three questions you bring.
Each one answered in brief here, and in full in the Body of Knowledge.
-
What do I actually have to build?
Five layers, in build order: see it, rule it, test it, contain it, prove it. A team of one builds a thin slice through all five: a registry the deploy writes to, one policy that blocks, one eval gate, an identity and a kill switch per agent, and a structured record from each.
-
Which tests count as evidence?
A test whose metrics and thresholds were fixed before the run, whose result is filed against the version it tested, and whose failure blocks the release. For high-risk systems the EU AI Act asks for testing against prior defined metrics and probabilistic thresholds (Art. 9(8)) 1.
-
What must the running system record?
Enough to reconstruct any decision: inputs, outputs, model and prompt versions, tool calls and approvals, as structured, timestamped events. A high-risk system must allow the automatic recording of events over its lifetime (Art. 12(1)), and its deployer keeps the logs it controls for at least six months (Art. 26(6)) 1.
Your route through the site.
In reading order: chapters at the section that matters, then the patterns, tools, templates, datasets and figures that turn them into work.
Orient
- The stack, layer by layer Reference · The five layers, the evidence each one produces and the tool categories that fill it.
- How to read the stack Chapter 04 · Why the build order runs see it, rule it, test it, contain it, prove it.
- The minimum viable stack Figure · The smallest version one engineer can run end to end.
Build
- The build as a chain of gates Chapter 14 · Every gate from use case to release, and the record each one leaves.
- AI system register entry Template · The registry row the deploy writes: owner, scope, tier and status.
- AI register entry builder Tool · Fill a register entry in the browser and export it as JSON.
- Policy Card Pattern · Policy as data the pipeline reads, not a document it ignores.
- Policy Card builder Tool · Turn one rule into a Policy Card, a Rego module with tests and the CI hook.
- Eval Gate in CI Pattern · The pipeline stage that fails the build when an eval fails.
- Fairness metric chooser Tool · Choose the fairness metric before the eval suite runs, with its caveats.
- Test plan schema Template · Metrics, thresholds and datasets frozen before the first run.
- Eval result schema Template · One record per run, filed against the version it tested.
- AIBOM Pattern · What the system is made of, as a bill of materials a scanner can read.
- Model card builder Tool · Draft a model card that points at its evidence, as a document you keep.
- Runtime Guardrail Pattern · Input, output and tool-call checks at the enforcement point.
- Governing AI agents Reference · If you ship agents: registry, identity, tool gateway, checkpoints and a kill switch.
- Agent control profile Tool · Write down the control profile of one agent, as a document you keep.
Prove
- Continuous Assurance Telemetry Pattern · Assurance produced from telemetry as the system runs, not from screenshots.
- Evidence record schema Template · The envelope every control writes its result into.
- Machine-Readable Evidence (OSCAL) Pattern · Evidence an auditor can query, in OSCAL.
- Obligations API (JSON) Dataset · Every obligation with a stable id, to key your evidence records to.
- Maturity self-check Tool · Read your function layer by layer and find the one move that raises the floor.
Start this week.
- Put one eval in CI that can fail the build, with its threshold in version control. Eval Gate in CI
- Have the deploy write a registry entry with an owner and a scope, and block deploys without one. AI system register entry
- Emit one structured evidence record per release: version, eval results, approver and date. Evidence record schema
- Key each record to an obligation id, so an auditor can query it. The obligation register
- Score your five layers with the self-check and pick one move. Maturity self-check
The obligations that matter most.
The duties a pipeline evidences directly: logging, testing, robustness, documentation, oversight by design, transparency, and the threat lists the evals run against.
| Obligation | Applies | Evidence |
|---|---|---|
| EU AI Act Art. 12 record-keeping and logging AIGE-OBL-EUAIA-ART12 | Deferred · | Structured, signed logs; OpenTelemetry traces; tamper-evident event store |
| EU AI Act Art. 15 accuracy, robustness and cybersecurity AIGE-OBL-EUAIA-ART15 | Deferred · | Eval gate; adversarial red-team suite; robustness and security controls; regression evals |
| EU AI Act Art. 9 risk management system AIGE-OBL-EUAIA-ART9 | Deferred · | Risk register as code; threat models; linkage to FRIA and eval results |
| EU AI Act Art. 11 technical documentation (Annex IV) AIGE-OBL-EUAIA-ART11 | Deferred · | AIBOM (CycloneDX ML-BOM, SPDX 3.0 AI); auto-generated technical documentation; model cards |
| EU AI Act Art. 14 human oversight AIGE-OBL-EUAIA-ART14 | Deferred · | Human-in-the-loop checkpoints; kill switch; override and escalation paths |
| EU AI Act Art. 50 transparency for certain AI systems AIGE-OBL-EUAIA-ART50 | In force · | Content labelling and machine-readable marking (e.g. C2PA-style); chatbot disclosure banner |
| OWASP LLM · Top 10 for LLM Applications 2026 AIGE-OBL-OWASP-LLM | Voluntary | Prompt-injection and output-handling controls; eval gate |
| OWASP Agentic · Top 10 for Agentic Applications 2026 AIGE-OBL-OWASP-AGENTIC | Voluntary | Agent threat model; adversarial evals; runtime guardrails; kill switch |
| NIST AI RMF · MEASURE AIGE-OBL-NISTRMF-MEASURE | Voluntary | Eval gates; adversarial red-team suite; metrics per failure mode |
Sources
- [1] Regulation (EU) 2024/1689 (Artificial Intelligence Act), consolidated text of 27 July 2026 (Arts. 9(8), 12(1) and 26(6)). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng (verified: primary)
Other routes.
- For CISOs and risk leads CISOs, heads of risk, model risk managers, internal audit and third-party risk managers.
- For legal counsel and DPOs In-house counsel, data protection officers, privacy and compliance leads, and contract managers.
- For executives and boards Board members, executive committees, and chief AI, data and technology officers.
- For the public sector Public bodies and operators of public services: CIOs, service owners, procurement and oversight.
- For SMEs and start-ups Small and medium-sized companies and start-ups, most of them buying more AI than they build.
- AIGP candidates The public AIGP body of knowledge read against this site. Not affiliated with or endorsed by IAPP.
- Certifications Certifications and assessments in AI governance, and what each one evidences.
Start at the top of the route.
Step 01 is The stack, layer by layer. Each step after it builds on the one before.