The AI Governance Engineering Thesis

Governance as a system to be built, run and measured, not a document to be signed.

  • v0.5.0
  • Updated
  • CC BY 4.0

Leer en español

Where AI governance engineering sits An architecture diagram generated by Archify. AI compliance / legal · interprets obligations · Governance functions AI compliance / legal interprets obligations Responsible AI / ethics · sets values & principles · Governance functions Responsible AI / ethics sets values & principles AI safety research · studies models in principle · Governance functions AI safety research studies models in principle GRC engineering · the parent discipline · Governance functions GRC engineering the parent discipline MLOps / LLMOps · builds, deploys, serves · Build & run MLOps / LLMOps builds, deploys, serves AI Governance Engineering · governance as engineered systems · Build & run AI Governance Engineering governance as engineered systems AI security engineering · the sibling, defends · Build & run AI security engineering the sibling, defends AI systems & agents · the object, in production · Architecture component AI systems & agents the object, in production Model risk management · validates models (SR 11-7) · Architecture component Model risk management validates models (SR 11-7) Audit-ready evidence · machine-readable proof · Architecture component Audit-ready evidence machine-readable proof Auditors / regulators · read & query the proof · Architecture component Auditors / regulators read & query the proof obligation → control values → controls consumes research AI specialisation defended system as input gates the pipeline builds & serves governs continuously validates the model emits runtime evidence audit is a query Governance functions Build & run Legend Backend Database Cloud Security External
Where AI governance engineering sitsWhere AI governance engineering sits among the adjacent roles (analyst, platform and assurance), turning governance intent into running controls and evidence. Generated from the Body of Knowledge.Open interactive diagram (opens in a new tab)

Version 0.5.0 · 2026-09-25 · Jorge García Aibar and Aurélie Pols


AI governance engineering is the application of engineering practice (systems thinking, product thinking and code) to the governance of AI systems. It treats governance not as a document to be signed but as a system to be built, run and measured, with the same rigour engineers already apply to the models and agents it governs.

It is more than “AI governance plus a few scripts.” It is a change in how the work is done. Legacy AI governance describes an AI system on paper and hopes the paper stays true. Engineering the governance means the policy is executable, the control runs in the pipeline, the evidence is produced as a by-product of the build, and the whole thing is judged by one test: did realised risk actually fall, and can a regulator or auditor read the proof? It is a capability, not a job title. Anyone close enough to the build can develop it. It is measured by realised risk reduction and by audit-ready evidence, never by how many frameworks appear on a slide.

The idea does not appear from nowhere. It inherits from a line of engineering movements that turned process into running systems: site reliability engineering, DevSecOps, policy-as-code and software supply-chain security. Most directly, it inherits from GRC engineering, which since about 2024 has turned governance, risk and compliance into a product built with code, tested in CI/CD and shipping evidence through APIs1. AI governance now needs the same step-change, because the thing being governed (models that retrain, prompts that change, agents that act on their own) moves faster than any document can follow.

Fundamental problems with legacy AI governance

1. Governance written for systems that no longer exist. Legacy AI governance runs on PDF policies and spreadsheet inventories that describe an AI system as it was on the day it was reviewed. But models are retrained, prompts are rewritten and agents acquire new tools by the day. The artefact is stale before it is signed. It is no accident that the function still sits mostly with privacy, legal and IT and only 5% with security3: far from the pipeline where the system actually changes.

2. Point-in-time review of a continuously changing thing. Annual assessments and committee sign-offs assume a system that holds still long enough to be judged. Frontier models and autonomous agents do not. Gartner expects more than 40% of agentic AI projects to be cancelled by the end of 2027, citing inadequate risk controls among the causes4, and predicts that by 2029 more than half of successful attacks on AI agents will exploit access-control weaknesses and prompt injection5: runtime failure modes that a once-a-year review is structurally blind to.

3. Governance as a gate at the end, not a property of the build. Governance arrives after the model is trained, as a checkpoint to clear before launch. Engineers experience it as a tax collected at the door, and nothing it produces is wired back into how the system is built. A policy that can only recommend cannot stop a bad release. An eval that can fail the build can. Governance placed at the end can only describe risk; governance built into the pipeline can prevent it.

4. Framework theatre. Mapping to NIST AI RMF or ISO/IEC 42001 becomes the end state instead of the starting point. A green mapping matrix is mistaken for a working control. Yet as of 2026-09-24 no harmonised standard is cited in the EU’s Official Journal, so even an ISO 42001 certificate confers no presumption of conformity with the AI Act6. Coverage is not assurance. A crosswalk proves you have read the framework, not that the control it points to actually fires; a green matrix over a broken control is “theatre with extra steps”2.

5. No runtime data path. The registry does not know what is running. Gartner asks AI governance platforms for “automated policy enforcement at runtime”11, yet one vendor’s comparison of the category, published by a competitor in it, finds that most of it “manages the program (inventories, assessments, framework mappings, evidence workflows) without any runtime data path”7. IBM’s own account of being named a Leader in Gartner’s first Magic Quadrant for AI Governance Platforms (June 2026) points the same way: it describes visibility into AI use cases and a roadmap of centralised AI asset inventory, lineage and use-case onboarding (the program layer), and does not mention runtime enforcement10. So the three questions that define the discipline (what AI is running, what is it allowed to do, what evidence proves it) go unanswered, because nothing is connected to production. Meanwhile a security vendor’s 2026 survey reports that roughly one in eight AI breaches involved agentic systems8: exactly the layer the paper registry cannot see.

Values

Eight affirmations. Each states what we build toward and, by naming it, what we build away from. Not all are new: values 1, 5 and 7 (governance as code, machine-readable evidence and measured risk reduction) are inherited from GRC engineering; values 2 and 4 (evals that fail the build, and agent identity and scope) are what AI forces us to add. Chapter 03 expands each with an in-practice example and the anti-pattern it rejects.

1. Governance is code, not a document. A policy in a PDF is a statement of intent a human must remember and apply; a policy as code is a control that executes, versioned in a repository and enforced without anyone remembering to. The document describes the rule; the code is the rule, and only what runs can be measured.

2. Evals fail builds; reviews only recommend. A review produces a recommendation someone may act on, later, or not. An eval produces a verdict with consequences: the model or agent passed or failed a defined test, and a failure blocks the release. We prefer controls that bite.

3. Evidence comes from runtime, not from a point-in-time attestation. An attestation says a control was in place when someone looked. Runtime evidence shows it working continuously, emitted by the system as it runs, because an AI system changes between reviews, and evidence gathered once decays immediately.

4. Every agent carries its own identity and scope. An actor on shared credentials is ungovernable: you cannot attribute its actions, revoke its access precisely, or bound what it may do. Identity is the precondition of accountability; scope is the precondition of containment, and both are established before the actor is allowed to act.

5. Evidence is machine-readable or it is not evidence. Evidence a human must produce, format and file by hand cannot be queried, diffed or verified at speed. Machine-readable artefacts (OSCAL, structured eval results, signed logs) turn the audit into a query and feed continuous assurance instead of a one-off binder.

6. Tooling must be inspectable and composable. You cannot trust a verdict you cannot trace. Tooling whose reasoning and data path you can open, bought or built, lets you follow a decision to the rule that produced it and the evidence it emitted, and wire it into the pipeline you already run rather than export into someone else’s.

7. Success is measured in realised risk reduction, not framework coverage. Mapping every control to a framework proves you have read it, not that any risk fell. We measure the thing itself: did the rate of the failure mode drop, did the blast radius shrink, did the incident get caught earlier? Coverage is an input; realised risk reduction is the outcome.

8. Governance is owned with engineering, not enforced from outside. Governance that sits apart and grants or denies passage is a bottleneck engineers route around. Owned jointly with engineering, built into the paved path, adopted because it is the easiest way to ship, it becomes part of how things are made, not a meeting one side dreads.

Principles

The values say what we prefer; these principles say what we commit to do. They are rules of action, not restatements of the preferences above.

Build the control at the earliest point it can block. Put every control where it can still stop the thing going wrong, and no later: in the repository, the build and the runtime, not in a review after the fact. The earliest enforceable point is the cheapest and the strongest, so that is where we put it.

Give every control teeth, or call it a signal. A control has to be able to change what happens next: block a merge, fail a deploy, revoke access. Anything that can only inform a committee is a signal, and we label it honestly as one rather than dress it up as a control.

Register and bound every actor before it acts. Nothing, human or non-human, gets to act until it has an owner, a declared scope and a way to be stopped. Autonomy is granted only where it can be attributed, contained and withdrawn, never by default.

Instrument the build to produce its own proof. Wire each control to emit its own record as it runs, so assurance falls out of the system instead of being assembled by hand. If demonstrating a control needs a screenshot, we have not finished building it.

Start from a named failure mode or a named harm. Design each control against a specific way the system fails (prompt injection, tool misuse, agent identity abuse, data exfiltration) or a specific harm to a person’s rights. If we cannot name the risk it answers, we do not build it.

Make the governed path the easiest path. Ship governance as tooling, templates and paved paths engineers adopt without asking permission, and measure adoption. If routing around the governance is easier than using it, we fix the product, not the people.

What AI governance engineers build

Not decks. Working artefacts, versioned in a repository and running in production:

  • Policy-as-code: governance rules as executable policy (OPA/Rego, Cedar, Policy Cards) that evaluate in CI/CD and at runtime.
  • An agent registry: the runtime-aware inventory of every model, service and agent, each with an owner, a scope and a status.
  • AIBOM and model/data cards: the bill of materials for an AI system (CycloneDX ML-BOM, SPDX 3.0 AI profile) and structured transparency documentation.
  • Eval gates in CI: adversarial and capability evals (Inspect, promptfoo, Garak, Giskard) wired into the pipeline so a failing eval blocks the release.
  • Runtime guardrails and kill switches: input/output controls, tool-call mediation and a tested way to stop an agent, at the point of action.
  • Continuous assurance telemetry: tracing and monitoring (OpenTelemetry, agent observability) that turns production behaviour into a live control signal.
  • Machine-readable evidence: OSCAL and signed, structured artefacts that make the audit a query instead of a scramble.
  • Incident pipelines: the plumbing to detect, triage and report serious incidents on the clock, including the EU AI Act’s Article 73 reporting for high-risk systems.
  • FRIA and DPIA templates as code: fundamental-rights and data-protection impact assessments maintained as versioned, reviewable artefacts, not one-off documents.

These artefacts map, layer by layer, onto the five-layer AI governance engineering stack: Govern-as- Code, Inventory & Transparency, Evals & Red Teaming as Evidence, Runtime Controls & Observability, and Assurance & Continuous Compliance.

We concede the inheritance plainly, because it is the honest defence against “this is just GRC with AI words.” Three of the five layers (Govern-as-Code, Inventory & Transparency, and Assurance & Continuous Compliance, which carry policy as code, the asset inventory and machine-readable evidence) are inherited from GRC engineering and carried across almost unchanged. Two are what AI forces us to add: evals and red-teaming as controls (layer 03), because the thing being governed is a model whose behaviour can only be established by testing it; and agent identity and runtime control (layer 04), because an autonomous actor has no analogue in classic GRC. The new work of the discipline concentrates in those two layers.

A discipline, distinct from its neighbours

AI governance engineering is not AI safety research, MLOps, model risk management, AI compliance or legal work, or Responsible AI ethics; it is the engineering that turns all of those into running controls and readable evidence. It is the AI-era sibling of AI security engineering, the direct descendant of GRC engineering. One name clash is worth flagging: some vendors use the same words, “AI governance engineering”, for the reverse problem, governing the AI tools that engineers use inside their own workflows9. That is governed AI engineering, not the discipline described here. Chapter 01 draws every one of these lines in full.

Authors

Jorge García Aibar (v0.1–v0.5.0), AI Governance & Privacy Engineer. LinkedIn: https://www.linkedin.com/in/jorgara

Aurélie Pols (v0.1–v0.5.0), Responsible AI (EU/Global), Privacy & Data Governance. LinkedIn: https://www.linkedin.com/in/aureliepols

Co-authors wanted. This is version 0.5.0: a public draft, deliberately incomplete. It was started by one practitioner and it needs many. If you build governance for AI systems (policy-as- code, agent registries, eval gates, runtime guardrails, continuous assurance) and you can bring a verified fact, a pattern that worked, or a sharper argument, you are invited to co-author. The discipline is a capability anyone can develop, and this text belongs to everyone who does the work.

Sign / get involved

Licence

This Thesis is licensed under CC BY 4.0. You may share and adapt it provided you give appropriate credit, link to the licence and indicate changes. Attribution: Jorge García Aibar and Aurélie Pols.

Sources

  1. [1] GRC Engineering Manifesto. grcengineering. ~2024. https://grc.engineering/ (verified: primary)
  2. [2] “What is GRC Engineering” (Ayoub Fandi). GRC Engineer. 2025. https://grcengineer.com/what-is-grc-engineering/ (verified: primary)
  3. [3] AI Governance Profession Report 2025. IAPP (with Credo AI). 2025-04-16. https://iapp.org/resources/article/ai-governance-profession-report/ (verified: primary)
  4. [4] “Gartner Predicts Over 40% of Agentic AI Projects Will Be Canceled by End of 2027”. Gartner. 2025-06-25. https://www.gartner.com/en/newsroom/press-releases/2025-06-25-gartner-predicts-over-40-percent-of-agentic-ai-projects-will-be-canceled-by-end-of-2027 (verified: primary)
  5. [5] “Gartner Forecasts the Market for Securing AI Will Reach Almost $5 Billion in 2027”. Gartner. 2026-08-26. https://www.gartner.com/en/newsroom/press-releases/2026-08-26-gartner-forecasts-the-market-for-securing-ai-will-reach-almost-5-billion-in-2027 (verified: primary)
  6. [6] Standardisation of the AI Act (no harmonised standard yet referenced in the Official Journal, so no Art. 40 presumption of conformity from any standard, ISO/IEC 42001 included; page last updated 2026-08-03; no Commission implementing decision citing one found in the Publications Office index on 2026-09-24). European Commission. 2026-08-03. https://digital-strategy.ec.europa.eu/en/policies/ai-act-standardisation (verified: primary)
  7. [7] “Best AI Governance Platforms in 2026: 14 Enterprise Vendors Compared” (vendor-published comparison of the 13 Magic Quadrant vendors plus its own product; runtime data path critique). Kosmoy. 2026-07-10. https://www.kosmoy.com/resources/blog/best-ai-governance-platforms-2026/ (verified: secondary)
  8. [8] 2026 AI Threat Landscape Report (vendor survey; key finding stated on the report page: one in eight breaches were agentic). HiddenLayer. 2026. https://www.hiddenlayer.com/report-and-guide/threatreport2026 (verified: primary)
  9. [9] “AI Governance Engineering” (governing AI used inside engineering workflows). Visure Solutions. 2026. https://visuresolutions.com/ai-engineering/ai-governance-engineering/ (verified: primary)
  10. [10] “IBM recognized as a Leader in the Gartner Magic Quadrant for AI Governance Platforms” (vendor announcement citing Gartner, Magic Quadrant for AI Governance Platforms, L. Kornutick et al., 17 June 2026, the first MQ for the category; visibility into AI use cases; roadmap: AI asset inventory and lineage, use-case onboarding). IBM. 2026-06-17. https://www.ibm.com/new/announcements/ibm-recognized-as-a-leader-in-gartner-magic-quadrant-for-ai-governance-platforms (verified: secondary)
  11. [11] “Global AI Regulations Fuel Billion-Dollar Market for AI Governance Platforms” (platforms should enable “automated policy enforcement at runtime”; AI governance spending USD 492M in 2026, over USD 1B by 2030). Gartner. 2026-02-17. https://www.gartner.com/en/newsroom/press-releases/2026-02-17-gartner-global-ai-regulations-fuel-billion-dollar-market-for-ai-governance-platforms (verified: primary)