Maturity, read layer by layer.

For each of the five stack layers, pick the highest observable criterion your running systems meet today. You get the per-layer profile, not a single score: the floor it sets and the single next move, with the pattern that builds it.

Indicative, not legal advice and not a conformity claim. Nothing you enter leaves your browser.

JavaScript is off or has not loaded, so the live result and the exports are not available. The questions, the criteria and the guide below still work as a worksheet you can read and fill in by hand.

Built from 07. Maturity model of the Body of Knowledge. Runs entirely in this page: no account and no upload.

Answer with the system, not the intention: each statement should be something you could show by querying the registry, running the gate or reading the evidence store. Each level assumes the ones below it. If no statement is true of a layer yet, pick "None of these yet".

A team, a platform or a quarter. It appears in the exports and in the link you copy.

Layer 1: Govern-as-Code

What this layer proves: A policy verdict tied to a commit, a pull request or a deploy, machine-readable and reproducible.

Layer 2: Inventory & Transparency

What this layer proves: A registry entry and its attached documents, ideally written by a deployment pipeline rather than typed by hand.

Layer 3: Evals & Red Teaming as Evidence

What this layer proves: A structured eval result: pass or fail against a threshold, versioned alongside the model it tested.

Layer 4: Runtime Controls & Observability

What this layer proves: A stream of runtime decisions and traces.

Layer 5: Assurance & Continuous Compliance

What this layer proves: A live assurance store that any of the lower layers writes into and an auditor can read from.

How to read the result

The model measures the running systems, not the paperwork, and it is read layer by layer [1]. Almost no real function sits at one clean level across all five layers; the usual picture is a ragged line, and that is the point of reading it by layer.

  • Profile. One reading per layer: the highest criterion in that layer's row that the systems meet today. Report the profile, not just the floor: it shows where the leverage is.
  • Floor. The overall level is the weakest layer, because a chain is as strong as its weakest link. You are at a level only when every layer has reached it.
  • Next move. The next cell to the right in the weakest layer's row. The smallest step to the next level is almost always to close the weakest layer, not to add a control to the strongest one. When several layers share the floor, the tool takes the first in build order (Layer 1 to Layer 5); the others still cap the floor until they move.

By hand, without JavaScript

  1. In each layer above, note the highest statement that is true today (Level 0 if none).
  2. The lowest of the five numbers is your floor; the layers that hold it set it.
  3. Your next move is the statement one level up in the first of those layers. Its metrics and checklist questions are listed below, and the pattern that builds it is in the patterns catalogue [2].

Metrics for each step up

Each step has metrics you can read off the systems [1]. Track the trend, not the single number. The most telling one across levels is evidence freshness: if your evidence ages in months, you are not yet continuous.

Level 1 Documented to Level 2 Inventoried

  • Percentage of AI systems and agents in the registry with a named owner and a class
  • Registry-to-production reconciliation gap (systems in production but not registered)

Level 2 Inventoried to Level 3 Tested

  • Percentage of registered systems with a versioned eval suite
  • Percentage with a recorded, timestamped eval result in the last release

Level 3 Tested to Level 4 Enforced

  • Percentage of releases passing through an eval gate (versus bypassing it)
  • Percentage of agents with a tested kill switch and a scoped, non-shared identity
  • Number of releases blocked with a logged reason

Level 4 Enforced to Level 5 Continuous

  • Mean time to detect an unauthorised agent action (an agent doing something outside its declared scope)
  • Evidence freshness (age of the most recent evidence artefact per control)
  • Percentage of controls whose status is answerable by a live query rather than a manual pull

Checklist questions, by level

Answer each with the system, not the intention. A "no" caps you at the level below [1]. Each level also has a typical failure that moves a function back down.

Level 1 Documented

  • Is every AI system covered by a written policy with a named owner?
  • Is there a risk register that a person maintains?

Typical failure: The document was last edited a quarter ago and no longer matches production; the artefact is stale before it is signed.

Level 2 Inventoried

  • Does the registry get an entry automatically at deploy, with owner, scope and status?
  • Can you list every model and agent running today, from the system of record, in under a minute?

Typical failure: Shadow AI. A system or agent reaches production without registering, so the inventory is complete only for the honest.

Level 3 Tested

  • Does every registered system have a versioned eval suite?
  • Are results stored with timestamps?
  • Do you run red-team evals against your agents?

Typical failure: The eval is run once before launch, pasted into a slide, and never re-run when the model or its prompts change.

Level 4 Enforced

  • Does a failing eval or policy check actually block a release?
  • Is an agent without an owner, scope and kill switch prevented from reaching production?
  • Can you show a release that was blocked, with the reason logged?

Typical failure: Brittle gates that engineers route around; a gate maintained by governance alone that engineering does not own; or a gate whose suite is trivial or unmaintained, so the block is real but the assurance is not.

Level 5 Continuous

  • Is runtime telemetry wired to control decisions, not just dashboards?
  • Is evidence emitted as machine-readable artefacts continuously?
  • Would an audit question be answered by a query rather than a collection sprint?

Typical failure: Telemetry that is collected but never wired to a decision; observability without enforcement decays back to Level 3 dressed up as Level 5.

What this is not

This self-check is not a certification, does not confer one and is not an audit. It reads the depth of the runtime data path; schemes such as ISO/IEC 42001 certification attest a management system, which is a different claim. The chapter sets out how the two relate in How this relates to certification and other assessments. The pattern links are illustrative, not a claim of conformity.

The criteria are the chapter's own, in its observable criteria table; the tool adds no criterion of its own.

Sources

  1. [1] 07. Maturity model (five levels): "Observable criteria, by layer and level", "Metrics per level", "Self-assessment checklist" and each level's typical failure, the text this tool renders. AI Governance Engineering Body of Knowledge. 2026-09-24. https://aigovernanceengineer.com/bok/maturity-model (verified: primary)
  2. [2] 05. Patterns: the catalogue each next move links to, one pattern per step. AI Governance Engineering Body of Knowledge. 2026-09-24. https://aigovernanceengineer.com/bok/patterns (verified: primary)