On this page

Pattern: Eval Gate in CI

An evaluation suite wired into CI so a model or agent ships only above a documented threshold: the eval run is the control, its result the evidence.

Layer 03 · Evals & Red Teaming as Evidence In the chapter 05 catalogue

Summary: Wire an evaluation suite into the CI/CD pipeline so that a model or agent must pass a defined test, above a documented threshold, before it can ship. The eval run is the control and its result is the evidence; a failing eval blocks the build.

Eval Gate in CI A workflow diagram generated by Archify. 01 / Development 02 / CI / CD Pipeline 03 / Eval Suite (Layer 03) 04 / Registry & Evidence EX / Blocked Release Change Run evals Gate + record Eval run in CI Below threshold Model / Prompt Change · retrained · new tool · Development › Change Model / Prompt Change retrained · new tool CI Pipeline · already runs tests · CI / CD Pipeline › Change CI Pipeline already runs tests Ship Release · deploy to prod · CI / CD Pipeline › Gate + record Ship Release deploy to prod Versioned Eval Suite · versioned with model · Eval Suite (Layer 03) › Eval run in CI › Run evals Versioned Eval Suite versioned with model Capability Eval · coverage floor · Eval Suite (Layer 03) › Eval run in CI › Run evals · e.g. promptfoo Capability Eval coverage floor e.g. promptfoo Adversarial Eval · injection-resistance · Eval Suite (Layer 03) › Eval run in CI › Gate + record · e.g. Garak Adversarial Eval injection-resistance e.g. Garak Threshold Gate · traces to failure mode · Eval Suite (Layer 03) › Gate + record · pass / fail Threshold Gate traces to failure mode pass / fail Structured Result · score · verdict · time · Registry & Evidence › Gate + record Structured Result score · verdict · time Registry Entry · evidence accrues · Registry & Evidence › Gate + record Registry Entry evidence accrues Pipeline Blocked · does not ship · Blocked Release › Below threshold › Gate + record Pipeline Blocked does not ship score score triggers CI file against entry load suite adversarial eval capability eval below threshold emit result above threshold Legend Agent logic Policy Context / trace External system
Eval Gate in CIHow a model or prompt change moves through an eval gate in CI: capability and adversarial evals against a versioned suite, a threshold that ships the release or blocks it and files the result against the registry. Generated from the Body of Knowledge.Open interactive diagram (opens in a new tab)

Objectives

Make “give every control teeth” concrete: give a testable property a consequence, so failure stops a release instead of filing a finding.

Target users

AI governance engineer, ML engineer, platform team.

Impacted stakeholders

Model owners, users exposed to the system, auditors.

Relevant principles

Give every control teeth; build the control at the earliest point it can block.

Context

A model or agent that changes (retrained, re-prompted, given a new tool) and a pipeline that already runs tests for functional correctness.

Problem

Evaluations run once before launch and pasted into a slide prove nothing after the next change. A review board that can only rate findings cannot stop a scheduled launch. Without a gate, evaluation is research, not control.

Solution

Version an eval suite alongside the model. Run at least one capability eval and one adversarial eval in CI (for example with Inspect, promptfoo, Garak or Giskard; illustrative). Set a threshold that traces to a named failure mode or obligation. Fail the pipeline below the threshold. Emit a structured result (suite id, model version, score, threshold, pass/fail, timestamp) filed against the registry entry.

Illustrative schema for the result:

{
  "suite_id": "injection-resistance.v4",
  "model_version": "csa-01@2026-09-18",
  "score": 0.982,
  "threshold": 0.95,
  "result": "pass",
  "timestamp": "2026-09-18T14:22:03Z"
}

Consequences

Regressions are caught before production and evidence accrues automatically. The trade-off is eval maintenance, run-time cost in CI, and the need to tune thresholds to avoid flaky gates.

Policy Card; Adversarial Red-Team Suite; Continuous Assurance Telemetry; Machine-Readable Evidence (OSCAL); Model Card as Control Evidence.

Maps to: EU AI Act Art. 15, Art. 55 · ISO/IEC 42001 · NIST AI RMF (Measure) · OWASP Agentic ASI01/ASI02 · Layer 03 Evals & Red Teaming as Evidence.

Threat IDs follow the OWASP Top 10 for Agentic Applications 2026 1 and function labels the NIST AI RMF2. Mappings are illustrative, not a claim of conformity.

Sources

  1. [1] Top 10 for Agentic Applications 2026 (ASI IDs). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
  2. [2] AI Risk Management Framework (AI RMF 1.0; Govern, Map, Measure, Manage). NIST. 2023-01-26. https://www.nist.gov/itl/ai-risk-management-framework (verified: primary)
Edit this page on GitHub
Cite this pattern

García Aibar, J. (2026). Pattern: Eval Gate in CI. In AI Governance Engineering: The Thesis & Body of Knowledge (v0.5.0), chapter 05, Patterns. https://doi.org/10.5281/zenodo.22956197. https://aigovernanceengineer.com/patterns/eval-gate-in-ci. CC BY 4.0

BibTeX

@misc{aige2026bok,
  author  = {Jorge García Aibar},
  title   = {{AI Governance Engineering: The Thesis \& Body of Knowledge}},
  chapter = {05. Patterns: Eval Gate in CI},
  year    = {2026},
  version = {0.5.0},
  doi     = {10.5281/zenodo.22956197},
  url     = {https://aigovernanceengineer.com/patterns/eval-gate-in-ci},
  note    = {Version 0.5.0}
}
Share on LinkedIn