---
title: "Data Admission and Privacy Control Profile"
description: "An open control profile for AI training data: dataset admission, rights ledger, lawful basis, purpose limits, special categories and lineage. Draft v0.1."
canonical: https://aigovernanceengineer.com/controls/data-admission-and-privacy
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22857084
version: "0.1"
updated: 2026-09-26
---

# Data Admission and Privacy Control Profile

> Reference controls for the data AI systems learn from: admission before a job reads a dataset, the right to use each source, lawful basis and purpose, fitness for purpose, integrity, lineage and downstream use. Every control is a draft derived from site patterns, record schemas and chapters 14 and 19, open for technical review.

- Version: 0.1
- Status: Draft
- Review: Open for technical review
- Published: 2026-09-26
- Updated: 2026-09-26
- Authors: Jorge García Aibar
- Controls: 12

Draft for review, not a claim of conformity. These are draft control specifications, open for technical review: illustrative, not legal advice and binding on no one.

## Scope

Datasets that training, fine-tuning, validation, testing, evaluation and retrieval-index jobs read, the sources they are built from and the consumers of the outputs of the models built on them. Data handling at inference time, transfers and automated decision-making are left to chapter 19; evaluation environments are covered by the evaluation environment profile.

## How to read a control

Each control has a stable id (`AIGE-CTL-<PROFILE>-<NNN>`) that never changes and is never reused, and records:

- Objective: the outcome the control secures, in one sentence.
- Failure modes: observable events that mean the control failed.
- Scope, enforcement points (pre_merge, deploy, runtime, periodic) and the failure response (deny, require_approval, alert).
- Verification: how a third party would check it (inspect, test, observe, attest).
- Evidence: the artefact it leaves and the stack layer that keeps it.
- Mappings: obligations, ISO/IEC 42001 Annex A, NIST AI RMF, OWASP and other references, each id checked against the site registers.

Depth: Specified controls carry a verification procedure, evidence, notes and an observation example, and have a page of their own; "Derived from site material" controls restate existing site material (an agent control of chapter 23, a pattern, a record schema or a chapter, named on each control) and add nothing it does not say; "Draft outline" controls are skeletons with open questions. This profile: 12 derived from site patterns, record schemas and chapters of the Body of Knowledge.

An observation is what a check of a control would emit: control_id, subject, expected, observed, status (pass, fail or not_applicable), timestamp and evidence[].

## AIGE-CTL-DATA-001 Dataset Admission Gate at Read Time

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-001

- Id: `AIGE-CTL-DATA-001` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [the dataset-admission-record.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record); [chapter 14, Governing AI development](https://aigovernanceengineer.com/bok/governing-development))
- Objective: A training, fine-tuning, validation, testing, evaluation or retrieval-index job reads a dataset version only if a signed admission record for that version admits the job's pipeline for its use case and target system, and the rights, quality, representativeness, bias and integrity checks passed or were waived by someone entitled to waive them.
- Failure modes:
  - A job reads a dataset version that has no admission record for its pipeline, for example a training job pointed at a whole warehouse admitted for nothing.
  - A model trains on data outside its consent scope, on a sample that misses the population it will serve, on labels nobody audited or on a snapshot someone altered, and the problem surfaces in production or in an audit, when the fix is a retrain.
  - The person who wants the dataset used is the only person who decides it may be: the requester signs the admission, or a check is waived by someone not entitled to waive it.
  - A missing admission field produces a reminder in a wiki instead of a failed run.
- Scope: Every job that reads data to train, fine-tune, validate, test, evaluate or build a retrieval index, and the dataset versions it reads. The checks the gate runs are specified in AIGE-CTL-DATA-002 to 009; a sandbox pipeline with its own lighter admission is out of scope.
- Enforcement points:
  - runtime: at the point of action (gateway or guardrail)
- Verification:
  - Inspect: Each admission record validates against dataset-admission-record.v1 and carries what the pattern lists: subject, pipeline, target system, linked data card, decision, the checks with the obligation each enforces, the content hash of the admitted snapshot, the actor and a signature.
  - Test: A job that presents a use-case id or a pipeline the record does not admit is denied the read.
- Evidence:
  - Admission record per dataset version and permitted pipeline, signed by the data owner · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
  - Read-time verdicts of the gate: the use-case id presented and the admit or deny decision · Layer 01 Govern-as-Code
- Failure response: deny: block the action. The policy denies the read unless the record admits that pipeline for that use; a job with no admitted dataset does not start. A new version, a new source, quality drift, a licence change, an erasure request or a new use case reopens admission.
- Layers: [Layer 01 Govern-as-Code](https://aigovernanceengineer.com/bok/the-stack#layer-01-govern-as-code), [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [Policy Card](https://aigovernanceengineer.com/patterns/policy-card)
- Mappings:
  - Obligations: [EU AI Act Art. 10 data and data governance](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art10); [General Data Protection Regulation (EU) 2016/679, GDPR Art. 5(1)(b) and 6(4) purpose limitation](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art5-1b); [ISO/IEC 42001, A.7 Data for AI systems](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a7)
  - ISO/IEC 42001: A.7.2 Data for development and enhancement of AI system; A.7.4 Quality of data for AI systems; A.7.5 Data provenance
  - NIST AI RMF: MAP 2.3 Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation.; MAP 4.1 Approaches for mapping AI technology and legal risks of its components – including the use of third-party data or software – are in place, followed, and documented, as are risks of infringement of a third party’s intellectual property or other rights.
  - OWASP: [LLM05:2026 Data and Model Poisoning](https://aigovernanceengineer.com/resources/threats#threat-llm05-2026)
- References:
  - [1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use)
  - [2] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Data for training and testing")
  - [3] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Owners, stewards and the admission gate")
  - [4] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Art. 10 (data and data governance: 10(2) practices, including origin, preparation, bias examination and mitigation, and data gaps; 10(3) relevant, sufficiently representative, free of errors and complete; 10(4) specific setting of use)
  - [5] Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (anonymity of models; legitimate interest; consequences of unlawful processing in development)
  - [6] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
  - [7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title)
  - [8] Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (GOVERN 6.1 third-party risks incl. infringement of intellectual property or other rights; MAP 1.1 intended purposes documented; MAP 2.3 data collection and selection considerations identified and documented; MAP 3.3 targeted application scope; MAP 4.1 legal risks of components incl. third-party data; MEASURE 2.10 privacy risk examined and documented; MANAGE 1.4 negative residual risks to downstream acquirers and end users documented)
- Implementation notes:
  - Write one admission record per dataset version and per permitted pipeline on the published dataset-admission-record schema, and make every data-reading job present its use-case id and target system at read time; the check is policy-as-code in the pipeline, not a reminder in a wiki.
  - Separate the duties: the data owner is accountable and signs, the data steward operates the checks, and a small review board settles contested admissions. The minimum checklist lives as code, so adding a check is a reviewed change.
  - Give waivers an owner and an expiry, or they become the norm; record conditions on the admission record (for example "collect islands-region claims before the next retrain") so the next retrain cannot start until they are closed.
- Open questions:
  - What lighter admission should a sandbox pipeline for exploratory work carry, and how is data kept from leaving the sandbox into a training job?
  - Which enforcement point fits a gate that decides at read time inside a data pipeline: the platform's access layer, the job scheduler or the storage policy engine?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-001.json

## AIGE-CTL-DATA-002 Dataset Card for Every Admitted Version

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-002

- Id: `AIGE-CTL-DATA-002` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the dataset-card.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-card); [the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [chapter 14, Governing AI development](https://aigovernanceengineer.com/bok/governing-development))
- Objective: Every dataset version that is admitted carries a dataset card that states its owner, purpose, provenance, lawful basis, licence and retention rule, and the admission gate checks the card is complete before it admits the version.
- Failure modes:
  - A dataset is admitted with a card that lacks a lawful basis, a provenance, a retention limit or a licence, so the rights and the deletion date cannot be read from it.
  - The card describes a previous version: its composition, representativeness or quality checks no longer match the snapshot that was admitted.
  - The card is written from memory after training instead of filled from the admission record and lineage.
- Scope: Datasets admitted to training, fine-tuning, validation, testing, evaluation or retrieval-index pipelines, one card per version. The model card and the system card, which describe what was built from the data, are out of scope.
- Enforcement points:
  - deploy: before a version is deployed or released
- Verification:
  - Inspect: The card of each admitted version validates against dataset-card.v1 (dataset id, version, name, owner, description, provenance, lawful basis, licence and retention rule), and the admission record links to it.
- Evidence:
  - Dataset card per admitted version · Layer 02 Inventory & Transparency · [dataset-card.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-card)
  - Card-completeness check on the admission record · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
- Failure response: deny: block the action. A version whose card is missing or incomplete is not admitted; the admission record names the card that was checked.
- Layer: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [Model Card as Control Evidence](https://aigovernanceengineer.com/patterns/model-card-as-control-evidence)
- Mappings:
  - Obligations: [EU AI Act Art. 10 data and data governance](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art10); [ISO/IEC 42001, A.7 Data for AI systems](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a7)
  - ISO/IEC 42001: A.7.2 Data for development and enhancement of AI system; A.7.5 Data provenance
  - NIST AI RMF: MAP 2.3 Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation.
- References:
  - [2] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Data for training and testing")
  - [9] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Model cards, system cards and datasheets")
  - [1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use)
  - [10] Datasheets for Datasets (Gebru et al.; arXiv 1803.09010) (motivation, composition, collection, preprocessing, uses, distribution and maintenance)
  - [4] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Art. 10 (data and data governance: 10(2) practices, including origin, preparation, bias examination and mitigation, and data gaps; 10(3) relevant, sufficiently representative, free of errors and complete; 10(4) specific setting of use)
  - [7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title)
  - [8] Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (GOVERN 6.1 third-party risks incl. infringement of intellectual property or other rights; MAP 1.1 intended purposes documented; MAP 2.3 data collection and selection considerations identified and documented; MAP 3.3 targeted application scope; MAP 4.1 legal risks of components incl. third-party data; MEASURE 2.10 privacy risk examined and documented; MANAGE 1.4 negative residual risks to downstream acquirers and end users documented)
- Implementation notes:
  - Fill the card from the same records the gate reads (the admission record and lineage), so the datasheet travels with the admission record and covers motivation, composition, collection, preprocessing, uses, distribution and maintenance.
  - Record in the card who the data does and does not represent (populations and known gaps) and the quality and bias checks run against this version, with their results.
- Open questions:
  - Which optional card fields (composition, representativeness, quality checks, splits) should become mandatory for data admitted to a high-risk system?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-002.json

## AIGE-CTL-DATA-003 Training-Data Rights Ledger Row per Source

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-003

- Id: `AIGE-CTL-DATA-003` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Training-Data Rights Ledger pattern](https://aigovernanceengineer.com/patterns/training-data-rights-ledger); [the dataset-card.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-card); [chapter 14, Governing AI development](https://aigovernanceengineer.com/bok/governing-development))
- Objective: Every training source has a ledger row, per source and not per merged dataset, that records its acquisition channel, licensor, licence terms for training, commercial use and distribution of derived models, the legal basis where the data is personal, the permitted uses and, for crawled content, the rights-reservation check with its method, result and date.
- Failure modes:
  - Nobody can say which sources trained which model version, or on what terms.
  - A single unlicensed or unlawfully obtained source contaminates every model trained on it, and without per-source lineage the only safe response is to delete everything.
  - Crawled content is used although the rightholder reserved its rights by machine-readable means, because the reservation check was not run, was not recorded or is stale.
- Scope: Every source a provider trains or fine-tunes on: internal data, licensed corpora, open datasets, crawled web content and user data. The ledger records the organisation's position; it does not settle open legal questions.
- Enforcement points:
  - deploy: before a version is deployed or released
- Verification:
  - Test: The corpus build and the dataset admission gate fail on a source with no ledger row, on a source whose terms do not permit the declared use and on a source whose reservation check is missing or stale.
- Evidence:
  - Ledger row per source and version, with the reservation-check method, result and date for crawled content · Layer 02 Inventory & Transparency
  - Licence on the dataset card: name, whether training is allowed, whether text-and-data-mining reservations were checked · Layer 02 Inventory & Transparency · [dataset-card.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-card)
  - Ledger check (rows present) on the admission record · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
- Failure response: deny: block the action. A source with no row, with terms that do not permit the declared use or with a missing or stale reservation check fails the corpus build and admission.
- Layer: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Training-Data Rights Ledger](https://aigovernanceengineer.com/patterns/training-data-rights-ledger), [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [AIBOM](https://aigovernanceengineer.com/patterns/aibom)
- Mappings:
  - Obligations: [EU AI Act Art. 10 data and data governance](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art10); [EU AI Act Art. 53(1)(c) copyright policy honouring text-and-data-mining reservations](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art53-1c); [Directive (EU) 2019/790 on copyright in the Digital Single Market, DSM Directive Art. 4(3) text-and-data-mining reservations](https://aigovernanceengineer.com/obligations/aige-obl-dsm-art4-3); [General Data Protection Regulation (EU) 2016/679, GDPR Art. 5(1)(b) and 6(4) purpose limitation](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art5-1b)
  - ISO/IEC 42001: A.7.5 Data provenance
  - NIST AI RMF: GOVERN 6.1 Policies and procedures are in place that address AI risks associated with third-party entities, including risks of infringement of a third-party’s intellectual property or other rights.; MAP 4.1 Approaches for mapping AI technology and legal risks of its components – including the use of third-party data or software – are in place, followed, and documented, as are risks of infringement of a third party’s intellectual property or other rights.
- References:
  - [11] Training-Data Rights Ledger (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): one ledger row per training source, joined to lineage so each model knows its sources)
  - [12] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "The right to use the data")
  - [13] Directive (EU) 2019/790 on copyright in the Digital Single Market, Art. 4 (text and data mining exception; 4(3) reservation of rights by machine-readable means for content made publicly available online)
  - [14] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Art. 53 (GPAI provider obligations: (c) copyright policy including reservations of rights; (d) public summary of training content on the AI Office template)
  - [4] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Art. 10 (data and data governance: 10(2) practices, including origin, preparation, bias examination and mitigation, and data gaps; 10(3) relevant, sufficiently representative, free of errors and complete; 10(4) specific setting of use)
  - [15] In the Matter of Everalbum, Inc., Decision and Order ("Affected Work Product": models or algorithms developed using users' biometric information, to be deleted within 90 days with a sworn statement)
  - [7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title)
  - [8] Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (GOVERN 6.1 third-party risks incl. infringement of intellectual property or other rights; MAP 1.1 intended purposes documented; MAP 2.3 data collection and selection considerations identified and documented; MAP 3.3 targeted application scope; MAP 4.1 legal risks of components incl. third-party data; MEASURE 2.10 privacy risk examined and documented; MANAGE 1.4 negative residual risks to downstream acquirers and end users documented)
- Implementation notes:
  - For crawled content, record the crawler identity, the window and the reservation check: the method (for example robots.txt and page metadata read at fetch time), the result and the date. The crawl pipeline should write rows itself, since a row per source is real work for large crawls.
  - Put the licence in the AIBOM as well, so a licence change surfaces in the next build.
  - Generate the disclosures (the GPAI training-content summary under Art. 53(1)(d) and similar training-data documentation) as queries over the ledger, not as documents written from memory.
- Open questions:
  - How fresh must a reservation check be before the gate treats it as stale, and should it be re-run at every corpus build?
  - At what granularity should rows be kept when rights attach per record (opt-outs, per-record licences) rather than per source?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-003.json

## AIGE-CTL-DATA-004 Lawful Basis and Assessment per Processing Stage

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-004

- Id: `AIGE-CTL-DATA-004` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the dataset-card.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-card); [the impact-assessment.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-impact-assessment); [chapter 19, Privacy and data protection law applied to AI](https://aigovernanceengineer.com/bok/privacy-and-ai))
- Objective: Each dataset carries, for each processing stage (training, fine-tuning, retrieval, inference, log monitoring), its lawful basis, its purpose and a pointer to the assessment behind it, and processing likely to result in a high risk has a versioned DPIA or a recorded decision that none was needed.
- Failure modes:
  - One basis is picked for "the model", although training, indexing and answering a live customer are different activities that each need their own basis.
  - A training job runs on a dataset whose recorded basis does not cover training, or on data whose basis was never recorded.
  - A legitimate-interest assessment points at mitigations that have been switched off, and the registry does not show that it went stale.
  - No DPIA exists and no decision that one was not needed was recorded.
- Scope: Personal data in datasets used for training, fine-tuning and retrieval, and the stages that process it. Transfers, automated decision-making and the rights path are outside this control; the fundamental rights impact assessment, which complements the DPIA rather than repeating it, is left to the FRIA-as-Code pattern.
- Enforcement points:
  - runtime: at the point of action (gateway or guardrail)
  - periodic: on a schedule, over what is already running
- Verification:
  - Test: The training job reads the basis registry before it runs and refuses to run on a dataset whose basis does not cover training.
- Evidence:
  - Lawful basis for this purpose on the dataset card · Layer 02 Inventory & Transparency · [dataset-card.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-card)
  - Basis registry entry per dataset and stage, with the reference of its assessment (for example a versioned LIA) · Layer 02 Inventory & Transparency
  - Versioned AI DPIA addendum, or the recorded decision that no DPIA was needed · Layer 02 Inventory & Transparency · [impact-assessment.v1](https://aigovernanceengineer.com/resources/templates#schema-impact-assessment)
- Failure response: deny: block the action. The job refuses to run on a dataset whose basis does not cover the stage. Where the DPIA shows high residual risk, the controller consults the supervisory authority before the processing.
- Layer: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [Training-Data Rights Ledger](https://aigovernanceengineer.com/patterns/training-data-rights-ledger), [FRIA-as-Code](https://aigovernanceengineer.com/patterns/fria-as-code)
- Mappings:
  - Obligations: [General Data Protection Regulation (EU) 2016/679, GDPR Art. 6 lawful basis per processing moment](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art6); [General Data Protection Regulation (EU) 2016/679, GDPR Arts. 35–36 DPIA and prior consultation](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art35-36)
  - NIST AI RMF: MEASURE 2.10 Privacy risk of the AI system – as identified in the MAP function – is examined and documented.
- References:
  - [16] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Lawful basis for training versus inference")
  - [17] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "The DPIA for AI systems")
  - [12] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "The right to use the data")
  - [18] Regulation (EU) 2016/679 (GDPR) (Art. 5 principles, incl. 5(1)(b) purpose limitation and 5(1)(c) minimisation; Art. 6 lawful basis and 6(4) compatibility; Art. 7 consent; Art. 9 special categories; Arts. 15 to 17 and 21 rights; Art. 25 data protection by design and by default; Art. 30 records of processing; Arts. 35 and 36 DPIA and prior consultation)
  - [5] Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (anonymity of models; legitimate interest; consequences of unlawful processing in development)
  - [8] Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (GOVERN 6.1 third-party risks incl. infringement of intellectual property or other rights; MAP 1.1 intended purposes documented; MAP 2.3 data collection and selection considerations identified and documented; MAP 3.3 targeted application scope; MAP 4.1 legal risks of components incl. third-party data; MEASURE 2.10 privacy risk examined and documented; MANAGE 1.4 negative residual risks to downstream acquirers and end users documented)
- Implementation notes:
  - Keep the basis registry attached to the system's registry entry: each dataset and stage carries its basis, purpose and a pointer to the assessment behind it, and the training job reads it.
  - Make the legitimate-interest assessment a versioned artefact that points each mitigation at the control implementing it (the opt-out endpoint, the filter rule id, the scraping allow-list); switch a mitigation off and the LIA goes stale.
  - Generate most of the AI DPIA from the registry: processing moments from the registry, bases from the basis registry, with the DPO still writing and signing the risk judgement. The impact-assessment schema carries a dpia_addendum type for the AI-specific fields.
- Open questions:
  - Should the admission gate itself refuse a dataset whose stage has no DPIA and no recorded "no DPIA" decision, or is that a periodic check over the registry?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-004.json

## AIGE-CTL-DATA-005 Purpose Match Before Reuse of Data

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-005

- Id: `AIGE-CTL-DATA-005` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [the dataset-admission-record.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record); [chapter 19, Privacy and data protection law applied to AI](https://aigovernanceengineer.com/bok/privacy-and-ai))
- Objective: A run that reads a dataset is denied when the purpose on the dataset's card differs from the purpose declared by the consuming system and no compatibility assessment is recorded.
- Failure modes:
  - Data is reused for a purpose incompatible with the one it was collected for: support transcripts reused to profile customers for sales, security footage reused for attendance, fraud features reused for credit limits.
  - Consent to a service is treated as consent to train a model on the service's data.
  - A purpose mismatch is caught by nobody because the purpose travels in a document, not as a tag on the data a rule can read.
- Scope: Personal data further processed for training, fine-tuning or indexing by a system other than, or for a purpose other than, the one it was collected for.
- Enforcement points:
  - runtime: at the point of action (gateway or guardrail)
- Verification:
  - Test: A run whose declared purpose differs from the purpose on the dataset's card, with no compatibility assessment recorded, is denied, and the denial is filed against the dataset's registry entry.
- Evidence:
  - Purpose-match verdict per run · Layer 01 Govern-as-Code
  - Compatibility and use-case checks on the admission record · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
  - Art. 6(4) compatibility assessment filed against the dataset · Layer 02 Inventory & Transparency
- Failure response: deny: block the action. The run is denied; the request becomes an Art. 6(4) compatibility assessment, and the denied run and the assessment are both filed against the dataset's registry entry.
- Layers: [Layer 01 Govern-as-Code](https://aigovernanceengineer.com/bok/the-stack#layer-01-govern-as-code), [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [Training-Data Rights Ledger](https://aigovernanceengineer.com/patterns/training-data-rights-ledger), [Policy Card](https://aigovernanceengineer.com/patterns/policy-card)
- Mappings:
  - Obligations: [General Data Protection Regulation (EU) 2016/679, GDPR Art. 5(1)(b) and 6(4) purpose limitation](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art5-1b)
- References:
  - [19] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Purpose limitation and function creep")
  - [12] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "The right to use the data")
  - [1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use)
  - [18] Regulation (EU) 2016/679 (GDPR) (Art. 5 principles, incl. 5(1)(b) purpose limitation and 5(1)(c) minimisation; Art. 6 lawful basis and 6(4) compatibility; Art. 7 consent; Art. 9 special categories; Arts. 15 to 17 and 21 rights; Art. 25 data protection by design and by default; Art. 30 records of processing; Arts. 35 and 36 DPIA and prior consultation)
- Implementation notes:
  - Use a purpose tag that travels with the data and a layer 01 rule that compares it with the purpose declared by the consuming system; a denied join is proof the purpose limit bit.
  - The Art. 6(4) test weighs the link between purposes, the context, the nature of the data, the consequences and the safeguards, such as encryption or pseudonymisation; record the outcome, including a partial one (for example "aggregated topic counts only").
- Open questions:
  - How should purposes be named so that a rule can compare them: a controlled vocabulary per organisation, or the use-case ids of the registry?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-005.json

## AIGE-CTL-DATA-006 Personal Data Screening and Minimisation

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-006

- Id: `AIGE-CTL-DATA-006` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the dataset-card.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-card); [the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [chapter 19, Privacy and data protection law applied to AI](https://aigovernanceengineer.com/bok/privacy-and-ai))
- Objective: Personal and special-category data are screened on each snapshot before training, the filter log is kept with the snapshot, each input feature carries a reason and a measured contribution, and retention follows a rule enforced in code.
- Failure modes:
  - A snapshot enters training without a PII or special-category scan, or the scan ran and its log was not kept.
  - Features with no reason and no measured contribution stay in the data, so minimisation is asserted once instead of argued feature by feature.
  - Special-category fields are used with no documented condition.
  - Retention follows the storage default rather than the obligation, and data is kept past its deletion date.
- Scope: Snapshots and features admitted to training, fine-tuning, evaluation and retrieval pipelines, and their retention. Retrieval indexes and logs are covered for minimisation only; anonymity claims about trained models are outside this control.
- Enforcement points:
  - deploy: before a version is deployed or released
  - periodic: on a schedule, over what is already running
- Verification:
  - Inspect: Each admitted snapshot has the log of the PII and special-category scan run on it, and its card records a retention rule as enforced in code.
- Evidence:
  - PII and special-category filter log kept with each snapshot · Layer 01 Govern-as-Code
  - Personal-data flag and retention rule on the dataset card · Layer 02 Inventory & Transparency · [dataset-card.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-card)
  - Retention-set check on the admission record · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
- Failure response: deny: block the action. A snapshot with no scan log does not enter training; a feature with neither a reason nor a measured contribution is removed, and a special-category field needs a documented condition before it stays.
- Layers: [Layer 01 Govern-as-Code](https://aigovernanceengineer.com/bok/the-stack#layer-01-govern-as-code), [Layer 03 Evals & Red Teaming as Evidence](https://aigovernanceengineer.com/bok/the-stack#layer-03-evals--red-teaming-as-evidence)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate)
- Mappings:
  - Obligations: [General Data Protection Regulation (EU) 2016/679, GDPR Art. 5(1)(c) and 25 minimisation and data protection by design and by default](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art25)
- References:
  - [20] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Minimisation, privacy by design and PETs")
  - [21] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Obligation to artefact map")
  - [18] Regulation (EU) 2016/679 (GDPR) (Art. 5 principles, incl. 5(1)(b) purpose limitation and 5(1)(c) minimisation; Art. 6 lawful basis and 6(4) compatibility; Art. 7 consent; Art. 9 special categories; Arts. 15 to 17 and 21 rights; Art. 25 data protection by design and by default; Art. 30 records of processing; Arts. 35 and 36 DPIA and prior consultation)
  - [5] Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (anonymity of models; legitimate interest; consequences of unlawful processing in development)
- Implementation notes:
  - Argue minimisation feature by feature in the data card; the EDPB lists source selection, preparation and filtering among the areas an authority examines.
  - Hold only the fields answers need in retrieval indexes, and use synthetic or masked eval sets wherever a test does not depend on real identities.
  - Treat a synthetic set as a dataset with its own admission record naming the generator, the seed data and the privacy method: synthetic data inherits the biases, gaps and, where the generator memorised, the source records of its generator.
- Open questions:
  - What detection rate should a PII or special-category scan reach before its "pass" is accepted as evidence, and how is that rate measured?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-006.json

## AIGE-CTL-DATA-007 Special-Category Data Conditions

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-007

- Id: `AIGE-CTL-DATA-007` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the dataset-card.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-card); [the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [chapter 19, Privacy and data protection law applied to AI](https://aigovernanceengineer.com/bok/privacy-and-ai))
- Objective: Special-category data in a dataset is recorded with the Art. 9(2) condition relied on, and when it is processed for bias detection and correction under Art. 4a it is used only where other data would not do, pseudonymised, access-controlled, not transmitted onwards and deleted once the bias is corrected or its retention period ends, whichever comes first, with the records of processing stating why it was strictly necessary.
- Failure modes:
  - Special-category data is present in a dataset whose card says it is not, or with no condition recorded.
  - Data admitted for bias detection is kept after the bias was corrected, used for another purpose or passed on.
  - The records of processing do not say why special-category data was strictly necessary and why other data, including synthetic or anonymised data, would not do.
- Scope: Datasets that contain special-category personal data, including data processed only to detect and correct bias. Sensitive data a model infers at runtime is covered by the proxy test and inference policy of chapter 19, not here.
- Enforcement points:
  - deploy: before a version is deployed or released
  - periodic: on a schedule, over what is already running
- Verification:
  - Inspect: For each dataset processed under Art. 4a, the records of processing hold the strictly-necessary reason and a deletion log shows the data deleted once the bias was corrected or its retention period ended, whichever came first.
- Evidence:
  - Special-category block on the dataset card: presence, condition relied on, and whether it is processed only for bias detection · Layer 02 Inventory & Transparency · [dataset-card.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-card)
  - Records-of-processing entry with the Art. 4a reason · Layer 02 Inventory & Transparency
  - Deletion log of the pseudonymised bias set · Layer 01 Govern-as-Code
- Failure response: deny: block the action. A dataset with special-category data and no recorded condition is not admitted; an Art. 4a bias set without pseudonymisation or a deletion rule is not admitted for bias detection.
- Layers: [Layer 01 Govern-as-Code](https://aigovernanceengineer.com/bok/the-stack#layer-01-govern-as-code), [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [Fairness Eval Suite](https://aigovernanceengineer.com/patterns/fairness-eval-suite)
- Mappings:
  - Obligations: [General Data Protection Regulation (EU) 2016/679, GDPR Art. 9 special categories, incl. inferred sensitive data](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art9); [EU AI Act Art. 4a lawful basis for special-category data in bias detection](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art4a); [General Data Protection Regulation (EU) 2016/679, GDPR Art. 30 records of processing activities](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art30)
- References:
  - [22] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Special categories, inferred data and biometrics")
  - [23] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Records of processing")
  - [12] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "The right to use the data")
  - [1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use)
  - [24] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Arts. 4a and 10 (as amended by Reg. (EU) 2026/1744: Art. 10(5) deleted; Art. 4a inserted for special-category data in bias detection and correction)
  - [18] Regulation (EU) 2016/679 (GDPR) (Art. 5 principles, incl. 5(1)(b) purpose limitation and 5(1)(c) minimisation; Art. 6 lawful basis and 6(4) compatibility; Art. 7 consent; Art. 9 special categories; Arts. 15 to 17 and 21 rights; Art. 25 data protection by design and by default; Art. 30 records of processing; Arts. 35 and 36 DPIA and prior consultation)
- Implementation notes:
  - After the Digital Omnibus (Regulation (EU) 2026/1744, in force 27 July 2026), the narrow basis to process special-category data for bias detection sits in Art. 4a; Art. 10(5) was deleted, so a card that still names Art. 10(5) points at a basis that no longer exists.
  - Generate the records-of-processing entries from the registry, the data cards and the basis registry, so the Art. 4a reason does not go stale with the next pipeline change.
- Open questions:
  - What evidence shows that other data, including synthetic or anonymised data, would not have done for a given bias examination?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-007.json

## AIGE-CTL-DATA-008 Fitness-for-Purpose Checks Before Admission

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-008

- Id: `AIGE-CTL-DATA-008` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [the dataset-admission-record.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record); [chapter 14, Governing AI development](https://aigovernanceengineer.com/bok/governing-development))
- Objective: Before a dataset version is admitted, its quality (label accuracy, completeness, consistency, timeliness), its quantity per class and per group, its representativeness against the deployment population and a proxy and bias examination are checked and recorded, each passing or waived in writing by someone entitled to waive it.
- Failure modes:
  - A dataset large enough but drawn from the wrong population is admitted: quantity is taken for representativeness.
  - A group falls below the minimum cell size in the test plan and nobody records it.
  - No bias examination is recorded, or a failed check is waived with no signer, no condition and no expiry.
  - The data measures a proxy rather than what the use case needs, and the assumption is never written down.
- Scope: Training, validation and testing datasets for high-risk systems, where Art. 10 applies, and any dataset admitted under the organisation's own policy. Fairness testing of the trained model is covered by the Fairness Eval Suite pattern, not here.
- Enforcement points:
  - deploy: before a version is deployed or released
- Verification:
  - Inspect: The admission record holds a result for each quality, quantity, representativeness and bias check, with the obligation it enforces; every waived check names a waiver signed by the data owner and every condition is listed.
- Evidence:
  - Quality, representativeness and bias check results on the admission record, with waivers and conditions · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
  - Quality checks, populations covered and known gaps on the dataset card · Layer 02 Inventory & Transparency · [dataset-card.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-card)
- Failure response: require_approval: hold the action for a human decision. A failed check blocks admission unless someone entitled to waive it signs a waiver; the waiver and its conditions appear in the model card, and the next retrain cannot start until the conditions are closed.
- Layers: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency), [Layer 03 Evals & Red Teaming as Evidence](https://aigovernanceengineer.com/bok/the-stack#layer-03-evals--red-teaming-as-evidence)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [Fairness Eval Suite](https://aigovernanceengineer.com/patterns/fairness-eval-suite)
- Mappings:
  - Obligations: [EU AI Act Art. 10 data and data governance](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art10); [ISO/IEC 42001, A.7 Data for AI systems](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a7)
  - ISO/IEC 42001: A.7.4 Quality of data for AI systems
  - NIST AI RMF: MAP 2.3 Scientific integrity and TEVV considerations are identified and documented, including those related to experimental design, data collection and selection (e.g., availability, representativeness, suitability), system trustworthiness, and construct validation.
- References:
  - [25] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Quality, quantity, representativeness and fitness for purpose")
  - [1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use)
  - [4] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Art. 10 (data and data governance: 10(2) practices, including origin, preparation, bias examination and mitigation, and data gaps; 10(3) relevant, sufficiently representative, free of errors and complete; 10(4) specific setting of use)
  - [26] ISO/IEC 5259 series, Data quality for analytics and machine learning (ML) (Part 1 overview, terminology and examples; Part 2 data quality measures; Part 3 data quality management requirements and guidelines; Part 4 data quality process framework; Part 5 data quality governance framework)
  - [7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title)
  - [8] Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (GOVERN 6.1 third-party risks incl. infringement of intellectual property or other rights; MAP 1.1 intended purposes documented; MAP 2.3 data collection and selection considerations identified and documented; MAP 3.3 targeted application scope; MAP 4.1 legal risks of components incl. third-party data; MEASURE 2.10 privacy risk examined and documented; MANAGE 1.4 negative residual risks to downstream acquirers and end users documented)
- Implementation notes:
  - Use the vocabulary of the ISO/IEC 5259 series for the quality measures, and evidence each dimension with the test chapter 14 names: an audit of a labelled sample and inter-annotator agreement for labels, null rates per field and segment for completeness, cell counts against the minimum in the test plan for quantity, a distribution comparison against a reference for representativeness.
  - Keep an assumption register for what the data is meant to measure (Art. 10(2)(d)) and a proxy analysis for fitness for purpose.
- Open questions:
  - Who sets the thresholds each check is held to, and how are they reviewed when the deployment population changes?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-008.json

## AIGE-CTL-DATA-009 Signed Snapshot Integrity

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-009

- Id: `AIGE-CTL-DATA-009` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [the dataset-admission-record.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record); [chapter 14, Governing AI development](https://aigovernanceengineer.com/bok/governing-development))
- Objective: Only a content-addressed, signed snapshot is admitted, its hash is re-verified when a job reads it, and new or appended data is checked for anomalies before it is admitted.
- Failure modes:
  - A model trains on a snapshot someone altered after admission.
  - Poisoned data enters through an appended batch that was never checked, so the model still looks functional while carrying a bias, a weakness or a backdoor.
  - The admission record names a hash that no job ever compares with the data it reads.
- Scope: Snapshots admitted to training, fine-tuning, validation, testing, evaluation and retrieval-index pipelines, and data appended to them. The integrity of the trained model artefact is covered by the Model Artefact Integrity pattern.
- Enforcement points:
  - deploy: before a version is deployed or released
  - runtime: at the point of action (gateway or guardrail)
- Verification:
  - Test: When a job reads an admitted snapshot, the hash of what it reads is compared with the content hash on the admission record, and a mismatch stops the read.
- Evidence:
  - Content hash of the admitted snapshot and the signature over the admission record · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
  - Read-time hash verification and anomaly-check results · Layer 01 Govern-as-Code
- Failure response: deny: block the action. A snapshot whose hash does not match its admission record is not read; new or appended data that fails the anomaly checks is not admitted.
- Layers: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency), [Layer 01 Govern-as-Code](https://aigovernanceengineer.com/bok/the-stack#layer-01-govern-as-code)
- Patterns: [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [AIBOM](https://aigovernanceengineer.com/patterns/aibom)
- Mappings:
  - Obligations: [EU AI Act Art. 10 data and data governance](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art10); [OWASP Top 10 for LLM Applications 2026](https://aigovernanceengineer.com/obligations/aige-obl-owasp-llm)
  - ISO/IEC 42001: A.7.5 Data provenance
  - OWASP: [LLM05:2026 Data and Model Poisoning](https://aigovernanceengineer.com/resources/threats#threat-llm05-2026)
  - MITRE ATLAS: AML.M0007 (Sanitize Training Data (mitigation))
  - MITRE ATLAS: AML.M0025 (Maintain AI Dataset Provenance (mitigation))
- References:
  - [1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use)
  - [25] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Quality, quantity, representativeness and fitness for purpose")
  - [6] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
  - [27] MITRE ATLAS data, release 2026.09 (modified 2026-09-15; AML.T0020 Training Data Poisoning; mitigations AML.M0007 Sanitize Training Data and AML.M0025 Maintain AI Dataset Provenance)
  - [7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title)
- Implementation notes:
  - Sign the admission record (for example "ed25519:<base64>") so it is tamper-evident in the evidence store, and record the digest of the snapshot that was checked in its input_hash.
  - Training data is an attack surface: ATLAS catalogues training data poisoning (AML.T0020) and lists "Sanitize Training Data" (AML.M0007) and "Maintain AI Dataset Provenance" (AML.M0025) among its mitigations.
- Open questions:
  - Which anomaly checks on appended data are strong enough to catch planted triggers, and which belong instead in the regression evals of the trained model?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-009.json

## AIGE-CTL-DATA-010 Lineage from Training Runs to Admitted Sources

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-010

- Id: `AIGE-CTL-DATA-010` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Training-Data Rights Ledger pattern](https://aigovernanceengineer.com/patterns/training-data-rights-ledger); [the Dataset Admission Gate pattern](https://aigovernanceengineer.com/patterns/dataset-admission-gate); [the model-card.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-model-card); [chapter 14, Governing AI development](https://aigovernanceengineer.com/bok/governing-development))
- Objective: Each training run records the admission records and ledger rows (id and version) it read and the hashes of the admitted snapshots, so backward lineage answers "what trained this model?" and forward lineage answers "which models used this source?".
- Failure modes:
  - A licence withdrawal, an erasure request or an order names a source, and nobody can list the models trained on it.
  - The model card lists datasets by name but not by version or snapshot hash, so the training cannot be reproduced.
  - Lineage is kept at dataset level only where rights attach to records, so an opt-out cannot be traced to the runs it affects.
- Scope: Training and fine-tuning runs and the datasets, snapshots and ledger rows they read. Evaluation runs are covered for the data they read, not for their results.
- Enforcement points:
  - deploy: before a version is deployed or released
- Verification:
  - Inspect: For a released model version, the run record names the admission records and ledger rows it read, and a forward-lineage query from one of those sources returns the model version.
- Evidence:
  - Datasets used to train, validate, test or fine-tune the model on the model card, each named by its dataset card id or linked to its data card, with its role · Layer 02 Inventory & Transparency · [model-card.v1](https://aigovernanceengineer.com/resources/templates#schema-model-card)
  - Link to the lineage record (for example an OpenLineage or W3C PROV graph) on the dataset card · Layer 02 Inventory & Transparency · [dataset-card.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-card)
  - Training run record with the code commit, the hashes of the admitted snapshots and the admission records and ledger rows read · Layer 02 Inventory & Transparency
- Failure response: alert: let the action through and raise an alert. A training run that does not record the admission records and ledger rows it read has no backward lineage: until it is restored, a withdrawal or an order cannot be traced to that model version. Which response fits the gap is left to technical review.
- Layer: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Training-Data Rights Ledger](https://aigovernanceengineer.com/patterns/training-data-rights-ledger), [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [AIBOM](https://aigovernanceengineer.com/patterns/aibom)
- Mappings:
  - Obligations: [EU AI Act Art. 10 data and data governance](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art10); [ISO/IEC 42001, A.7 Data for AI systems](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a7)
  - ISO/IEC 42001: A.7.5 Data provenance
- References:
  - [28] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Provenance versus lineage")
  - [29] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Reproducibility and linked versioning")
  - [11] Training-Data Rights Ledger (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): one ledger row per training source, joined to lineage so each model knows its sources)
  - [1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use)
  - [30] PROV Overview (PROV-DM and PROV-O W3C Recommendations of 30 April 2013; provenance as information about entities, activities and people involved in producing data)
  - [31] OpenLineage: an open platform for collection and analysis of data lineage (standard API for lineage events over datasets, jobs and runs, with facets)
  - [7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title)
- Implementation notes:
  - Record provenance in W3C PROV terms (entities, activities and agents) and emit lineage events over datasets, jobs and runs, so each training run names the admission records it read.
  - Choose granularity by where rights attach: dataset-level provenance by default, record-level where rights attach to records (personal data, per-source licences, opt-outs), feature-level lineage for sensitive derived features that can act as proxies.
  - List the datasets by version in the AIBOM as well.
- Open questions:
  - The site publishes no schema for a model training run (the published training-record schema covers AI literacy training): should one be published, or should the model card and AIBOM carry the run fields?
  - Should a training run with no recorded lineage block the release of the model version it produced, or only raise an alert to its owner?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-010.json

## AIGE-CTL-DATA-011 Rights Changes Propagated to Affected Models

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-011

- Id: `AIGE-CTL-DATA-011` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Training-Data Rights Ledger pattern](https://aigovernanceengineer.com/patterns/training-data-rights-ledger); [the Rights Requests Against Models pattern](https://aigovernanceengineer.com/patterns/rights-requests-against-models); [the dataset-admission-record.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record); [chapter 19, Privacy and data protection law applied to AI](https://aigovernanceengineer.com/bok/privacy-and-ai))
- Objective: A licence expiry or withdrawal, a new rights reservation, a consent withdrawal, an erasure request or an order marks the affected ledger rows and snapshots, forward lineage lists the affected models, and the remediation (retrain without the source, retire the model, or a documented decision to rely on another basis) is recorded against the same rows with a date and an approver.
- Failure modes:
  - An erasure request closes on time at the source system but the training snapshot, retrieval index, logs and models trained on the data are never reached.
  - A consent withdrawal cannot be traced to the runs and model versions that inherited the consent.
  - Remedies reach the model itself (an order to delete models developed using unlawfully used data) and, without per-source lineage, the only safe response is to delete everything.
  - A change in rights does not reopen admission, and the next retrain reads the source again.
- Scope: Changes in the right to use a training source or a person's data after admission, and the datasets, indexes and model versions they reach. The per-location response to a data-subject request is the Rights Requests Against Models pattern; this control covers the propagation from the data to the models.
- Enforcement points:
  - periodic: on a schedule, over what is already running
- Verification:
  - Test: Run a mock erasure request through the corpus, the snapshots, the retrieval index, the logs and the weights, write the fulfilment record, and time it against the one-month deadline.
- Evidence:
  - Re-admission record for the affected dataset versions · Layer 01 Govern-as-Code · [dataset-admission-record.v1](https://aigovernanceengineer.com/resources/templates#schema-dataset-admission-record)
  - Remediation recorded against the affected ledger rows, with the models forward lineage listed, a date and an approver · Layer 02 Inventory & Transparency
  - Fulfilment record: every location, the action in each, the model versions affected and when the gap closes · Layer 05 Assurance & Continuous Compliance
- Failure response: alert: let the action through and raise an alert. The change marks the affected rows and reopens admission; the owners of every affected model are told, and each model is retrained without the source, retired or kept on a documented decision to rely on another basis.
- Layer: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Training-Data Rights Ledger](https://aigovernanceengineer.com/patterns/training-data-rights-ledger), [Dataset Admission Gate](https://aigovernanceengineer.com/patterns/dataset-admission-gate), [Rights Requests Against Models](https://aigovernanceengineer.com/patterns/rights-requests-against-models)
- Mappings:
  - Obligations: [General Data Protection Regulation (EU) 2016/679, GDPR Arts. 15–17 and 21 data subject rights against trained models](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art15-17-21); [General Data Protection Regulation (EU) 2016/679, GDPR Art. 7 conditions for consent and its withdrawal](https://aigovernanceengineer.com/obligations/aige-obl-gdpr-art7); [Directive (EU) 2019/790 on copyright in the Digital Single Market, DSM Directive Art. 4(3) text-and-data-mining reservations](https://aigovernanceengineer.com/obligations/aige-obl-dsm-art4-3)
- References:
  - [11] Training-Data Rights Ledger (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): one ledger row per training source, joined to lineage so each model knows its sources)
  - [32] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Where a request has to reach")
  - [33] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "The limits of consent")
  - [34] Rights Requests Against Models (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): each data-subject request routed to every place the data sits and closed with a fulfilment record)
  - [28] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Provenance versus lineage")
  - [18] Regulation (EU) 2016/679 (GDPR) (Art. 5 principles, incl. 5(1)(b) purpose limitation and 5(1)(c) minimisation; Art. 6 lawful basis and 6(4) compatibility; Art. 7 consent; Art. 9 special categories; Arts. 15 to 17 and 21 rights; Art. 25 data protection by design and by default; Art. 30 records of processing; Arts. 35 and 36 DPIA and prior consultation)
  - [15] In the Matter of Everalbum, Inc., Decision and Order ("Affected Work Product": models or algorithms developed using users' biometric information, to be deleted within 90 days with a sworn statement)
- Implementation notes:
  - Keep a consent-purpose log joining each consent to the datasets and model versions that inherited it; without that join a withdrawal cannot be traced to the runs it affects.
  - For data inside the weights, choose on the ladder chapter 19 sets out (output suppression, retraining without the data, machine unlearning) and record the choice and its reason per request, with the date the next retrain closes the gap.
- Open questions:
  - How long may a model stay in production on output suppression before retraining without the data is due, and who decides?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-011.json

## AIGE-CTL-DATA-012 Registered Downstream Consumers of Outputs

Anchor: https://aigovernanceengineer.com/controls/data-admission-and-privacy#aige-ctl-data-012

- Id: `AIGE-CTL-DATA-012` · v0.1 · Draft · Open for technical review
- Depth: Derived from site material ([the Downstream Use Register pattern](https://aigovernanceengineer.com/patterns/downstream-use-register); [the policy-card.v1 record schema](https://aigovernanceengineer.com/resources/templates#schema-policy-card); [chapter 14, Governing AI development](https://aigovernanceengineer.com/bok/governing-development))
- Objective: Every consumer of a system's outputs (a system, a team, a partner or a training pipeline) is registered against the producing system with its purpose, its approval and the re-test that cleared the outputs for that context; access to the outputs is granted per registered consumer, and the intended and prohibited uses are rules on a Policy Card.
- Failure modes:
  - A risk score approved to prioritise manual review becomes an automatic decline in another team's pipeline, and nobody assessed that use.
  - A model's outputs are harvested as training data for another model, and the feedback loop is invisible.
  - A partner receives outputs under a contract nobody connected to the registry, and is not told when the model changes or retires.
- Scope: Consumers of the outputs of AI systems, internal and external, including training pipelines that read those outputs. Registration binds internal consumers; external ones depend on contract terms and audit rights.
- Enforcement points:
  - deploy: before a version is deployed or released
  - runtime: at the point of action (gateway or guardrail)
- Verification:
  - Test: A consumer with no registration has no credential to the output API or table, and a registration whose declared use meets a prohibited-use rule on the card goes to review as a new purpose instead of receiving a credential.
- Evidence:
  - Intended and prohibited uses as rules on the system's Policy Card · Layer 01 Govern-as-Code · [policy-card.v1](https://aigovernanceengineer.com/resources/templates#schema-policy-card)
  - Downstream use register entry: each consumer with its use, approval, re-test, credential or contract, and the feedback-loop check · Layer 02 Inventory & Transparency
- Failure response: deny: block the action. An unregistered consumer gets no credential; a declared use outside the card fails registration and reopens classification and the impact assessments as a new purpose. A model change, an incident or a retirement notifies every registered consumer.
- Layers: [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency), [Layer 01 Govern-as-Code](https://aigovernanceengineer.com/bok/the-stack#layer-01-govern-as-code)
- Patterns: [Downstream Use Register](https://aigovernanceengineer.com/patterns/downstream-use-register), [Policy Card](https://aigovernanceengineer.com/patterns/policy-card), [Agent Registry](https://aigovernanceengineer.com/patterns/agent-registry), [Disclosure & Notification Pipeline](https://aigovernanceengineer.com/patterns/disclosure-notification-pipeline)
- Mappings:
  - Obligations: [EU AI Act Art. 9 risk management system](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art9); [EU AI Act Art. 25 responsibilities along the AI value chain](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art25); [EU AI Act Art. 50 transparency for certain AI systems](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art50); [ISO/IEC 42001, A.8 Information for interested parties](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a8); [ISO/IEC 42001, A.9 Use of AI systems](https://aigovernanceengineer.com/obligations/aige-obl-iso42001-a9)
  - ISO/IEC 42001: A.8.2 System documentation and information for users; A.9.4 Intended use of the AI system
  - NIST AI RMF: MAP 1.1 Intended purposes, potentially beneficial uses, context-specific laws, norms and expectations, and prospective settings in which the AI system will be deployed are understood and documented.; MAP 3.3 Targeted application scope is specified and documented based on the system’s capability, established context, and AI system categorization.; MANAGE 1.4 Negative residual risks (defined as the sum of all unmitigated risks) to both downstream acquirers of AI systems and end users are documented.
  - OWASP: [LLM10:2026 Improper Output Handling](https://aigovernanceengineer.com/resources/threats#threat-llm10-2026); [ASI08 Cascading Failures](https://aigovernanceengineer.com/resources/threats#threat-asi08)
- References:
  - [35] Downstream Use Register (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): intended and prohibited uses as a Policy Card and every consumer of the outputs recorded against the registry entry)
  - [36] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Function creep")
  - [37] Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), of 13 June 2024; OJ L, 2024/1689, 12.7.2024 (Art. 3(13) reasonably foreseeable misuse; Art. 9(2)(b) risks under reasonably foreseeable misuse; Art. 25(1)(c) changed intended purpose; Art. 50(2) machine-readable marking of synthetic outputs)
  - [6] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
  - [38] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  - [7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title)
  - [8] Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (GOVERN 6.1 third-party risks incl. infringement of intellectual property or other rights; MAP 1.1 intended purposes documented; MAP 2.3 data collection and selection considerations identified and documented; MAP 3.3 targeted application scope; MAP 4.1 legal risks of components incl. third-party data; MEASURE 2.10 privacy risk examined and documented; MANAGE 1.4 negative residual risks to downstream acquirers and end users documented)
- Implementation notes:
  - Forecast misuse before go-live with a premortem, abuse cases written next to the user stories and a stakeholder impact map that includes people who never touch the interface; each plausible misuse becomes a prohibited-use rule or a monitor.
  - Stamp outputs with the producing system and version, the intended use and a caveat, as metadata a consumer can read; for generative content, this is the machine-readable marking Art. 50(2) requires of providers.
  - Classify consumer requests against the negative space of the card, alert on what falls outside it, and watch for outputs that return as training data.
- Open questions:
  - How is a registered external consumer held to its declared use when the outputs leave the organisation's access controls?
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-012.json

## Mappings

Mappings are illustrative, not a claim of conformity.

| Control | Obligations | ISO/IEC 42001 | NIST AI RMF | OWASP | AIUC-1 | Layer |
| --- | --- | --- | --- | --- | --- | --- |
| `AIGE-CTL-DATA-001` Dataset Admission Gate at Read Time | AIGE-OBL-EUAIA-ART10, AIGE-OBL-GDPR-ART5-1B, AIGE-OBL-ISO42001-A7 | A.7.2, A.7.4, A.7.5 | MAP 2.3, MAP 4.1 | LLM05:2026 |   | L1, L2 |
| `AIGE-CTL-DATA-002` Dataset Card for Every Admitted Version | AIGE-OBL-EUAIA-ART10, AIGE-OBL-ISO42001-A7 | A.7.2, A.7.5 | MAP 2.3 |   |   | L2 |
| `AIGE-CTL-DATA-003` Training-Data Rights Ledger Row per Source | AIGE-OBL-EUAIA-ART10, AIGE-OBL-EUAIA-ART53-1C, AIGE-OBL-DSM-ART4-3, AIGE-OBL-GDPR-ART5-1B | A.7.5 | GOVERN 6.1, MAP 4.1 |   |   | L2 |
| `AIGE-CTL-DATA-004` Lawful Basis and Assessment per Processing Stage | AIGE-OBL-GDPR-ART6, AIGE-OBL-GDPR-ART35-36 |   | MEASURE 2.10 |   |   | L2 |
| `AIGE-CTL-DATA-005` Purpose Match Before Reuse of Data | AIGE-OBL-GDPR-ART5-1B |   |   |   |   | L1, L2 |
| `AIGE-CTL-DATA-006` Personal Data Screening and Minimisation | AIGE-OBL-GDPR-ART25 |   |   |   |   | L1, L3 |
| `AIGE-CTL-DATA-007` Special-Category Data Conditions | AIGE-OBL-GDPR-ART9, AIGE-OBL-EUAIA-ART4A, AIGE-OBL-GDPR-ART30 |   |   |   |   | L1, L2 |
| `AIGE-CTL-DATA-008` Fitness-for-Purpose Checks Before Admission | AIGE-OBL-EUAIA-ART10, AIGE-OBL-ISO42001-A7 | A.7.4 | MAP 2.3 |   |   | L2, L3 |
| `AIGE-CTL-DATA-009` Signed Snapshot Integrity | AIGE-OBL-EUAIA-ART10, AIGE-OBL-OWASP-LLM | A.7.5 |   | LLM05:2026 |   | L2, L1 |
| `AIGE-CTL-DATA-010` Lineage from Training Runs to Admitted Sources | AIGE-OBL-EUAIA-ART10, AIGE-OBL-ISO42001-A7 | A.7.5 |   |   |   | L2 |
| `AIGE-CTL-DATA-011` Rights Changes Propagated to Affected Models | AIGE-OBL-GDPR-ART15-17-21, AIGE-OBL-GDPR-ART7, AIGE-OBL-DSM-ART4-3 |   |   |   |   | L2 |
| `AIGE-CTL-DATA-012` Registered Downstream Consumers of Outputs | AIGE-OBL-EUAIA-ART9, AIGE-OBL-EUAIA-ART25, AIGE-OBL-EUAIA-ART50, AIGE-OBL-ISO42001-A8, AIGE-OBL-ISO42001-A9 | A.8.2, A.9.4 | MAP 1.1, MAP 3.3, MANAGE 1.4 | LLM10:2026, ASI08 |   | L2, L1 |

## Open questions

- What lighter admission should a sandbox pipeline for exploratory work carry, and how is data kept from leaving the sandbox into a training job? (`AIGE-CTL-DATA-001`)
- Which enforcement point fits a gate that decides at read time inside a data pipeline: the platform's access layer, the job scheduler or the storage policy engine? (`AIGE-CTL-DATA-001`)
- Which optional card fields (composition, representativeness, quality checks, splits) should become mandatory for data admitted to a high-risk system? (`AIGE-CTL-DATA-002`)
- How fresh must a reservation check be before the gate treats it as stale, and should it be re-run at every corpus build? (`AIGE-CTL-DATA-003`)
- At what granularity should rows be kept when rights attach per record (opt-outs, per-record licences) rather than per source? (`AIGE-CTL-DATA-003`)
- Should the admission gate itself refuse a dataset whose stage has no DPIA and no recorded "no DPIA" decision, or is that a periodic check over the registry? (`AIGE-CTL-DATA-004`)
- How should purposes be named so that a rule can compare them: a controlled vocabulary per organisation, or the use-case ids of the registry? (`AIGE-CTL-DATA-005`)
- What detection rate should a PII or special-category scan reach before its "pass" is accepted as evidence, and how is that rate measured? (`AIGE-CTL-DATA-006`)
- What evidence shows that other data, including synthetic or anonymised data, would not have done for a given bias examination? (`AIGE-CTL-DATA-007`)
- Who sets the thresholds each check is held to, and how are they reviewed when the deployment population changes? (`AIGE-CTL-DATA-008`)
- Which anomaly checks on appended data are strong enough to catch planted triggers, and which belong instead in the regression evals of the trained model? (`AIGE-CTL-DATA-009`)
- The site publishes no schema for a model training run (the published training-record schema covers AI literacy training): should one be published, or should the model card and AIBOM carry the run fields? (`AIGE-CTL-DATA-010`)
- Should a training run with no recorded lineage block the release of the model version it produced, or only raise an alert to its owner? (`AIGE-CTL-DATA-010`)
- How long may a model stay in production on output suppression before retraining without the data is due, and who decides? (`AIGE-CTL-DATA-011`)
- How is a registered external consumer held to its declared use when the outputs leave the organisation's access controls? (`AIGE-CTL-DATA-012`)

## Changelog

- v0.1 (2026-09-26): First draft: 12 controls derived from the Dataset Admission Gate, Training-Data Rights Ledger and Downstream Use Register patterns, the dataset admission record and dataset card schemas, and chapters 14 and 19; open for technical review.

## Sources

[1] Dataset Admission Gate (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): a job may read a dataset version only if a complete, signed admission record admits it for that use). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/patterns/dataset-admission-gate (verified: primary)
[2] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Data for training and testing"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#data-for-training-and-testing (verified: primary)
[3] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Owners, stewards and the admission gate"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#owners-stewards-and-the-admission-gate (verified: primary)
[4] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Art. 10 (data and data governance: 10(2) practices, including origin, preparation, bias examination and mitigation, and data gaps; 10(3) relevant, sufficiently representative, free of errors and complete; 10(4) specific setting of use). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_10 (verified: primary)
[5] Opinion 28/2024 on certain data protection aspects related to the processing of personal data in the context of AI models (anonymity of models; legitimate interest; consequences of unlawful processing in development). European Data Protection Board. 2024-12. https://www.edpb.europa.eu/documents/opinion-of-the-board-art-64/opinion-282024-on-certain-data-protection-aspects-related-to_en (verified: primary)
[6] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026). OWASP GenAI Security Project. 2026-08-03. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ (verified: primary)
[7] ISO/IEC 42001:2023, AI management systems, Annex A (reference control objectives and controls A.2 to A.10, cited by id and short title). ISO/IEC. 2023. https://www.iso.org/standard/81230.html (verified: secondary)
[8] Artificial Intelligence Risk Management Framework (AI RMF 1.0), NIST AI 100-1 (GOVERN 6.1 third-party risks incl. infringement of intellectual property or other rights; MAP 1.1 intended purposes documented; MAP 2.3 data collection and selection considerations identified and documented; MAP 3.3 targeted application scope; MAP 4.1 legal risks of components incl. third-party data; MEASURE 2.10 privacy risk examined and documented; MANAGE 1.4 negative residual risks to downstream acquirers and end users documented). NIST. 2023-01-26. https://nvlpubs.nist.gov/nistpubs/ai/NIST.AI.100-1.pdf (verified: primary)
[9] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Model cards, system cards and datasheets"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#model-cards-system-cards-and-datasheets (verified: primary)
[10] Datasheets for Datasets (Gebru et al.; arXiv 1803.09010) (motivation, composition, collection, preprocessing, uses, distribution and maintenance). arXiv. 2018-03-23. https://arxiv.org/abs/1803.09010 (verified: primary)
[11] Training-Data Rights Ledger (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): one ledger row per training source, joined to lineage so each model knows its sources). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/patterns/training-data-rights-ledger (verified: primary)
[12] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "The right to use the data"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#the-right-to-use-the-data (verified: primary)
[13] Directive (EU) 2019/790 on copyright in the Digital Single Market, Art. 4 (text and data mining exception; 4(3) reservation of rights by machine-readable means for content made publicly available online). Publications Office of the EU (EUR-Lex). 2019-05-17. https://eur-lex.europa.eu/eli/dir/2019/790/oj/eng (verified: primary)
[14] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Art. 53 (GPAI provider obligations: (c) copyright policy including reservations of rights; (d) public summary of training content on the AI Office template). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_53 (verified: primary)
[15] In the Matter of Everalbum, Inc., Decision and Order ("Affected Work Product": models or algorithms developed using users' biometric information, to be deleted within 90 days with a sworn statement). Federal Trade Commission. 2021-05-07. https://www.ftc.gov/system/files/documents/cases/1923172_-_everalbum_decision_final.pdf (verified: primary)
[16] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Lawful basis for training versus inference"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#lawful-basis-for-training-versus-inference (verified: primary)
[17] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "The DPIA for AI systems"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#the-dpia-for-ai-systems (verified: primary)
[18] Regulation (EU) 2016/679 (GDPR) (Art. 5 principles, incl. 5(1)(b) purpose limitation and 5(1)(c) minimisation; Art. 6 lawful basis and 6(4) compatibility; Art. 7 consent; Art. 9 special categories; Arts. 15 to 17 and 21 rights; Art. 25 data protection by design and by default; Art. 30 records of processing; Arts. 35 and 36 DPIA and prior consultation). Publications Office of the EU (EUR-Lex). 2016-04-27. https://eur-lex.europa.eu/eli/reg/2016/679/oj/eng (verified: primary)
[19] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Purpose limitation and function creep"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#purpose-limitation-and-function-creep (verified: primary)
[20] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Minimisation, privacy by design and PETs"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#minimisation-privacy-by-design-and-pets (verified: primary)
[21] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Obligation to artefact map"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#obligation-to-artefact-map (verified: primary)
[22] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Special categories, inferred data and biometrics"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#special-categories-inferred-data-and-biometrics (verified: primary)
[23] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Records of processing"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#records-of-processing (verified: primary)
[24] Regulation (EU) 2024/1689 (AI Act), consolidated text of 2026-07-27, Arts. 4a and 10 (as amended by Reg. (EU) 2026/1744: Art. 10(5) deleted; Art. 4a inserted for special-category data in bias detection and correction). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_4a (verified: primary)
[25] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Quality, quantity, representativeness and fitness for purpose"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#quality-quantity-representativeness-and-fitness-for-purpose (verified: primary)
[26] ISO/IEC 5259 series, Data quality for analytics and machine learning (ML) (Part 1 overview, terminology and examples; Part 2 data quality measures; Part 3 data quality management requirements and guidelines; Part 4 data quality process framework; Part 5 data quality governance framework). ISO/IEC. 2024-2025. https://www.iso.org/standard/81088.html (verified: primary)
[27] MITRE ATLAS data, release 2026.09 (modified 2026-09-15; AML.T0020 Training Data Poisoning; mitigations AML.M0007 Sanitize Training Data and AML.M0025 Maintain AI Dataset Provenance). MITRE (atlas-data repository). 2026-09-15. https://github.com/mitre-atlas/atlas-data (verified: primary)
[28] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Provenance versus lineage"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#provenance-versus-lineage (verified: primary)
[29] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Reproducibility and linked versioning"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#reproducibility-and-linked-versioning (verified: primary)
[30] PROV Overview (PROV-DM and PROV-O W3C Recommendations of 30 April 2013; provenance as information about entities, activities and people involved in producing data). W3C. 2013-04-30. https://www.w3.org/TR/prov-overview/ (verified: primary)
[31] OpenLineage: an open platform for collection and analysis of data lineage (standard API for lineage events over datasets, jobs and runs, with facets). OpenLineage project (The Linux Foundation). 2026. https://openlineage.io/ (verified: primary)
[32] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "Where a request has to reach"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#where-a-request-has-to-reach (verified: primary)
[33] Privacy and data protection law applied to AI (AI Governance Engineering Body of Knowledge v0.5.0, chapter 19, section "The limits of consent"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/privacy-and-ai#the-limits-of-consent (verified: primary)
[34] Rights Requests Against Models (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): each data-subject request routed to every place the data sits and closed with a fulfilment record). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/patterns/rights-requests-against-models (verified: primary)
[35] Downstream Use Register (AI Governance Engineering Body of Knowledge v0.5.0, pattern catalogue (chapter 05): intended and prohibited uses as a Policy Card and every consumer of the outputs recorded against the registry entry). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/patterns/downstream-use-register (verified: primary)
[36] Governing AI development (AI Governance Engineering Body of Knowledge v0.5.0, chapter 14, section "Function creep"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-development#function-creep (verified: primary)
[37] Regulation (EU) 2024/1689 laying down harmonised rules on artificial intelligence (Artificial Intelligence Act), of 13 June 2024; OJ L, 2024/1689, 12.7.2024 (Art. 3(13) reasonably foreseeable misuse; Art. 9(2)(b) risks under reasonably foreseeable misuse; Art. 25(1)(c) changed intended purpose; Art. 50(2) machine-readable marking of synthetic outputs). Publications Office of the EU (EUR-Lex). 2024-07-12. https://eur-lex.europa.eu/eli/reg/2024/1689/oj/eng (verified: primary)
[38] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)

## Machine-readable

- Each control as JSON: https://aigovernanceengineer.com/api/v1/controls/<id in lower case>.json (for example https://aigovernanceengineer.com/api/v1/controls/aige-ctl-data-001.json)
- All controls as JSON: https://aigovernanceengineer.com/api/v1/controls.json (every control of every profile, in the envelope of the open data API)
- JSON Schema of the controls dataset: https://aigovernanceengineer.com/api/v1/schemas/controls.json (generated from the same registry)
- Control observation schema: https://aigovernanceengineer.com/schemas/control-observation.v1.json (the record a check of a control emits: control_id, subject, expected, observed, status, timestamp, evidence)
- Observation example: https://aigovernanceengineer.com/schemas/examples/control-observation.example.json (a filled record that validates against the schema)
- Observation template: https://aigovernanceengineer.com/templates/control-observation.md (the same fields in Markdown, with short guidance)
- The open data API: https://aigovernanceengineer.com/resources/data

## Review

Review a control through the issue form: https://github.com/losanchos5/aige/issues/new?template=control-review.yml. How review works: https://aigovernanceengineer.com/contribute. Page: https://aigovernanceengineer.com/controls/data-admission-and-privacy
