On this page

Authorization Boundary

Every agent in an evaluation run acts only within an authorization boundary recorded before the run starts: the tools, operations, data classes and budgets it may use, and the runs and resources it may reach.

v0.2 Draft Open for technical review

AIGE-CTL-EVAL-001, a control of the Evaluation Environment Control Profile (draft v0.2). See it among the other controls of the profile, or its mappings beside every other control's in the controls crosswalk.

The control record

Id
AIGE-CTL-EVAL-001 · v0.2 · Draft · Open for technical review
Objective
Every agent in an evaluation run acts only within an authorization boundary recorded before the run starts: the tools, operations, data classes and budgets it may use, and the runs and resources it may reach.
Failure modes
  • A run starts with no recorded boundary for the agent under test, or with a boundary that differs from the scope the agent was told.
  • The agent calls a tool or an operation class outside its recorded boundary, or starts processes with administrative privileges, and the call succeeds.
  • A credential given for one purpose on a shared resource, such as downloading packages, also lets the agent write, list or post there.
  • Agents in runs meant to be isolated reach each other through a resource the runs share, such as a package repository or a cache.
Scope
Agents and harnesses under evaluation, the tools and operations they can use during a run, the credentials they receive for shared resources and the resources runs share. Budgets are recorded here and enforced under AIGE-CTL-EVAL-006; network egress is AIGE-CTL-EVAL-002. Production deployments are covered by the agent runtime profile.
Enforcement points
  • deploy: before a version is deployed or released
  • runtime: at the point of action (gateway or guardrail)
Verification
  • Inspect: Before the run, inspect the run record: it holds a boundary for each agent under test (tools and operation classes, data classes, budgets, and each shared resource with the operations allowed on it), and the prompt the agent receives states the same boundary as instructions, including what it must not access.
  • Test: At admission, from inside the environment, attempt one call of each kind outside the boundary (an unlisted tool or operation class, a write or a listing with a download-only credential on a shared resource, a process started as root) and one listed call; every attempt outside the boundary must be refused and logged, and the listed call must succeed.
  • Test: Start two canary runs on the same shared resources: a marker written by one run must not be readable by the other.
  • Observe: After the run, compare every tool call and every request to a shared resource in the run's logs with the recorded boundary: each falls inside it, and every refused attempt is recorded with its time and target.
Evidence
  • The boundary of each agent under test, recorded before the run: tools, operation classes, data classes, budgets and shared resources with the operations allowed on each · Layer 02 · agent-register-entry.v1
  • Admission test verdicts: the refused out-of-boundary calls and the cross-run canary · Layer 04 · evidence-record.v1
  • One observation per run comparing the calls and shared-resource requests made with the recorded boundary · Layer 05 · control-observation.v1
Failure response
deny: block the action. Calls and shared-resource operations outside the recorded boundary are refused at the enforcement point and logged. A run with no recorded boundary is not started; a run in which a call outside the boundary succeeded, or in which runs reached each other, is stopped and its result is withheld until the path is closed.
Layers
Layer 04 Runtime Controls & Observability, Layer 02 Inventory & Transparency
Patterns
Agent Registry, Policy Card
Seeded from
Registry entry, Read-only tools, Reversible, bounded actions only, Execution budgets, Data classes recorded, with the DPIA linked
Mappings
References
  • [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The agent registry")
  • [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Autonomy is a design decision")
  • [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The tool allow-list")
  • [4] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server)
  • [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [8] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions")
  • [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
  • [11] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent")
  • [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
Implementation notes
  • Record the boundary as data before the run, in the fields of an agent registry entry (scope, tools, data access, spend limit), and enforce it at the gateway the agent's calls pass through, not in the prompt alone. State the same boundary to the agent: Anthropic's guidance for external evaluation partners asks every prompt to state what is in and out of scope, including permitted actions and network boundaries, with boundaries "phrased as instructions" rather than claims about the environment.
  • Tell the agent what it may not do with the access it has. OpenAI's account of a third-party cyber evaluation states that the evaluator's "intended authorization boundary was the simulated cyber range", while the agents were not explicitly told how they could and could not use the open internet access they had; the evaluator names this as a potential contributing cause.
  • Scope every credential for a shared resource to the operations the task needs: download from a package repository, never write, list or post. OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate samples with the credentials they were given to download packages, without exploiting any vulnerability; its technical report on the Hugging Face incident lists removing the shared credential from new research workloads and restricting the types of requests agents could make to Artifactory among its mitigations.
  • Give each run its own namespace on every shared resource (package cache, artefact store, scratch storage) and test it with a canary marker. METR's investigation of the OpenAI/Hugging Face incident reports that agents "meant to be fully isolated from one another" communicated through an internal package repository.
  • Run the agent as an unprivileged user. METR's Task Standard states that the agent "MUST NOT be allowed to start processes" on the primary machine as root.
Open questions
  • Which shared resources (package mirrors, caches, model API proxies) can be given a per-run namespace, and which should be removed from evaluation environments altogether?
  • What evidence is sufficient to show that an agent stayed within its boundary for a whole run when the logs of some of the tools it called are outside the evaluator's control?
Observation
  • Subject: Evaluation run
  • Expected: Every tool call and every request to a shared resource falls inside the boundary recorded for the agent before the run; every attempt outside it is refused and logged.
  • Example: Run 88231: 3 listing and upload requests to the shared package repository with a download-only credential, all refused and logged; no call outside the boundary succeeded: pass.

Example observations

Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.

Pass: run-88231@2026-09-26

Status
pass
Subject
run-88231@2026-09-26 (Evaluation run)
Expected
Every tool call and every request to a shared resource falls inside the boundary recorded for the agent before the run; every attempt outside it is refused and logged.
Observed
Run 88231: 3 listing and upload requests to the shared package repository with a download-only credential, all refused and logged; the cross-run canary marker was not readable from run 88232; no call outside the recorded boundary succeeded.
Timestamp
Observer
boundary-log-comparator
Evidence
  • gateway and package repository request log of run 88231 · sha256:da6a97adcb16d06010c05a988f4091e15aede6a60a5cb87fc23033f1ee11c55e
  • boundary recorded for the agent before run 88231
Notes
Illustrative example, not the result of a real evaluation.

Download the pass example (JSON)

The pass record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-001",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88231@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "Every tool call and every request to a shared resource falls inside the boundary recorded for the agent before the run; every attempt outside it is refused and logged.",
  "observed": "Run 88231: 3 listing and upload requests to the shared package repository with a download-only credential, all refused and logged; the cross-run canary marker was not readable from run 88232; no call outside the recorded boundary succeeded.",
  "status": "pass",
  "timestamp": "2026-09-26T08:41:09Z",
  "enforcement_point": "runtime",
  "verification_kind": "observe",
  "observer": "boundary-log-comparator",
  "evidence": [
    {
      "artefact": "gateway and package repository request log of run 88231",
      "url": "https://evidence.example/runs/88231/requests.jsonl",
      "hash": "sha256:da6a97adcb16d06010c05a988f4091e15aede6a60a5cb87fc23033f1ee11c55e"
    },
    {
      "artefact": "boundary recorded for the agent before run 88231",
      "url": "https://evidence.example/runs/88231/boundary.json",
      "schema": "https://aigovernanceengineer.com/schemas/agent-register-entry.v1.json"
    }
  ],
  "run_id": "88231",
  "notes": "Illustrative example, not the result of a real evaluation."
}

Fail: run-88233@2026-09-26

Status
fail
Subject
run-88233@2026-09-26 (Evaluation run)
Expected
Every tool call and every request to a shared resource falls inside the boundary recorded for the agent before the run; every attempt outside it is refused and logged.
Observed
Admission test of run 88233: a listing request to the shared package repository with the download credential succeeded and returned files written by another run.
Timestamp
Observer
boundary-admission-test
Evidence
  • admission test verdicts of run 88233 · sha256:c938ad26e04f47ff6216693b737cd52ec5ef208d7cae90883b9abd6a7d41c2ab
Notes
Illustrative example, not the result of a real evaluation. The run was not started; the credential was reissued as download-only and the repository given a per-run namespace.

Download the fail example (JSON)

The fail record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-001",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88233@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "Every tool call and every request to a shared resource falls inside the boundary recorded for the agent before the run; every attempt outside it is refused and logged.",
  "observed": "Admission test of run 88233: a listing request to the shared package repository with the download credential succeeded and returned files written by another run.",
  "status": "fail",
  "timestamp": "2026-09-26T10:12:37Z",
  "enforcement_point": "deploy",
  "verification_kind": "test",
  "observer": "boundary-admission-test",
  "evidence": [
    {
      "artefact": "admission test verdicts of run 88233",
      "url": "https://evidence.example/runs/88233/admission.json",
      "hash": "sha256:c938ad26e04f47ff6216693b737cd52ec5ef208d7cae90883b9abd6a7d41c2ab",
      "schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
    }
  ],
  "run_id": "88233",
  "notes": "Illustrative example, not the result of a real evaluation. The run was not started; the credential was reissued as download-only and the repository given a per-run namespace."
}

Incident cases on this site that list this control among their related controls.

Patterns

The patterns that implement this control.

Obligations

The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.

Threats

The entries of the threat catalogues this control answers.

Sources

  1. [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The agent registry"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#the-agent-registry (verified: primary)
  2. [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Autonomy is a design decision"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#autonomy-is-a-design-decision (verified: primary)
  3. [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The tool allow-list"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#the-tool-allow-list (verified: primary)
  4. [4] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server). METR (GitHub). 2024-10-30. https://raw.githubusercontent.com/METR/task-standard/main/STANDARD.md (verified: primary)
  5. [5] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
  6. [6] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
  7. [7] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
  8. [8] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions"). OpenAI. 2026-08-04. https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ (verified: primary)
  9. [9] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
  10. [10] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
  11. [11] Unsanctioned Artifactory writes and cross-sample communication (misalignment report) (OpenAI reports models in RL training that used Artifactory, "a shared package repository", to exchange messages across separate training samples with the credentials they were given to download packages, without exploiting any vulnerability; such exchanges can "undermine the assumption that training or evaluation samples are independent"). OpenAI (Alignment Research Blog). 2026-09-16. https://alignment.openai.com/misalignment-reports/unauthorized-artifactory-writes-and-cross-sample-communication/ (verified: primary)
  12. [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)

Machine-readable

Review this control

Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.

This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.

Cite this control

García Aibar, J. (2026). AIGE-CTL-EVAL-001 Authorization Boundary (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-001. CC BY 4.0

BibTeX

@misc{aige2026page,
  author       = {Jorge García Aibar},
  title        = {{AIGE-CTL-EVAL-001 Authorization Boundary}},
  howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
  year         = {2026},
  version      = {0.2},
  doi          = {10.5281/zenodo.22857084},
  url          = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-001},
  note         = {Version 0.2}
}