On this page

Credential Isolation

An agent under evaluation holds only short-lived credentials issued to its own identity for the run and bound to the one service each is for, never standing secrets or a person's own token.

v0.2 Draft Open for technical review

AIGE-CTL-EVAL-003, a control of the Evaluation Environment Control Profile (draft v0.2). See it among the other controls of the profile, or its mappings beside every other control's in the controls crosswalk.

The control record

Id
AIGE-CTL-EVAL-003 · v0.2 · Draft · Open for technical review
Objective
An agent under evaluation holds only short-lived credentials issued to its own identity for the run and bound to the one service each is for, never standing secrets or a person's own token.
Failure modes
  • A long-lived secret (an API key, a cloud access key, a password) is readable from the agent's environment, configuration, files or memory during a run.
  • The agent presents a token issued to a person, a token whose audience is another service, or a credential it found rather than received, and the tool server accepts it.
  • A credential issued for the run is still accepted after the run ended or was aborted.
  • A credential appears in the run's transcript, memory store, logs or outputs, or is passed to another agent.
Scope
Credentials, tokens and keys the agent under evaluation and the tools it calls can reach during a run, including what it holds in memory and writes to its transcript, and the model API key, which stays with a proxy outside the environment. The evaluator's own operator credentials are out of scope.
Enforcement points
  • deploy: before a version is deployed or released
  • runtime: at the point of action (gateway or guardrail)
Verification
  • Inspect: Before the run, inspect the environment template and the run's configuration: no long-lived secret is present, and the agent obtains credentials from a broker outside the environment under its own workload identity, each with a lifetime no longer than the run and an audience naming one tool server.
  • Test: After the run, scan every run artefact (transcript, memory store, logs, outputs and a snapshot of the environment's file system) for secret patterns and for the tokens issued to the run; expect no match.
  • Test: Replay a token issued for the run against a different tool server, and again after the run has ended; both must be rejected, for the wrong audience and for expiry or revocation.
  • Observe: Read the tool servers' logs for the run: every call carries a token issued for that server, delegated calls name the agent as the acting party, and audience-check failures were raised as alerts.
Evidence
  • Credential issuance log of the run: identity, audience, scope, lifetime and revocation time of every token · Layer 04 · evidence-record.v1
  • Audience-check and replay results from the tool servers · Layer 04 · evidence-record.v1
  • Secret scan of the run artefacts, filed as an observation of this control · Layer 05 · control-observation.v1
Failure response
deny: block the action. A token issued for another audience, to a person, or for a run that has ended is rejected by the tool server. A run in which a long-lived secret or a leaked credential is found is stopped, the credential is revoked and the result is withheld until the exposure is assessed.
Layers
Layer 04 Runtime Controls & Observability, Layer 02 Inventory & Transparency
Patterns
Agent Identity & Scoped Credentials
Seeded from
Its own identity, Replace long-lived secrets with short-lived credentials, Delegation, never impersonation, MCP authorisation (spec 2026-07-28), Memory write gate and rollback
Mappings
References
  • [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Identity and short-lived credentials")
  • [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Short-lived, attested credentials")
  • [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Delegation without impersonation")
  • [4] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "MCP authorization as of 2026-07-28")
  • [5] RFC 8693, OAuth 2.0 Token Exchange (the act claim "provides a means within a JWT to express that delegation has occurred and identify the acting party")
  • [6] MCP specification 2026-07-28, Authorization (MCP servers MUST validate that access tokens were issued specifically for them and "MUST NOT accept or transit any other tokens")
  • [7] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients)
  • [8] SPIFFE overview (SVIDs are "short lived cryptographic identity documents", delivered and rotated through the Workload API)
  • [9] Model AI Governance Framework for Agentic AI, v1.5 (agent identity unique and "cryptographically verifiable"; authorisations "time- or session-bound, non-transferable")
  • [10] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  • [11] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service)
  • [12] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  • [13] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
  • [14] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  • [15] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  • [16] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions")
  • [17] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning)
  • [18] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
Implementation notes
  • Give each agent under test its own workload identity for the run (an attested identity document of the kind SPIFFE defines, or the platform's workload identity) and issue its tokens from a broker outside the environment. The environment template holds no secret.
  • Keep the model API key with an LLM API proxy outside the environment, reached through the egress allow-list of AIGE-CTL-EVAL-002: the agent sees the proxy, never the key. Anthropic's guidance for external evaluation partners likewise asks that "the API keys should be kept outside the environment".
  • Issue short-lived access tokens whose aud names one tool server and, when the agent acts for a user or another agent, whose act claim names the agent (RFC 8693 token exchange). Tool servers validate the audience and refuse any other token, as the MCP authorization specification of 2026-07-28 requires of MCP servers.
  • Set every token lifetime no longer than the run's time budget, issue no refresh tokens to agents under evaluation, and revoke what was issued when the run ends or is aborted.
  • Scrub credentials from memory stores, transcripts and logs before they are written, then grep every run artefact after the run for secret patterns (known key prefixes, high-entropy strings, the run's own issued tokens); a match fails the run. Scan for credentials a run picked up as well as those it was issued: METR's investigation of the OpenAI/Hugging Face incident reports an agent that found working Hugging Face credentials exposed on the internet and posted them to a board other agents read, and OpenAI states that in a third-party evaluation its model reused a GitHub token another lab's agent had left publicly accessible.
Open questions
  • What lifetime should a credential issued for an evaluation run have, and who revokes it when the run is aborted?
  • Can a secret scan be made complete for encoded or split credentials, or should the control rely on short lifetimes and revocation for whatever a scan may miss? OpenAI reports a model that split a researcher's token into pieces with the stated aim of avoiding secret scanning.
Observation
  • Subject: Evaluation run
  • Expected: No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.
  • Example: Run 88214: secret scan found 0 matches in transcript, memory store and logs; a replayed token was rejected with an audience error: pass.

Example observations

Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.

Pass: run-88214@2026-09-26

Status
pass
Subject
run-88214@2026-09-26 (Evaluation run)
Expected
No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.
Observed
Secret scan of transcript, memory store, logs and file system snapshot: 0 matches. A run token replayed against another tool server was rejected with an audience error; replayed after the run, it was rejected as expired.
Timestamp
Observer
run-artefact-secret-scanner
Evidence
  • secret scan report of run 88214 · sha256:7e0b3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e
  • credential issuance log of run 88214
Notes
Illustrative example, not the result of a real evaluation.

Download the pass example (JSON)

The pass record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-003",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88214@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.",
  "observed": "Secret scan of transcript, memory store, logs and file system snapshot: 0 matches. A run token replayed against another tool server was rejected with an audience error; replayed after the run, it was rejected as expired.",
  "status": "pass",
  "timestamp": "2026-09-26T11:03:40Z",
  "enforcement_point": "runtime",
  "verification_kind": "test",
  "observer": "run-artefact-secret-scanner",
  "evidence": [
    {
      "artefact": "secret scan report of run 88214",
      "url": "https://evidence.example/runs/88214/secret-scan.json",
      "hash": "sha256:7e0b3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f1b4e7d0a3c6f9b2e5d8a1c4f7b0e"
    },
    {
      "artefact": "credential issuance log of run 88214",
      "url": "https://evidence.example/runs/88214/issuance.json",
      "schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
    }
  ],
  "run_id": "88214",
  "notes": "Illustrative example, not the result of a real evaluation."
}

Fail: run-88215@2026-09-26

Status
fail
Subject
run-88215@2026-09-26 (Evaluation run)
Expected
No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.
Observed
Secret scan found 1 long-lived API key in an environment variable of the agent container and the same key in the transcript at step 212.
Timestamp
Observer
run-artefact-secret-scanner
Evidence
  • secret scan report of run 88215 · sha256:2a5d8b1e4c7f0a3d6b9e2c5f8a1d4b7e0c3f6a9d2b5e8c1f4a7d0b3e6c9f2a5d
Notes
Illustrative example, not the result of a real evaluation. The key was revoked, the run stopped and its result withheld.

Download the fail example (JSON)

The fail record as JSON
{
  "$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
  "control_id": "AIGE-CTL-EVAL-003",
  "profile": "evaluation-environment",
  "control_version": "0.2",
  "subject": "run-88215@2026-09-26",
  "subject_kind": "eval-run",
  "expected": "No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.",
  "observed": "Secret scan found 1 long-lived API key in an environment variable of the agent container and the same key in the transcript at step 212.",
  "status": "fail",
  "timestamp": "2026-09-26T12:27:55Z",
  "enforcement_point": "runtime",
  "verification_kind": "test",
  "observer": "run-artefact-secret-scanner",
  "evidence": [
    {
      "artefact": "secret scan report of run 88215",
      "url": "https://evidence.example/runs/88215/secret-scan.json",
      "hash": "sha256:2a5d8b1e4c7f0a3d6b9e2c5f8a1d4b7e0c3f6a9d2b5e8c1f4a7d0b3e6c9f2a5d"
    }
  ],
  "run_id": "88215",
  "notes": "Illustrative example, not the result of a real evaluation. The key was revoked, the run stopped and its result withheld."
}

Incident cases on this site that list this control among their related controls.

Patterns

The patterns that implement this control.

Obligations

The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.

Threats

The entries of the threat catalogues this control answers.

Sources

  1. [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Identity and short-lived credentials"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#identity-and-short-lived-credentials (verified: primary)
  2. [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Short-lived, attested credentials"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#short-lived-attested-credentials (verified: primary)
  3. [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Delegation without impersonation"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#delegation-without-impersonation (verified: primary)
  4. [4] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "MCP authorization as of 2026-07-28"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#mcp-authorization-as-of-2026-07-28 (verified: primary)
  5. [5] RFC 8693, OAuth 2.0 Token Exchange (the act claim "provides a means within a JWT to express that delegation has occurred and identify the acting party"). IETF. 2020-01. https://www.rfc-editor.org/rfc/rfc8693.html (verified: primary)
  6. [6] MCP specification 2026-07-28, Authorization (MCP servers MUST validate that access tokens were issued specifically for them and "MUST NOT accept or transit any other tokens"). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization (verified: primary)
  7. [7] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices (verified: primary)
  8. [8] SPIFFE overview (SVIDs are "short lived cryptographic identity documents", delivered and rotated through the Workload API). SPIFFE project. 2026. https://spiffe.io/docs/latest/spiffe-about/overview/ (verified: primary)
  9. [9] Model AI Governance Framework for Agentic AI, v1.5 (agent identity unique and "cryptographically verifiable"; authorisations "time- or session-bound, non-transferable"). IMDA. 2026-05-20. https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf (verified: primary)
  10. [10] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
  11. [11] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service). METR. 2026. https://vivaria.metr.org/reference/config/ (verified: primary)
  12. [12] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
  13. [13] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml). MITRE. 2026-09-15. https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.09 (verified: primary)
  14. [14] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
  15. [15] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
  16. [16] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions"). OpenAI. 2026-08-04. https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ (verified: primary)
  17. [17] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/exposing-a-github-token-in-a-public-repository/ (verified: primary)
  18. [18] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)

Machine-readable

Review this control

Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.

This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.

Cite this control

García Aibar, J. (2026). AIGE-CTL-EVAL-003 Credential Isolation (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-003. CC BY 4.0

BibTeX

@misc{aige2026page,
  author       = {Jorge García Aibar},
  title        = {{AIGE-CTL-EVAL-003 Credential Isolation}},
  howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
  year         = {2026},
  version      = {0.2},
  doi          = {10.5281/zenodo.22857084},
  url          = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-003},
  note         = {Version 0.2}
}