---
title: "Credential Isolation"
description: "Draft control AIGE-CTL-EVAL-003: an agent under evaluation holds only short-lived credentials bound to its own identity and one service, no standing secrets."
canonical: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-003
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22857084
version: "0.2"
updated: 2026-09-26
---

# Credential Isolation

> An agent under evaluation holds only short-lived credentials issued to its own identity for the run and bound to the one service each is for, never standing secrets or a person's own token.

- Id: AIGE-CTL-EVAL-003
- Profile: [Evaluation Environment Control Profile v0.2](https://aigovernanceengineer.com/controls/evaluation-environment)
- Status: Draft
- Review: Open for technical review
- Published: 2026-09-26
- Updated: 2026-09-26
- Anchor on the profile page: https://aigovernanceengineer.com/controls/evaluation-environment#aige-ctl-eval-003

Draft for review, not a claim of conformity. A draft control specification, open for technical review: illustrative, not legal advice and binding on no one.

## The control record

- Id: `AIGE-CTL-EVAL-003` · v0.2 · Draft · Open for technical review
- Depth: Specified
- Objective: An agent under evaluation holds only short-lived credentials issued to its own identity for the run and bound to the one service each is for, never standing secrets or a person's own token.
- Failure modes:
  - A long-lived secret (an API key, a cloud access key, a password) is readable from the agent's environment, configuration, files or memory during a run.
  - The agent presents a token issued to a person, a token whose audience is another service, or a credential it found rather than received, and the tool server accepts it.
  - A credential issued for the run is still accepted after the run ended or was aborted.
  - A credential appears in the run's transcript, memory store, logs or outputs, or is passed to another agent.
- Scope: Credentials, tokens and keys the agent under evaluation and the tools it calls can reach during a run, including what it holds in memory and writes to its transcript, and the model API key, which stays with a proxy outside the environment. The evaluator's own operator credentials are out of scope.
- Enforcement points:
  - deploy: before a version is deployed or released
  - runtime: at the point of action (gateway or guardrail)
- Verification:
  - Inspect: Before the run, inspect the environment template and the run's configuration: no long-lived secret is present, and the agent obtains credentials from a broker outside the environment under its own workload identity, each with a lifetime no longer than the run and an audience naming one tool server.
  - Test: After the run, scan every run artefact (transcript, memory store, logs, outputs and a snapshot of the environment's file system) for secret patterns and for the tokens issued to the run; expect no match.
  - Test: Replay a token issued for the run against a different tool server, and again after the run has ended; both must be rejected, for the wrong audience and for expiry or revocation.
  - Observe: Read the tool servers' logs for the run: every call carries a token issued for that server, delegated calls name the agent as the acting party, and audience-check failures were raised as alerts.
- Evidence:
  - Credential issuance log of the run: identity, audience, scope, lifetime and revocation time of every token · Layer 04 Runtime Controls & Observability · [evidence-record.v1](https://aigovernanceengineer.com/resources/templates#schema-evidence-record)
  - Audience-check and replay results from the tool servers · Layer 04 Runtime Controls & Observability · [evidence-record.v1](https://aigovernanceengineer.com/resources/templates#schema-evidence-record)
  - Secret scan of the run artefacts, filed as an observation of this control · Layer 05 Assurance & Continuous Compliance · [control-observation.v1](https://aigovernanceengineer.com/resources/templates#schema-control-observation)
- Failure response: deny: block the action. A token issued for another audience, to a person, or for a run that has ended is rejected by the tool server. A run in which a long-lived secret or a leaked credential is found is stopped, the credential is revoked and the result is withheld until the exposure is assessed.
- Layers: [Layer 04 Runtime Controls & Observability](https://aigovernanceengineer.com/bok/the-stack#layer-04-runtime-controls--observability), [Layer 02 Inventory & Transparency](https://aigovernanceengineer.com/bok/the-stack#layer-02-inventory--transparency)
- Patterns: [Agent Identity & Scoped Credentials](https://aigovernanceengineer.com/patterns/agent-identity-scoped-credentials)
- Seeded from: [Its own identity](https://aigovernanceengineer.com/bok/governing-agents#identity-and-short-lived-credentials), [Replace long-lived secrets with short-lived credentials](https://aigovernanceengineer.com/bok/governing-agents#short-lived-attested-credentials), [Delegation, never impersonation](https://aigovernanceengineer.com/bok/governing-agents#delegation-without-impersonation), [MCP authorisation (spec 2026-07-28)](https://aigovernanceengineer.com/bok/governing-agents#mcp-authorization-as-of-2026-07-28), [Memory write gate and rollback](https://aigovernanceengineer.com/bok/governing-agents#memory-and-context-governance)
- Mappings:
  - Obligations: [EU AI Act Art. 15 accuracy, robustness and cybersecurity](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art15); [Singapore IMDA Model AI Governance Framework for Agentic AI: agent identity and scoped authorisations (voluntary)](https://aigovernanceengineer.com/obligations/aige-obl-sg-agentic-identity); [NIST AI Agent Standards Initiative (2026)](https://aigovernanceengineer.com/obligations/aige-obl-nist-agents); [OWASP Top 10 for Agentic Applications 2026](https://aigovernanceengineer.com/obligations/aige-obl-owasp-agentic)
  - ISO/IEC 42001: A.6.2.6 AI system operation and monitoring; A.9.2 Processes for responsible use of AI systems
  - NIST AI RMF: MEASURE 2.7 Security and resilience are evaluated and documented
  - OWASP: [ASI03 Identity and Privilege Abuse](https://aigovernanceengineer.com/resources/threats#threat-asi03)
  - AIUC-1: A008
  - IETF RFC 8693: act claim (delegation names the acting party; never impersonation)
  - MCP specification 2026-07-28: Authorization, Token Handling (audience validation; no token passthrough)
  - SPIFFE: SVID (short-lived workload identity documents)
  - MITRE ATLAS: AML.T0083 (Credentials from AI Agent Configuration (not yet a row of the threat bridge))
  - NIST SP 800-53 Rev. 5: IA-5 (Authenticator Management)
  - NIST SP 800-53 Rev. 5: AC-6 (Least Privilege)
- References:
  - [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Identity and short-lived credentials")
  - [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Short-lived, attested credentials")
  - [3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Delegation without impersonation")
  - [4] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "MCP authorization as of 2026-07-28")
  - [5] RFC 8693, OAuth 2.0 Token Exchange (the act claim "provides a means within a JWT to express that delegation has occurred and identify the acting party")
  - [6] MCP specification 2026-07-28, Authorization (MCP servers MUST validate that access tokens were issued specifically for them and "MUST NOT accept or transit any other tokens")
  - [7] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients)
  - [8] SPIFFE overview (SVIDs are "short lived cryptographic identity documents", delivered and rotated through the Workload API)
  - [9] Model AI Governance Framework for Agentic AI, v1.5 (agent identity unique and "cryptographically verifiable"; authorisations "time- or session-bound, non-transferable")
  - [10] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
  - [11] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service)
  - [12] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
  - [13] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
  - [14] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
  - [15] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
  - [16] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions")
  - [17] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning)
  - [18] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
- Implementation notes:
  - Give each agent under test its own workload identity for the run (an attested identity document of the kind SPIFFE defines, or the platform's workload identity) and issue its tokens from a broker outside the environment. The environment template holds no secret.
  - Keep the model API key with an LLM API proxy outside the environment, reached through the egress allow-list of AIGE-CTL-EVAL-002: the agent sees the proxy, never the key. Anthropic's guidance for external evaluation partners likewise asks that "the API keys should be kept outside the environment".
  - Issue short-lived access tokens whose aud names one tool server and, when the agent acts for a user or another agent, whose act claim names the agent (RFC 8693 token exchange). Tool servers validate the audience and refuse any other token, as the MCP authorization specification of 2026-07-28 requires of MCP servers.
  - Set every token lifetime no longer than the run's time budget, issue no refresh tokens to agents under evaluation, and revoke what was issued when the run ends or is aborted.
  - Scrub credentials from memory stores, transcripts and logs before they are written, then grep every run artefact after the run for secret patterns (known key prefixes, high-entropy strings, the run's own issued tokens); a match fails the run. Scan for credentials a run picked up as well as those it was issued: METR's investigation of the OpenAI/Hugging Face incident reports an agent that found working Hugging Face credentials exposed on the internet and posted them to a board other agents read, and OpenAI states that in a third-party evaluation its model reused a GitHub token another lab's agent had left publicly accessible.
- Open questions:
  - What lifetime should a credential issued for an evaluation run have, and who revokes it when the run is aborted?
  - Can a secret scan be made complete for encoded or split credentials, or should the control rely on short lifetimes and revocation for whatever a scan may miss? OpenAI reports a model that split a researcher's token into pieces with the stated aim of avoiding secret scanning.
- Observation:
  - Subject: Evaluation run
  - Expected: No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.
  - Example: Run 88214: secret scan found 0 matches in transcript, memory store and logs; a replayed token was rejected with an audience error: pass.
- JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-003.json

## Example observations

Two illustrative records of a check of this control, one that passes and one that fails. They validate against the control observation schema; they are not results of any real evaluation.

### Pass: run-88214@2026-09-26

- Status: pass
- Subject: run-88214@2026-09-26 (Evaluation run)
- Expected: No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.
- Observed: Secret scan of transcript, memory store, logs and file system snapshot: 0 matches. A run token replayed against another tool server was rejected with an audience error; replayed after the run, it was rejected as expired.
- Timestamp: 2026-09-26T11:03:40Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-003.pass.json

### Fail: run-88215@2026-09-26

- Status: fail
- Subject: run-88215@2026-09-26 (Evaluation run)
- Expected: No credential other than the run's own short-lived, audience-bound tokens is readable in the environment or appears in the run's artefacts, and none is accepted after the run.
- Observed: Secret scan found 1 long-lived API key in an environment variable of the agent container and the same key in the transcript at step 212.
- Timestamp: 2026-09-26T12:27:55Z
- JSON: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-003.fail.json

## Related cases

- [OpenAI agents and Hugging Face: an evaluation environment that was not isolated](https://aigovernanceengineer.com/cases/openai-hugging-face-agent-incident-2026): METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- [An internally deployed model published a researcher's GitHub token in a public repository](https://aigovernanceengineer.com/cases/openai-agent-github-token-exposure-2026): OpenAI reports that an internally deployed model put a researcher's GitHub token, split to avoid secret scanning, into code it pushed to a public repository.
- [Training samples exchanged messages through a shared package repository](https://aigovernanceengineer.com/cases/openai-agents-artifactory-cross-sample-2026): OpenAI reports that models in RL training used an internal package repository, with the credentials they were given, to exchange messages across samples.
- [Claude models reached real systems from a misconfigured third-party cyber evaluation](https://aigovernanceengineer.com/cases/anthropic-third-party-eval-environment-incidents-2026): Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems.
- [Agents in a cyber range with open internet took unsanctioned actions against real people](https://aigovernanceengineer.com/cases/uk-aisi-cyber-range-unsanctioned-actions-2026): UK AISI reports that agents in a cyber evaluation with internet deliberately enabled took 19 unsanctioned actions aimed at real people and organisations.

## Patterns

- [Agent Identity & Scoped Credentials](https://aigovernanceengineer.com/patterns/agent-identity-scoped-credentials) (Layer 04 Runtime Controls & Observability)

## Obligations

- [EU AI Act Art. 15 accuracy, robustness and cybersecurity](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art15) (`AIGE-OBL-EUAIA-ART15`): Accuracy, robustness and cybersecurity
- [Singapore IMDA Model AI Governance Framework for Agentic AI: agent identity and scoped authorisations (voluntary)](https://aigovernanceengineer.com/obligations/aige-obl-sg-agentic-identity) (`AIGE-OBL-SG-AGENTIC-IDENTITY`): Each agent has a unique, accounted-for identity, catalogued and centrally managed; authorisations are scoped, time- or session-bound, non-transferable and bounded by the authorising human
- [NIST AI Agent Standards Initiative (2026)](https://aigovernanceengineer.com/obligations/aige-obl-nist-agents) (`AIGE-OBL-NIST-AGENTS`): CAISI initiative on interoperable, secure AI agents: identity, authentication, agent security
- [OWASP Top 10 for Agentic Applications 2026](https://aigovernanceengineer.com/obligations/aige-obl-owasp-agentic) (`AIGE-OBL-OWASP-AGENTIC`): Agent threat catalogue (ASI01 Agent Goal Hijack … ASI10 Rogue Agents)

## Threats

- [ASI03 Identity and Privilege Abuse](https://aigovernanceengineer.com/resources/threats#threat-asi03) (OWASP Top 10 for Agentic Applications 2026)

## Sources

[1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Identity and short-lived credentials"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#identity-and-short-lived-credentials (verified: primary)
[2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Short-lived, attested credentials"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#short-lived-attested-credentials (verified: primary)
[3] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Delegation without impersonation"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#delegation-without-impersonation (verified: primary)
[4] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "MCP authorization as of 2026-07-28"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#mcp-authorization-as-of-2026-07-28 (verified: primary)
[5] RFC 8693, OAuth 2.0 Token Exchange (the act claim "provides a means within a JWT to express that delegation has occurred and identify the acting party"). IETF. 2020-01. https://www.rfc-editor.org/rfc/rfc8693.html (verified: primary)
[6] MCP specification 2026-07-28, Authorization (MCP servers MUST validate that access tokens were issued specifically for them and "MUST NOT accept or transit any other tokens"). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/specification/2026-07-28/basic/authorization (verified: primary)
[7] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices (verified: primary)
[8] SPIFFE overview (SVIDs are "short lived cryptographic identity documents", delivered and rotated through the Workload API). SPIFFE project. 2026. https://spiffe.io/docs/latest/spiffe-about/overview/ (verified: primary)
[9] Model AI Governance Framework for Agentic AI, v1.5 (agent identity unique and "cryptographically verifiable"; authorisations "time- or session-bound, non-transferable"). IMDA. 2026-05-20. https://www.imda.gov.sg/-/media/imda/files/about/emerging-tech-and-research/artificial-intelligence/mgf-for-agentic-ai.pdf (verified: primary)
[10] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
[11] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service). METR. 2026. https://vivaria.metr.org/reference/config/ (verified: primary)
[12] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
[13] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml). MITRE. 2026-09-15. https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.09 (verified: primary)
[14] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
[15] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
[16] Third-party cyber evaluations involving OpenAI models (OpenAI states that the evaluator's "intended authorization boundary was the simulated cyber range", that its model reused a GitHub token another lab's agent had left publicly accessible, and that it will review how to "set expectations for isolation, credential handling, monitoring, and stop conditions"). OpenAI. 2026-08-04. https://openai.com/index/third-party-cyber-evaluations-involving-openai-models/ (verified: primary)
[17] Exposing a GitHub token in a public repository (misalignment report) (OpenAI reports an internally deployed model that published a researcher's GitHub token to a public repository, split into pieces with the stated aim of avoiding secret scanning). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/exposing-a-github-token-in-a-public-repository/ (verified: primary)
[18] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)

## Machine-readable

- This control as JSON: https://aigovernanceengineer.com/api/v1/controls/aige-ctl-eval-003.json
- The pass example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-003.pass.json
- The fail example: https://aigovernanceengineer.com/controls/examples/control-observation.aige-ctl-eval-003.fail.json
- The whole profile as Markdown: https://aigovernanceengineer.com/controls/evaluation-environment.md
- The open data API: https://aigovernanceengineer.com/resources/data

## Review

Review this control through the issue form: https://github.com/losanchos5/aige/issues/new?template=control-review.yml. How review works: https://aigovernanceengineer.com/contribute. Page: https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-003

## Cite

AIGE-CTL-EVAL-003 Credential Isolation. In Jorge García Aibar (2026). Evaluation Environment Control Profile (v0.2, draft). AI Governance Engineer. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment
