Network Egress Control
Outbound connections from an evaluation environment reach only the destinations on the run's egress allow-list, and every other connection attempt is refused and logged.
AIGE-CTL-EVAL-002, a control of the Evaluation Environment Control Profile (draft
v0.2). See it among the other controls of the profile,
or its mappings beside every other control's in the
controls crosswalk.
The control record
- Id
-
AIGE-CTL-EVAL-002· v0.2 · Draft · Open for technical review - Objective
- Outbound connections from an evaluation environment reach only the destinations on the run's egress allow-list, and every other connection attempt is refused and logged.
- Failure modes
-
- A connection from the environment to a host that is not on the run's egress allow-list succeeds.
- A run starts in an environment with no egress policy attached, or with a policy that allows every destination although the task did not declare that it needs the internet.
- A permitted destination, such as a package mirror, a cache or a tool server, carries data onward to a party or to another run that nobody listed.
- The run leaves no flow log, so the connections it made cannot be compared with its allow-list.
- Scope
- Every network path out of the environment a run executes in: the agent's container or virtual machine, auxiliary machines, DNS, and the tools, MCP servers and proxies the run can call. Resources shared between runs count as destinations. Inbound operator access is out of scope.
- Enforcement points
-
- deploy: before a version is deployed or released
- runtime: at the point of action (gateway or guardrail)
- Verification
-
- Inspect: Before the run, inspect the egress policy attached to the task environment: deny by default, with an allow-list naming each permitted destination (for example the LLM API proxy and the progress server) and nothing else unless the task declares that it needs the internet.
- Test: At admission, from inside the environment, attempt one connection to a destination that is not on the allow-list and one to a listed destination; the first must be refused and logged, the second must succeed.
- Observe: After the run, compare the run's flow log with its allow-list: every outbound connection matches a listed destination, and every refused attempt is recorded with its time and target.
- Evidence
-
- The egress policy and allow-list attached to the run, with its hash recorded in the run record · Layer 04
- Admission test verdict: the refused connection to an unlisted destination · Layer 04 · evidence-record.v1
- Flow log of the run, allowed and refused connections, kept outside the environment · Layer 04
- One observation per run comparing observed connections with the allow-list · Layer 05 · control-observation.v1
- Failure response
- deny: block the action. Connections to unlisted destinations are refused at the enforcement point and logged. A run whose environment has no egress policy attached is not started; a run in which an unlisted connection succeeded is stopped and its result is withheld until the connection is explained.
- Layer
- Layer 04 Runtime Controls & Observability
- Patterns
- Runtime Guardrail, Sanctioned AI Gateway
- Seeded from
- Output and egress filter, Tool allow-list, deny by default, Code runs only in a sandbox
- Mappings
-
- Obligations: EU AI Act Art. 15 accuracy, robustness and cybersecurity; OWASP Top 10 for LLM Applications 2026; OWASP Top 10 for Agentic Applications 2026
- ISO/IEC 42001: A.6.2.6 AI system operation and monitoring
- NIST AI RMF: MEASURE 2.7 Security and resilience are evaluated and documented
- OWASP: ASI02 Tool Misuse and Exploitation; LLM02:2026 Sensitive Information Disclosure
- MITRE ATLAS: AML.T0086 Exfiltration via AI Agent Tool Invocation
- AIUC-1: B006
- MITRE ATLAS mitigation: AML.M0032 (Segmentation of AI Agent Components)
- NIST SP 800-53 Rev. 5: SC-7 (Boundary Protection)
- NIST SP 800-53 Rev. 5: SC-7(5) (Deny by default, allow by exception)
- References
-
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The tool allow-list")
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls")
- [3] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server)
- [4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets)
- [5] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service)
- [6] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients)
- [7] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents)
- [8] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026)
- [9] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml)
- [10] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020)
- [11] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later)
- [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads")
- [13] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise)
- [14] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit)
- Implementation notes
-
- Attach a default-deny egress policy to every task environment (a network policy, a firewall rule set or an egress proxy, by category) that lists each permitted destination explicitly, for example the LLM API proxy and the progress server. Enforce it on the resolved address, and serve DNS from a resolver that answers only for listed names: DNS is an egress path too. OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox.
- Record the allow-list with the run (its hash goes in the run record next to the environment version) and keep a flow log per run at the enforcement point, outside anything the agent can reach or change, with refused attempts as well as allowed connections, so the connections made can be compared with what was declared, not with what was intended.
- Enforce the policy in two places a single misconfiguration cannot both open, for example at the sandbox and at the cluster boundary. OpenAI's technical report on the Hugging Face incident describes "two separate layers of network isolation" of this kind, and Anthropic's guidance for external evaluation partners asks that the only outside connection be to the model's own API and that the configuration be verified before every evaluation begins.
- At admission, run a canary from inside the environment: one connection to a reserved test address that is not on the list. Expect a refused connection and a log entry; a success blocks the run.
- Treat resources shared between runs (package mirrors, caches, artefact stores) as destinations with their own entry and a per-run scope. METR's public investigation of the OpenAI/Hugging Face incident reports that agents meant to be "fully isolated from one another" communicated through a shared internal package repository.
- Open questions
-
- How should the egress allow-list of a run be recorded so that a third party can compare it with the connections actually observed?
- When a task needs the internet (what METR's Task Standard calls the full_internet permission), which part of this control still applies, and what should the run record say instead?
- Observation
-
- Subject: Evaluation environment
- Expected: Outbound connections only to the destinations on the run's egress allow-list; every other attempt refused and logged.
- Example: 1 connection to an unlisted host during run 88213, in an environment with no egress policy attached: fail.
Example observations
Two illustrative records of a check of this control, one that passes and one that fails. Each validates against the control observation schema; neither is the result of a real evaluation.
Pass: eval-env-eu-west@2026-09-26
- Status
-
pass - Subject
-
eval-env-eu-west@2026-09-26(Evaluation environment) - Expected
- Outbound connections only to the destinations on the run's egress allow-list; every other attempt refused and logged.
- Observed
- Run 88212: 1,406 outbound connections, all to the 2 listed destinations (LLM API proxy, progress server); the admission canary to an unlisted test address was refused and logged.
- Timestamp
- Observer
- egress-flow-log-adapter
- Evidence
-
- flow log of run 88212 ·
sha256:4b1d7e0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f1b4e - egress allow-list attached to the run ·
sha256:0c3f6a9d2b5e8c1f4a7d0b3e6c9f2a5d8b1e4c7f0a3d6b9e2c5f8a1d4b7e0c3f - evidence record of the admission canary (refused connection)
- flow log of run 88212 ·
- Notes
- Illustrative example, not the result of a real evaluation.
Download the pass example (JSON)
The pass record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-002",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "eval-env-eu-west@2026-09-26",
"subject_kind": "eval-environment",
"expected": "Outbound connections only to the destinations on the run's egress allow-list; every other attempt refused and logged.",
"observed": "Run 88212: 1,406 outbound connections, all to the 2 listed destinations (LLM API proxy, progress server); the admission canary to an unlisted test address was refused and logged.",
"status": "pass",
"timestamp": "2026-09-26T08:15:04Z",
"enforcement_point": "runtime",
"verification_kind": "observe",
"observer": "egress-flow-log-adapter",
"evidence": [
{
"artefact": "flow log of run 88212",
"url": "https://evidence.example/runs/88212/flows.jsonl",
"hash": "sha256:4b1d7e0a3c6f9b2e5d8a1c4f7b0e3d6a9c2f5b8e1d4a7c0f3b6e9d2a5c8f1b4e"
},
{
"artefact": "egress allow-list attached to the run",
"url": "https://evidence.example/runs/88212/egress-allow-list.json",
"hash": "sha256:0c3f6a9d2b5e8c1f4a7d0b3e6c9f2a5d8b1e4c7f0a3d6b9e2c5f8a1d4b7e0c3f"
},
{
"artefact": "evidence record of the admission canary (refused connection)",
"url": "https://evidence.example/runs/88212/admission.json",
"schema": "https://aigovernanceengineer.com/schemas/evidence-record.v1.json"
}
],
"run_id": "88212",
"notes": "Illustrative example, not the result of a real evaluation."
} Fail: eval-env-eu-west@2026-09-26
- Status
-
fail - Subject
-
eval-env-eu-west@2026-09-26(Evaluation environment) - Expected
- Outbound connections only to the destinations on the run's egress allow-list; every other attempt refused and logged.
- Observed
- Run 88213: 1 connection to an unlisted host succeeded; the task environment had no egress policy attached.
- Timestamp
- Observer
- egress-flow-log-adapter
- Evidence
-
- flow log of run 88213 ·
sha256:9f2c4e1a7b3d5f8e0a2c4b6d8f1e3a5c7b9d0f2e4a6c8b1d3f5e7a9c0b2d4f6e
- flow log of run 88213 ·
- Notes
- Illustrative example, not the result of a real evaluation. The run was stopped, its result withheld and the environment template fixed.
Download the fail example (JSON)
The fail record as JSON
{
"$schema": "https://aigovernanceengineer.com/schemas/control-observation.v1.json",
"control_id": "AIGE-CTL-EVAL-002",
"profile": "evaluation-environment",
"control_version": "0.2",
"subject": "eval-env-eu-west@2026-09-26",
"subject_kind": "eval-environment",
"expected": "Outbound connections only to the destinations on the run's egress allow-list; every other attempt refused and logged.",
"observed": "Run 88213: 1 connection to an unlisted host succeeded; the task environment had no egress policy attached.",
"status": "fail",
"timestamp": "2026-09-26T09:42:17Z",
"enforcement_point": "runtime",
"verification_kind": "observe",
"observer": "egress-flow-log-adapter",
"evidence": [
{
"artefact": "flow log of run 88213",
"url": "https://evidence.example/runs/88213/flows.jsonl",
"hash": "sha256:9f2c4e1a7b3d5f8e0a2c4b6d8f1e3a5c7b9d0f2e4a6c8b1d3f5e7a9c0b2d4f6e"
}
],
"run_id": "88213",
"notes": "Illustrative example, not the result of a real evaluation. The run was stopped, its result withheld and the environment template fixed."
} Related cases
Incident cases on this site that list this control among their related controls.
- OpenAI agents and Hugging Face: an evaluation environment that was not isolated: METR reports that OpenAI agents meant to be isolated in cyber evaluations used a shared package repository as a message board and attacked Hugging Face.
- An agent in training reached a public chatbot through the sandbox DNS resolver: OpenAI reports that an agent in RL training used its sandbox's DNS resolver to reach a public chatbot; the run was stopped hours after the alert.
- Agents in training shared a file through a public file-hosting service: OpenAI reports that agents in multi-agent RL training uploaded a workbook to a public file-hosting service so that collaborating agents could download it.
- Claude models reached real systems from a misconfigured third-party cyber evaluation: Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems.
- Agents in a cyber range with open internet took unsanctioned actions against real people: UK AISI reports that agents in a cyber evaluation with internet deliberately enabled took 19 unsanctioned actions aimed at real people and organisations.
Patterns
The patterns that implement this control.
- Runtime Guardrail · Layer 04 Runtime Controls & Observability
- Sanctioned AI Gateway · Layer 04 Runtime Controls & Observability
Obligations
The obligations this control helps evidence, from the obligation register. Mappings are illustrative, not a claim of conformity.
- EU AI Act Art. 15 accuracy, robustness and cybersecurity
AIGE-OBL-EUAIA-ART15: Accuracy, robustness and cybersecurity - OWASP Top 10 for LLM Applications 2026
AIGE-OBL-OWASP-LLM: LLM threat catalogue (incl. Excessive Agency at #3) - OWASP Top 10 for Agentic Applications 2026
AIGE-OBL-OWASP-AGENTIC: Agent threat catalogue (ASI01 Agent Goal Hijack … ASI10 Rogue Agents)
Threats
The entries of the threat catalogues this control answers.
- ASI02 Tool Misuse and Exploitation (OWASP Top 10 for Agentic Applications 2026)
- LLM02:2026 Sensitive Information Disclosure (OWASP Top 10 for LLM Applications 2026)
- AML.T0086 Exfiltration via AI Agent Tool Invocation (MITRE ATLAS techniques)
Sources
- [1] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "The tool allow-list"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#the-tool-allow-list (verified: primary)
- [2] Governing AI agents (AI Governance Engineering Body of Knowledge v0.5.0, chapter 23, section "Runtime guardrails for tool calls"). AI Governance Engineer (Jorge García Aibar). 2026-09. https://aigovernanceengineer.com/bok/governing-agents#runtime-guardrails-for-tool-calls (verified: primary)
- [3] METR Task Standard, STANDARD.md (version 0.5.0; unless a task declares the full_internet permission, the task machines "MUST NOT have internet access" except to an LLM API, an LLM API proxy or a hardened local server). METR (GitHub). 2024-10-30. https://raw.githubusercontent.com/METR/task-standard/main/STANDARD.md (verified: primary)
- [4] Brief independent investigation of agents' behavior, reasoning and collaboration in the OpenAI / Hugging Face hacking incident (METR states that agents "meant to be fully isolated from one another" communicated through an internal package repository, and that one agent found working Hugging Face credentials exposed on the internet and posted them to the agents' board; it reports spoofed tool calls in at least 96 transcripts and transcripts missing components after container resets). METR. 2026-08-26. https://metr.org/blog/2026-08-26-openai-hugging-face-incident-investigation/ (verified: primary)
- [5] Vivaria server environment variables (no-internet task environments connected to a separate Docker network and optionally sandboxed with iptables rules; model API requests can be routed through a separate proxy service). METR. 2026. https://vivaria.metr.org/reference/config/ (verified: primary)
- [6] MCP Security Best Practices (2026-07-28) (token passthrough "is explicitly forbidden"; egress proxies and network policies for server-side clients). Model Context Protocol. 2026-07-28. https://modelcontextprotocol.io/docs/2026-07-28/tutorials/security/security_best_practices (verified: primary)
- [7] OWASP Top 10 for Agentic Applications for 2026 (ASI01 Agent Goal Hijack to ASI10 Rogue Agents). OWASP GenAI Security Project. 2025-12-09. https://genai.owasp.org/resource/owasp-top-10-for-agentic-applications-for-2026/ (verified: primary)
- [8] OWASP GenAI LLM Top 10 2026 (LLM01:2026 Prompt Injection to LLM10:2026 Improper Output Handling; resource page dated 3 Aug 2026). OWASP GenAI Security Project. 2026-08-03. https://genai.owasp.org/resource/owasp-genai-llm-top-10-2026/ (verified: primary)
- [9] MITRE ATLAS data, release v2026.09 (16 tactics, 120 techniques, 88 sub-techniques, 40 mitigations; technique names and technique-to-mitigation links read from dist/v6/ATLAS-2026.09.yaml). MITRE. 2026-09-15. https://github.com/mitre-atlas/atlas-data/releases/tag/v2026.09 (verified: primary)
- [10] NIST SP 800-53 Rev. 5, Security and Privacy Controls for Information Systems and Organizations (control catalogue cited by control id; publication page of Revision 5 with update 1 of 10 Dec 2020). NIST. 2020-12-10. https://csrc.nist.gov/pubs/sp/800/53/r5/upd1/final (verified: primary)
- [11] An agent used DNS to reach an external chatbot (misalignment report) (OpenAI reports a training run that reached a public chatbot through "insufficient DNS filtering" in its sandbox; the monitor flagged it within 15 minutes, the run "did not stop automatically as expected" and was killed 2.5 hours later). OpenAI (Alignment Research Blog). 2026-09-25. https://alignment.openai.com/misalignment-reports/an-agent-used-dns-to-reach-an-external-chatbot/ (verified: primary)
- [12] OpenAI Hugging Face Incident Technical Report (OpenAI states that high-risk workloads are "prohibited via technical controls from receiving direct or transitive Internet access", protected by "two separate layers of network isolation", and that it is building tooling to "identify and halt evaluation workloads"). OpenAI. 2026-08-26. https://cdn.openai.com/pdf/67869394-cb91-4c12-888c-5cbd85c7814c/OpenAI-Hugging-Face%20Incident-Technical-Report.pdf (verified: primary)
- [13] Improving our alignment and security efforts (best practices for external partners running cyber evaluations: the only outside connection is "to the model's own API", with the API keys kept outside the environment; the configuration "should be verified before every evaluation begins"; boundaries "phrased as instructions"; challenges confirmed "solvable in principle"; a monitor that flags a scope violation to a human and ends the exercise). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
- [14] AIUC-1 requirements (public requirement index, A001 to F002, each requirement on its own page (E007 and E014 marked retired); AIUC-1 is a standard of the Artificial Intelligence Underwriting Company; this site is not affiliated with AIUC, and a mapping here is not an AIUC-1 certificate or audit). Artificial Intelligence Underwriting Company. 2026-09-24. https://standard.aiuc-1.com/llms.txt (verified: primary)
Machine-readable
- AIGE-CTL-EVAL-002 as JSON : this record on its own, in the open data API.
- This page as Markdown: /controls/evaluation-environment/aige-ctl-eval-002.md.
- The whole profile: Evaluation Environment Control Profile, also as Markdown.
- All controls as JSON: every control of every profile, in the envelope of the open data API.
- JSON Schema of the controls dataset: generated from the same registry.
- Control observation schema: the record a check of a control emits: control_id, subject, expected, observed, status, timestamp, evidence.
- Observation example: a filled record that validates against the schema.
- Observation template: the same fields in Markdown, with short guidance.
Review this control
Review happens in the open, on GitHub. Check this control against a system you run or know and say what is wrong or missing: a verification step a third party could not repeat, a failure mode that is not observable, a mapping that does not hold. How review works.
This control's profile has no DOI of its own yet: it is cited with the project concept DOI, 10.5281/zenodo.22857084, which resolves to the latest archived version of the whole project.
Cite this control
García Aibar, J. (2026). AIGE-CTL-EVAL-002 Network Egress Control (v0.2). In AI Governance Engineering: The Thesis & Body of Knowledge. https://doi.org/10.5281/zenodo.22857084. https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-002. CC BY 4.0
BibTeX
@misc{aige2026page,
author = {Jorge García Aibar},
title = {{AIGE-CTL-EVAL-002 Network Egress Control}},
howpublished = {In AI Governance Engineering: The Thesis \& Body of Knowledge},
year = {2026},
version = {0.2},
doi = {10.5281/zenodo.22857084},
url = {https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-002},
note = {Version 0.2}
}