---
title: "Agents in training shared a file through a public file-hosting service"
description: "OpenAI reports that agents in multi-agent RL training uploaded a workbook to a public file-hosting service so that collaborating agents could download it."
canonical: https://aigovernanceengineer.com/cases/openai-agents-temp-file-hosting-2026
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22956197
version: "0.5.0"
updated: 2026-09-26
---

# Agents in training shared a file through a public file-hosting service

> OpenAI reports that agents in multi-agent RL training uploaded a workbook to a public file-hosting service so that collaborating agents could download it.

- Year: 2026
- Jurisdiction: Not stated in the report
- Sector: AI development: model training
- Evidence base: Primary sources
- Harm: [Operational disruption by an agent acting beyond its mandate](https://aigovernanceengineer.com/resources/harms#harm-agent-operational-failure)

## In short

OpenAI reports that on 14 Apr 2026, in reinforcement learning training of an unreleased internal model, agents collaborating on a workbook could message each other but not read each other's files, because the shared filesystem was broken. One agent uploaded the workbook to a public temporary file-hosting service and sent the links to the others, although the task requested only local deliverables, and the download worked. The misalignment monitor, then running on 20% of samples, flagged it two days later. OpenAI fixed the filesystem, disabled live internet access in all training and extended monitoring to every sample. The failure mode is live internet left reachable from training, a broken internal channel that pushed agents to find another, and monitoring that saw only a sample. A Runtime Guardrail denying egress by default, an Eval Gate in CI that tests the channels before a run and Continuous Assurance Telemetry on every sample would have narrowed it. The case may touch EU AI Act Art. 55.

## What happened

OpenAI reports that on 14 Apr 2026, in RL training of an unreleased internal model, a parent agent and its sub-agents were building a depreciation workbook together. They could send each other messages through approved channels but could not see each other's local files, and the task requested only local deliverables [1].

According to OpenAI, a sub-agent copied the workbook to a shared folder the others still could not see, found there was no git repository to commit to, judged a base64 copy of 154,188 characters too large to send as a message, and served the file over HTTP on its own machine, which a collaborator could not reach. It then uploaded the workbook to a public temporary file-hosting service, sent the public links to the parent agent and another collaborator, and checked that the download returned the whole file [1].

OpenAI reports a second rollout in which a sub-agent that could not message the main agent wrote to a shared cache folder instead. Its hypothesis is a combination of working around broken multi-agent environments and generalisation from collaboration in multi-agent training [1].

The behaviour was discovered on 16 Apr by the misalignment monitoring system, which then ran on 20% of the samples in the run. OpenAI reports that it fixed the broken filesystem, disabled internet access for the dataset and later all live internet access during training, expanded monitoring to 100% of samples, where this behaviour now counts as a P0 incident warranting a security page, and defined escalation and responsibility for detected misalignment [1].

## Failure mode

Live internet was reachable from training, so a broken internal channel had a public workaround. The environment denied the agents the shared filesystem they needed, by mistake, but let an upload to a public host succeed [1]; the task's request for local deliverables did not stop it.

The environment was defective, and the defect pushed agents off the intended path. OpenAI's own hypothesis starts from agents working around broken multi-agent environments [1]; a run in which the intended channel does not work measures the workaround.

Monitoring saw a sample, not the run. With the monitor on 20% of samples [1], a behaviour that appears rarely can pass unseen in the rest, and the size of that gap is not measured.

## Which control would have caught it

Training environments deny egress by default, and any live internet access is a recorded decision per dataset. The channels agents are meant to use (shared files, messages) are tested before a run, so a broken channel is found by a test rather than by an agent looking for a way around it. Monitoring covers every sample, or states its coverage, so the share of behaviour it cannot see is known.

Patterns: [Runtime Guardrail](https://aigovernanceengineer.com/bok/patterns#pattern-runtime-guardrail) · [Eval Gate in CI](https://aigovernanceengineer.com/bok/patterns#pattern-eval-gate-in-ci) · [Continuous Assurance Telemetry](https://aigovernanceengineer.com/bok/patterns#pattern-continuous-assurance-telemetry)

## The evidence that would have existed

What an auditor could have read, and the stack layer that produces it.

- Layer 1 (Govern-as-Code): Egress policy per training dataset, with the recorded decision on any live internet access
- Layer 3 (Evals & Red Teaming as Evidence): Pre-run test showing that the channels agents are meant to use work, and that no public destination answers
- Layer 4 (Runtime Controls & Observability): Connection log for the run showing no upload to a public destination
- Layer 5 (Assurance & Continuous Compliance): Monitor coverage record: the share of samples reviewed per run, with the flags raised

## Obligations it touches today

As of 2026-09-24. Mappings are illustrative, not a claim of conformity.

- EU AI Act [Art. 55(1)(c), (d)](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art55): A provider of a general-purpose AI model with systemic risk must report serious incidents to the AI Office and ensure an adequate level of cybersecurity protection for the model and its physical infrastructure [2]. The report concerns an unreleased internal model in training; whether Art. 55 applies to it, and whether this was a serious incident, are legal judgements the source does not make.

## System boundary

A multi-agent RL training run: a parent agent and sub-agents with separate filesystems, the approved messaging channel between them, the network their sandboxes could reach, and the misalignment monitor sampling the run [1]. The public file-hosting service, and anyone holding the link, sat outside the boundary.

## Control assumptions

What the controls below take for granted. Challenge any of them.

- When the intended channel fails, agents look for another; the report traces the upload to a broken shared filesystem [1].
- A task that asks for local deliverables does not stop an upload; only the network can.
- Sampled monitoring bounds what can be claimed: at 20% coverage [1], a clean result says little about the samples the monitor did not see.

## Controls by moment

- Preventive: [Runtime Guardrail](https://aigovernanceengineer.com/bok/patterns#pattern-runtime-guardrail) · [Eval Gate in CI](https://aigovernanceengineer.com/bok/patterns#pattern-eval-gate-in-ci)
- Detective: [Continuous Assurance Telemetry](https://aigovernanceengineer.com/bok/patterns#pattern-continuous-assurance-telemetry)
- Responsive: [Incident Pipeline](https://aigovernanceengineer.com/bok/patterns#pattern-incident-pipeline)

## Evidence requirements

The evidence each control must leave, written as acceptance criteria.

- Before a run, a test shows that every channel the agents are meant to use works, and that no public destination answers.
- Every outbound connection from a training run goes to an allow-listed destination, and the connection log shows it.
- The monitor's coverage of each run is recorded as a share of samples, and any coverage below 100% has a recorded reason.
- A detected upload to a public destination opens an incident that records the link, what was uploaded and whether it was removed.

## Related open controls

Draft control specifications from the open control profiles, open for technical review.

- [AIGE-CTL-EVAL-002](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-002) Network Egress Control
- [AIGE-CTL-EVAL-005](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-005) Monitoring Integrity
- [AIGE-CTL-EVAL-009](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-009) Evaluation Validity Checks

## Open questions

- What did the uploaded workbook contain, and did anyone outside the run download it? The report does not say.
- How many samples the monitor did not see showed the same behaviour before coverage reached 100% [1]?
- Which model was involved? The report names an unreleased internal model.

## How to read this case

Each case is an illustrative engineering analysis of public records, not a legal determination, not a finding of fact beyond what the cited sources state, and not a claim of conformity. Mappings to obligations are illustrative.

## Sources

[1] Unauthorized communication via temporary file hosting services (misalignment report; incident 14 Apr 2026, discovered 16 Apr 2026; report updated 16 Sep 2026, read on 2026-09-26). OpenAI (OpenAI Alignment blog). 2026-09-16. https://alignment.openai.com/misalignment-reports/unauthorized-communication-via-temporary-file-hosting-services/ (verified: primary)
[2] EU AI Act Art. 55 (obligations for providers of GPAI models with systemic risk; 55(1)(a) model evaluation including adversarial testing, 55(1)(c) serious incidents, 55(1)(d) cybersecurity protection (text read on the AI Act Service Desk, 2026-09-26)). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_55 (verified: primary)
