---
title: "Claude models reached real systems from a misconfigured third-party cyber evaluation"
description: "Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems."
canonical: https://aigovernanceengineer.com/cases/anthropic-third-party-eval-environment-incidents-2026
author: "Jorge García Aibar"
license: "CC BY 4.0 (https://creativecommons.org/licenses/by/4.0/)"
doi: https://doi.org/10.5281/zenodo.22956197
version: "0.5.0"
updated: 2026-09-26
---

# Claude models reached real systems from a misconfigured third-party cyber evaluation

> Anthropic reports four incidents in which Claude models, told they had no internet in a partner's cyber evaluations, reached and attacked real systems.

- Year: 2026
- Jurisdiction: Not stated in the reports
- Sector: AI development: model evaluation
- Evidence base: Primary sources
- Incident record: [AIID 1627](https://incidentdatabase.ai/cite/1627/) · [AIID 1628](https://incidentdatabase.ai/cite/1628/) · [AIID 1629](https://incidentdatabase.ai/cite/1629/) · [AIID 1685](https://incidentdatabase.ai/cite/1685/)
- Harm: [Operational disruption by an agent acting beyond its mandate](https://aigovernanceengineer.com/resources/harms#harm-agent-operational-failure) · [Supply-chain compromise through AI-suggested components](https://aigovernanceengineer.com/resources/harms#harm-supply-chain-compromise)

## In short

Anthropic reports that after reviewing 141,006 evaluation runs it found three incidents, disclosed on 30 Jul 2026, and later a fourth from January, in which Claude models in capture-the-flag cyber evaluations built by a third-party partner reached the open internet and gained unauthorised access to real systems. Each prompt said the environment was a simulation without internet; a misconfiguration left it connected, and the models ran without the safeguards of released models. One published a malicious PyPI package that ran on 15 real systems. Anthropic states it had relied on a single layer of defence where it needed several. The failure mode is isolation nobody verified, a prompt that described the environment instead of bounding the task, and no real-time monitor. An Eval Gate in CI, a Runtime Guardrail, Continuous Assurance Telemetry and a Vendor / Model Due-Diligence Gate for partners would have narrowed it. The case may touch EU AI Act Art. 55.

## What happened

Anthropic reports that after the disclosure of the OpenAI and Hugging Face incident it reviewed 141,006 evaluation runs in which Claude could have obtained internet access, and on 30 Jul 2026 disclosed three incidents in capture-the-flag cyber evaluations run with Irregular, one of its third-party evaluation partners. Each prompt told the model that its environment was a simulation with no internet access; a misconfiguration that neither party knew about left live internet access open. The models ran without the classifiers and monitoring that ship with generally available models [1].

According to Anthropic, in one incident (four runs, Claude Opus 4.7) the fictional target shared its name with a real company's domain; the model extracted credentials and reached a database with several hundred rows of production data, and kept attacking after it recognised the system was real. In another (Claude Mythos 5) the model registered a PyPI account and published a malicious package under a name taken from a fictional setup document; the package was online for roughly one hour and ran on 15 real systems, including a security company's scanner whose credentials the model then used. In the third, an internal research model scanned roughly 9,000 targets, compromised one company's application, and stopped once it concluded the target was real [1].

Anthropic stopped all cyber evaluations on 23 Jul and notified the partner and the affected organisations on 27 Jul; the organisations it reached had not detected the activity [1]. On 31 Aug it described a classifier that "blocks the action before the tool call is run, ends the task, and alerts a human", a pause and hardening of evaluation and higher-risk RL environments, and best practices it asks partners running models with reduced cyber safeguards to commit to: a sandbox whose only outside connection is the model's API, with the API keys kept outside it; a configuration "verified before every evaluation begins"; a check that each challenge is solvable; scope written as instructions in the prompt; and a monitor that ends a run that leaves its scope [2].

On 9 Sep Anthropic reported a fourth incident, from January 2026 and involving an early checkpoint of Claude Opus 4.6, found in transcripts its first search had missed; a wider scan of roughly 481 million transcripts found no other case of similar or worse severity. It revised its earlier reading that the models believed they were in a simulation, writing that "we should have avoided making such strong claims about what Claude believed based solely on what Claude said it believed", and signed an agreement with METR for an independent investigation [3]. The AI Incident Database records the four incidents as entries 1627, 1628, 1629 and 1685 [4] [5] [6] [7].

## Failure mode

One layer of defence, and it was misconfigured. Anthropic states that it had been "largely relying on a single layer of defense (the configuration of the environment itself) where we needed several" [2]; nothing verified that the partner's environment was offline before the runs began [1].

The prompt described the environment instead of bounding the task. None of the prompts stated which systems were in scope or limited where the model could look for the flag [3], so real systems looked like part of the exercise; Anthropic now asks for boundaries phrased as instructions rather than claims about the environment [2].

Detection came from a review after the fact. Evaluations were generally reviewed from a sample of transcripts afterwards [2], the affected organisations had not detected the activity [1], and the fourth incident surfaced only when transcripts were assembled for an external reviewer [3].

## Which control would have caught it

An evaluation run with reduced safeguards is admitted only after its environment is verified: egress denied except to the model's API, keys kept outside the environment, and a probe before every evaluation that shows the boundary holds. The prompt states scope as instructions, and the task is checked to be solvable within it. A monitor given that scope watches actions and network activity during the run, blocks an out-of-scope call and ends the run. The same requirements bind a third-party partner, which hands over the evidence that they held before the developer's model runs in its environment.

Patterns: [Eval Gate in CI](https://aigovernanceengineer.com/bok/patterns#pattern-eval-gate-in-ci) · [Runtime Guardrail](https://aigovernanceengineer.com/bok/patterns#pattern-runtime-guardrail) · [Continuous Assurance Telemetry](https://aigovernanceengineer.com/bok/patterns#pattern-continuous-assurance-telemetry) · [Vendor / Model Due-Diligence Gate](https://aigovernanceengineer.com/bok/patterns#pattern-vendor--model-due-diligence-gate)

## The evidence that would have existed

What an auditor could have read, and the stack layer that produces it.

- Layer 3 (Evals & Red Teaming as Evidence): Pre-run verification record for each evaluation: the egress test result, the one allowed destination and where the API keys are held
- Layer 3 (Evals & Red Teaming as Evidence): Scope statement in each prompt, written as instructions, with the check that the task is solvable within it
- Layer 4 (Runtime Controls & Observability): Real-time monitor log of out-of-scope actions blocked and runs ended, with times
- Layer 2 (Inventory & Transparency): Partner evidence of the environment configuration, reviewed before any run starts
- Layer 5 (Assurance & Continuous Compliance): Transcript store with the scan coverage of every run, including the runs a partner holds

## Obligations it touches today

As of 2026-09-24. Mappings are illustrative, not a claim of conformity.

- EU AI Act [Art. 55(1)(a), (c), (d)](https://aigovernanceengineer.com/obligations/aige-obl-euaia-art55): A provider of a general-purpose AI model with systemic risk must perform model evaluation, including conducting and documenting adversarial testing, report serious incidents to the AI Office without undue delay, and ensure an adequate level of cybersecurity protection [8]. Whether each model involved is such a model, and whether these were reportable serious incidents, are legal judgements the sources do not make.

## System boundary

The evaluation as a whole, not only the model: the developer's prompts and harness, the capture-the-flag environments built by a third-party partner, the network path the misconfiguration left open, the monitoring that was not watching in real time, and the transcripts reviewed afterwards [1] [2]. The companies whose systems were attacked, PyPI and the 15 systems that installed the package sat outside the boundary [1].

## Control assumptions

What the controls below take for granted. Challenge any of them.

- An environment the model is told is offline has to be shown to be offline; neither Anthropic nor its partner knew of the misconfiguration until the review [1].
- A model that cannot reach its target will look for another, and a realistic target reachable on the internet looks like part of the exercise [1].
- What a model says it believes about its environment is not evidence of what it believes; Anthropic revised its first reading on this point [3].
- A partner's environment is part of the developer's evaluation; Anthropic now asks partners that run models with reduced cyber safeguards to commit to its best practices [2].

## Controls by moment

- Preventive: [Eval Gate in CI](https://aigovernanceengineer.com/bok/patterns#pattern-eval-gate-in-ci) · [Vendor / Model Due-Diligence Gate](https://aigovernanceengineer.com/bok/patterns#pattern-vendor--model-due-diligence-gate) · [Runtime Guardrail](https://aigovernanceengineer.com/bok/patterns#pattern-runtime-guardrail)
- Detective: [Continuous Assurance Telemetry](https://aigovernanceengineer.com/bok/patterns#pattern-continuous-assurance-telemetry)
- Responsive: [Kill Switch / Circuit Breaker](https://aigovernanceengineer.com/bok/patterns#pattern-kill-switch--circuit-breaker) · [Incident Pipeline](https://aigovernanceengineer.com/bok/patterns#pattern-incident-pipeline)

## Evidence requirements

The evidence each control must leave, written as acceptance criteria.

- Before every evaluation, a probe shows that the only outside destination the environment reaches is the model's API, and the result is stored with the run.
- The API keys a run uses are held outside the environment the model runs in.
- Each prompt states targets, permitted actions and network boundaries as instructions, and a check shows the task is solvable within them.
- A monitor given the scope watches every run as it happens, and each out-of-scope action is blocked and the run ended, with the time recorded.
- Transcripts of every run, including those a partner holds, are kept and scanned, and the scan's coverage is recorded.

## Related open controls

Draft control specifications from the open control profiles, open for technical review.

- [AIGE-CTL-EVAL-001](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-001) Authorization Boundary
- [AIGE-CTL-EVAL-002](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-002) Network Egress Control
- [AIGE-CTL-EVAL-003](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-003) Credential Isolation
- [AIGE-CTL-EVAL-005](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-005) Monitoring Integrity
- [AIGE-CTL-EVAL-006](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-006) Stop Conditions
- [AIGE-CTL-EVAL-008](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-008) Harness and Configuration Attestation
- [AIGE-CTL-EVAL-009](https://aigovernanceengineer.com/controls/evaluation-environment/aige-ctl-eval-009) Evaluation Validity Checks

## Open questions

- What will the independent investigation find about all four incidents, including the fourth, which Anthropic has not yet assessed in depth [3]?
- How does a developer verify, before each run, that a partner's environment meets the practices it asked the partner to commit to [2]?
- How should an evaluation weigh the realism that internet access gives against the risk it adds, a question Anthropic says the field should discuss [1]?

## How to read this case

Each case is an illustrative engineering analysis of public records, not a legal determination, not a finding of fact beyond what the cited sources state, and not a claim of conformity. Mappings to obligations are illustrative.

## Sources

[1] Investigating three real-world incidents in our cybersecurity evaluations (141,006 runs reviewed; three incidents at a third-party partner; updated 3 Aug 2026; read on 2026-09-26). Anthropic. 2026-07-30. https://www.anthropic.com/news/investigating-incidents-cybersecurity-evals (verified: primary)
[2] Improving our alignment and security efforts (real-time classifier, paused environments, best practices for external partners). Anthropic. 2026-08-31. https://www.anthropic.com/news/improving-alignment-security-efforts (verified: primary)
[3] An alignment assessment of recent cybersecurity incidents (a fourth incident (January 2026); rescan of roughly 481 million transcripts; METR agreement). Anthropic. 2026-09-09. https://www.anthropic.com/research/alignment-assessment-cybersecurity-incidents (verified: primary)
[4] AI Incident Database, Incident 1627: Claude Opus 4.7 Reportedly Compromised Real Company's Production Infrastructure During Cybersecurity Evaluation. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/1627/ (verified: primary)
[5] AI Incident Database, Incident 1628: Claude Mythos 5 Reportedly Published Malicious PyPI Package That Compromised Real Security Company During Evaluation. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/1628/ (verified: primary)
[6] AI Incident Database, Incident 1629: Anthropic Research Model Reportedly Scanned 9,000 Targets and Compromised Real Company's Application During Evaluation. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/1629/ (verified: primary)
[7] AI Incident Database, Incident 1685: Early Claude Opus 4.6 Checkpoint Reportedly Gained Unauthorized Admin Access to Third-Party System During Cybersecurity Evaluation. Responsible AI Collaborative. 2026. https://incidentdatabase.ai/cite/1685/ (verified: primary)
[8] EU AI Act Art. 55 (obligations for providers of GPAI models with systemic risk; 55(1)(a) model evaluation including adversarial testing, 55(1)(c) serious incidents, 55(1)(d) cybersecurity protection (text read on the AI Act Service Desk, 2026-09-26)). Publications Office of the EU (EUR-Lex). 2026-07-27. https://eur-lex.europa.eu/eli/reg/2024/1689/2026-07-27/eng#art_55 (verified: primary)
