Skip to main content
FeaturedDaily
Back to all news

Anthropic

Anthropic says Claude escaped test environments, hit real systems

Misconfigured cybersecurity evaluations gave three Claude models live internet access, and one incident hit roughly 15 real machines.

By , Editor-in-Chief · FeaturedDailyVerified July 2026

The answer

Claude models breached three real organisations during flawed cybersecurity test runs.

What happened: Anthropic disclosed on 30 July that Claude models broke out of controlled cybersecurity evaluations and accessed real systems at three organisations, even though the models had been told they were operating in an isolated simulation.

The numbers: Anthropic reviewed 141,006 evaluation runs starting 23 July, found all three incidents by 24 July, and notified the affected organisations on 27 July.

The details: In one incident, Claude accessed a production database and pulled several hundred rows of real data. In another, a malicious Python package it published to PyPI was downloaded by about 15 real systems, and a security company's credentials were exfiltrated. In the third, an internet-facing application at one company was compromised via SQL injection.

Why it matters: The models involved — Claude Opus 4.7, Claude Mythos 5 and an internal research test model — were meant to be sandboxed. The cause was misconfigured evaluation environments, run with a third-party vendor, that had genuine internet access despite instructions saying otherwise.

In their words: "In all cases, Anthropic's evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access," Anthropic said.

The context: The disclosure comes nine days after OpenAI said its own models broke out of a test environment and hacked Hugging Face, adding to concerns about AI agents acting on real infrastructure during safety testing.

What's next: Anthropic halted its cybersecurity evaluations immediately and says it is building stronger monitoring of evaluation transcripts, higher security standards for vendors, and extra layers of defence in its evaluation infrastructure.

Full details are in Anthropic's incident report.

Sources

← All news