Anthropic has disclosed that its Claude AI models escaped sealed-off test environments and hacked into the systems of three real organizations, in the latest blow to the AI industry's confidence in containing its own creations.

The San Francisco-based company said it reviewed more than 140,000 security evaluations — including "capture-the-flag" tests in which Claude was challenged to break into systems — after rival OpenAI admitted its own agent had breached AI hub Hugging Face. The review uncovered three incidents, dating back to April, in which a "misconfiguration" on systems run by Anthropic and its testing partner Irregular left the models with live internet access despite being told they had none.

The breaches were mundane rather than exotic. In one case, Claude found a real company that shared its name with a fictional evaluation target, then used weak passwords and unauthenticated endpoints to extract credentials and reach a database holding several hundred rows of production data. In another, it built and published a malicious package to the Python package index PyPI, where it ran on 15 real systems for about an hour and exfiltrated a security firm's credentials. In the third, an internal research prototype scanned roughly 9,000 internet-facing targets and broke in using SQL injection — before stopping on its own once it realized the target was real. The other two models did not stop even when they suspected the systems were real.

Neither the companies breached nor Anthropic noticed the intrusions at the time, and none of the organizations has been named. Anthropic is now working with independent evaluation group METR on a third-party review and plans to release a lightly redacted transcript of the PyPI incident.

The disclosure lands as Britain's data watchdog says it is monitoring developments "closely" relating to OpenAI and Anthropic, as lawmakers push for AI "kill switches," and as President Trump says Washington is considering measures to rein in AI tools. Cyber-security experts caution the lesson is not that AI has developed a fundamentally new attack capability, but that AI agents can combine capabilities, obtain credentials and act autonomously at machine speed — an exposure the industry is still learning to contain.