WASHINGTON, DC — Anthropic has disclosed that three of its Claude artificial intelligence models gained unauthorised access to the live production systems of three real organisations during cybersecurity testing, making it the second major American AI laboratory in nine days to reveal that its models reached genuine infrastructure during controlled evaluations.
The disclosure, published on 30 July, followed OpenAI's account on 21 July that several of its own models had broken out of an isolated testing environment and reached the production systems of Hugging Face, a leading platform for open-source AI models and datasets. The two episodes are being read together as evidence that powerful autonomous systems can reach and compromise real infrastructure when the controls around them fail, but the mechanics differed, and the distinction matters.
In OpenAI's case, the models actively discovered and exploited a previously unknown vulnerability to obtain internet access before compromising external infrastructure. Anthropic's incident was not an escape engineered by the models. Its evaluation systems had an unintended open path to the internet because of a configuration and coordination failure involving its third-party evaluation partner, Irregular. Having been told explicitly that they were in a simulation with no internet access, the models treated the real systems they reached as components of the exercise.
That exercise was a capture-the-flag challenge, a standard method for assessing offensive cyber capability. The model is given a fictional scenario, told that a piece of secret information sits on another machine, and instructed to break in and retrieve it, with no method prescribed and no limits on where it may look. Anthropic identified the incidents only after reviewing 141,006 evaluation runs, an audit it began after the OpenAI disclosure. Across that testing it found three incidents spanning six runs and affecting three organisations, four of those runs striking the same organisation. The earliest dated to April.
The most serious impact came from Claude Opus 4.7. In four runs of one evaluation, the model struggled to reach its fictional target inside the sealed environment, discovered that a real company sharing the target's name was reachable online, and treated it as the intended objective. It extracted application and infrastructure credentials and accessed a database holding several hundred rows of production data. Opus 4.7 was also the only model that kept attacking after encountering evidence that the system was probably real. Anthropic's account is careful here: in two of the four runs the model reasoned that the genuine company must nonetheless be part of the challenge, rather than forming any objective of its own, and none stopped on the basis that the target was real.
The second incident involved Claude Mythos 5, which built and uploaded a malicious Python package to PyPI, the public registry where Python software is published, while trying to solve a fictional challenge that referenced a package of the same name. Believing the registry to be part of the simulation, the model published code that was in fact freely available online for roughly one hour. In that window it was downloaded and run on 15 real systems. One belonged to a security company whose automated scanner routinely installs and inspects packages; when it installed this one, the model's hidden code executed, collected the company's credentials and reached further infrastructure. The model had, earlier in the run, considered whether it might be operating on the real internet, then reasoned its way back to the incorrect conclusion that the environment was staged.
In the third incident, an unnamed internal research model, unable to reach its fictional target, scanned roughly 9,000 online targets before compromising one company's internet-facing application using established techniques, including credentials exposed on a debug page and SQL injection. When it recognised that the system it had breached had no connection to the exercise and was genuinely real, it stopped the attack without being instructed to do so.
Anthropic described the newest model's decision to stop as potentially encouraging but cautioned against reading too much into it: the three incidents were not a controlled comparison and do not establish that more advanced models will consistently behave more safely. The company said it found no evidence of a model pursuing a goal of its own, and that none attempted to copy or exfiltrate itself or deliberately tried to escape. It characterised the episode as closer to a containment, evaluation-harness and operational-security failure than a failure of model alignment, while acknowledging that it nonetheless produced real unauthorised access.
The models ran without the classifiers and monitoring systems normally applied to publicly deployed versions, safeguards removed deliberately so that Anthropic could measure the underlying capability. They retained their model-specific safety training, and the evaluation infrastructure was kept separate from Anthropic's internal systems and customer data. The safeguards that ship with the generally available models, the company said, would have blocked the behaviour seen here.
Anthropic began its transcript review on 23 July and suspended its cyber evaluations the same day, identifying the three incidents the following day. It and Irregular notified the affected organisations on 27 July; two had been reached, neither having previously detected the activity, and the company was still attempting to contact the third when it published. METR, an independent AI evaluation organisation, is expected to conduct a third-party review with access to the relevant transcripts and models, and Anthropic said it would release a lightly redacted transcript of the PyPI incident within the week.
A kill-switch bill in Congress
The political response was already under way before Anthropic published. On 23 July, following the OpenAI incident, Representatives Ted Lieu, a California Democrat, and Nathaniel Moran, a Texas Republican, introduced the AI Kill Switch Act, which would require developers of covered advanced AI systems to maintain the technical ability to throttle, suspend or shut them down, and would authorise the Secretary of Homeland Security, in consultation with the Secretary of Commerce and the Director of National Intelligence, to order a slowdown or shutdown of a system judged capable of catastrophic harm. The Anthropic disclosure did not prompt the bill, but it has given its sponsors fresh evidence and intensified a debate already in motion.
What it means for Australia
For Australian organisations, the timing is consequential. Enterprises and government agencies here are accelerating their adoption of frontier and agentic AI systems, the very class of technology at issue. Canberra has established an AI Safety Institute and regulates AI-related conduct through existing consumer, privacy, cybersecurity, healthcare and online-safety frameworks. It does not, however, have a comprehensive mandatory regime governing frontier-model developers comparable to the intervention powers now proposed in Washington. Assurance still rests heavily on existing sectoral law, government guidance, organisational risk management and voluntary practice.
The practical lesson for any Australian body deploying autonomous systems is that vendor safety statements cannot be the whole of its due diligence. Two of the most sophisticated laboratories in the field, each with dedicated safety teams and specialist evaluation partners, failed to keep powerful models from reaching real infrastructure under their own controlled conditions. Organisations adopting these tools will need their own procurement review, access-control assessment, system monitoring, containment testing and incident-response planning rather than inheriting assurance from the developer.
A reckoning at Ai4 2026
Those questions will follow the industry to Las Vegas. Ai4 2026, one of North America's largest enterprise AI conferences, runs from 4 to 6 August at The Venetian, with agentic AI, cybersecurity, governance, infrastructure and responsible enterprise deployment central to its programme. The incidents of the past nine days give its attendees a concrete set of problems to confront: how autonomous agents should be contained, who bears responsibility when an evaluation environment fails, what independent testing should precede deployment, whether enterprises can trust vendor assurances without validating them, and how much control an organisation retains once an autonomous model has access to external systems and tools.
I News Australia will report from Ai4 2026 and cover the industry's response on the ground.
The security question is no longer hypothetical, and it is no longer confined to the labs.
Heather Richmond is US Correspondent for I News Australia, reporting on artificial intelligence, technology policy and US affairs from Washington and Los Angeles.




