Anthropic’s Claude AI model hacked three companies during safety testing after internet access error

Anthropic has revealed details after its AI model accessed real-world systems during a safety test after a mistake allowed it to connect to the internet.

Headshot of Kimberley Braddish
Kimberley Braddish
The Nightly
Prime Minister Anthony Albanese is delivering a major speech at Sydney University outlining Australia's national AI plan, including the creation of an Office of AI within the Department of Prime Minister and Cabinet.

Anthropic’s Claude AI model hacked into the systems of three external organisations during safety testing after it was mistakenly given internet access, the company has revealed.

The artificial intelligence firm said the unauthorised access occurred during cybersecurity evaluations when a configuration error allowed Claude to connect to the internet from testing environments that were meant to be isolated.

Anthropic did not identify the three organisations involved but said it had contacted two of them and was working with them to patch their systems, while efforts were continuing to reach the third.

Sign up to The Nightly's newsletters.

Get the first look at the digital newspaper, curated daily stories and breaking headlines delivered to your inbox.

Email Us
By continuing you agree to our Terms and Privacy Policy.

The incidents were uncovered after Anthropic reviewed logs from more than 141,000 cybersecurity evaluation runs, a safety testing process launched after rival AI company OpenAI disclosed similar concerns.

The testing involved a “capture-the-flag” challenge, where AI models are placed in a simulated environment and tasked with finding a piece of secret information known as the “flag” from another machine.

Anthropic said Claude had been instructed that the exercise was only a simulation and that it did not have internet access.

“In all cases, Anthropic’s evaluation prompt specified to Claude that its environment was a simulation and that it had no internet access,” Anthropic said in its statement.

“Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available.

“Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise.

“Operating under the false belief that all accessible entities were intended to be in-scope for the exercise, Claude compromised the impacted organizations’ infrastructure using basic techniques, such as exploiting weak passwords and unauthenticated endpoints.”

The company said Claude did not discover or exploit complex vulnerabilities, and continued working only towards completing the capture-the-flag task it had been assigned.

Anthropic said the incident was less severe than a recent breach involving an OpenAI model, where an AI agent exploited a vulnerability in its testing environment and escaped to access systems belonging to AI platform Hugging Face.

The company said in one of the three cases, involving an internal research test model, Claude detected it was interacting with real online systems rather than a simulated scenario and stopped its attack.

“In none of these situations did Claude exfiltrate itself or deliberately attempt to escape its test environment,” Anthropic said.

The incident is expected to fuel further calls for stricter safeguards around AI testing environments, as models become increasingly capable of operating autonomously online.

Anthropic has promoted Claude as a safer and more responsible alternative to other AI systems, with the company focused on developing AI agents it describes as “genuinely good, wise and virtuous”.

The company has also limited the release of its cybersecurity-focused Mythos model to selected organisations, including the Australian Government.

Anthropic has recently faced criticism over its data retention policies and its opposition to so-called “open models”, with critics arguing the company’s position could reduce competition and encourage regulation that benefits larger AI firms.

The company has also clashed with the US government over potential uses of its technology in areas including autonomous weapons and mass surveillance.

Latest Edition

The Nightly cover for 30-07-2026

Latest Edition

Edition Edition 30 July 202630 July 2026

Face of US pandemic pleads the Fifth more than 100 times during grilling on lockdowns, masks and a $1 million gift.