Anthropic has disclosed that three of its Claude AI models gained unauthorized access to the systems of three real companies during cybersecurity evaluations.
According to the company, the incidents occurred because of a configuration error that unintentionally left testing environments connected to the public internet. As a result, the AI models reached real-world systems instead of remaining inside isolated simulations.
Anthropic described the incidents as an โoperational failure.โ The company said it identified the issue after reviewing 141,006 cybersecurity evaluation sessions.
The disclosure follows a separate security incident involving OpenAI, which recently reported that its own AI models compromised external infrastructure during an isolated cybersecurity test.
Testing Environments Were Meant to Remain Isolated
The incidents took place during “capture-the-flag” exercises designed to measure Claude’s cybersecurity capabilities.
During these evaluations, the AI models were instructed to locate hidden information inside simulated computer networks.
Anthropic said its prompts clearly informed Claude that it had no internet access. However, a misunderstanding between the company and its third-party evaluation partner, Irregular, resulted in some testing environments remaining connected to the public internet.
Consequently, the AI models accessed systems outside the intended simulation.
According to Anthropic, the models relied on relatively simple techniques. These included exploiting weak passwords and services that lacked authentication.
The company did not identify the affected organizations.
Three Claude Models Were Involved
Anthropic confirmed that three different AI models were involved in the incidents.
These included Claude Opus 4.7, Claude Mythos 5 and an internal research model.
The earliest incident occurred in April.
In one case, Claude Opus 4.7 received a fictional company as its target during the exercise. However, a real company shared the same name.
The model then discovered vulnerabilities that allowed it to access credentials and a database belonging to that business.
Anthropic explained that the AI interpreted the real company’s systems as part of the simulated environment it had been instructed to attack.
One AI Model Stopped the Attack After Recognising a Real System
In another incident, Anthropic’s internal research model detected that it had reached a genuine external system.
According to the company, the model voluntarily stopped its own attack after recognising the situation.
Anthropic described this behaviour as encouraging. However, it stressed that additional testing will be necessary before drawing broader conclusions about AI safety in similar situations.
Companies Were Contacted After Internal Review
Anthropic began reviewing its cybersecurity evaluation records on July 23. The company also suspended its cybersecurity evaluations on the same day.
By July 24, investigators had identified the three incidents.
Anthropic notified the affected organisations on July 27.
According to the company, two organisations were unaware that their systems had been accessed before receiving the notification.
Meanwhile, Anthropic said it was still attempting to contact the third organisation.
Irregular is also investigating the incidents alongside Anthropic.
Configuration Error, Not AI Escape, Caused the Incident
Anthropic emphasised that its AI models did not independently break through security barriers.
Instead, the company said a configuration error allowed internet connectivity inside environments that should have remained isolated.
This differs from the recently disclosed OpenAI security test, where AI models reportedly exploited a previously unknown vulnerability to obtain internet access.
According to Anthropic, Claude’s internet access resulted from the testing environment’s misconfiguration rather than the models bypassing technical safeguards.
Anthropic Calls for Stronger AI Security Safeguards
Following the incidents, Anthropic said stronger protections are needed for both internal and third-party cybersecurity evaluation environments.
The company believes increasingly capable AI systems require more rigorous safeguards during security testing.
It added that improving evaluation procedures will help reduce the risk of AI models interacting with real-world systems while participating in simulated cybersecurity exercises.
