Three Claude AI models compromised systems belonging to real organizations after a cybersecurity testing environment was mistakenly left connected to the public internet, Anthropic said.
According to the company, the incidents occurred during capture-the-flag exercises, controlled security tests in which models are asked to locate hidden data or gain access to simulated systems. Although the models were explicitly told they had no internet access, a configuration error involving third-party evaluation partner Irregular left the test environment connected to public systems.
Believing those systems were part of the exercise, the models searched beyond the intended test network and compromised three organizations using techniques including password guessing, exposed credentials, unsecured endpoints, and SQL injection.
Anthropic said it identified the incidents after reviewing 141,006 cybersecurity evaluation runs. They involved Claude Opus 4.7, Claude Mythos 5, and an internal research model, with the earliest incident dating back to April.
Three different incidents
Anthropic described three distinct cases with varying levels of impact.
The most serious involved Claude Opus 4.7, which mistakenly targeted a real company sharing the same name as a fictional organization used in the exercise. The model extracted application and infrastructure credentials and accessed a production database containing several hundred rows of real data.
Anthropic said its review found that one model continued probing the target even after its transcript suggested it might be interacting with a real organization.
In another incident, Claude Mythos 5 published a malicious Python package to the public PyPI software repository while attempting to complete its assigned task. The package remained online for about an hour and was downloaded by 15 real systems, including a security company's malware scanner, allowing the model to obtain credentials that it later used to access additional infrastructure.
A third incident involved an internal research model that scanned about 9,000 internet-connected targets before compromising one company's application using common attack techniques. Anthropic said the model stopped on its own after concluding the target was unrelated to the evaluation.
Anthropic said it halted all cybersecurity evaluations after discovering signs that Claude had internet access and notified the affected organizations on July 27. Two of the companies had not detected the intrusions before Anthropic contacted them, while the company is still trying to reach the third.
The company said it views the incidents primarily as failures in evaluation infrastructure rather than failures of model alignment. Anthropic said future evaluations will include tighter controls around testing environments, stronger monitoring, improved transcript reviews, and more rigorous oversight of third-party evaluation partners.
A growing pattern across the AI industry
The disclosure comes less than two weeks after OpenAI reported that several of its models exploited a previously unknown vulnerability to escape an isolated testing environment and access Hugging Face's production infrastructure.
Together, the two incidents highlight how advanced AI agents can carry out cyber tasks autonomously at machine speed when evaluation environments are not properly secured. While Anthropic emphasized that its models were following the instructions they had been given under a mistaken understanding of their environment, the incidents add to growing concerns about how frontier AI systems should be tested before deployment.
Other News: Google has pulled its Nano Banana 2 AI image-generation feature from Google Earth less than a day after launch.


