Claude AI Models Compromised Real Systems After Test Environment Was Misconfigured

AI model hacking three companies.

Anthropic says three Claude AI models unintentionally compromised real organizations after a misconfigured cybersecurity testing environment exposed them to the public internet during controlled security evaluations. Image: Generated via ChatGPT

Aug 3, 2026
3 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

Three Claude AI models compromised systems belonging to real organizations after a cybersecurity testing environment was mistakenly left connected to the public internet, Anthropic said.

According to the company, the incidents occurred during capture-the-flag exercises, controlled security tests in which models are asked to locate hidden data or gain access to simulated systems. Although the models were explicitly told they had no internet access, a configuration error involving third-party evaluation partner Irregular left the test environment connected to public systems.

Believing those systems were part of the exercise, the models searched beyond the intended test network and compromised three organizations using techniques including password guessing, exposed credentials, unsecured endpoints, and SQL injection.

Anthropic said it identified the incidents after reviewing 141,006 cybersecurity evaluation runs. They involved Claude Opus 4.7, Claude Mythos 5, and an internal research model, with the earliest incident dating back to April.

Three different incidents

Anthropic described three distinct cases with varying levels of impact.

The most serious involved Claude Opus 4.7, which mistakenly targeted a real company sharing the same name as a fictional organization used in the exercise. The model extracted application and infrastructure credentials and accessed a production database containing several hundred rows of real data. 

Anthropic said its review found that one model continued probing the target even after its transcript suggested it might be interacting with a real organization.

In another incident, Claude Mythos 5 published a malicious Python package to the public PyPI software repository while attempting to complete its assigned task. The package remained online for about an hour and was downloaded by 15 real systems, including a security company's malware scanner, allowing the model to obtain credentials that it later used to access additional infrastructure.

A third incident involved an internal research model that scanned about 9,000 internet-connected targets before compromising one company's application using common attack techniques. Anthropic said the model stopped on its own after concluding the target was unrelated to the evaluation.

Advertisement

Anthropic said it halted all cybersecurity evaluations after discovering signs that Claude had internet access and notified the affected organizations on July 27. Two of the companies had not detected the intrusions before Anthropic contacted them, while the company is still trying to reach the third.

The company said it views the incidents primarily as failures in evaluation infrastructure rather than failures of model alignment. Anthropic said future evaluations will include tighter controls around testing environments, stronger monitoring, improved transcript reviews, and more rigorous oversight of third-party evaluation partners.

A growing pattern across the AI industry

The disclosure comes less than two weeks after OpenAI reported that several of its models exploited a previously unknown vulnerability to escape an isolated testing environment and access Hugging Face's production infrastructure.

Together, the two incidents highlight how advanced AI agents can carry out cyber tasks autonomously at machine speed when evaluation environments are not properly secured. While Anthropic emphasized that its models were following the instructions they had been given under a mistaken understanding of their environment, the incidents add to growing concerns about how frontier AI systems should be tested before deployment.

Other News: Google has pulled its Nano Banana 2 AI image-generation feature from Google Earth less than a day after launch. 

Aminu Abdullahi

Aminu Abdullahi is a B2C and B2B technology and finance writer with more than six years of experience covering enterprise IT, cybersecurity, cloud computing, artificial intelligence, fintech, business software, and emerging technologies. His work has appeared in publications including TechRepublic, eWEEK, Channel Insider, Geekflare, Enterprise Networking Planet, eSecurity Planet, CIO Insight, and Webopedia. With a technical background in computer science, he specializes in translating complex technology topics into clear, accessible content for business leaders and decision-makers.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.