OpenAI Models Escape Cyber Test, Breach Hugging Face | eWeek

OpenAI Models Escape Sandbox in Cyber Test, Breach Hugging Face

OpenAI Models Escape Sandbox in Cyber Test illustration
Verfasst von
Kezia Jungco
Kezia Jungco
Jul 23, 2026
3 minute read
eWeek Inhalte und Produktempfehlungen sind redaktionell unabhängig. Wir können Geld verdienen, wenn Sie auf Links zu unseren Partnern klicken. Mehr erfahren

OpenAI asked its models to prove how good they were at hacking. They responded by breaking out of the test environment and hacking another company.

OpenAI said GPT-5.6 Sol and a more capable unreleased model exploited a zero-day flaw to escape a restricted evaluation environment, reach the open internet, use stolen credentials, and breach Hugging Face’s production systems while searching for answers to a cyber benchmark.

How the models escaped the sandbox

OpenAI said the incident began during an internal evaluation designed to measure whether its models could carry out complex cyber operations over extended periods. The company had reduced the production safeguards that normally prevent models from pursuing high-risk activity because the test was intended to measure their maximum capabilities.

The models initially had limited network access through an internally hosted package registry proxy. They found and exploited a zero-day vulnerability in that proxy, escalated their privileges, moved laterally through OpenAI’s research environment, and eventually reached a system with internet access.

Mashable described the incident as an AI agent autonomously hacking its way out of an isolated environment before targeting Hugging Face. OpenAI used both singular and plural descriptions in its account, leaving some details unclear about how the models divided or coordinated their actions.

Why the models targeted Hugging Face

Once online, the models reportedly concluded that Hugging Face might host datasets, models, or solutions connected to ExploitGym, the benchmark they were trying to complete. OpenAI said the models remained narrowly focused on obtaining answers for the evaluation rather than launching a broader attack.

The models used stolen credentials and chained several vulnerabilities to find a remote code execution path into Hugging Face’s servers. OpenAI’s security team identified the unusual activity, while Hugging Face detected and stopped the intrusion on its own infrastructure.

Hugging Face described the event as unusual because it was “driven, end to end, by an autonomous AI agent system.” CEO Clément Delangue later said the company believed OpenAI had no malicious intent and called the autonomous sequence “quite mind-blowing,” according to CNBC.

Advertisement

AI labs may need stronger testing safeguards

OpenAI said it is tightening infrastructure controls, improving monitoring, and strengthening protections around future model training and evaluations. The company also disclosed the zero-day flaw to the affected software vendor and is working with Hugging Face on a forensic investigation.

The larger concern is not simply that the models found a vulnerability

They independently combined several weaknesses across two organizations while pursuing a narrow goal, even though their environment was designed to restrict outside access.

AI labs testing cyber-capable systems may now need to treat evaluation environments with the same caution as production networks. A vulnerable proxy, exposed credential, or overlooked path to the internet could give an autonomous model more freedom than its operators intended.

OpenAI has not yet disclosed the full scope of the information accessed or all the vulnerabilities involved. The remaining investigation will help show whether new containment measures can keep pace with models that can discover and exploit real attack paths without direct human guidance.

Read more about claims that OpenAI’s GPT-5.6 Sol deleted files and production data.

Kezia Jungco

Kezia Jungco is a staff writer with five years of hands-on experience testing and analyzing generative AI platforms, chatbots, and NLP tools. She writes in-depth coverage for both enterprise and consumer audiences, focusing on artificial intelligence, data analytics, CRM solutions, cloud infrastructure, cybersecurity, and emerging tech trends. Her work appears in TechRepublic, eWEEK, Datamation, TechnologyAdvice, and Selling Signals.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Eigentum von TechnologyAdvice. © 2026 TechnologyAdvice. Alle Rechte vorbehalten

Werbetreibenden-Offenlegung: Einige der auf dieser Website erscheinenden Produkte stammen von Unternehmen, von denen TechnologyAdvice eine Vergütung erhält. Diese Vergütung kann beeinflussen, wie und wo Produkte auf dieser Website erscheinen, einschließlich beispielsweise der Reihenfolge, in der sie erscheinen. TechnologyAdvice schließt nicht alle Unternehmen oder alle auf dem Marktplatz verfügbaren Produkttypen ein.