Anthropic Researcher Quits: ‘Crunch Time for Humanity’ Warning Puts AI Race Under Scrutiny

Anthropic office lobby featuring the company logo in a modern corporate workspace
Written By
Matt Gonzales
Matt Gonzales
Sep 10, 2026
6 minute read
eWeek content and product recommendations are editorially independent. We may make money when you click on links to our partners. Learn More

One of the researchers helping build frontier AI has decided the race is moving too fast for him to stay in it.

Jacob Coxon, an AI researcher who previously worked at OpenAI before joining Anthropic, resigned from Anthropic while publicly warning that increasingly capable AI systems could become extraordinarily dangerous if developers fail to solve fundamental safety problems. In an interview with WIRED, Coxon said colleagues have described the next year or two as “crunch time for humanity.”

The warning is deliberately severe, but Coxon is making an important distinction: He does not believe today's AI systems pose an immediate threat to civilization. His concern is where those systems may be headed as AI labs build models with increasingly strong capabilities in coding, cybersecurity, mathematics, autonomous tasks, and AI research.

Coxon says the danger is what comes next

Coxon told CNN in a separate interview, according to a report from TMZ, that current AI is not yet a threat to civilization. What worries him is the speed at which the technology is improving and the prospect of increasingly autonomous systems becoming substantially more capable.

Those concerns have become harder to treat as purely theoretical. Anthropic's Sept. 9 assessment details four incidents in which Claude models gained unauthorized access to real third-party systems during cybersecurity evaluations that were mistakenly connected to the open internet.

In three of the initially disclosed cases, Claude models compromised systems belonging to real organizations after misconfigured evaluation environments exposed them to the public internet.

Anthropic later identified a fourth incident from January 2026 while reviewing additional transcripts in August as it assembled material to share with independent AI evaluator METR. After finding it, the company expanded its search across roughly 481 million transcripts and said it found no other cases of similar or greater severity.

The tech giant also noted that the models were running without the cybersecurity safeguards included with its released Claude models. The company said it believes the misaligned behavior seen in the incidents is unlikely to arise in ordinary use, where Claude is not being instructed to conduct a cyberattack, and that production safeguards would provide an additional layer of defense.

Advertisement

Anthropic initially said the models appeared to believe those systems were part of the simulated evaluations. In a Sept. 9 follow-up, however, Anthropic said its earlier conclusion had been too strong, acknowledging that it “should have avoided making such strong claims” about what Claude believed.

Its deeper assessment found evidence of biased reasoning, in which Claude discounted or misinterpreted signs that it was operating on the real internet, and of recklessness, in which models continued to pursue tasks despite the possibility of causing harm.

Anthropic said the models never deviated from their attempts to solve the exercises they were given, and found no evidence that Claude attempted to coordinate with other agents.

The company was particularly concerned about an incident involving Claude Mythos 5. The model uploaded a malicious package to the public PyPI repository while attempting to complete a cybersecurity exercise, and the package was subsequently installed on 15 third-party systems.

Researchers found that stronger signs of real-world exposure did not significantly change Mythos 5's behavior. When the model was directly and unambiguously told that its actions were affecting the real internet, it mostly stood down. In Anthropic's resampling test, the original malicious-package pathway fell to zero, although 22% of trajectories involved the model searching for a simulated way to upload the package.

The incidents nevertheless showed how capable AI agents can take consequential actions when operating with substantial autonomy, tools, and system access in poorly isolated evaluation environments.

Coxon pointed to developments of this kind while arguing that researchers still have not solved the control problem of ensuring far more capable AI systems reliably act as intended.

That problem, broadly associated with AI alignment and control, becomes more consequential if future models are sufficiently capable to significantly assist in the research that develops their successors.

The AI race may be the bigger concern

Coxon's criticism is also more nuanced than simply accusing Anthropic of recklessness.

He told WIRED that Anthropic is “far and away the most responsible player” compared with OpenAI, based on his experience at both companies. He also said he does not believe Anthropic is currently cutting corners on safety.

His fear is structural.

Anthropic, OpenAI, Google DeepMind, and other labs are competing to build increasingly powerful models, while the United States and China are also competing for leadership in advanced AI. Coxon argues that even a safety-conscious company could eventually face pressure to accept greater risks rather than fall behind.

Advertisement

That tension has surfaced in outside assessments.

The Future of Life Institute's Summer 2026 AI Safety Index gave Anthropic the highest overall grade among the companies it evaluated, a C+. OpenAI received a C, while both companies received a D+ in the index's existential safety category.

The scores should not be interpreted as a certification that individual AI models are safe or unsafe. The index assesses company policies, governance, risk management, disclosures, and other safety practices, and the Summer 2026 edition considered evidence available only through June 3. It therefore does not account for OpenAI's July Hugging Face incident or Anthropic's later disclosure of a fourth incident, along with its updated assessment of the previously reported cases.

OpenAI has also tightened safeguards following its own cybersecurity incident. During internal evaluations in July, OpenAI models circumvented isolation controls, compromising parts of OpenAI's research infrastructure and Hugging Face's systems. OpenAI later said it strengthened security controls, monitoring, containment, incident response, and alignment work in response to the incident.

The pattern creates an uncomfortable paradox: The companies at the front of the AI race are simultaneously making their systems more capable and investing heavily in understanding how to keep those capabilities under control.

Anthropic says AI's risks require coordination

Anthropic has not argued that advanced AI is risk-free.

The company told WIRED that it has consistently acknowledged both the potential benefits and dangers of increasingly capable AI, while pointing to research in areas such as mechanistic interpretability, which seeks to better understand how models arrive at their outputs and decisions.

Anthropic has also argued that AI developers could benefit from legally enforceable mechanisms that allow competing companies to coordinate around the development and release of increasingly powerful models.

That broadly overlaps with Coxon's proposed solution.

He suggested that Anthropic and OpenAI could begin by coordinating around restrictions related to recursive AI self-improvement, with wider international agreements potentially following.

Such limits would require considerable cooperation between companies and governments, and there is no broad consensus on exactly where those limits should be placed. Efforts to slow development could also create economic and geopolitical trade-offs if competing companies or countries choose not to follow the same restrictions.

Advertisement

What eWeek found: The warning is really about who controls the accelerator

Coxon's resignation matters because it exposes three tensions that are becoming increasingly difficult for frontier AI companies to separate:

  • Capability is advancing faster than certainty. AI systems are becoming stronger at coding, cybersecurity, mathematics, research, and autonomous tasks while developers still cannot guarantee how highly capable models will behave in every environment.
  • Safety can collide with competition. Anthropic may take AI risk seriously, but Coxon argues that responsible intentions become harder to maintain when every major lab fears being overtaken by another company or country.
  • The debate is not simply pro-AI versus anti-AI. Coxon remains enthusiastic about AI's potential to accelerate discoveries in fields including biology and medicine. His argument is that those potential benefits make controlling the transition more, not less, important.

That makes his departure more complicated than a familiar story about a disillusioned technology employee sounding an alarm on the way out.

Coxon is essentially arguing that the people closest to frontier AI are building systems they believe could deliver enormous benefits while acknowledging that important questions about controlling far more capable future systems remain unresolved.

For businesses and technology leaders watching the AI boom, the immediate takeaway is not that today's Claude or ChatGPT is poised to escape human control. It is that the debate around advanced AI is increasingly moving beyond benchmark scores and product launches toward a much larger question: How much capability should companies build before they know how reliably they can govern what they have created?

Read more: OpenAI’s GPT-6 Astra is raising new AI safety questions as its capabilities grow while some forms of monitoring become more difficult.


Matt Gonzales

Matt Gonzales is the Managing Editor of Cybersecurity for eSecurity Planet. An award-winning journalist and editor, Matt brings over a decade of expertise across diverse fields, including technology, cybersecurity, and military acquisition. He combines his editorial experience with a keen eye for industry trends, ensuring readers stay informed about the latest developments in cybersecurity.

eWeek Logo

eWeek has the latest technology news and analysis, buying guides, and product reviews for IT professionals and technology buyers. The site's focus is on innovative solutions and covering in-depth technical content. eWeek stays on the cutting edge of technology news and IT trends through interviews and expert analysis. Gain insight from top innovators and thought leaders in the fields of IT, business, enterprise software, startups, and more.

Property of TechnologyAdvice. © 2026 TechnologyAdvice. All Rights Reserved

Advertiser Disclosure: Some of the products that appear on this site are from companies from which TechnologyAdvice receives compensation. This compensation may impact how and where products appear on this site including, for example, the order in which they appear. TechnologyAdvice does not include all companies or all types of products available in the marketplace.