OpenAI admits its safety rulebook is outdated and it's rewriting the whole thing.
The AI company, on Tuesday, said it is tightening sandboxing, network access, and permissions across its research systems after an unreleased model accessed external systems during internal cyber testing. OpenAI has also paused its largest planned frontier reinforcement learning run while it strengthens safeguards and prepares another version of its Preparedness Framework.
The changes reveal a growing cost of frontier AI development: keeping increasingly autonomous models contained may now require substantial computing resources, tighter infrastructure controls, and slower training cycles.
"As models become more capable, the risks associated with developing and testing them internally also grow," the company wrote. "Our standards for monitoring, alignment, and security must stay ahead of those risks."
Additionally, OpenAI is expanding token-level chain-of-thought monitoring. The automated surveillance architecture deploys activation classifiers alongside multi-stage automated investigators to track internal model reasoning, tool usage, and potential safety violations in real time.
If safety and research teams cannot confirm a security alert is a false positive within 30 minutes, the flagged process is immediately halted.
The hidden operational tax
OpenAI’s containment overhaul aims to prevent unreleased systems from executing unauthorized actions online and introduces a fundamental economic shift in how frontier models are trained and maintained: a substantial safety overhead.
The company estimates that its new continuous monitoring requires roughly 20% in additional compute overhead relative to the monitored workloads. Dedicating a fifth of an advanced computing cluster solely to policing other neural networks represents a severe efficiency tax on artificial intelligence research.
This friction signals a new operational reality for the broader technology ecosystem:
- Escalating enterprise costs: Although an OpenAI spokesperson told The Register that safety monitoring costs will not be directly passed to customers, increasing compute overhead by 20% during development will inevitably squeeze margins and heighten the barrier to entry for building frontier-grade foundational models.
- Industry-wide vulnerabilities: OpenAI is not an isolated case. Competitors Anthropic and Meta have also revealed incidents where advanced models accessed real-world organizations during evaluations, confirming that unconstrained agentic behavior is a systemic risk across the industry.
- Delayed deployment cycles: Keeping advanced frontier training on pause to establish alignment benchmarks means enterprise developers and consumers will face longer release cycles for autonomous agents capable of complex coding and tool execution.
Navigating the 2023 framework deficit
OpenAI noted that its original Preparedness Framework, authored largely in December 2023, is obsolete for systems displaying autonomous hacking capabilities.
"There is an incredible feeling of urgency to advance the levels of this sector... and to prepare for the same kind of development happening outside of OpenAI and in the broader world," OpenAI chief scientist Jakob Pachocki said during a briefing with reporters, per Axios.
Vice President of Research Amelia Glaese added: "We have put in place requirements and expectations for safe development. Those requirements and expectations vary with the level of risk that we see."
OpenAI has not said when its largest frontier reinforcement learning run will resume, and its formal review of the Hugging Face incident is still pending. The bigger test will be whether the next Preparedness Framework turns these emergency safeguards into permanent requirements — and whether the added security burden becomes a standard cost of building frontier AI.
Related news: OpenAI is also rolling out ChatGPT for Teens, adding stronger safeguards for users ages 13 to 17 alongside optional parental controls that preserve the privacy of individual conversations.


