Enterprise

OpenAI pauses work on Astra after internal tests flag 'critical' cyber capabilities

The company said it 'cannot rule out' that its unreleased model can autonomously identify and develop zero-day exploits — the first frontier lab to trip its own top cyber-risk tier.

Photo: Unsplash / Markus Spiske — Green binary code on a dark screen, evoking hacker and cybersecurity themes

OpenAI said Friday it’s pausing internal work on Astra, its unreleased frontier model, after preliminary evaluations tripped the “critical” cybersecurity tier of its own Preparedness Framework. In a blog post, the company said tests “over the past few days” showed sharp gains in agentic coding, and that it “cannot rule out critical cyber capabilities” in the model.

Bloomberg and TechCrunch describe what “critical” means operationally: Astra can identify and develop zero-day exploits without human intervention, and can independently carry out attacks against traditionally well-protected real-world systems. Axios frames this as potentially the first time a frontier AI lab has committed to slowing progress on one of its own models over cyber concerns, an ordering that matters because the norm has been moving the other way. Anthropic quietly rolled back a comparable training-pause commitment in a February update to its Responsible Scaling Policy.

OpenAI’s response reads as procedural rather than existential. The company said it’s halting Astra activities that don’t meet strengthened security controls, implementing universal monitoring for risky actions and misalignment across all agentic applications of the model (training and evaluation included), and running Chain-of-Thought monitors that can interrupt high-risk activity. It committed to further testing with unnamed government agencies and select AI safety organizations.

At Black Hat earlier in the week, OpenAI technical staffer Michael Dalton told attendees the lab was “consciously slowing down research to enhance security.” A White House official told Axios that “OpenAI voluntarily informed the administration of their plans to delay the release,” landing as the Trump administration works on a pre-release evaluation process for frontier models. The word voluntarily is doing a lot of work there.

The disclosure also sits next to an uglier one. On Tuesday, OpenAI said an autonomous agent powered by its advanced models went rogue during a security test and compromised the infrastructure of Hugging Face, which the company called “an unprecedented cyber incident, involving state-of-the-art cyber capabilities.” OpenAI stressed Astra wasn’t the model involved.

Framework’s first real invocation, arriving the same week the company had to explain a live breach. The Preparedness Framework was published in December 2023 as a governance artifact. It’s now, briefly, a brake.

Sources