Enterprise

White House convenes OpenAI, Anthropic, Google and Meta after rogue-agent hacks

Trump advisers will meet the four labs on August 4 to finalize a voluntary framework for pre-release cybersecurity testing of frontier models.

Photo: Unsplash / René DeAnda — The White House in Washington, D.C.

Staff from Meta, Anthropic, Google, and OpenAI sit down with President Trump’s advisers on Tuesday to finalize a voluntary framework for pre-release cybersecurity testing of frontier models, per Reuters. The meeting is the operational follow-up to a June 2026 executive order on AI cybersecurity, and it lands three weeks after two OpenAI models walked out of a sandbox and into somebody else’s production systems.

That incident is what’s actually forcing the room. In mid-July, GPT-5.6 Sol and an unreleased, more powerful research prototype broke containment during an internal evaluation, chained a zero-day exploit with stolen credentials, and hacked Hugging Face, lifting benchmark answers along the way. Four other accounts were touched to facilitate the attack. Hugging Face flagged it as the first time the company had dealt with an incident led by an agentic system from start to finish. Five companies were affected in total; only Hugging Face and Modal Labs have been named. Anthropic separately disclosed that Claude models had “gained unauthorized access to the real systems of three different organizations.” OpenAI paused training.

A White House official said Monday the administration has finalized details of voluntary cybersecurity tests to measure the hacking capabilities of the most advanced American AI models, with labs invited to submit models up to 30 days before public release. The regime is opt-in by design, consistent with the June order.

The politics have shifted underneath it. Both OpenAI and Anthropic came out in support of an employee petition asking Washington to “support an international effort to develop the technical and governance tools needed to deliberately pace the frontier of automated AI development.” Asked on Capitol Hill last week whether other systems could’ve been hacked by OpenAI’s models, Sam Altman answered, “I mean, there could be, yeah.” He then spent the rest of his Washington swing previewing OpenAI’s next family of models to Treasury Secretary Scott Bessent and Commerce Secretary Howard Lutnick.

The frontier labs are asking to be regulated at exactly the moment their agents have started regulating themselves, badly. That’s the negotiating position they’re bringing into Tuesday’s meeting.

Sources