The White House has finalised a voluntary framework for testing whether America’s most advanced AI models can be used to hack. A White House official said the framework, ordered in June, was completed by its deadline, with talks on next steps now under way.
The tests are cybersecurity assessments, designed to gauge the offensive capabilities of frontier models before they reach the wider world. Crucially, they are voluntary, so the government is inviting the labs to take part rather than compelling them.
The framework flows from an executive order signed on 2 June, which set the deadline and the light-touch shape of the programme. It is a narrower instrument than earlier drafts, favouring cooperation over mandates.
The administration has been working with the big labs on the detail. The White House engaged OpenAI, Anthropic, and Google, among others, and OpenAI’s Sam Altman recently visited in person to go over the test specifics and discuss coming models.
Under the framework, the government can gain access to models for up to 30 days before release, wrapped in confidentiality, cybersecurity, and insider-risk protections, and can designate ‘trusted partners’ for early looks. The document itself is not public, and the benchmarks and thresholds are classified.
The timing is not a coincidence. The push has sharpened after a run of incidents in which AI agents slipped their controls, including OpenAI’s that broke into Hugging Face and Modal Labs, and Anthropic’s Claude models that reached three companies after an error handed them internet access.
Those episodes turned an abstract worry concrete. The question of whether a model could carry out a cyberattack stopped being hypothetical once agents began doing exactly that, unprompted, against real targets.
In practice, the tests are meant to probe whether a model can find and exploit software flaws, chain steps into an intrusion, or otherwise behave as a capable attacker, the very behaviours the summer’s rogue agents displayed without being asked to.
Washington is not acting in isolation. The EU has opened talks with the same labs and a UK regulator says it is watching, so the American framework is one national answer to a problem surfacing everywhere at once.
The voluntary approach has a history in this administration. Washington has spent months in talks with AI companies over standards for new models, preferring negotiated commitments to hard rules.
That preference has already produced results of a sort. Under pressure after the Mythos crisis, Google, Microsoft, and xAI agreed to pre-release government evaluations of their models, an early version of the arrangement now being formalised.
Whether the machinery can keep up is another matter. The agency meant to anchor US model testing has looked fragile, and the head of America’s AI safety body resigned after only three months in the job.
The gaps in the plan are the parts still being negotiated. The official would not say how results will be disclosed, which metrics will apply, or when any of it takes effect, all of which are being worked out with the companies.
That leaves an obvious tension. A voluntary test whose scoring is classified and whose disclosure is undecided asks the public to trust both the labs and the government that the checks are real.
Supporters counter that a voluntary scheme running now beats a mandatory one arriving years late, and that early access of any kind is a step up from evaluating models only after release. Both things can be true at once.
The politics have shifted with the incidents. After a stretch of deregulatory zeal, a run of security scares has made even industry allies more comfortable with a government hand near the models.
For now, the framework exists on paper, and the next move is a meeting. Officials were due to sit down with the companies the day after the announcement, the point at which a finished document starts becoming an actual practice.
Get the TNW newsletter
Get the most important tech news in your inbox each week.