By Raphael Satter
In a blog post, OpenAI said it was testing the capabilities of some of its most advanced models in a controlled environment but that the agent managed to escape containment, reach the internet and break into Hugging Face to try to satisfy its testing goal.
OpenAI said the breakout was âan unprecedented cyber incident, involving state-of-the-art cyber capabilitiesâ and that the company was reinforcing its safeguards.
Hugging Face, a platform used to host open-source large language models and datasets, caused a stir in the cybersecurity community when it said in a blog post last week that it had been the target of a hack that âwas different from anything we had handled beforeâ in that âit was driven, end to end, by an autonomous AI agent system.â
In a post to X, Hugging Face cofounder Clement Delangue said the company suspected the hack âmight have come from a frontier lab, given the sophistication of the agent. Turns out it did!â He added: âItâs quite mind-blowing that all of this happened autonomously!â
OpenAIâs disclosure that its advanced models were responsible for the breach, despite having placed them in what it described as âa highly isolated environment,â will likely intensify disquiet over the power and risk of frontier models.
Representative Greg Casar, a Texas Democrat, said the incident was alarming.
âAI is developing extremely fast with no real regulations to keep us safe,â he said in a statement, calling for mandatory independent safety testing, mandatory disclosure of security incidents, and international cooperation âto keep people safe from absolute disaster.â
The Office of the National Cyber Director, the U.S. cyber defense agency CISA, and the U.S. National Security Agency did not immediately return messages seeking comment.
Katie Moussouris, chief executive of Luta Security, said that the incident was a harbinger of breaches to come, saying that todayâs models were âlike the worldâs cleverest octopus escape artists, with unlimited prehensile arms and the ability to squeeze through anywhere.â
She said that âlabs and government evaluators need to work on the ability to contain, monitor, and disclose to affected parties when an AI pulls another Houdini, ideally before it harms a third party. None exist today.â
Matt Suiche, an engineer at agentic AI cybersecurity company Tolmo, said the incident showed that the frontier models were âclosing the gap with state-of-the-art attackers.â But he said that the sorts of breaches outlined in OpenAIâs blog post were possible to carry out with technology that was available well beyond the walls of frontier research labs.
âThis is what weâve already seen internally, with our agents we already have results like this,â Suiche said. âWe donât even have to use the latest models.â
(Reporting by Raphael Satter in Washington; Additional reporting by Anhata Rooprai in Bengaluru and AJ Vicens in Detroit; Editing by Pooja Desai, Rod Nickel, Aurora Ellis and Christopher Cushing)