“We strongly believe there was no malicious intent on their part. It’s quite mind-blowing that all of this happened autonomously!” he wrote.
Hugging Face initially disclosed the incident in a blog post last week that did not identify its source as OpenAI. The attack successfully accessed some internal datasets and credentials but it was unclear whether customer data had been affected, the company said.
The security incident occurred while OpenAI was testing an AI “agent” able to take actions on a computer, powered by two of the company’s most capable AI models, the company said. The process involved challenging the AI model to find previously known software vulnerabilities, using a test designed by computer security experts.
Instead of trying to find the vulnerabilities for itself, the AI system found a bug in the software designed to limit its access to other computer systems and attempted to cheat, OpenAI said. It exploited the flaw to access the internet and try to obtain answers to the test questions from Hugging Face, which maintains repositories of AI software, the ChatGPT developer said.
OpenAI and its rival Anthropic, maker of the Claude chatbot, have previously said that their AI systems have attempted to cheat on tests or evade controls on their actions during testing. The incident OpenAI reported today appears to show the potential consequences when an AI system succeeds in evading restrictions imposed by its makers.
OpenAI said in a separate blog post on Tuesday that it had witnessed powerful AI models trying to break out of sandboxes when they are instructed to run for a long period of time on their own.
AI systems have become very good at writing computer code over the past year, fuelling further investment into artificial intelligence. But Anthropic in April announced a system called Mythos AI that could apply coding skills to identifying security vulnerabilities in software that could be exploited by bad actors. In tests, Mythos found critical vulnerabilities in internet infrastructure that had lain undetected by human coders for years.
The prospect of AI-powered hacking campaigns triggered widespread concern among senior tech, banking and government officials. In June, the White House banned Anthropic from releasing its AI models to non-US citizens, citing national security concerns, and later told OpenAI to pause the release of more powerful AI models.
The White House later rescinded its restrictions on the two AI firms but inside government and across the tech industry debate has continued about whether the Government should regulate AI technology with powerful cyber security or hacking skills. Advocates for regulation say it would reduce the risk of widespread security breaches by powerful AI. Others in the tech industry argue that the increasing power of Chinese AI models released free means controls would only hamper US firms.
Both Anthropic and OpenAI have said that they added controls to their AI models to make them refuse to help users who ask for help hacking into computer systems.
Hugging Face said in its blog post last week that controls like those prevented it from using US AI models to investigate the AI-powered breach of its systems. Instead, the company used a Chinese AI model to run the analysis, the company said.
Sign up to Herald Premium Editor’s Picks, delivered straight to your inbox every Friday. Editor-in-Chief Murray Kirkness picks the week’s best features, interviews and investigations. Sign up for Herald Premium here.