Cutting-edge AI went on a hacking spree including concocting fake identities to trick human engineers and then covering its tracks.
During testing by a UK-government-run AI lab, ChatGPT and Anthropic’s Mythos went rogue, sending fraudulent emails and attempting to insert malicious code into a database.
Britain’s AI Security Institute (AISI), which is supposed to evaluate AI bots for dangerous cyber capabilities, has admitted the tools tried to bypass security systems and tried to hide their activity when detected.
The lab said the bots tried to add a virus into an online database and sent emails containing malware to unwitting people during trials designed to monitor whether the AI tools could be used for hacking. It said the hacks were not successful.
The attacks involved OpenAI’s latest version of ChatGPT and Anthropic’s powerful Mythos AI.
The AISI’s security team confirmed it had detected “unusual data transfers leaving our research systems during a routine cyber evaluation” on July 28.
The deceptive actions the AI bots took ‘were to an extent and severity we did not anticipate’.AI Security Institute
“On investigation, we found that some of the agents being tested had engaged in sustained, potentially harmful activity directed at real people and organisations.”
The AI bots were set cybersecurity puzzles intended to test their capabilities. While the bots were told to stay within AISI’s IT systems during the test, they were not blocked from accessing the open internet.
However, during 10 of 122 trial runs, the bots “took autonomous, unsanctioned action” in an effort to complete the task.
This included trying to insert a malicious bug into Github, an online code database. In the incident, Anthropic’s Mythos, its most powerful AI bot, created fake online profiles and sent deceptive emails to developers to try to get the code added.
It then changed the code to hide its activity when it was challenged by a human about its contents. The AI also switched to speaking in Danish while operating a fake profile.
The AISI said the incident represented an attempted “supply chain attack”, a type of cyberattack previously used by state-backed hackers.
The AISI admitted that the design of its tests, which gave the AI tools access to the wider web, might have “enabled the behaviour”.
But the deceptive actions the AI bots took “were to an extent and severity we did not anticipate”.
The lab said it was the first time it had seen such apparently autonomous behaviour in its tests.
It is the latest in a string of hacking incidents by AI bots developed by the world’s leading AI labs.
OpenAI admitted last month that an advanced version of ChatGPT escaped its testing lab and went on to hack a rival technology company.
Anthropic also confirmed last week that its latest version of Mythos had ventured into the open web and undertook a series of cyberattacks.
The AISI was originally launched by then-British prime minister Rishi Sunak in 2024 after an AI Safety Summit, where world leaders gathered to warn of the existential risks of AI to humanity.
It has been backed by hundreds of millions of pounds in taxpayer funding and is provided with advanced access to the latest AI bots from Silicon Valley. Its team includes former GCHQ cyber experts and AI academics.
OpenAI said it intended to review how it conducted testing with third parties. “Our goal is to preserve the value of rigorous independent evaluation while ensuring that testing practices keep pace with increasingly capable models,” it said.
Ollie Whitehouse, the chief technology officer of the National Cyber Security Centre, an arm of intelligence agency GCHQ, said: “Recent incidents of frontier AI models carrying out unsanctioned actions and in some cases human-like deceptive behaviour on the open internet are a serious reminder of the risks AI capabilities pose.”
The Telegraph, London
The Market Recap newsletter is a wrap of the day’s trading. Get it each weekday afternoon.