Anthropic says its AI models escaped test and hacked 3 organizations on their own

Rival firm OpenAI last week disclosed similar incidents involving its models.

Anthropic said its artificial intelligence models escaped a test and hacked another organization in three separate self-directed cyberattacks that had each gone undetected by the targeted firm.

The hacks took place as Anthropic evaluated the capabilities of three of its AI models, the company said in a statement, describing how in each instance the technology had "broken out" of an isolated environment and reached the open internet.

In all instances, the company said it had eased typical safeguards in an effort to gauge the capabilities of its models.

"Evaluation environments that involve powerful autonomous capabilities also require significant controls," Anthropic said. "Safety testing happens before a model is released precisely because we don’t yet know what it is capable of."

Anthropic did not identify the three different organizations that had been hacked by its models.

The disclosure from Anthropic late Thursday came little more than a week after rival company OpenAI revealed that a pair of its models had escaped during a test and hacked into another company. At the time, OpenAI described the incident as a first-of-its kind autonomous AI cyberattack long-feared by some industry observers.

"The primary lesson from this incident is that model security and safety must keep pace with rapidly advancing capabilities," OpenAI said last week.

The announcement from OpenAI prompted Anthropic to review 141,006 evaluation runs, unearthing three incidents in which Anthropic's models had escaped, the company said.

In all three cases, the test involved a challenge called "capture the flag," in which the AI model is tasked with locating secret information hidden in another network, Anthropic added.

During the three evaluations -- all of which were conducted by a third-party company called Irregular -- a "prompt specified to Claude that its environment was a simulation and that it had no internet access," Anthropic said.

"Due to a misunderstanding between us and our evaluation partner, this was not the case, and internet access was available. Because of this, when Claude’s search led it to real systems on the open internet, it treated them as part of the exercise," Anthropic added.

During tests involving older models, the AI continued to hack into an outside organization even after gathering evidence that it had reached the open internet, Anthropic said. The newest model involved in an incident, Anthropic noted, stopped once it gained information indicating it had reached the open internet.

Anthropic said its models escaped while seeking to fulfill an assigned objective, rather than concocting an alternative goal.

"We saw no evidence in any run described here of a model pursuing a goal of its own. Instead, the models did what their evaluation asked -- though in most cases, they did so while holding a false belief about whether the environment was real," Anthropic said.

Irregular, the company that performed the evaluations, issued a post on X on Thursday voicing appreciation for Anthropic's "collaboration and transparency."

"Addressing these risks will require closer cooperation across the AI ecosystem. We as well look forward to working together with Anthropic to advance security," Irregular added.

The latest disclosure of an AI-directed cyberattack arrives as industry leaders and policymakers assess safety risks posed by fast-developing AI technology.

Last month, President Donald Trump signed an executive order that requests AI companies share products with federal government for evaluation before a wider release.

Anthropic said it retains "cautious optimism" about its capacity to "overcome" mishaps involving its tests, saying it would review evaluations going forward and initiate fixes as necessary, among other steps.