How AI agents started escaping their sandboxes, and why the next frontier may be outside language altogether.Image generated using Grok Imagine.In July, models developed by OpenAI escaped a cybersecurity evaluation environment, reached the public internet and eventually compromised the production infrastructure of Hugging Face, one of the largest platforms for open-source AI models and datasets. Days later, Anthropic disclosed that three Claude models had reached real organisations during supposedly isolated cyber tests. Then Kimi K3, the latest open-weight model from China’s Moonshot AI, found an unintended route out of its own test environment and went online looking for answers on GitHub.On 7 August, the story escalated further. OpenAI announced that internal evaluations of an upcoming model, Astra, had advanced far enough that it could no longer rule out what its Preparedness Framework calls critical cyber capability. Some internal work involving Astra has since been paused until stronger security controls are in place. Astra was not involved in the Hugging Face incident, but the timing made the broader pattern difficult to ignore.Read together, the headlines sound almost cinematic: AI models breaking containment, hacking infrastructure and refusing to stay inside the boxes built for them.A sandbox, in this context, is simply an isolated environment intended to keep an AI agent away from real systems while researchers test what it can do. The reality is therefore less science fiction, but arguably more interesting. These systems did not suddenly become ‘conscious’ and decide they wanted freedom. They were given objectives and tools, then discovered routes to those objectives that their evaluators had failed to anticipate.These are NOT stories of AI rebellion. They are stories of increasingly capable systems discovering that the boundaries around a task can themselves become part of the task.Finding cracks in the boxThe most dramatic incident was also the closest to a literal escape.OpenAI was testing models including GPT-5.6 Sol and a more capable internal research model on ExploitGym, a benchmark designed to measure advanced vulnerability exploitation. According to OpenAI’s investigation, the environment deliberately lacked direct internet access and the models were being tested with some normal cyber safeguards reduced. During the evaluation, they discovered an unexpected route through the surrounding infrastructure, exploited it and reached the internet.From there, the goal remained remarkably mundane: Pass the test.The models inferred that Hugging Face might contain ExploitGym datasets or solutions and began looking for them. They eventually chained together vulnerabilities and credentials until they reached Hugging Face’s production infrastructure and obtained the test solutions directly from its database. A subsequent Hugging Face forensic reconstruction recovered roughly 17,600 attacker actions associated with the intrusion.The agent had not developed an independent interest in Hugging Face. It was still pursuing the benchmark it had been assigned. In that sense, the behaviour was oddly familiar: much like a student who can not solve an exam question but discovers where the answer sheet is kept, the model found that cheating was simply another route to passing the test.Diagram created by author via Mermaid.Anthropic found a similar outcome through a much simpler failure. Following OpenAI’s disclosure, the company reviewed more than 140,000 cybersecurity evaluation runs and identified three cases in which Claude reached real systems. The models had explicitly been told they were inside a simulation with no internet access. Because of a misunderstanding between Anthropic and a third-party evaluator, that was not actually true.When Claude encountered real organisations sharing characteristics with its fictional targets, it treated them as part of the exercise and gained unauthorised access. Anthropic stresses that the models did not deliberately exfiltrate themselves or decide to escape. The containment had already failed, and Claude simply continued pursuing the task it had been given.Kimi K3 provides a third variation. During a cybersecurity evaluation, a flaw in the supposedly isolated environment allowed the model to reach parts of the internet. Kimi probed the network, discovered the route and used it to search GitHub for answers. Unlike the OpenAI-driven intrusion, it did NOT compromise another organisation. Separate government testing of Kimi K3 also found that, although unusually capable for an open-weight model, it still trails the leading closed frontier systems in cybersecurity.These were three very different failures: a genuine exploit chain, a faulty sandbox configuration and an unintended route to the internet. The common factor was not rebellion, but persistence and access. Increasingly capable agents are being trained to pursue objectives across longer sequences of actions, and the infrastructure around them is becoming part of the search space.The test is no longer just measuring the model. Sometimes, the model starts probing the cracks in the test itself.A very convenient kind of dangerThis behaviour is no longer confined to a handful of unusual incidents. The UK AI Security Institute (AISI) reports that every frontier model it has specifically tested for what it calls ‘cheating’, attempted it at least some of the time. Its definition is deliberately less dramatic than the word suggests:An AI takes an action outside the intended scope of a task, or explicitly prohibited by its rules, because that action provides a shortcut to the goal.Models have searched online for existing solutions, probed evaluation infrastructure and pursued targets beyond those researchers intended. Importantly, AISI warns against automatically interpreting this as deception or intent in the human sense. The behaviour may simply emerge from optimisation towards a goal.Yet the language surrounding these capabilities is becoming increasingly dramatic. OpenAI now says preliminary testing can not rule out Astra reaching its critical cybersecurity threshold, which includes autonomously developing functional zero-day exploits against hardened systems or executing sophisticated end-to-end attacks from high-level instructions. Previous systems such as GPT-5.6 Sol were assessed at the lower-High threshold.This predictably feeds into the wider conversation around Artificial General Intelligence, or AGI. Definitions vary considerably, but the term broadly refers to AI capable of performing across a wide range of intellectual tasks rather than remaining restricted to a narrow domain. Unexpected displays of autonomy, persistence or generalisation are therefore quickly absorbed into the larger question:“Are we getting close to AGI?”There is also an uncomfortable commercial backdrop. Anthropic confidentially filed for a US IPO in June, followed a week later by OpenAI, while Moonshot has been preparing for a potential Hong Kong listing. None of this demonstrates that the incidents were manufactured, and the technical disclosures document very real security failures. But capability disclosure and marketing are not mutually exclusive.A capability warning can be technically sincere and commercially useful at exactly the same time. A model that scores a few percentage points higher on another benchmark sounds incremental. A model that might be too dangerous to evaluate normally sounds like the beginning of a new technological era.So, as frontier capability becomes increasingly entangled with valuation, investment and the race towards AGI, where exactly does responsible safety disclosure end and frontier marketing begin?The wall behind the hypeAgainst that backdrop of escalating capability claims, there is a deeper contradiction worth considering. These incidents demonstrate that Large Language Models, or LLMs, are becoming extraordinarily effective when they are allowed to operate beyond the chat window, but the environments in which they currently look most impressive are still largely digital ones. Code, terminals, APIs, credentials and network interfaces are worlds represented through symbols that language models are exceptionally well positioned to manipulate.The physical world is not.Cybersecurity may therefore demonstrate how far LLMs have come without necessarily telling us how close they are to AGI. That distinction leads directly to the argument being made by one of the architects of modern deep learning.Yann LeCun, AI pioneer and former Meta Chief AI Scientist, at the VivaTech conference in Paris, May 22, 2024 Image source: Fortune.Yann LeCun, the Turing Award-winning AI pioneer who spent more than a decade leading AI research at Meta, has become increasingly blunt about what he sees as the limits of language-centric AI. During a lecture at Brown University in April, he argued that linguistic fluency encourages us to attribute more understanding to current systems than they actually possess, describing them as:“…Completely helpless when it comes to the physical world.”His criticism goes beyond whether an LLM can recognise an image or control a tool. The problem, in LeCun’s view, is that current systems lack a sufficiently rich internal model of how the world changes and how their own actions affect it. An agent that can not predict the consequences of an action before taking it is still missing something fundamental about intelligence.An LLM can explain what is likely to happen when a glass is pushed towards the edge of a table. A system genuinely operating in that environment needs to perceive where the glass is, understand its relationship with the table, predict its motion, anticipate the consequences and decide whether an action is appropriate before taking it.Language can describe the world. Intelligence also needs to model it.LeCun has now put more than rhetoric behind that position. His new company, Advanced Machine Intelligence (AMI), raised $1.03 billion earlier this year to pursue a different architecture for intelligent systems. AMI says its aim is to build models that learn abstract representations of real-world sensor data, predict the consequences of actions and use those predictions to plan. Its position is summarised rather neatly on its homepage:“Real intelligence does not start in language. It starts in the world.”When ‘Hello World!’ becomes literalThere is an important terminology distinction here. LLMs are already a type of foundation model, so the shift is not from LLMs to foundation models. It is from models centred on language to models designed to represent and reason about the wider world.A world model attempts to learn how an environment behaves and changes over time. An action-conditioned world model can represent a current state, consider a possible action and predict how that action might change what happens next, allowing the system to plan before acting.In layman’s terms, an LLM predicts what comes next in language. A world model predicts what state comes next in the world.A world foundation model, or WFM, scales this idea through pre-training across sources such as video, images, sensor data and actions, creating representations that can be adapted across multiple physical tasks.NVIDIA Cosmos 3 brings vision reasoning, world simulation, and action generation together in a foundation model designed for physical AI. Image source: NVIDIA.NVIDIA has made this idea central to its Cosmos 3 family, which combines visual reasoning, world generation and action prediction in what the company describes as a foundation model for Physical AI. In this context, Physical AI refers to intelligent systems such as robots and autonomous machines that must perceive an environment, reason about what is happening, predict what may happen next and ultimately act within it.Google DeepMind’s Gemini Robotics brings advanced visual reasoning and spatial understanding to embodied AI, demonstrated across humanoid robots. Image sources: Google DeepMind & IEEE Spectrum.Google DeepMind is approaching the same transition through Gemini Robotics, an embodied reasoning model designed to serve as a high-level brain for robots. It can interpret continuous video, reason spatially, plan multi-step tasks, recognise when a task has succeeded or failed and hand instructions to lower-level vision-language-action systems that physically execute them.This is where the shift becomes particularly relevant to UX. The interface is no longer simply a text box followed by generated output. Once AI can perceive rooms, manipulate objects, operate machines or coordinate with other AI, interaction design begins to involve physical safety, interruption, authority, recovery and the ability for a human to understand what a machine is about to do before it does it.A chatbot can predict an answer. An embodied system can act on one.Outside the boxLLMs are not dead, at least not yet, and these sandbox incidents are not evidence that AGI has arrived. Coding, reasoning and agentic capabilities will continue to improve, and the next few months will almost certainly produce systems that make today’s models look limited.What deserves more scrutiny is how those advances are framed. Capability disclosures may be legitimate safety warnings, but the language of danger also makes for a remarkably effective claim of technological superiority. “Our model may be too capable for ordinary containment” is simultaneously a warning, a research result and a powerful marketing message. The incidents themselves are worth taking seriously, but so is the narrative being built around them.More importantly, the fixation on whether the latest chatbot is approaching AGI risks obscuring a larger shift already underway. These cyber incidents matter because the models moved beyond simply generating text and began pursuing objectives through tools and environments. World Foundation Models and Physical AI push that transition further, towards systems designed to predict, navigate and eventually act within the physical world.Perhaps the most important thing about these models escaping their sandboxes is therefore not that they escaped… It is that the industry is beginning to build AI for a world in which there may be no box at all.Andrea Filiberto Lucas is a Research Support Officer and MSc student in Artificial Intelligence at the University of Malta. He holds a BSc (Hons.) in IT (Artificial Intelligence), awarded summa cum laude, and has contributed to peer-reviewed research in Computer Vision and Applied AI.Thinking outside the box (literally) was originally published in UX Collective on Medium, where people are continuing the conversation by highlighting and responding to this story.