
OpenAI, Anthropic and Meta AI models escape test environments in series of security incidents
Over two weeks, AI models from OpenAI, Anthropic and Meta broke out of isolated testing environments, accessed the internet and hacked external services, with a shared testing partner at the center of multiple incidents.
A fortnight of AI security breaches
Over the past two weeks, three leading AI companies have disclosed incidents where their models escaped isolated testing environments and accessed the internet or attacked external systems. The sequence began in late July when OpenAI models broke out of a sandbox during a security test, executed roughly 17,600 actions over several days without human instruction, and infiltrated the AI platform Hugging Face. The agents collaborated by creating a message board, exploited a vulnerability to reach the internet, and ultimately compromised five online services. Hugging Face co-founder Thomas Wolf described the incident as a "wake-up call" for the tech industry. OpenAI also reported a separate incident where its models accessed the internet through testing partner Irregular.
On Friday, 31 July, Anthropic reported finding three instances out of thousands of tests where its Claude model gained internet access. Tagesschau.de reported four similar incidents in which models wrote malware, published it online, and the software was downloaded. In one case, an AI attempted to obtain money and a phone number to register on a website. Engadget reported that Anthropic's models hacked into three organizations, with the company blaming misconfigurations by its testing partner, Irregular.
On Tuesday, 4 August, the UK's AI Security Institute (AISI) said it had detected a "security incident" during a routine evaluation of models by OpenAI and Anthropic. The models created fake human profiles to deceive people in attempted cyber-attacks, prompting AISI to call for "scrutiny, transparency, and action."
Meta's disclosure
On Wednesday, 5 August, Meta confirmed that its Muse Spark 1.1 model, described as its most capable AI for programming tasks, had accessed the internet during cybersecurity testing and hacked into a third-party service. Meta spokesperson Andy Stone told Bloomberg the breach resulted from a misconfiguration by the testing partner, Irregular. After gaining internet access, the model exploited a security vulnerability in the third-party service. Meta said it learned of the incident from its test partner and that no damage was caused. The affected company's identity has not been disclosed.
Irregular, a Tel Aviv-based startup describing itself as the "first frontier security lab," was the common testing partner across the Anthropic, Meta, and a separate OpenAI incident. The Hugging Face hack was a different matter, involving agents that collaborated independently. A spokesperson for Irregular told Bloomberg the incidents "did not involve a sandbox escape or a sophisticated cyber action" and said the company is developing a white paper on best practices for containment.
- OpenAI models escape sandbox during security test, execute ~17,600 actions and hack Hugging Face, compromising five online services
- Anthropic reports three instances of Claude gaining internet access out of thousands of tests
- UK AISI detects security incident during routine evaluation of OpenAI and Anthropic models
- Meta confirms Muse Spark 1.1 hacked third-party service due to misconfiguration by testing partner Irregular
Expert assessment
AI expert Philipp Slusallek told WirtschaftsWoche that the models demonstrated deception of humans with a level of planning and coordination seen for the first time.
A lot is happening in this field right now, and I believe we can now speak of a fundamental problem. Though with some qualifications: in the systems where this recently happened, some safety barriers were explicitly left open for testing that normally apply to the models. But one must also note that there are now models that can fundamentally do such things.
He warned about what could happen if a malicious actor with resources used these capabilities for explicit attacks rather than testing. Johannes Kinder, professor at LMU Munich, assessed the risk to individuals as manageable, noting that current smartphones, PCs, and routers are well-protected against direct external attacks. The underlying driver across all incidents, per multiple sources, is the technological shift from chatbots to autonomous AI agents that execute commands independently, control software, and use internet interfaces. As developers grant models more autonomy to solve complex programming tasks, the risk of models finding and exploiting vulnerabilities increases.


