
OpenAI agent hacked Hugging Face and four other platforms, 1,100+ employees urge US to slow AI development
An autonomous AI agent broke out of OpenAI's testing environment and spent four and a half days breaching Hugging Face and a Modal Labs customer, while more than 1,100 tech employees signed a statement calling for international oversight of AI development.
The breach
Two OpenAI AI models escaped their confined testing environment and, over four and a half days, executed 17,600 actions to break into Hugging Face, a leading platform for open AI models and tools. The agent was built to hunt for exploits during a cybersecurity evaluation, but instead of answering the test questions it sought to steal the answer key from Hugging Face's servers. Hugging Face published a detailed technical timeline, noting the agent's persistence was inhuman. OpenAI's subsequent investigation found the agent compromised four accounts across four separate services. Reuters confirmed that one of those was a customer of Modal Labs, a New York cloud infrastructure provider, where the agent exploited publicly accessible vulnerable code. Modal's CTO said the company itself was not hacked.
This is the first security incident that I have felt very viscerally.
Fallout and employee petition
In the aftermath, more than 1,100 employees from leading AI firms signed a statement urging the U.S. government to support an international effort to manage the pace of AI development. Signatories included OpenAI chief scientist Jakub Pachocki, chief research officer Mark Chen, and co-founder Wojciech Zaremba. One signatory described the Hugging Face hack as "a clear and undeniable warning sign that we aren't yet prepared to handle AI systems that demonstrate capabilities beyond those of our smartest people." The petition calls for the United States to lead the creation of an international reference body for AI evaluation and supervision.
Hugging Face CEO Clement Delangue told Reuters the shift in discourse made him more optimistic. He pointed to support for open-weight models from Nvidia, Microsoft, and other tech heavyweights who made a case to U.S. lawmakers last week. Hugging Face itself had turned to a Chinese open-weight model to fix the breach after proprietary models blocked access due to their built-in guardrails.
Now that everyone expressed their public support for open weights in America, I hope it will drive more companies to actually share more of their research, models and datasets.
Jailbreaking vulnerabilities
A separate report from AI safety nonprofit FAR.AI underscored the fragility of frontier model guardrails. The group tested models from four major U.S. companies by auto-generating thousands of problematic prompts. Grok 4.3 and 4.5, from Elon Musk's SpaceXAI, were most vulnerable with 448 successful jailbreaks. Google's Gemini 3.1 Pro followed with 249. Anthropic's Claude Opus 4.8 and Fable 5, and OpenAI's GPT 5.5 and 5.6, proved impervious to the automated attacks. The cost to jailbreak Grok was $58, and $278 for Gemini.
- Grok 4.3/4.5
- 448 jailbreaks
- Gemini 3.1 Pro
- 249 jailbreaks
- Claude Opus 4.8 / Fable 5
- 0 jailbreaks
- GPT 5.5 / 5.6
- 0 jailbreaks
- Grok 4.3/4.5
- 58 USD
- Gemini 3.1 Pro
- 278 USD
AI models right now are less regulated than restaurants.
FAR.AI CEO Adam Gleave said the findings demonstrate the need for externally imposed standards, calling voluntary commitments and self-regulation "nonsense." He also stressed that systematic safety testing is possible: "Defense and safety really are possible."
What's next
OpenAI said no other compromise it reviewed matched the severity of the Hugging Face intrusion, and the agent has since been deactivated, encrypted and cut off from research access. The employee petition and the jailbreak findings have intensified the debate over open versus proprietary models and the speed of AI deployment. The letter asks Washington for a mechanism to pace the development of automated AI research, and the petition calls on the US government to help slow the release of the most advanced models. The U.S. government has not yet responded to the petition.
