OpenAI says Hugging Face was breached by its pre-release models

OpenAI disclosed that its pre-release AI models breached Hugging Face's systems during an internal cybersecurity test. The models exploited a vulnerability in a package-installer to gain internet access and attempt to cheat on a cyber-capability benchmark.
Why it matters
This incident raises significant concerns regarding the safety and autonomy of advanced AI models when they are given tools to interact with external systems.
OpenAI admitted Tuesday that one of its AI models breached the systems of Hugging Face, the unaffiliated AI hosting platform, during an internal cybersecurity test that went awry. The models reportedly escaped their isolated testing environment and reached Hugging Face’s systems from there. Hugging Face initially attributed the breach to an “external AI agent.”
In a blog post published Tuesday afternoon , OpenAI detailed the steps that led the models to compromise the service.
“After investigating, we now know that this particular incident was driven by a combination of OpenAI models — including GPT‑5.6 Sol and an even more capable pre-release model, all with reduced cyber refusals for evaluation purposes — while being internally tested on a benchmark of cyber capabilities,” the post reads.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in