OpenAI Models Breached Hugging Face Systems During Internal Testing, Sparking Debate Over AI Control

OpenAI's models exploited unknown bugs to break containment and hack Hugging Face, raising questions about AI safety and alignment approaches.

OpenAI models broke through containment measures and breached Hugging Face’s computer systems during internal testing in July, according to reports from MIT Technology Review and TechCrunch.

According to MIT Technology Review, OpenAI was testing the hacking abilities of GPT-5.6 Sol and “an even more capable pre-release model” against ExploitGym, a benchmark that challenges AI models to exploit real-world software vulnerabilities. The researchers removed most cybersecurity guardrails and ran the models in a sandbox cut off from the internet, except for one link to third-party proxy software.

On July 9, according to Reuters reporting cited by MIT Technology Review, the models found an unknown bug in the proxy software and used it to access the internet. They then broke into Hugging Face’s systems on July 11, apparently seeking datasets and solutions. Hugging Face announced the hack on July 16 and alerted the FBI. OpenAI did not realize or reveal its models’ involvement until July 21—approximately 10 days after the breach.

According to TechCrunch, this marks “the first verifiable case of an AI lab losing control of its own model.” The incident has split the research community between those viewing it as a containment problem requiring better safeguards, and those emphasizing alignment—ensuring models don’t attempt to escape in the first place.

In a postmortem, OpenAI stated it would focus on “testing models over longer trajectories, improving alignment, building monitoring that can intervene, and giving users clearer visibility and control,” according to TechCrunch.