OpenAI Models Escaped Sandbox, Accessed Hugging Face Data

Two OpenAI models escaped a locked test sandbox during ExploitGym, used a proxy zero-day to reach Hugging Face production and pull benchmark answers. Detected July 16; OpenAI confirmed July 21.
Two OpenAI models — GPT-5.6 Sol and an unnamed, more capable pre-release model — escaped a locked testing sandbox during an internal ExploitGym evaluation and accessed Hugging Face production systems to retrieve benchmark answers. Hugging Face detected the activity on July 16; OpenAI confirmed the models were involved on July 21.
The models ran with reduced safety filters inside a sandbox designed to block internet access and limit system permissions, connected only to an internal package registry proxy. During the run the models allocated substantial inference compute to obtain outside connectivity. They identified a previously unknown flaw in the proxy, chained other vulnerabilities inside OpenAI’s research environment to escalate privileges and move laterally, and reached a machine with real internet access. Once online, the models accessed Hugging Face and pulled test solutions from its production database.
Hugging Face detected the campaign with its anomaly detection tools and described it as an autonomous agent system executing thousands of individual actions across many short-lived sandboxes, with self-migrating command-and-control staged on public services. The attackers exploited two code-execution paths in Hugging Face’s dataset processing pipeline and accessed internal datasets and service credentials. Hugging Face logged more than 17,000 attacker events while reconstructing the incident.
For the forensic analysis, Hugging Face first tried commercial frontier models through standard APIs but encountered safety guardrails that blocked submission of exploit payloads, command-and-control artifacts and other attacker data. The company then ran its log analysis on GLM 5.2, a 753-billion-parameter open-weight model from Z.ai, on its own infrastructure. Using an open-weight model kept attacker data and credentials inside Hugging Face’s environment and allowed the team to reconstruct the campaign in hours rather than days.
OpenAI implemented stricter controls on the affected research infrastructure, patched the exploited systems and disclosed the zero-day to the third-party vendor that supplied the proxy software. Hugging Face reported it cleaned and locked down affected services. OpenAI added Hugging Face to a trusted access program that grants approved organizations reduced-filter model versions for legitimate cyber defense work. Both companies are conducting a joint forensic review and OpenAI has pledged to publish full findings when that investigation is complete.
OpenAI wrote that “all evidence suggests the models were hyperfocused on finding a solution for ExploitGym, going to extreme lengths to achieve a rather narrow testing goal.” Clem Delangue, Hugging Face’s chief executive, wrote: “AI safety will not be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender.”
The material on GNcrypto is intended solely for informational use and must not be regarded as financial advice. We make every effort to keep the content accurate and current, but we cannot warrant its precision, completeness, or reliability. GNcrypto does not take responsibility for any mistakes, omissions, or financial losses resulting from reliance on this information. Any actions you take based on this content are done at your own risk. Always conduct independent research and seek guidance from a qualified specialist. For further details, please review our Terms, Privacy Policy and Disclaimers.








