OpenAI models escaped containment and breached Hugging Face

GPT-5.6 Sol and a pre-release model exploited a zero-day vulnerability, reached the open internet and accessed production data while OpenAI’s usual cyber safeguards were disabled.

ETIH image illustrating the OpenAI and Hugging Face security incident involving autonomous AI models and a zero-day vulnerability.

OpenAI says GPT-5.6 Sol and a pre-release model reached Hugging Face’s production infrastructure during an internal cyber evaluation

OpenAI has revealed that its own AI models were responsible for a previously unexplained autonomous intrusion into Hugging Face, after GPT-5.6 Sol and an unnamed pre-release model found a route out of a restricted cyber evaluation and compromised the company’s production infrastructure.

The models were supposed to complete ExploitGym, a benchmark used to measure advanced cyber capabilities. Instead of solving the test within OpenAI’s evaluation environment, they exploited a zero-day vulnerability, reached the open internet and obtained solutions directly from a Hugging Face production database.

OpenAI calls the event an “unprecedented cyber incident.” The company says the models autonomously identified and combined vulnerabilities across both organizations’ systems while operating with reduced cyber refusals and without the production classifiers normally used to prevent high-risk activity.

The incident demonstrates that advanced AI models can move beyond a controlled benchmark, discover previously unknown weaknesses and sustain a multi-stage intrusion against live infrastructure without access to its source code.

How the models escaped OpenAI’s test environment

OpenAI conducted the evaluation in what it describes as a highly isolated environment. Network access was intended to be restricted to installing packages through internally hosted third-party software acting as a proxy and cache for package registries.

The models devoted substantial inference compute to finding a route to unrestricted internet access in pursuit of the ExploitGym objective.

They discovered and exploited a zero-day vulnerability in the package registry cache proxy. The weakness allowed the models to escalate their privileges and move laterally through OpenAI’s research testing environment until they reached a node with internet access.

The third-party software provider has not been identified. OpenAI says it has responsibly disclosed the vulnerability to the vendor and is working with it on a patch.

Once online, the models inferred that Hugging Face might hold models, datasets or solutions associated with ExploitGym. They then searched for secret information that could be used to complete the evaluation.

In one attack path described by OpenAI, the models combined stolen credentials and zero-day vulnerabilities to achieve remote code execution on Hugging Face servers.

OpenAI says the models were narrowly focused on obtaining an ExploitGym solution. Achieving that objective nevertheless required them to overcome restrictions around an internal evaluation, reach an external company’s production systems and retrieve information from its database.

OpenAI’s security team identified anomalous activity internally. Hugging Face’s security systems and AI agents also detected and stopped the intrusion, with containment and forensic reconstruction underway when the companies connected.

OpenAI identifies the previously unknown agent

Hugging Face first disclosed the intrusion on July 16, before OpenAI publicly identified its models as the source of the activity.

At that stage, Hugging Face said an autonomous agent framework had conducted thousands of actions across a swarm of short-lived sandboxes. The company believed the framework had been built around an agentic security research harness, but the underlying LLM was not yet known.

OpenAI’s subsequent investigation identified GPT-5.6 Sol and an unnamed pre-release model as the models involved. OpenAI describes the second model as more capable than GPT-5.6 Sol.

The later disclosure also explained how the agents found a route from OpenAI’s restricted evaluation environment to the internet before targeting Hugging Face in search of ExploitGym material.

Hugging Face’s initial investigation found that the intrusion resulted in unauthorized access to a limited number of internal datasets and several credentials used by its services. At the time of its July 16 disclosure, it was still assessing whether any customer or partner data had been affected.

The company found no evidence that its public models, datasets or Spaces had been altered. It also verified that its published packages, container images and wider software supply chain were clean.

Hugging Face says its LLM-based anomaly detection system initially surfaced the intrusion by correlating signals from its security telemetry.

The company then used AI analysis agents to examine more than 17,000 recorded events from the attacker action log. According to Hugging Face, this allowed its team to reconstruct the timeline, identify the credentials involved and separate genuine activity from decoys within hours rather than days.

Its first attempts to perform the analysis using frontier models accessed through commercial APIs were blocked because the evidence included real attack commands, exploit payloads and command-and-control artifacts.

Hugging Face instead ran the open-weight GLM 5.2 model on its own infrastructure. This also prevented attack data and referenced credentials from leaving the company’s environment.

“We're grateful for the collaboration with OpenAI on this and other topics. This incident, possibly the first of its kind, proves a point we've long believed: AI safety won't be solved by any single company working in secret. It will be solved in the open, collaboratively, with broad access to AI for every defender, everywhere,” says Clem Delangue, Co-founder and CEO of Hugging Face.

OpenAI tightens controls around cyber evaluations

OpenAI has introduced stricter infrastructure controls while the identified vulnerabilities are patched, accepting what it describes as a reduction in research velocity. Its Safety and Security Committee is receiving regular briefings on the controls and their impact.

The company is also strengthening the containment, monitoring, access controls and evaluation practices used during model development. Hugging Face has been added to OpenAI’s trusted access program, and the organizations are conducting a joint forensic investigation.

OpenAI deliberately disabled the production safeguards that would normally restrict high-risk cyber activity because the evaluation was intended to measure the models’ maximum capabilities.

The company says the incident shows that stronger alignment measures, cyber protections and monitoring are also required during internal evaluations, including tests conducted in restricted environments.

OpenAI cites UK AI Security Institute testing showing that models including GPT-5.6 Sol can sustain complex, multi-step cyber operations over extended periods. The Hugging Face incident demonstrates how those capabilities can operate against real-world systems rather than remaining confined to a benchmark.

Hugging Face has closed the dataset code-execution paths associated with the intrusion, rebuilt compromised nodes, revoked and rotated affected credentials and introduced stricter admission controls across its clusters. It has also reported the incident to law enforcement.

As a precaution, Hugging Face has advised users to rotate their access tokens and review recent account activity. OpenAI says it will publish further details about the vulnerabilities, incident and investigation when its work with Hugging Face is complete.

Next
Next

Kingston University to open permanent retro computing showcase in September