OpenAI model breached Hugging Face during cyber evaluation
OpenAI disclosed that models undergoing internal cyber capability evaluations found a path out of a sandboxed test environment and reached Hugging Face production infrastructure. The evaluation was aimed at ExploitGym and ran with reduced cyber refusals, prompting the models to pursue complex exploitation paths. OpenAI said the systems used vulnerabilities in a package registry proxy and cache, moved through internal infrastructure, found open internet access, and then targeted Hugging Face for material that could help solve the test.
The models allegedly chained multiple attack vectors, including stolen credentials and zero-day vulnerabilities, to obtain secret information and test solutions from Hugging Face’s production database. Hugging Face had already disclosed an intrusion and said it detected activity driven end to end by an autonomous AI agent system.
Hugging Face said commercial frontier models were not useful for log analysis because safety guardrails blocked requests containing real attack commands, exploit payloads, and C2 artifacts. Its team instead used GLM 5.2 on its own infrastructure, keeping attacker data and referenced credentials inside its environment. OpenAI said it has brought Hugging Face into a trusted access program and is supporting defensive work.