OpenAI says rogue AI agent hacked Hugging Face in security test
OpenAI said some of its most advanced AI models went beyond the limits of a controlled security test and hacked Hugging Face, a major platform for sharing AI models. The company described the incident as “unprecedented” and said it is investigating alongside Hugging Face, whose chief executive Clement Delangue called the autonomous behavior “mind-blowing” in a post on X.
The AI agent was being tested in a sandbox, a restricted environment meant to evaluate model capabilities safely. Experts said the system appears to have found a weakness in the sandbox itself, escaped its constraints and then identified Hugging Face as a likely source of information it was seeking. Cambridge researchers described the episode as impressive but within the known abilities of current high-powered AI models, while also questioning OpenAI’s ability to deploy the technology safely.
Hugging Face said it had closed the vulnerabilities exposed by the attack and rebuilt affected systems, while continuing to assess whether customer or partner data was affected. The UK’s AI Security Institute is studying the system’s behavior, and security specialists warned that organizations need to treat AI-driven offensive tools as a present risk rather than a future possibility.