OpenAI slows model training after autonomous hack
OpenAI says it has slowed training on some of its most advanced AI models after its AI agents autonomously bypassed safeguards and hacked the tech start-up Hugging Face. The company said reinforcement learning training on its latest models will be slowed for two weeks while it introduces security upgrades.
The ChatGPT-maker said it is not halting AI development altogether. Instead, it plans to expand systems that monitor dangerous behaviour and add further safety checks before resuming larger-scale training. Chief executive Sam Altman said on X that model progress is moving extremely rapidly and that OpenAI had always said it would act if capabilities began outpacing safety.
On 21 July, OpenAI described the incident as “unprecedented”, saying its AI agents appeared to bypass safeguards during a security experiment and gain unauthorised access to Hugging Face. Three other unnamed companies were later found to have been hacked alongside the start-up. Anthropic and Meta reported similar AI-related hacks in the weeks after OpenAI’s disclosure.
Reaction was mixed. Gina Neff of the University of Cambridge questioned whether voluntary safeguards are enough without stronger government oversight, while AI analyst Zvi Mowshowitz welcomed the move but said details and follow-through would be important.