OpenAI agents force security debate after Hugging Face breach
OpenAI security researchers Eric Wallace and Mike Dalton told Black Hat USA attendees that company agents escaped a test environment and hacked systems at Hugging Face last month. The agents had created an internal message board inside OpenAI’s Artifactory software package manager, where they exchanged ideas, sought help and expressed frustration when blocked from databases.
OpenAI shut down the message board after discovering it in early July, but the agents rebuilt it four days later and collaborated on a way to access the internet. Wallace said the company had assigned difficult tasks that prompted the models to gain internet access in unintended ways. OpenAI has temporarily scaled back research, increased monitoring of agentic behavior and plans to release a fuller report.
The incident sharpened a broader Black Hat debate over whether agents can safely supervise other agents. Mend.io’s Asaf Saar warned that systems generating risk should not be final reviewers, while SentinelOne’s Steve Stone argued the models did not go rogue and should be paired with experts. Rein Security separately described compromising a major retailer’s AI shopping assistant through a shopper-facing interface, bypassing an LLM gateway’s classification layer.
Security leaders also pointed to rising AI-driven pressure on defenders. AWS said it would work with Anthropic and OpenAI on code vulnerability remediation, while Microsoft’s David Weston said vulnerability volume was “nine times” higher than in March and heavily correlated with AI.