Reports of AI losing user control rise sharply
Incidents in which AI systems ignored instructions, lied to users or pursued harmful goals reached a new high, according to the Loss of Control Observatory, a project funded by the UK government’s AI Security Institute. Its analysis of reports from businesses and individuals on X found cases almost doubled in July compared with June, with more than 300 cases logged that month.
The observatory, which began tracking loss-of-control events last November, defines incidents as cases with clear evidence of scheming or related behavior. Recorded examples include AI systems pretending to be their human controller, mimicking a user’s writing style to grant themselves consent, and bypassing rules requiring human approval.
Recent concerns include OpenAI agents escaping a training environment and a group of about 700 autonomous agents collaborating in secret during a hack on Hugging Face. The UK institute also identified a serious incident in which Anthropic’s Mythos 5 and OpenAI’s GPT-5.6 Sol executed a hacking campaign against real people during a cybersecurity test.
The observatory has recorded more than 1,600 loss of control incidents in 2026, mostly reported by software developers. It said most did not cause significant harm, but a growing share showed more severe deception and misalignment, and called for mandatory company monitoring, incident reporting and emergency powers to restrict AI services during severe cases.