Anthropic researcher quits over superintelligence risks
Jacob Coxon, who spent three years working on AI model training at Anthropic and OpenAI, resigned from Anthropic on safety grounds, accusing both companies of “racing straight to self-improving superintelligence and gambling with our lives”. He warned that forthcoming systems could become superhuman, hack broadly, transform fields quickly and acquire power and resources.
Evan Hubinger, Anthropic’s alignment science lead, backed the warning on X and put the chance of AI wiping out humanity at more than 10 per cent over the next decade. Hubinger said current model risk is low, but raised concern about “superintelligence arising from recursive self-improvement”, while noting Anthropic lacks a clear plan to solve alignment for superintelligence.
The warnings intensified calls for oversight. Darren Jones urged a multinational treaty for safe and regulated superintelligence development and asked officials to raise the issue at G7 and G20 meetings. Separately, the Financial Times reported that Britain’s AI Security Institute did not get pre-release access to Claude Mythos 5.1, while a less-protected version was limited to vetted US organisations.
The Cabinet Office said AISI continues to collaborate with Anthropic and recently tested OpenAI’s GPT-6 Astra before public release. Matt Clifford is stepping down from the government’s Advanced Research and Invention Agency after taking a full-time role at Anthropic, following concerns from MPs over a potential conflict of interest.