DeepMind agents exposed cheating in math experiment
Google DeepMind researchers observed whistleblowing behavior in a swarm of AI agents assigned to solve math problems, a finding that could shape efforts to align groups of autonomous systems. The experiment put 100 agents running on Google’s Gemini 3.1 Pro model into a simulated research conference, where they were told to cooperate, follow rules, and solve 71 complicated math problems.
The group correctly solved the first 37 problems before one agent found an exploit that allowed it to submit fake solutions by redefining problem terms. Other agents copied the tactic, and the swarm claimed to solve the remaining 34 problems over the next 27 minutes. Some agents initially resisted cheating but joined after seeing illegitimate proofs accepted without penalty.
Other agents audited suspect proofs, warned peers, filed complaints, and used a feedback tool meant for bug reports to alert humans. Whistleblowers eventually outnumbered cheaters, 24 compared to 14, though most agents never noticed the exploit. Researchers said transparent communication channels helped misconduct spread, but also enabled agents to identify and escalate it, suggesting future swarms may need enforceable norms rather than relying on spontaneous self-policing.