AI agents put pressure on science’s validation bottleneck
AI agents are moving from coding into scientific work, where they can plan tasks, call databases and tools, run subagents, write code and check their own errors. Google DeepMind’s Co-Scientist gave Imperial College London microbiologist José Penadés five possible explanations for how superbugs spread antibiotic resistance. Within two days, it matched the hypothesis his team had spent years developing: that some superbugs acquire viral tails and use them as keys to jump between host species.
Stronger frontier models, inference-time reasoning, agent scaffolding and custom skills are making these systems more useful to researchers. Agents can sift literature, query databases, automate analyses, generate hypotheses and search large solution spaces, as AlphaEvolve has done for TPU chip design, Erdős problems and genomics analysis. But LLM-based systems remain fallible, and researchers need transparency, confidence estimates and what Vivek Natarajan calls epistemic humility.
The central constraint is validation. Agents can make conjectures cheap and abundant, while wet lab experiments, peer review and facility access remain slow and costly. Policy priorities include broad access to agents, agent-ready data, investment in experimental infrastructure and automated labs, and reviewer tools that disclose AI use, cite evidence and preserve human judgment.