AI agents fall short on open-ended research
A Princeton-led group of researchers found that AI agents are not yet capable of conducting open-ended AI research at the level needed for recursive self-improvement. The study used a method called shadow evaluation, asking Anthropic’s Claude Opus 4.8, running on OpenClaw, to answer research questions from two unpublished papers submitted to NeurIPS 2026.
The agents were given six days, Anthropic API credits, GPU resources, virtual computers, and web access to produce conference-quality papers. The original paper authors rejected both submissions. The agents reviewed literature, ran hundreds of experiments, and compiled results, but they pursued weak approaches too quickly, struggled to rethink failing plans, wrote unclearly, and produced no novel contribution.
The findings suggest a gap between AI systems’ strength in narrow, checkable engineering tasks and the broader judgment required for original research. The study had limitations, including its small scope and the fact that graders knew the submissions were AI-generated, but it may temper claims that rapid recursive self-improvement is imminent.