Historical AI models struggle with Einstein-style discovery
Scientists are testing whether language models trained only on historical knowledge can rediscover major scientific breakthroughs. Demis Hassabis of Google DeepMind suggested using data available before 1911 to see whether an LLM could reproduce general relativity, framing the exercise as a possible benchmark for artificial general intelligence.
Early efforts suggest current systems are better at pattern matching than at the abductive reasoning behind paradigm shifts. Researchers argue that breakthroughs such as relativity require a creative leap from sparse or anomalous evidence into a durable world model, not just statistical extrapolation from large data sets. An MIT orbital-mechanics model trained on synthetic planetary systems failed to infer Newtonian gravitation, instead producing a different incorrect law for each system.
Independent researcher Michael Hla trained Machina Mirabilis on pre-1900 data and reported only limited “glimpses of intuition” after prompts related to quantum mechanics and relativity. Other teams found that historical training sets are noisy and leaky: a model intended to use only knowledge up to 1930 still answered later questions correctly. The Ranke-4B project has built models with cut-offs of 1913, 1929, 1933, 1939 and 1946, aiming to detect “sparks of genius” rather than full breakthroughs.