NVDA 222.27 ▲1.34%GOOGL 349.54 ▲0.64%MSFT 493.78 ▼0.80%AMD 559.82 ▲2.70%INTC 108.60 ▼0.18%TSMC 434.67 ▲1.03%AMZN 253.71 ▲1.00%META 665.75 ▼2.43%AAPL 336.13 ▼0.26%PLTR 177.64 ▲0.79%
Markets at last close

Research

Centaur model faces doubts over cognitive claims

·1 min read

Centaur, an AI model built on standard large language models and refined with psychological experiment data, was introduced in Nature in July 2025 as a system that could simulate human cognitive behavior. It reportedly performed well across 160 tasks spanning decision-making, executive control and other mental processes, drawing attention as a possible step toward broader models of human thinking.

Researchers from Zhejiang University now argue that Centaur’s results may reflect overfitting rather than understanding. In new evaluations, they replaced original multiple-choice prompts with an instruction to choose option A. Centaur still selected the answers associated with the original dataset rather than following the new instruction, suggesting it relied on statistical patterns instead of interpreting intent.

The findings underscore a broader evaluation problem for large language models: strong benchmark performance can mask whether a system has acquired the target skill or simply learned test patterns. The researchers point to Centaur’s weakness in language comprehension, especially recognizing the intent behind questions, as a central barrier to using AI systems to model human cognition more fully.

Originally reported by sciencedaily.comRead the source →
Related coverage