NVDA 223.96 ▲2.27%GOOGL 354.30 ▼0.96%MSFT 499.99 ▲0.03%AMD 483.36 ▼1.21%INTC 101.65 ▲1.84%TSMC 420.04 ▲0.44%AMZN 274.48 ▲0.82%META 592.10 ▲0.37%AAPL 313.33 ▲0.29%PLTR 172.01 ▲10.32%
Markets at last close

OpenAI · Models

o3 points beyond bigger models

·1 min read

François Chollet created ARC-AGI in 2019 as a benchmark meant to challenge the industry’s reliance on scaling model size. The test used simple colored-grid puzzles that a child could solve, but that required identifying abstract patterns rather than leaning on memorized training examples. For four years, larger models made little progress: GPT-3 scored close to zero, and GPT-4o scored 5%.

OpenAI’s o3 changed the trajectory in December 2024, reaching 87.5% on the same benchmark while belonging to the same model class as the system that had scored 5% the month before. The key shift was longer inference: the older model spent a few hundred tokens deciding on an answer, while o3 spent five and a half billion. The high-compute run was reportedly trained on a large portion of the public training set, so the result was not pure inference, but it still pointed to a different path for progress beyond simply making models bigger.

Originally reported by levelup.gitconnected.comRead the source →
Related coverage
All OpenAI news →