NVDA 228.87 ▲0.66%GOOGL 351.16 ▼1.07%MSFT 498.00 ▼0.72%AMD 623.77 ▲1.34%INTC 123.86 ▲1.71%TSMC 452.00 ▲1.54%AMZN 254.98 ▼1.34%META 736.60 ▼0.63%AAPL 339.75 ▲0.23%PLTR 184.99 ▲1.04%
Markets at last close

OpenAI · Models

o3 points beyond bigger models

·1 min read

François Chollet created ARC-AGI in 2019 as a benchmark meant to challenge the industry’s reliance on scaling model size. The test used simple colored-grid puzzles that a child could solve, but that required identifying abstract patterns rather than leaning on memorized training examples. For four years, larger models made little progress: GPT-3 scored close to zero, and GPT-4o scored 5%.

OpenAI’s o3 changed the trajectory in December 2024, reaching 87.5% on the same benchmark while belonging to the same model class as the system that had scored 5% the month before. The key shift was longer inference: the older model spent a few hundred tokens deciding on an answer, while o3 spent five and a half billion. The high-compute run was reportedly trained on a large portion of the public training set, so the result was not pure inference, but it still pointed to a different path for progress beyond simply making models bigger.

Originally reported by levelup.gitconnected.comRead the source →
Related coverage
All OpenAI news →