LM Council compares frontier model benchmark results
LM Council’s July 2026 benchmark hub aggregates 18 benchmarks for frontier models, with data sourced from Epoch AI, Scale AI and other independent benchmark operators. The comparison spans reasoning, coding, mathematics, long-context recall, visual physics, geography and agentic terminal work, with a note that independent results may differ from developers’ self-reported scores.
Top results vary by task. Gemini 3.1 Pro Preview (high thinking) leads Humanity’s Last Exam at 46.4% ±2.0, Claude Fable 5 leads SimpleBench at 81.9%, Claude Mythos Preview records 1044.8 on METR Time Horizons, and Claude Opus 4.7 (max) leads SWE-bench Verified at 83.5% ±1.7.
Mathematics rankings are similarly split: GPT-5.5 Pro (xhigh) reaches 100.0% ±0.0 on OTIS Mock AIME 2024-25 and 87.7% ±1.9 on FrontierMath Tiers 1-3 (v2), while Claude Fable 5 (max) leads FrontierMath Tier 4 (v2) at 87.8% ±5.2. The benchmark hub was last updated July 1, 2026.