NVDA 223.96 ▲2.27%GOOGL 354.30 ▼0.96%MSFT 499.99 ▲0.03%AMD 483.36 ▼1.21%INTC 101.65 ▲1.84%TSMC 420.04 ▲0.44%AMZN 274.48 ▲0.82%META 592.10 ▲0.37%AAPL 313.33 ▲0.29%PLTR 172.01 ▲10.32%
Markets at last close

Anthropic · Models

Claude models lead Vellum’s 2026 LLM rankings

·1 min read

Vellum’s LLM leaderboard, updated 1 Jul 2026, compares model versions released after April 2024 using provider results, Vellum evaluations and open-source community tests. It excludes saturated benchmarks such as MMLU and emphasizes newer measures across broad capability, reasoning, coding, automation, web browsing, computer use and terminal tasks.

Claude Mythos 5 leads Humanity’s Last Exam with 64.5%, ahead of Claude Opus 4.8 at 57.9% and Claude Sonnet 5 at 57.4%. Claude Sonnet 5 tops GPQA Diamond with 96.2%, while Claude Mythos 5 ranks first in SWE Bench with 95.5% and Terminal-Bench 2.1 with 88%.

Claude Fable 5 leads several tool-use categories, scoring 17.4% on AutoBench, 85% on OSWorld and 88% on BrowseComp. The speed and pricing tables list Llama 4 Scout at 2600 t/s, GPT-5.3 Codex at 0.003s TTFT and Nova Micro at $0.04 / $0.14 per 1M tokens.

Originally reported by vellum.aiRead the source →
Related coverage
All Anthropic news →