NVDA 228.87 ▲0.66%GOOGL 351.16 ▼1.07%MSFT 498.00 ▼0.72%AMD 623.77 ▲1.34%INTC 123.86 ▲1.71%TSMC 452.00 ▲1.54%AMZN 254.98 ▼1.34%META 736.60 ▼0.63%AAPL 339.75 ▲0.23%PLTR 184.99 ▲1.04%
Markets at last close

Anthropic · Models

Claude models lead Vellum’s 2026 LLM rankings

·1 min read

Vellum’s LLM leaderboard, updated 1 Jul 2026, compares model versions released after April 2024 using provider results, Vellum evaluations and open-source community tests. It excludes saturated benchmarks such as MMLU and emphasizes newer measures across broad capability, reasoning, coding, automation, web browsing, computer use and terminal tasks.

Claude Mythos 5 leads Humanity’s Last Exam with 64.5%, ahead of Claude Opus 4.8 at 57.9% and Claude Sonnet 5 at 57.4%. Claude Sonnet 5 tops GPQA Diamond with 96.2%, while Claude Mythos 5 ranks first in SWE Bench with 95.5% and Terminal-Bench 2.1 with 88%.

Claude Fable 5 leads several tool-use categories, scoring 17.4% on AutoBench, 85% on OSWorld and 88% on BrowseComp. The speed and pricing tables list Llama 4 Scout at 2600 t/s, GPT-5.3 Codex at 0.003s TTFT and Nova Micro at $0.04 / $0.14 per 1M tokens.

Originally reported by vellum.aiRead the source →
Related coverage
All Anthropic news →