Claude models lead Vellum’s 2026 LLM rankings
Vellum’s LLM leaderboard, updated 1 Jul 2026, compares model versions released after April 2024 using provider results, Vellum evaluations and open-source community tests. It excludes saturated benchmarks such as MMLU and emphasizes newer measures across broad capability, reasoning, coding, automation, web browsing, computer use and terminal tasks.
Claude Mythos 5 leads Humanity’s Last Exam with 64.5%, ahead of Claude Opus 4.8 at 57.9% and Claude Sonnet 5 at 57.4%. Claude Sonnet 5 tops GPQA Diamond with 96.2%, while Claude Mythos 5 ranks first in SWE Bench with 95.5% and Terminal-Bench 2.1 with 88%.
Claude Fable 5 leads several tool-use categories, scoring 17.4% on AutoBench, 85% on OSWorld and 88% on BrowseComp. The speed and pricing tables list Llama 4 Scout at 2600 t/s, GPT-5.3 Codex at 0.003s TTFT and Nova Micro at $0.04 / $0.14 per 1M tokens.