Claude Mythos 5 tops SWE-bench Verified ranking
Claude Mythos 5 leads BenchLM.ai’s SWE-bench Verified leaderboard with 95.5% as of July 6, 2026, ahead of Claude Fable 5 at 95% and Claude Opus 4.8 at 88.6%. Anthropic holds the top five positions, while GPT-5.3 Codex ranks sixth with 85% and Ornith-1.0-397B follows at 82.4%.
SWE-bench Verified evaluates models on resolving real GitHub issues from open-source Python repositories including Django, Flask, and scikit-learn. The benchmark includes 500 verified issues and focuses on code patch generation, requiring models to understand codebases, write fixes, and pass test suites.
BenchLM.ai lists 57 models on the leaderboard and categorizes the benchmark under Coding. The Coding category carries a 20% weight in BenchLM.ai’s overall scoring system, while SWE-bench Verified contributes 16% of that category score, making performance on the benchmark relevant to broader model rankings.