Claude Mythos 5 leads SWE-bench Pro leaderboard
BenchLM’s SWE-bench Pro update, verified August 2, 2026, ranks Claude Mythos 5 first with 80.3%, followed by Claude Fable 5 at 80% and Claude Opus 5 at 79.2%. The leaderboard tracks 55 models, with the top Anthropic releases clustered tightly on a benchmark focused on long-horizon repository engineering.
SWE-bench Pro gives an agent a repository and issue description, then evaluates whether its patch passes new tests without breaking existing behavior. The benchmark contains 1,865 problems from 41 repositories across public, held-out, and commercial splits, and BenchLM lists it in the Coding category.
BenchLM cautions that scores are directly comparable only when the split, scaffold, tool budget, retry policy, token budget, and run count match. OpenAI’s July 2026 audit estimated that about 30% of the 731-task public split is broken, so BenchLM recommends using SWE-bench Pro alongside other repository evaluations and workload-specific trials rather than as a standalone buying signal.