Qwen3.7 Max leads the MMLU ranking
Qwen3.7 Max ranks first on the MMLU benchmark as of August 15, 2026, scoring 93.7%. GPT-5 follows at 93.5%, while o3 ranks third at 93.1%.
MMLU, or Massive Multitask Language Understanding, tests knowledge across 57 subjects and is categorized as a general knowledge benchmark. The ranking includes 77 models, with a best score of 93.7, an average score of 82.4, and a standard deviation of 12.7.
The dataset pairs benchmark results with provider pricing, including input and output costs per million tokens. Qwen3.7 Max is listed at $1.250 input and $3.750 output, while GPT-5 is listed at $1.250 input and $10.000 output.
Originally reported by pricepertoken.comRead the source →
Related coverage
Alibaba breaks out AI revenue for the first time
2 days ago
Bloomberg technology agenda focuses on China chips and AI deals
6 days ago
Chinese AI models gain ground in Russia
1 week ago
JadeProx deploys TriBack Loader in government and healthcare attacks
4 weeks ago