Qwen3.7 Max leads the MMLU ranking
Qwen3.7 Max ranks first on the MMLU benchmark as of August 15, 2026, scoring 93.7%. GPT-5 follows at 93.5%, while o3 ranks third at 93.1%.
MMLU, or Massive Multitask Language Understanding, tests knowledge across 57 subjects and is categorized as a general knowledge benchmark. The ranking includes 77 models, with a best score of 93.7, an average score of 82.4, and a standard deviation of 12.7.
The dataset pairs benchmark results with provider pricing, including input and output costs per million tokens. Qwen3.7 Max is listed at $1.250 input and $3.750 output, while GPT-5 is listed at $1.250 input and $10.000 output.
Originally reported by pricepertoken.comRead the source →
Related coverage
Qwen3.8-Flash-Next shows strong agentic results
3 days ago
Qwen previews Qwen4 architecture with Qwen3.8-Flash-Next
5 days ago
Alibaba’s Qwen3.8 open models focus on agentic coding
2 weeks ago
New facts can alter LLM behavior in narrow contexts
2 weeks ago