NVDA 214.72 ▼0.98%GOOGL 344.82 ▲1.22%MSFT 483.24 ▲0.43%AMD 473.25 ▲0.81%INTC 90.07 ▼2.24%TSMC 418.95 ▲0.71%AMZN 258.63 ▼0.57%META 549.90 ▲0.75%AAPL 309.35 ▼0.63%PLTR 179.94 ▲3.44%
Markets at last close

Alibaba · Models

Qwen3.7 Max leads the MMLU ranking

·1 min read

Qwen3.7 Max ranks first on the MMLU benchmark as of August 15, 2026, scoring 93.7%. GPT-5 follows at 93.5%, while o3 ranks third at 93.1%.

MMLU, or Massive Multitask Language Understanding, tests knowledge across 57 subjects and is categorized as a general knowledge benchmark. The ranking includes 77 models, with a best score of 93.7, an average score of 82.4, and a standard deviation of 12.7.

The dataset pairs benchmark results with provider pricing, including input and output costs per million tokens. Qwen3.7 Max is listed at $1.250 input and $3.750 output, while GPT-5 is listed at $1.250 input and $10.000 output.

Originally reported by pricepertoken.comRead the source →
Related coverage
All Alibaba news →