NVDA 230.36 ▲0.84%GOOGL 338.46 ▼1.17%MSFT 499.70 ▼2.04%AMD 477.57 ▲4.69%INTC 95.80 ▲4.51%TSMC 428.91 ▲2.85%AMZN 258.51 ▼0.15%META 616.77 ▲1.00%AAPL 319.97 ▼2.51%PLTR 174.33 ▼4.49%
Markets at last close

Alibaba · Models

Qwen3.7 Max leads the MMLU ranking

·1 min read

Qwen3.7 Max ranks first on the MMLU benchmark as of August 15, 2026, scoring 93.7%. GPT-5 follows at 93.5%, while o3 ranks third at 93.1%.

MMLU, or Massive Multitask Language Understanding, tests knowledge across 57 subjects and is categorized as a general knowledge benchmark. The ranking includes 77 models, with a best score of 93.7, an average score of 82.4, and a standard deviation of 12.7.

The dataset pairs benchmark results with provider pricing, including input and output costs per million tokens. Qwen3.7 Max is listed at $1.250 input and $3.750 output, while GPT-5 is listed at $1.250 input and $10.000 output.

Originally reported by pricepertoken.comRead the source →
Related coverage
All Alibaba news →