Vera Rubin makes MLPerf debut with strong GB300 gains
Nvidia’s Vera Rubin NVL72 made its first public appearance in MLPerf Inference v6.1, published on September 16, 2026, as a preview submission rather than a generally available product. Nvidia said the rack delivered up to 2.5x the DeepSeek-R1 throughput and up to 3.7x the Qwen3-VL throughput of its current GB300 NVL72 system, with the largest gains showing up in latency-sensitive interactive scenarios.
The results come with caveats. Vera Rubin’s scores were submitted in MLCommons’ preview category, meaning the hardware and software stack are still pre-launch and may change before commercial availability. Nvidia has not disclosed a launch date, pricing, memory capacity, transistor count, or TDP for Vera Rubin NVL72 alongside the benchmark data.
AMD used the same round to emphasize scale and software progress, including a 512-GPU Instinct MI355X cluster submitted by Crusoe and ROCm-driven gains on unchanged MI355X hardware. AMD reported a 28% jump in offline throughput and a 38% jump in server-mode throughput on gpt-oss 120B, plus a 70% gain on a Wan 2.2 video-generation test.
The broader MLPerf round drew 30 organizations, 120 systems, and 486 individual results, reflecting a shift toward rack-scale AI infrastructure, reasoning workloads, multimodal models, retrieval-augmented generation, and agentic edge tests.