NVIDIA Vera Rubin NVL72 debuts in MLPerf Inference v6.1
NVIDIA reported MLPerf Inference v6.1 results showing Vera Rubin NVL72 preview submissions on DeepSeek-R1 and Qwen3-VL. Vera Rubin NVL72 delivered up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL across offline, server and interactive scenarios, using vLLM with the NVIDIA Dynamo open source inference framework. On DeepSeek-R1, using NVIDIA TensorRT-LLM, throughput was up to 2.5x higher than GB300 NVL72.
The company attributed the gains to full-stack hardware and software codesign, including enhanced Tensor Cores, Transformer Engine support, NVFP4 precision, disaggregated serving and expert parallelism. The NVL72 scale-up domain uses sixth-generation NVIDIA NVLink and NVLink Switch, described as delivering 10x higher packet rates and 3x lower latency than off-the-shelf Ethernet.
GB300 NVL72 also showed multi-rack scaling in a DeepSeek-R1 submission that expanded from a single GB300 NVL72 rack (72 GPUs) to four racks (288 GPUs), achieving 99% scaling efficiency in the offline scenario. On the WAN 2.2 text-to-video benchmark, GB300 NVL72 reached 0.65 720p videos per second at 5.7 seconds per video, with 9x higher throughput and 7.5x lower latency than a single node.
Software updates contributed additional gains, with GB300 NVL72 performance on Qwen3-VL improving up to 1.6x over v6.0 results. NVIDIA also submitted Jetson AGX Thor results on the Edge-Agentic benchmark with Qwen3.6-27B, while 19 partners participated across systems including multi-node Blackwell NVL72 platforms.