NVIDIA reported MLPerf Inference v6.1 results on September 16, 2026, marking the Vera Rubin NVL72 system's first benchmark submission. According to NVIDIA, the platform delivered up to 3.7x higher throughput than GB300 NVL72 on Qwen3-VL and up to 2.5x on DeepSeek-R1. A four-rack, 288-GPU GB300 NVL72 configuration achieved 99% scaling efficiency on DeepSeek-R1. Software optimizations alone produced up to 1.6x performance gains over v6.0. Results are vendor-submitted preview figures; MLCommons published official v6.1 data at mlcommons.org. Nineteen ecosystem partners also submitted results.

Source: NVIDIA News. Original announcement: September 16, 2026.