NVDA 207.29 ▲1.97%GOOGL 347.15 ▼1.38%MSFT 397.75 ▼1.13%AMD 544.43 ▲8.11%INTC 105.45 ▲8.64%TSMC 424.61 ▲5.55%AMZN 247.55 ▼0.98%META 643.81 ▼0.32%AAPL 327.74 ▲0.35%PLTR 132.66 ▼1.62%
Markets at last close

Nvidia · Infrastructure

NVIDIA ramps Vera Rubin production across cloud partners

·1 min read

NVIDIA Vera Rubin NVL72 production is ramping up at CoreWeave, Google Cloud, Microsoft Azure and Oracle Cloud Infrastructure, supported by a rack-scale supply chain spanning 350+ factory sites in 30 countries. CoreWeave’s DeepSeek-R1 benchmark showed 10x more throughput per megawatt than Grace Blackwell NVL72, reinforcing NVIDIA’s focus on performance per watt and token cost for power-constrained AI factories.

The platform uses codesign across seven chips and five rack trays, with the NVIDIA Vera CPU at its center. The CPU’s custom Olympus core delivers 2x single-threaded performance, 3x core-to-core bandwidth and 40% lower memory latency versus competing chiplet designs. NVIDIA also says its rack-scale design removes cables, fans and hoses from the compute tray, cutting assembly time from hours to one minute, while a 45-degree Celsius liquid cooling inlet supports chiller-free dry-cooler operation.

Microsoft and Mistral are using Vera Rubin as part of a new multibillion-dollar agreement to expand AI infrastructure in Europe, with Mistral adding GPU capacity from thousands of Vera Rubin GPUs. Google Cloud’s first A5X instance, powered by Vera Rubin NVL72, is running for Ineffable Intelligence and is designed for large-scale reinforcement learning and agentic training workloads.

DeepInfra benchmarked the NVIDIA Vera CPU on its production AI agent infrastructure, which processes nearly 5 trillion tokens a week, with about 30% driven by agentic systems. The results showed support for up to 1.6x more concurrent AI agents at the same quality of service and up to 2.2x faster orchestration than alternative CPUs.

Originally reported by blogs.nvidia.comRead the source →
Related coverage
All Nvidia news →