NVDA 212.17 ▲0.57%GOOGL 344.98 ▼1.26%MSFT 497.12 ▼1.64%AMD 504.20 ▲2.19%INTC 97.14 ▼0.05%TSMC 413.75 ▼1.02%AMZN 248.42 ▼2.02%META 670.24 ▲0.70%AAPL 331.34 ▼0.52%PLTR 172.56 ▼0.43%
Markets at last close

Nvidia · Infrastructure

NVIDIA emphasizes tokens per megawatt for AI factories

·1 min read

At AI Infra Summit, NVIDIA framed AI infrastructure around energy efficiency and tokens per megawatt as agentic workloads increase demand for performance, scale and power-aware operations. Ian Buck, vice president of hyperscale and high-performance computing, outlined collaborations with Amazon’s Annapurna Labs on NVHBM custom high-bandwidth memory and d-Matrix on NVLink Fusion, combining NVIDIA Vera CPUs with d-Matrix Raptor XPUs for low-latency inference.

NVIDIA positioned Vera Rubin systems, Dynamo inference software, NeMo libraries and networking technologies including NVLink, Spectrum-X Ethernet, ConnectX SuperNICs and BlueField as a full-stack AI factory platform. DSX MaxLPS can deliver up to 1.4x more tokens per megawatt through factory-wide power optimization, while Lambda reported 23% better performance per watt after running 19 nodes within a power budget typically used for 16 full-power nodes.

Emerald AI and NVIDIA demonstrated automated load reduction with Silicon Valley Power, with DSX Flex responding to demand signals while protecting priority AI workloads. Pinterest is using NVIDIA Blackwell and Dynamo for conversational AI in visual discovery, while Vera Rubin NVL72 results on SemiAnalysis AgentX showed up to 30x higher throughput per megawatt than NVIDIA GB300 NVL72 and up to 45x lower cost per million tokens.

Additional demonstrations highlighted Groq 3 LPX for deterministic ultralow-latency inference, startup benchmarks for the Vera CPU across agent and data workloads, and NVLink 6 resiliency features designed to keep very large GPU deployments operating reliably.

Originally reported by blogs.nvidia.comRead the source →
Related coverage
All Nvidia news →