NVDA 206.84 ▼0.92%GOOGL 319.74 ▲0.65%MSFT 381.70 ▲0.03%AMD 521.95 ▼3.29%INTC 92.32 ▼7.89%TSMC 403.41 ▼2.93%AMZN 232.11 ▼0.66%META 595.19 ▼1.80%AAPL 333.02 ▲3.53%PLTR 122.92 ▼0.36%
Markets at last close

Nvidia · Infrastructure

NVIDIA centers AI factory efficiency on performance per watt

·1 min read

Power capacity is presented as the central limit for AI factories, where revenue and profitability depend on how many tokens can be generated within fixed energy budgets. As agentic AI increases token demand, performance per watt becomes the key efficiency metric for infrastructure designed to scale under power constraints.

Frontier models are described as increasingly built around mixture-of-experts architectures that benefit from larger GPU domains connected by ultrafast scale-up interconnects. NVIDIA says its Hopper generation used an eight-GPU domain, while Blackwell NVL72 moves to a 72-GPU domain. Across leading open models, NVIDIA GB300 NVL72 is claimed to deliver up to 25x performance per watt compared with Hopper, with model-specific results including up to 20x on GLM5.1 and up to 10x for Kimi K2.6.

The gains are tied to full-stack rack-scale design, including NVLink Switch, inference software such as NVIDIA Dynamo, TensorRT LLM, SGLang and vLLM, and optimizations including NVFP4 quantization, disaggregated serving, expert parallelism and KV-aware routing. DSX MaxLPS is described as shifting power between GPUs and racks in real time, supporting warm-water liquid cooling and enabling operators to run up to 40% more GPUs within the same power budget.

NVIDIA also points to production deployments by Anthropic, OpenAI, SpaceXAI, CoreWeave, Perplexity and Fireworks AI. CoreWeave has deployed Kimi K2.6 on NVIDIA GB300 NVL72, Perplexity runs Qwen3 235B and Qwen3.5-397B-A17B on NVIDIA GB200 NVL72, and Fireworks AI deploys GLM 5.2 on Blackwell for customers including Cursor and Factory AI.

Originally reported by blogs.nvidia.comRead the source →
Related coverage
All Nvidia news →