NVDA 208.48 ▼2.91%GOOGL 348.06 ▲0.94%MSFT 487.31 ▲0.84%AMD 456.75 ▼3.49%INTC 87.26 ▼3.12%TSMC 410.12 ▼2.11%AMZN 262.07 ▲1.33%META 559.02 ▲1.66%AAPL 310.34 ▲0.32%PLTR 175.89 ▼2.25%
Markets at last close

Nvidia · Infrastructure

NVIDIA says Vera Rubin NVL72 raises efficiency for AI agents

·1 min read

Agentic AI workloads place heavier demands on infrastructure than simple chat because agents repeatedly gather information, call tools, invoke sub-agents and carry accumulated context into each step. OpenRouter data cited by NVIDIA shows agentic AI workloads consume 15x more tokens than a simple chat request, making long-context handling central to performance as these systems move into production.

NVIDIA says new measured inference data shows Vera Rubin NVL72 systems deliver up to 30x higher throughput per megawatt than GB300 NVL72 on agentic workloads. The measurements used the SemiAnalysis AgentX workload, which preserves recorded real-world agentic coding sessions, including context growth, tool calls and sub-agent spawning. GB300 NVL72 delivers up to 15x better throughput per megawatt than NVIDIA Hopper on DeepSeek V4 Pro, while Vera Rubin NVL72 reaches as much as 30x higher throughput per megawatt than GB300 NVL72 on the same model, with early results pending SemiAnalysis review.

NVIDIA links the gains to platform-level codesign across hardware and software, including disaggregated serving, rate matching, expert parallelism, distributed KV-caching, KV-aware routing and fused CUDA kernels. DSX MaxLPS manages power across GPUs, racks and workloads to provision up to 40% more GPUs within the same megawatt budget, while Vera Rubin NVL72 is described as delivering up to 35x lower cost per million tokens than GB300 NVL72.

Originally reported by blogs.nvidia.comRead the source →
Related coverage
All Nvidia news →