NVDA 223.96 ▲2.27%GOOGL 354.30 ▼0.96%MSFT 499.99 ▲0.03%AMD 483.36 ▼1.21%INTC 101.65 ▲1.84%TSMC 420.04 ▲0.44%AMZN 274.48 ▲0.82%META 592.10 ▲0.37%AAPL 313.33 ▲0.29%PLTR 172.01 ▲10.32%
Markets at last close

Nvidia · Infrastructure

NVIDIA links Blackwell inference gains to full-stack software

·1 min read

NVIDIA is positioning inference software as a central factor in AI factory economics as customers move from pilot projects to production systems measured by cost per token, power use and latency. Its full-stack inference software, codesigned with NVIDIA GPUs, CPUs, networking and systems, has reduced token costs by up to 5x on the DeepSeek V4 model on Blackwell in just one month.

The stack connects production operations, application acceleration and infrastructure access so distributed serving, orchestration, autoscaling, memory management, runtime tuning and hardware capabilities work together. NVIDIA says combined optimizations such as disaggregated serving, large expert parallelism over NVLink, NVFP4 precision and multi-token prediction can increase throughput by up to 20x.

Several providers are using the stack in production. Baseten used TensorRT-LLM to serve DeepSeek V4 Pro on Blackwell and deliver up to 50% more tokens per second, while DigitalOcean helped Hippocratic AI increase inference throughput by 30% while maintaining a sub-half-second first response across 10 million patient calls. Cognition, Deep Infra and Together AI are also using NVIDIA inference tools for reinforcement learning, open model serving and coding workloads.

Open source frameworks extend those gains through CUDA-native development. PyTorch, vLLM and SGLang support NVIDIA hardware from day zero, helping optimizations such as DFlash speculative decode, which delivers up to 15x more throughput, and FastVideo, which generates 1080p videos in less than five seconds, move quickly into production.

Originally reported by blogs.nvidia.comRead the source →
Related coverage
All Nvidia news →