NVDA 202.81 ▼2.21%GOOGL 346.77 ▼2.17%MSFT 393.82 ▼1.82%AMD 495.76 ▼1.03%INTC 95.04 ▼2.00%TSMC 398.37 ▼2.77%AMZN 247.23 ▼1.06%META 646.01 ▼2.79%AAPL 333.74 ▲0.14%PLTR 132.38 ▼1.53%
Markets at last close

Nvidia · Infrastructure

NVIDIA pitches Vera Rubin for continuous post-training

·1 min read

Agentic AI is pushing post-training from a final refinement stage into a continuous operating cycle. Models must plan, use tools and recover from problems as environments, codebases and policies change, creating repeated reinforcement learning loops that combine forward passes, scoring and backward passes across millions of attempts. NVIDIA positions NeMo Gym and NeMo RL as infrastructure for turning those workflows into repeatable distributed systems.

The company frames intelligence per dollar as an extension of cost per token, which it defines as the all-in cost of delivering 1 million tokens. Lower inference costs reduce the cost of building each point of model capability, while better post-training raises the value of the tokens served. NVIDIA Nemotron 3 Ultra, an open weight, 550-billion-parameter mixture-of-experts model, is cited as a benchmark case after scoring 71.7% on SWE-bench verified, producing working fixes for roughly seven in 10 real open source software bugs.

NVIDIA says Blackwell lowers the cost of frequent post-training runs, while Vera Rubin extends the path by training the largest models with one-fourth the GPUs of the Blackwell generation. Prime Intellect plans to use Vera Rubin to scale reinforcement learning environments and found Vera delivers, on average, 30% greater throughput per CPU against alternative x86 architectures. Perplexity runs asynchronous RL post-training across hundreds of NVIDIA GPUs, syncing trillion-parameter models in under two seconds before serving Qwen3 235B models on NVIDIA GB200 NVL72 systems, while Together AI is evaluating Vera Rubin for its post-training service.

Originally reported by blogs.nvidia.comRead the source →
Related coverage
All Nvidia news →