NVIDIA Vera targets faster agent workloads
Agentic AI workloads put CPUs on the critical path because models repeatedly hand off work such as tool calls, code execution, data processing, KV-cache handling and result analysis. NVIDIA argues that conventional data center CPUs have prioritized high core counts and lower cost per rentable core, often at the expense of single-threaded performance, memory bandwidth per core and predictable latency.
Vera is NVIDIA’s answer to that bottleneck, built around the custom Olympus CPU core for agent loops that depend on sequential steps. Olympus delivers 50% higher instructions per cycle than NVIDIA Grace, while Vera pairs the cores with up to 1.2TB/s of LPDDR5X memory bandwidth at less than 40 watts of memory power. A monolithic compute die provides 3.4TB/s of core-to-core bandwidth, described as 3x greater than any other data center CPU, and supports all 88 cores with full memory performance.
NVIDIA says Vera delivers 1.8x the sustained per-core performance of x86 in loaded CPU workloads representing agentic execution. Perplexity tested Vera on a coding workflow involving repository cloning and sandboxed tests, completing the job about 1.5x faster than x86 and starting concurrent sandboxes up to 1.9x faster. Partners also measured 3x faster large-scale SQL analytics with Starburst and up to 6x lower latency on real-time streaming with Redpanda against leading x86 server CPUs.