NVDA 230.86 ▲1.09%GOOGL 338.24 ▼1.70%MSFT 512.80 ▼0.02%AMD 615.73 ▲0.65%INTC 120.00 ▼0.19%TSMC 459.20 ▲0.66%AMZN 248.23 ▼0.37%META 725.93 ▲0.10%AAPL 330.32 ▼0.81%PLTR 190.04 ▲1.60%
Markets at last close

Microsoft · Infrastructure

Agent workloads put server CPUs back on the critical path

·1 min read

AI agents are shifting server CPU demand from a supporting role into a core infrastructure constraint. Unlike chatbot requests that mostly wait on GPU inference, agent workflows loop through reasoning, tool calls, code execution, API requests, verification, and orchestration, much of which runs on the CPU while accelerators sit idle.

Microsoft Azure production studies put the host CPU on the critical path. In a 24-hour fleet trace, tool execution matched or exceeded LLM inference time for more than 27% of requests. Across 13.5 million GitHub Copilot sessions, LLM and tool calls ran 1:1, while 92% of tool wall-clock time sat on the critical path. Those findings point to requirements including loaded per-core latency, cache and predictor isolation, role-specific cores, memory capacity per core, and hardware-assisted scheduling.

NVIDIA Vera is described as the best-aligned current architecture because of its big-core design, statically partitioned threads, and single compute die with no NUMA. Intel is strongest on memory uniformity and priority cores, Arm on memory latency and partitioning, and AMD on density and power management, but none includes the scheduling engine Microsoft prioritized.

Futurum models the 2030 server CPU market at $245.9 billion, with standalone AI CPUs at $164.7 billion, or 67%. The winning architecture has not been built yet, and catching up depends more on design choices than process nodes.

Originally reported by futurumgroup.comRead the source →
Related coverage
All Microsoft news →