AI inference chip market projected to reach USD 923.72 billion by 2035
The global AI inference and accelerator chips market reached USD 115.60 billion in 2025 and is projected to reach USD 923.72 billion by 2035, with a CAGR of 23.1% during 2026 To 2035. Growth is tied to rising use of generative AI, agentic AI, large language models and real-time intelligent applications in data centers, cloud platforms and edge environments, where buyers are seeking faster inference, lower latency, better energy efficiency and workload-specific computing.
Recent activity points to a shift toward custom silicon and disaggregated inference infrastructure. OpenAI and Broadcom unveiled Jalapeño in June 2026 for LLM inference, AMD and Cerebras announced a July 2026 partnership combining AMD Helios with Cerebras’ Wafer-Scale Engine, and AMD agreed in August 2026 to acquire Taalas. NVIDIA said Groq 3 LPX entered full production in August 2026 as an interactive inference accelerator for fast token generation and latency-sensitive agentic AI workloads.
Competition is concentrated among NVIDIA, AMD, Intel, Google, Amazon Web Services, Broadcom, Qualcomm, Cerebras Systems and Groq, with differentiation based on compute performance, memory bandwidth, latency, power efficiency, software ecosystems and specialized acceleration. The United States remains a core market for hyperscale data centers and custom AI silicon, while Japan, Germany and South Korea are seeing demand tied to robotics, automotive AI, industrial automation, edge computing, memory and semiconductor manufacturing.