NVIDIA revives Rubin CPX with HBM4 memory
NVIDIA has reportedly restarted development of its Rubin CPX AI accelerator after previously putting the project on hold. Supply chain analyst Ming Chi Kuo reported that the revised design replaces the earlier GDDR7 plan with HBM4 memory. The accelerator, derived from the Rubin GPU family, was originally aimed at large-scale agentic AI workloads and had been slated for 128 GB of GDDR7 memory, but the revived version is now expected to include 168 GB of HBM4 memory.
The Rubin CPX is designed for prefill workloads, including input context handling for LLMs and KV cache processing, while regular Rubin GPUs would handle decode tasks. NVIDIA is expected to house the SKU in separate racks with configurations ranging from 64 to 256 standalone CPX GPUs per rack. A rack tray with eight CPX GPUs can support 1.34 TB of long-context prefill and the related KV cache. The shift to HBM4 has reportedly required a package redesign, likely using TSMC’s CoWoS-S or CoWoS-L packaging.