NVDA 207.29 ▲1.97%GOOGL 347.15 ▼1.38%MSFT 397.75 ▼1.13%AMD 544.43 ▲8.11%INTC 105.45 ▲8.64%TSMC 424.61 ▲5.55%AMZN 247.55 ▼0.98%META 643.81 ▼0.32%AAPL 327.74 ▲0.35%PLTR 132.66 ▼1.62%
Markets at last close

Nvidia · Chips

NVIDIA details Rubin GPU architecture

·1 min read

NVIDIA showcased its Rubin GPU architecture as an accelerator designed to scale computing across multiple racks, systems, and data centers. The design is capable of delivering 50 PetaFLOPS of compute power at sparse NVFP4 operations, while the full configuration includes up to 224 streaming multiprocessors and 288 GB of HBM4 memory. The chip integrates 336 billion transistors and 896 Tensor Cores with a third-generation Transformer Engine.

Rubin organizes GPU resources into Graphics Processor Clusters with a centralized L2 cache, while the GigaThread Engine coordinates workflows and improves resource utilization. MIG Control partitions allow the GPU to be divided into multiple virtual GPUs. NVIDIA is using 12-high HBM4 modules with dedicated controllers and physical layers, reaching up to 22 TB/s peak memory bandwidth. NVLink 6 manages scaling and GPU communication, with all-to-all GPU links operating at 3,600 GB/s of fabric bandwidth, NVLink-C2C supporting CPU-GPU communication at 1,800 GB/s, and a PCIe Gen 6 switch providing 256 GB/s for external devices.

Originally reported by techpowerup.comRead the source →
Related coverage
All Nvidia news →