NVDA 228.45 ▲1.80%GOOGL 342.48 ▲1.59%MSFT 510.12 ▲2.68%AMD 456.16 ▼0.20%INTC 91.67 ▲1.80%TSMC 417.01 ▲0.36%AMZN 258.90 ▲1.54%META 610.68 ▲3.01%AAPL 328.21 ▲1.00%PLTR 182.53 ▲7.71%
Markets at last close

Micron · Models

Micron explores NAND closer to GPUs for larger LLMs

·1 min read

Micron is reportedly investigating high-endurance NAND Flash modules designed to sit closer to GPUs instead of remaining in distant storage pools behind multiple protocols. The concept, described as near-GPU NAND, would target workloads that do not need the latency and bandwidth of HBM or regular DRAM, while still requiring faster access than conventional system storage can provide.

The architecture would give GPUs access to a NAND Flash pool with lower density than conventional TLC or QLC NAND, prioritizing I/O speed, bandwidth, and read performance. A potential implementation could place hundreds of gigabytes of storage directly on the GPU package or PCB, in a role similar to DRAM technologies such as HBM or GDDR7/LPDDR6.

The added tier would sit between GPU memory and regular system storage, functioning as a fast buffer for inference on massive LLMs. By reducing memory constraints, systems could potentially run larger models with fewer GPUs, using additional storage and memory tiers to support the required inference size. The NAND would also be significantly cheaper than HBM or regular DRAM and sized to system specifications.

Originally reported by techpowerup.comRead the source →
Related coverage
All Micron news →