NVDA 203.28 ▲0.23%GOOGL 351.99 ▲1.51%MSFT 402.29 ▲2.15%AMD 503.57 ▲1.58%INTC 97.06 ▲2.13%TSMC 402.30 ▲0.99%AMZN 249.99 ▲1.12%META 645.85 ▼0.02%AAPL 326.59 ▼2.14%PLTR 134.85 ▲1.87%
Markets at last close

Apple · Research

Apple’s on-device AI push raises memory questions

·1 min read

Apple is exploring ways to run larger AI models directly on iPhones, with attention focused on PrismML’s work to compress large language models for local use. PrismML has managed to shrink Alibaba’s open-source Qwen 3.6 model to run entirely on an iPhone 17 Pro, despite the model having 27 billion parameters.

The technical debate centers on whether extreme compression can make local inference practical without unacceptable losses in accuracy or performance. One view is that reducing neural network parameters to one bit per parameter could let Apple’s 10B-parameter on-device model fit in just over 1GB of RAM, making it feasible for a 6 GB iPhone, while the iPhone 15 has only 6GB of RAM.

Memory remains the main constraint for larger on-device models, followed by bandwidth and compute cores. Some locally runnable 4-bit quantized models range up to ~20B parameters and can fit in 8-12GB, but tighter memory limits can increase compromises in accuracy and hallucination rates. Larger context windows, personalization through user data, and multimodal inputs also add memory demands.

Supporters see on-device AI as a way to improve privacy, speed, and energy efficiency compared with cloud processing. Skeptics argue that newer models will continue to outpace phone hardware and that highly compressed approaches may require too much processing power, risking battery drain on current iPhones.

Originally reported by forums.macrumors.comRead the source →
Related coverage
All Apple news →