NVIDIA pushes local AI on RTX systems at IFA 2026
NVIDIA, Microsoft and partners used IFA 2026 to expand local AI support across RTX and DGX hardware, with simplified setup for Hermes Agent, OpenClaw and Perplexity Portable Computer. The updates are aimed at reducing manual model selection, inference server setup and tuning, while keeping more agent workflows on user devices.
Perplexity Portable Computer runs on NVIDIA RTX GPUs with at least 24GB VRAM on Linux, with Windows support coming soon, and can escalate selected tasks to 15+ frontier models in the cloud with user permission. Hermes Agent adds one-click local model setup on Windows for RTX and DGX systems, while OpenClaw, described as the largest AI project on GitHub with more than 380K stars, is getting an optimized Windows setup for RTX GPUs with at least 24GB of VRAM.
NVIDIA also highlighted inference gains for local workloads, including llama.cpp throughput improvements of up to 1.9x on a GeForce RTX 5090, plus vLLM gains of 1.2x on RTX PRO 6000 Blackwell Workstation Edition and up to 1.4x on two DGX Spark clusters. A new open source NVIDIA Personal AI Router beta can distribute inference requests across compatible PCs on a local network.
RTX Spark Windows PCs are scheduled to arrive in October 2026, with Acer showing a compact desktop concept and Lenovo announcing Yoga Pro 9n and Yoga 9n 2-in-1 systems. NVIDIA described RTX Spark as using a 1 Petaflop RTX Blackwell GPU, up to 128GB of unified memory and a 20-core Grace CPU for creators, gamers and local agents.