NVIDIA announced a set of local AI tools and speed-ups at IFA 2026 in Berlin. Hermes Agent, OpenClaw and Perplexity Portable Computer are gaining simplified setup for local models on NVIDIA GPUs. Each is built on llama.cpp and carries NVIDIA's inference tuning. NVIDIA says new llama.cpp and vLLM work makes local inference faster, and the gains reach users through LM Studio and Ollama. It also released PAIR, a free open-source tool that spreads inference requests across other PCs on a home network. PAIR is in beta for Windows, macOS and Linux. RTX Spark Windows PCs arrive in October, with designs from Lenovo and Acer. CyberLink's PhotoDirector AI PC Mode will add on-device image editing when RTX Spark launches.
What changed
Running a local agent meant picking a model, finding an inference server and tuning quantization by hand.
What it unlocks
One-click setup of a local model inside Hermes Agent, plus spreading inference jobs across idle PCs on a home network.
- up to 1.9x faster on llama.cpp
- 1.2x faster on RTX PRO 6000
- 24GB VRAM minimum for local setup
- 128GB unified memory on RTX Spark
What you need to act on it
- an NVIDIA RTX GPU with at least 24GB VRAM
- Windows for the one-click agent setup
- GeForce RTX 20 Series or newer for PAIR
Sources