Release2026-09-03

NVIDIA announced a set of local AI tools and speed-ups at IFA 2026 in Berlin. Hermes Agent, OpenClaw and Perplexity Portable Computer are gaining simplified setup for local models on NVIDIA GPUs. Each is built on llama.cpp and carries NVIDIA's inference tuning. NVIDIA says new llama.cpp and vLLM work makes local inference faster, and the gains reach users through LM Studio and Ollama. It also released PAIR, a free open-source tool that spreads inference requests across other PCs on a home network. PAIR is in beta for Windows, macOS and Linux. RTX Spark Windows PCs arrive in October, with designs from Lenovo and Acer. CyberLink's PhotoDirector AI PC Mode will add on-device image editing when RTX Spark launches.

What changed

Running a local agent meant picking a model, finding an inference server and tuning quantization by hand.

What it unlocks

One-click setup of a local model inside Hermes Agent, plus spreading inference jobs across idle PCs on a home network.

  • up to 1.9x faster on llama.cpp
  • 1.2x faster on RTX PRO 6000
  • 24GB VRAM minimum for local setup
  • 128GB unified memory on RTX Spark

What you need to act on it

  • an NVIDIA RTX GPU with at least 24GB VRAM
  • Windows for the one-click agent setup
  • GeForce RTX 20 Series or newer for PAIR

Sources

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.