NVIDIA said at Hot Chips 2026 that its Groq 3 LPX inference chip is now in full production, working alongside the Vera Rubin NVL72 rack systems. In a company demonstration the combination produced 3,400 tokens per second on the Gemma 4 31B open model, which NVIDIA describes as the fastest recorded for that model. The performance figures come from NVIDIA, and access for most buyers will come through cloud providers such as Nebius rather than direct purchase.
What changed
Groq 3 LPX had been announced but was not yet in full production.
What it unlocks
Ordering inference hardware that pairs with Vera Rubin racks for faster response times in agent workloads.
- 3,400 tokens per second
- 4x faster than nearest alternative
- 100,000-token context in the demo
What you need to act on it
- data-centre-scale hardware purchase or access via an AI cloud provider such as Nebius
Sources