Release2026-08-24

NVIDIA said at Hot Chips 2026 that its Groq 3 LPX inference chip is now in full production, working alongside the Vera Rubin NVL72 rack systems. In a company demonstration the combination produced 3,400 tokens per second on the Gemma 4 31B open model, which NVIDIA describes as the fastest recorded for that model. The performance figures come from NVIDIA, and access for most buyers will come through cloud providers such as Nebius rather than direct purchase.

What changed

Groq 3 LPX had been announced but was not yet in full production.

What it unlocks

Ordering inference hardware that pairs with Vera Rubin racks for faster response times in agent workloads.

  • 3,400 tokens per second
  • 4x faster than nearest alternative
  • 100,000-token context in the demo

What you need to act on it

  • data-centre-scale hardware purchase or access via an AI cloud provider such as Nebius

Send this to someone who needs it

Shares the story and its sources. Nothing about you.

What does this mean for your job?

This is the story as everyone gets it. Once a week we send you the version written for your role — what changed, why it matters for the work you actually do, and one thing to try. Free while we tune it.