NVIDIA published its own measurements showing its new Vera Rubin NVL72 systems do far more AI work per unit of electricity than the previous GB300 generation, with a matching drop in the cost of producing text. The figures come from recorded real-world coding sessions run by AI agents rather than short chat requests. The results are NVIDIA's own, are still pending review by the analyst firm whose test workload was used, and do not yet include the system's new CPU.
What changed
GB300 NVL72 was NVIDIA's top system for agentic inference, itself up to 15x Hopper's throughput per megawatt.
What it unlocks
Running long, multi-step AI agent sessions continuously within a fixed data centre power budget.
- 30x throughput per megawatt vs GB300
- 35x lower cost per million tokens
- up to 40% more GPUs per megawatt
What you need to act on it
- access to Vera Rubin NVL72 hardware or a cloud provider offering it
Sources