RadixArk and Google Cloud said they are working together to make SGLang, a widely used open-source system for serving AI models, run on Google's TPU chips as well as on graphics processors. The JAX-based version, SGL-JAX, is available now and supports model families including Gemma, Qwen, DeepSeek and Grok. A PyTorch-native TPU version is planned for later in 2026, and the partners say new open models will be supported on TPUs the same day as on GPUs.
What changed
SGLang ran mainly on GPUs, so teams wanting to serve models on Google's own AI chips had to switch to different serving software.
What it unlocks
Running the same open-source model-serving setup and its API on Google Cloud TPUs instead of GPUs, using the JAX version available today.
- more than 30,000 GitHub stars
- over 1,700 contributors
What you need to act on it
- Google Cloud TPU access
- SGL-JAX today; the PyTorch-native TPU backend arrives later in 2026
- lmsys.org2026-07-30