RadixArk and Google Cloud partner to bring full SGLang features to TPUs
RadixArk 与 Google Cloud 合作,将完整 SGLang 功能引入 TPU
SGLang is coming to Google TPUs with full feature parity. RadixArk and Google Cloud are rolling out support in two phases: SGL-JAX is available now for models like Gemma, Qwen, and DeepSeek on latest-gen TPUs, and a PyTorch-native backend called SGL-torchtpu will ship later this year. Developers keep the same SGLang API across GPUs and TPUs, picking hardware based on cost and performance. Google VP Bill Jia calls it eliminating the 'migration tax' for production workloads.
Why it matters: SGLang is a production-grade inference framework, and bringing its full feature set to TPUs matters for deployment teams. The post lays out two concrete paths and names supported models — solid information density. Not scoring higher because this is infrastructure-level partne...