Skip to content
Synced · WeChat

Zhipu deploys ZCube, raising inference throughput 15% on the same GPUs

推翻二十年组网逻辑,智谱落地ZCube,让同样的GPU多干15%的活

Zhipu deployed ZCube in a thousand-GPU GLM-5.1 production inference cluster, replacing ROFT while keeping GPUs, software stack, and business code unchanged; throughput rose by over 15%, TTFT P99 fell 40.6%, and switch plus optical module costs dropped by one third.

Why it matters: HKR-H/K/R all pass: Zhipu reports ZCube in a GLM-5.1 1k-GPU production inference cluster with +15% throughput and 40.6% lower TTFT P99. Single-source infra optimization keeps it below major model-release weight.

Read the original ↗Export Markdown