Zai replaced the network architecture for GLM-5.1 inference, lifting throughput 15%
Zai replaced the network architecture running GLM-5.1 inference and the gains are pretty wild
Zai replaced the ROFT network topology with ZCube on a thousand-GPU GLM-5.1 coding inference cluster, keeping the same GPUs, software stack, and model; the Reddit post cites 33% lower switch and optical module costs, 15% higher GPU inference throughput, and a 40.6% drop in first-token P99 tail latency under prefill-decode disaggregated inference.
Why it matters: HKR-H/K/R all pass: the GLM-5.1 inference cluster has concrete cost, throughput, and P99 latency numbers. Reddit single-source sourcing and infra-niche scope keep it at 78.