The Technology Behind GLM-5.1 Reaching 400 Tokens/s
GLM-5.1 达到 400 tokens/s 背后的技术:当推理速度成为新的 Scaling Law
Zhipu GLM-5.1 high-speed API claims 400 tokens/s, and the post says TileRT reconstructs GPU inference at the execution-model level; the RSS snippet does not disclose benchmark conditions, hardware, pricing, or latency distribution.
Why it matters: HKR-H/K/R all pass: 400 tokens/s is a concrete hook, TileRT adds mechanism, and latency/cost resonates with builders. It stays at 78 because the speed is claimed, with no independent test or pricing condition disclosed.