Skip to content
Computing Life · Share · Yage

The Technology Behind GLM-5.1 Reaching 400 Tokens/s

GLM-5.1 达到 400 tokens/s 背后的技术:当推理速度成为新的 Scaling Law

Zhipu GLM-5.1 high-speed API claims 400 tokens/s, and the post says TileRT reconstructs GPU inference at the execution-model level; the RSS snippet does not disclose benchmark conditions, hardware, pricing, or latency distribution.

Why it matters: HKR-H/K/R all pass: 400 tokens/s is a concrete hook, TileRT adds mechanism, and latency/cost resonates with builders. It stays at 78 because the speed is claimed, with no independent test or pricing condition disclosed.

Read the original ↗Export Markdown