Skip to content
AI HOT (Curated Pool)

SGLang adds Waterfill and LPLB to improve DeepEP MoE load balancing

SGLang 引入 Waterfill 与 LPLB 提升 DeepEP MoE 负载均衡

LMSYS and NVIDIA introduced two dispatch-time load balancers in SGLang for DeepEP MoE inference. Waterfill handles the shared expert by sending it to the least-loaded GPU; it boosts throughput by 1.48–4.66% on DeepSeek V3/R1 and lifts DeepSeek V4 from 49,253 tok/s to 51,677 tok/s (+4.92%). LPLB uses linear programming to route tokens across redundant expert replicas, adding 0.84–7.34% throughput on the same benchmarks. Both methods preserve model accuracy. The post reports throughput gains but does not disclose latency impact.

Why it matters: Joint release from LMSYS and NVIDIA with concrete throughput gains and two named methods. Useful for anyone running DeepSeek-family models in production. Narrow audience and purely engineering-layer keeps it at the featured threshold.

Read the original ↗Export Markdown