Skip to content
AI HOT (Curated Pool)

SGLang Integrates DSpark: Confidence-Driven, Variable-Length Speculative Decoding

SGLang 集成 DSpark 推测解码:置信度驱动的可变长度验证

LMSYS integrated DSpark speculative decoding into SGLang. The draft model's confidence head decides how many tokens to verify per request, instead of always verifying a full block. On H200 with DeepSeek-V4-Flash, it beats both MTP and non-spec baselines in throughput and per-user decode speed. The key engineering move is a ragged, variable-length verify under full CUDA graphs: when the scheduler trims the verify budget, the engine replays a genuinely smaller graph, not a padded one. A cost-table profiler lets the scheduler size each request's budget online, and a cap-accept mode exposes the acceptance ceiling hidden by trimming.

Why it matters: LMSYS integrated DSpark into SGLang with real H200 benchmarks—not a paper announcement but an engine-level landing. H and K are solid, but the topic is inference optimization, which doesn't resonate broadly, so R is absent and the score stays at the featured threshold of 78.

Read the original ↗Export Markdown