LMSYS releases DFlash speculative decoding, boosting Qwen 3.5 397B throughput by 4.3x
下一代投机解码:DFlash 与 Spec V2
LMSYS, Z Lab, and Modal jointly released DFlash speculative decoding with the Spec V2 engine. DFlash uses a block diffusion draft model to generate an entire block of candidate tokens in parallel, instead of one-by-one like traditional drafters. It injects hidden states from the target model's middle layers directly into the draft model's KV cache, so the drafter skips full context modeling and focuses on predicting the next block. On the HumanEval coding benchmark, Qwen 3.5 397B-A17B with DFlash achieves 4.3x the throughput of the baseline and 1.5x that of native MTP. Draft models are released in triplicate on Hugging Face and runnable with a single SGLang command.
Why it matters: LMSYS, Z Lab, and Modal ship DFlash + Spec V2 with KV injection and block-diffusion parallel drafting — a concrete inference-acceleration improvement with an SGLang implementation. But pure low-level inference optimization has a high bar for general readers and lacks broad res...