SpecForge v0.3: LMSYS releases a disaggregated speculative decoding training stack and new open draft models
SpecForge v0.3.0 发布:统一解耦与共置投机解码栈,新增开放 SpecBundle 草稿模型
SpecForge v0.3 decouples target-model inference from draft-model training. Patched SGLang servers capture features, Mooncake transports tensors, and trainer workers consume them independently. On an 8×H20 testbed, 3 servers + 5 trainers deliver ~10% higher end-to-end training throughput than the previous colocated design. The runtime now supports six speculative decoding families—EAGLE3, DFlash, Domino, DSpark, and more—and ships community-contributed draft models trained entirely on open data.
Why it matters: LMSYS disaggregated speculative decoding training into three independent contracts — capture, delivery, lifecycle — and showed ~10% throughput gain on 8×H20 with 3 servers + 5 training nodes. The architecture is clean and the numbers are concrete, but the audience is narrow (i...