Liquid AI adds a 280M speculative decoding drafter to its 3B vision model, hitting 3.13× decode speedup on-device
Accelerating vision-language models with LFM2.5-VL-DSpark
Liquid AI released LFM2.5-VL-DSpark, an experimental speculative decoding drafter for its LFM2.5-VL-3B vision-language model. The drafter adds only 280M parameters (8.9% of the 3B target), leaves output quality unchanged, and delivers up to 3.13× decode speedup on-device and 2.66× on an H100; end-to-end gains reach 2.62× and 2.27×. It taps hidden states from intermediate layers of the target model to draft candidate tokens—image patches and text tokens are projected into a shared representation beforehand, so the inference algorithm stays identical to the text-only version. Day-one integrations include llama.cpp, MLX-VLM, and SGLang. The post does not disclose training data size, absolute latency numbers, or speedup variation across batch sizes.
Why it matters: Liquid AI shipped a speculative decoding module for its 3B vision model, hitting 3.13x on-device and 2.66x on H100 — concrete, reproducible numbers. But Liquid AI's ecosystem is small, so this reads more like a technical proof than an industry event, landing right at the featu...