Skip to content
Trending storyWatching

Liquid AI adds a 280M speculative decoding drafter to its 3B vision model, hitting 3.13× decode speedup on-device

1 report1 sourceupdated 2 days ago

What happened

Summary

Liquid AI 放出了一个实验性的“投机解码”加速模块,专门配给自家的 3B 视觉语言模型 LFM2.5-VL-3B。这个模块只多了 2.8 亿参数,占原模型的 8.9%,但能让纯解码部分在端侧跑快 3.13 倍,在 H100 上快 2.66 倍;算上图编码等环节,端到端最快分别能到 2.62 倍和 2.27 倍,而且输出质量不变。它的原理是偷看大...

Coverage

Follow the reports to see the story from different sides.

Sep 24
  1. Hugging Face BlogPick
    Liquid AI adds a 280M speculative decoding drafter to its 3B vision model, hitting 3.13× decode speedup on-device

    Liquid AI released LFM2.5-VL-DSpark, an experimental speculative decoding drafter for its LFM2.5-VL-3B vision-language model. The drafter adds only 280M parameters (8.9% of the 3B target), leaves output quality unchanged, and delivers up to 3.13× decode speedup on-device and 2.66× on an H100; end-to-end gains reach 2.62× and 2.27×. It taps hidden states from intermediate layers of the target model to draft candidate tokens—image patches and text tokens are projected into a shared representation beforehand, so the inference algorithm stays identical to the text-only version. Day-one integrations include llama.cpp, MLX-VLM, and SGLang. The post does not disclose training data size, absolute latency numbers, or speedup variation across batch sizes.

Heat over time

Not enough continuous observations to draw a trend yet.