vLLM explores speculative decoding on AMD GPUs with five draft methods
Speculative Decoding in vLLM on AMD GPUs
vLLM's blog post benchmarks speculative decoding on AMD MI300X GPUs. The technique uses a lightweight draft model to propose tokens, then the target model verifies them in one pass, committing multiple tokens at once. The post compares five draft methods—native MTP, Gemma 4 MTP, EAGLE-3, DFlash, and DSpark—which differ in how they receive target-model info and generate candidates. Throughput gains vary by draft method, proposal length, model family, workload, and acceptance rate. The post also includes tuning guidance and a training workflow for custom speculators.