MTP speculative decoding with Gemma 4: assistant model choice makes or breaks speed gains
Not All MTP Assistants Are Created Equal
A user tested MTP speculative decoding with Gemma 4 Heretic models in llama.cpp and found assistant model selection is everything. A 26B Q8 jumped from 30 t/s to 62 t/s; a 12B Q4 went from 12 t/s to 54 t/s. Two GGUFs with the same name aren't always identical. Unquantized assistants consistently beat Q4/Q8 assistants by roughly 10 t/s. Draft count of 1 gave the best results across the board. Always check logs to confirm MTP actually initialized—otherwise you're benchmarking the base model by accident.
Why it matters: Solid benchmarks with concrete numbers: 26B Q8 went from 30 to 62 tok/s, 12B Q4 from 12 to 54 tok/s. Actionable for local inference users. Downside: single Reddit post with no cross-source verification, and Gemma 4 has a narrower audience than Llama/DeepSeek.