Skip to content
r/LocalLLaMA

llama.cpp prompt lookup drafting gets 42x faster

42x Faster Prompt Lookup Drafting in llama.cpp

The prompt lookup speculative decoding implementation in llama.cpp was rewritten and is now 42x faster. The post only provides a title and a preview image; it doesn't disclose the optimization approach, tested models, or hardware. I'd hold off until the full blog post is available.

Read the original ↗Export Markdown