llama.cpp prompt lookup drafting gets 42x faster
42x Faster Prompt Lookup Drafting in llama.cpp
The prompt lookup speculative decoding implementation in llama.cpp was rewritten and is now 42x faster. The post only provides a title and a preview image; it doesn't disclose the optimization approach, tested models, or hardware. I'd hold off until the full blog post is available.