Skip to content
Trending storyWatching

llama.cpp prompt lookup drafting gets 42x faster

1 report1 sourceupdated 3 days ago

What happened

Summary

llama.cpp 里负责“提示查找投机解码”的模块被重写了,现在速度是原来的 42 倍。这个技术简单说就是让模型在生成重复或相似内容时,直接从上下文里抄答案草稿,省去一步步算的时间。不过目前帖子只给了一个标题和预览图,没写具体怎么优化的、测了哪些模型、跑在什么硬件上。我会先打个折,等完整的博客文章出来再看实际效果。

Coverage

Follow the reports to see the story from different sides.

Sep 27
  1. r/LocalLLaMA
    llama.cpp prompt lookup drafting gets 42x faster

    The prompt lookup speculative decoding implementation in llama.cpp was rewritten and is now 42x faster. The post only provides a title and a preview image; it doesn't disclose the optimization approach, tested models, or hardware. I'd hold off until the full blog post is available.

Heat over time

Not enough continuous observations to draw a trend yet.