# llama.cpp 的提示查找投机解码重写后快了 42 倍

> 原标题：42x Faster Prompt Lookup Drafting in llama.cpp

- 来源：r/LocalLLaMA
- 发布时间：2026-09-27T00:23:41.000Z
- AX AI 日报：https://ai-daily.ax0x.ai/items/59161
- 原文：https://www.reddit.com/r/LocalLLaMA/comments/1wr5ylm/42x_faster_prompt_lookup_drafting_in_llamacpp/

## 摘要

llama.cpp 里负责“提示查找投机解码”的模块被重写了，现在速度是原来的 42 倍。这个技术简单说就是让模型在生成重复或相似内容时，直接从上下文里抄答案草稿，省去一步步算的时间。不过目前帖子只给了一个标题和预览图，没写具体怎么优化的、测了哪些模型、跑在什么硬件上。我会先打个折，等完整的博客文章出来再看实际效果。
