Trending storyPast story
llama.cpp merges tiled mul_mat for k-quants, CPU inference speedup expected
1 report1 sourceupdated 4 days ago
What happened
Summary
llama.cpp 合并了一个 PR,给 CPU 推理加上了分块矩阵乘法(tiled mul_mat)。简单说就是把大矩阵切成小块算,让 CPU 缓存更友好、内存带宽压力更小——这对跑本地大模型是实打实的优化。不过正文被 Reddit 屏蔽了,没披露具体提速多少、内存省了多少。如果是真的,本地推理延迟应该能降一截。
Coverage
Follow the reports to see the story from different sides.
Sep 26
- r/LocalLLaMAllama.cpp merges tiled mul_mat for k-quants, CPU inference speedup expected
PR #27851 by jbooth adds tiled matrix multiplication for k-quants in llama.cpp's ggml-cpu backend. The post body is blocked by Reddit, so no speedup or memory numbers are disclosed. Tiled mul_mat improves CPU cache utilization and reduces memory bandwidth pressure—a real win for local LLM inference.