Skip to content
Computing Life · Share · Yage

DeepSeek Engram: Moving static knowledge out of GPU via lookup tables to free up reasoning capacity

把知识搬出 GPU:DeepSeek Engram 与模型的第二条稀疏轴

DeepSeek V4.1 Flash assigns 196B parameters to Engram, a conditional memory module stored in host RAM instead of GPU VRAM. Lookup keys are built from the last few tokens, so addresses are known ahead of time; RDMA prefetch hides the transfer latency behind computation. In the paper's self-reported results, reasoning gains outpace knowledge gains: BBH +5.0, needle-in-a-haystack retrieval jumps from 84.2 to 97.0. The mechanism: offloading static local mappings frees up early-layer compute and attention budget for multi-step reasoning and long-range dependencies. The team also introduces 'sparsity allocation'—experiments suggest ~20–25% of sparse capacity going to Engram works best, though no independent replication exists yet. Qwen3.8 Flash-Next adopts a similar design, signaling that external static memory is entering the mainstream.

Why it matters: DeepSeek packed a 196B-parameter lookup module called Engram into V4.1 Flash—no matmuls, no GPU memory residency, using hash keys and RDMA prefetch to decouple knowledge retrieval from compute. The self-reported gains are stronger on reasoning than on knowledge QA, which is co...

Read the original ↗Export Markdown