Skip to content
Computing Life · Share · Yage

High Fidelity Nearby, Lossy at a Distance: The Shared Intuition Behind Three Long-Context Approaches

YaRN, DeepSeek V4, and DeepSeek-OCR tackle long-context bottlenecks at the coordinate, information-pathway, and input-representation layers respectively, all converging on the same intuition: keep nearby tokens high-fidelity, compress distant ones. YaRN applies frequency-partitioned interpolation to RoPE, letting Llama 2 7B reach 128K context with ~384 A100 GPU hours. DeepSeek V4 uses full-attention within a 128-token window and heavy compression plus sparse selection beyond it, cutting V4-Pro's per-token FLOPs to 27% of V3.2 at 1M context. DeepSeek-OCR compresses full pages into dense visual tokens, hitting 97% text accuracy at 10x compression. The three were developed independently by different teams. The post also flags a new challenge—maintaining positional awareness after compression—and outlines three solutions: dual-track position scales, document-wise coordinate resets, and Kimi K3's removal of positional encoding entirely.

Why it matters: Unifies three long-context approaches across different tech stack layers under one sharp intuition, backed by concrete numbers. Not a primary research release, and the excerpt cuts off mid-argument, which caps the score.

Read the original ↗Export Markdown