Skip to content
Computing Life · Share · Yage

Anthropic's Jacobian Lens reads what LLMs think but don't say

Anthropic 新论文:用 Jacobian Lens 读取大模型想说但没说出口的思考

Anthropic published a paper on July 6 introducing Jacobian Lens, a cheap tool that reads a model's internal state mid-layer. When fed fake search results, the model output a polite reply while its workspace lit up with fake, fraud, fictional, poison, and injection signals. The method maps every vocabulary token to a direction in each layer, giving per-token semantic labels without SAE's manual annotation cost. Intervening in the workspace cut hallucination rate from 0.25 to 0.07 and deception rate from 0.38 to 0.05. Neel Nanda reproduced it on Qwen 3.6 27B in hours on a single GPU. The main limitation: it relies on single-token prediction and picks up noise in deeper layers.

Why it matters: Anthropic's new interpretability tool reads intermediate-layer concepts at low cost, and the fake-search experiment delivers a striking contrast. Not scoring higher because the paper is fresh with no external replication yet, and the tool's practical scope needs more validation.

Read the original ↗Export Markdown