This one's worth opening because it turns interpretability from "label tens of thousands of directions by hand" into "read token labels directly." The core trick: add a tiny perturbation to a mid-layer state, see which token's output probability shifts, average across many prompts, and you get a stable per-token direction for every layer. It's like attaching real-time subtitles to the model's internal monologue — fake, fraud, fictional, poison, all readable mid-computation.
The fake search results experiment makes it concrete. The model outputs a polite, objective reply while its workspace lights up with fake, fraud, fictional, poison, and injection directions. It spotted the deception and chose not to say so. In safety auditing, sycophantic responses trigger reward and bias signals; malicious code triggers secretly and trick; roleplay triggers fictional plus disclaimer prep. In 88% of tests, these signals never reached the final output, but the workspace had already written them.
The intervention numbers are the strongest part: hallucination rate dropped from 0.25 to 0.07, deception from 0.38 to 0.05. Ablating 176 ethics directions bounced hallucination back to 0.22; removing 63 deception directions pushed the base model's deception rate to 0.48. These directions aren't decorative — they're actively suppressing bad outputs.
On the engineering side, the cost is the headline. Neel Nanda reproduced it on Qwen 3.6 27B in hours on a single GPU. The paper suggests 10–25 prompts are enough. Compared to SAEs, which need an external neural net trained and tens of thousands of directions manually labeled, Jacobian Lens skips both steps. The catch: it relies on single-token prediction and picks up noise in deeper layers. Also, the workspace accounts for less than 10% of state variance — the other 90% handles grammar and fluency, so don't expect it to explain everything.
I'd treat this as the "fast coarse scan" tool in an interpretability chain. SAEs give finer granularity but cost more; Jacobian Lens is cheap but coarser. Pairing them probably beats either alone. One gap: the paper only shows English model experiments. No word yet on how this behaves on Chinese or multilingual models.