Skip to content
r/LocalLLaMA

You can now read Gemma 3's mind

Anthropic released NLA research to explain Gemma 3 27B Instruct activations for each generated token. The post links Auto Verbalizer and Activation Reconstructor weights on Hugging Face. Neuronpedia hosts an interactive page; the post does not disclose evaluation scores.

Why it matters: HKR-H/K/R all pass: Anthropic interpretability research ships reproducible weights and a Neuronpedia UI. No eval scores are disclosed, so it stays in the 78–84 band, not P1.

Read the original ↗Export Markdown