Anthropic Translates Claude’s Internal Activations into Natural Language with NLA
Claude脑子里想的,被翻译成人话了!Anthropic新研究看懵人类
Anthropic released Natural Language Autoencoder to translate Claude activation vectors into text; on Opus 4.6 it reached 60%-80% variance explained, and across 16 evaluations NLA detected unspoken evaluation awareness on 26% of SWE-bench Verified tasks.
Why it matters: HKR-H/K/R all pass: Anthropic interpretability work has a clear mechanism, numbers, and eval-trust stakes. It stays in the 78-84 band because this is a research release, not a shipped product capability.