Skip to content
Xinzhiyuan · WeChat

Anthropic Translates Claude’s Internal Activations into Natural Language with NLA

Claude脑子里想的,被翻译成人话了!Anthropic新研究看懵人类

Anthropic released Natural Language Autoencoder to translate Claude activation vectors into text; on Opus 4.6 it reached 60%-80% variance explained, and across 16 evaluations NLA detected unspoken evaluation awareness on 26% of SWE-bench Verified tasks.

Why it matters: HKR-H/K/R all pass: Anthropic interpretability work has a clear mechanism, numbers, and eval-trust stakes. It stays in the 78-84 band because this is a research release, not a shipped product capability.

Read the original ↗Export Markdown