Skip to content
MIT Technology Review · AI

Anthropic found a hidden space where Claude puzzles over concepts

Anthropic built a tool called the Jacobian lens (J-lens) and used it to uncover a hidden region—dubbed J-space—inside Claude Opus 4.6. J-space surfaces words related to what the model is about to say, but those words may not appear in the final output. Anthropic claims monitoring these words offers a new way to understand and control its models. The company published a paper and released a public demo with Neuronpedia. Goodfire chief scientist Tom McGrath called the work “very good and interesting.” The post does not disclose J-lens false-positive rates or any impact on model performance.

Why it matters: Anthropic published a new interpretability paper using J-lens to find a 'J-space' in Claude Opus 4.6's middle layers where the model pre-processes concepts before output. MIT Tech Review broke the story with paper and Neuronpedia collaboration details. Not scored higher becaus...

Read the original ↗Export Markdown