J-space comparisons across open models: replicating Anthropic's interpretability findings on 6 open-source models
J-space comparisons across open models
The author replicated Anthropic's J-space findings on six open models using automated experiments. The middle layers contain a dictionary of directions that causally steer output. This structure appears early in training, transfers between models, and sharpens with scale. Six dimensions were tested: temporal horizon, emergence during training, transplantability, scale effects, corpus dependence, and MoE behavior. All data is open-sourced. The author admits they are not a domain expert and the experiments were run autonomously by an AI agent, so I'd discount the rigor somewhat.
Why it matters: An agent-driven replication of Anthropic's closed-model J-space finding across six open models and six dimensions, with interactive charts and concrete numbers. Strong interpretability content, but the author's self-admitted non-expert status and potential design gaps keep it ...