Meet the New Biologists Treating LLMs Like Aliens
Meet the new biologists treating LLMs like aliens
MIT Technology Review reports that Anthropic, OpenAI, and Google DeepMind are using mechanistic interpretability to study LLMs; as a scale reference, a 200B-parameter model in 14-point print would cover 46 square miles. The post says Anthropic uses sparse autoencoders to mimic target models, linked a Claude 3 Sonnet region to the Golden Gate Bridge in 2024, and in a July experiment found Claude used different internal paths for “bananas are yellow” versus “bananas are red.” The key point for practitioners is that weak internal coherence constrains alignment and predictability.
Why it matters: Strong HKR-H/K/R: the framing is novel, and the piece includes concrete mech-interpretability examples rather than vague opinion. I score it as featured but below the top band because this is a high-quality reported synthesis, not a fresh model launch or a single new breakthrough