Can LLMs notice what's inside their own activations? This paper tests with injected representations
Emergent Introspective Awareness in Large Language Models
Jack Lindsey bypasses conversation-based introspection tests by injecting known concept representations into model activations and measuring whether the model can report them. Claude Opus 4 and 4.1 lead most experiments: they notice injected concepts, recall prior internal states, and distinguish their own outputs from human-written prefills. The capacity is real but highly unreliable and context-dependent, with trends sensitive to post-training choices. Models can also modulate their activations when told or incentivized to 'think about' a concept. The paper calls this functional introspective awareness, not anything like human self-reflection.
Why it matters: Novel method — not conversation testing but direct activation manipulation — with Claude Opus 4 standing out and concrete experimental findings. But the capability is unstable and context-dependent, and it's a single paper without cross-source discussion yet, so it stays below...