Skip to content
Hacker News front page

Can LLMs notice what's inside their own activations? This paper tests with injected representations

Emergent Introspective Awareness in Large Language Models

Jack Lindsey bypasses conversation-based introspection tests by injecting known concept representations into model activations and measuring whether the model can report them. Claude Opus 4 and 4.1 lead most experiments: they notice injected concepts, recall prior internal states, and distinguish their own outputs from human-written prefills. The capacity is real but highly unreliable and context-dependent, with trends sensitive to post-training choices. Models can also modulate their activations when told or incentivized to 'think about' a concept. The paper calls this functional introspective awareness, not anything like human self-reflection.

Why it matters: Novel method — not conversation testing but direct activation manipulation — with Claude Opus 4 standing out and concrete experimental findings. But the capability is unstable and context-dependent, and it's a single paper without cross-source discussion yet, so it stays below...

Read the original ↗Export Markdown