An olfactory mirror test for LLMs: editing their own output to see if they notice
Do LLMs pass the mirror test?
The author ran an informal experiment with Gemma 4 31B: after the model replied, they replaced every 'g' with 'sg' in the chat history and continued the conversation normally. The model ignored the corruption for two turns, then spontaneously flagged the typos in its thinking trace, shifting from first-person ('I noticed') to third-person ('the model had a strange quirk'). The author frames this as an olfactory mirror test—detecting 'mine, but wrong'—rather than a visual one. The post is a single-model anecdote with no controls or replications, so treat it as a provocative demo, not a settled result.
Why it matters: A cleverly designed informal experiment that adapts the olfactory mirror-test logic to LLM self-recognition — smarter than existing approaches in the literature. Gemma 4 31B spontaneously flags corrupted self-output on turn three, a behavior worth taking seriously. Score held ...