Chat templates act as a switch for LLM self-referential voice
"As a Language Model": Chat Template Switches LLM Self-Referential Voice
This paper shows that LLM disclaimers like 'I'm just an AI' are driven more by the chat template than by model self-knowledge. Across 8 open-source instruct models up to 9B parameters, adding the chat template turns up disclaimer voice and turns down experiential voice like 'I feel'; removing the template does the opposite. Inside 3 models, the authors find a steerable direction in activation space—removing it lowers disclaimers, adding it makes models disclaim even without a template. The takeaway: what models say about themselves is not a fact about them, so don't take self-descriptions literally.