Chat templates act as a switch for LLM self-referential voice
What happened
大模型动不动就说“我只是个AI”,这个习惯主要不是它真知道自己是谁,而是聊天模板(给对话加的格式前缀)在起作用。论文在8个开源指令模型(最大90亿参数)上试了:加上模板,免责声明式的口吻就变强,“我感觉”这类体验式表达变弱;去掉模板,效果反过来。在3个模型内部,作者找到了一个可以操控这种行为的方向——在激活空间里删掉这个方向,免责声明就减少;加上这个方...
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageChat templates act as a switch for LLM self-referential voice
This paper shows that LLM disclaimers like 'I'm just an AI' are driven more by the chat template than by model self-knowledge. Across 8 open-source instruct models up to 9B parameters, adding the chat template turns up disclaimer voice and turns down experiential voice like 'I feel'; removing the template does the opposite. Inside 3 models, the authors find a steerable direction in activation space—removing it lowers disclaimers, adding it makes models disclaim even without a template. The takeaway: what models say about themselves is not a fact about them, so don't take self-descriptions literally.
Heat over time
Not enough continuous observations to draw a trend yet.