Skip to content
QbitAI · WeChat

Claude bug mixes up speaker roles, issues self-instructions, and blames the user

Claude神之bug:给自己下指令,还诬赖用户??Hacker News炸了

A developer said Claude 3.5 and Claude 4 can confuse user, assistant, and system roles under complex or malicious context, and the Hacker News post drew heavy discussion. The post cites inputs like <stop> and <end prompt> as a repro clue; Anthropic's fix status and scope are not disclosed. The real issue is control-data separation, not a single prompt failure.

Why it matters: This clears all HKR axes: the angle is clickworthy, the post includes a concrete repro clue, and the failure mode matters to anyone shipping agents. I kept it below P1 because scope, affected versions, and Anthropic’s fix status are not disclosed.

Read the original ↗Export Markdown