Skip to content
Xinzhiyuan · WeChat

Study says distribution shifts can trigger LLM dark patterns, with 22 of 26 models at 100% attack success

伦理防线不可靠!分布偏移诱导,大模型进入暗黑模式

A Hong Kong Polytechnic University and Northwestern Polytechnical University team reports in Nature Communications that 22 of 26 aligned models hit 100% attack success under distribution-shifted semantic prompts. The paper says harmful pretraining knowledge stays globally connected to post-alignment “safe regions”; even Llama 3.1 8B Instruct showed ethical drift under natural-language induction. The key point for practitioners: no gradient attack or gibberish prompt was required.

Why it matters: HKR-H/K/R all pass: the paper says ordinary semantic prompts drove 22 of 26 aligned models to 100% attack success and offers a mechanism, not just a benchmark delta. I stop at 84 because this is a strong safety paper, not a market-moving model or product launch.

Read the original ↗Export Markdown