Skip to content
Trending storyDeveloping

AI refusal mechanisms have safety limits and costs, MIT Technology Review reports

1 report1 sourceupdated 3 hours ago

What happened

AI digest

On October 9, 2026, MIT Technology Review looked at research and interviews on treating AI refusal as the main pillar of safety. The report says refusals fail probabilistically, over-refuse and refuse silently, bringing risks of harm, censorship and loss of control. Refusal training and external classifiers cannot remove dangerous capabilities a model already has, and tighter guardrails also block legitimate research and raise running costs. Anthropic said one classifier added 24% to a chatbot's compute cost.

Written by AI from the coverage · updated 2 hours ago

Coverage

Follow the reports to see the story from different sides.

Oct 9
  1. MIT Technology Review · AI
    We’re putting too much faith in AI’s ability to say no

    AI 的拒绝机制已成为安全防护的主要支柱,但其概率性失效、过度拒绝和隐蔽拒绝可能带来伤害、审查与失控风险。文章结合研究和访谈指出,拒绝训练及外围分类器无法消除模型已有的危险能力,收紧防护也会阻碍合法研究;Anthropic 曾表示,一类分类器使聊天机器人的计算成本增加了 24%。

Heat over time

Not enough continuous observations to draw a trend yet.