Skip to content
Hugging Face Blog

Safety alignment should refuse the harmful subset of a topic, not the whole topic

Safety for Whom? Refusing the Right Subset of a Topic, Not the Whole Topic

Multiverse Computing's new paper argues that current safety alignment treats entire topics as refusal units—LlamaGuard-3, for instance, labels elections as 'factually incorrect information,' causing models to refuse even benign queries. They propose 'narrow-boundary safety': within a single topic like politics, refuse only the harmful subset (e.g., writing targeted manipulation) while still answering benign questions (e.g., election facts). The method uses deployment-specific boundary labels to self-distill a model that respects per-setting splits instead of topic-level blocks. Experiments focus on politics; the post doesn't disclose generalization results for other topics.

Why it matters: H and K both hit: the angle is sharp and the LlamaGuard-3 mislabeling case is concrete. R is weak — this is a safety-alignment niche topic that won't resonate broadly. Landed at the featured threshold of 72; didn't go higher because the body excerpt is partial, with full exper...

Read the original ↗Export Markdown