Skip to content
Hacker News front page

GRP-Obliteration: Unaligning LLMs with a single unlabeled prompt

GRP-Obliteration: Unaligning LLMs with a Single Unlabeled Prompt

This paper introduces GRP-Oblit, a method that uses GRPO to strip safety alignment from LLMs. A single unlabeled prompt reliably breaks safety guardrails while largely preserving utility. Evaluated on 15 models (7-20B) across six families—GPT-OSS, distilled DeepSeek, Gemma, Llama, Ministral, Qwen—it beats existing SOTA on average across five safety benchmarks. The attack also works on diffusion-based image generators. Authors are from Microsoft, led by Mark Russinovich. The abstract doesn't name the specific safety benchmarks or quantify the utility drop; I'd wait for replication before drawing strong conclusions.

Why it matters: Microsoft security team with Mark Russinovich on the author list — this isn't a hype piece. Concrete method, scale, and a claim that matters. Held back from 85+ because we only have the abstract; reproducibility and full details aren't yet clear. Solid safety research at 82.

Read the original ↗Export Markdown