Skip to content
r/LocalLLaMA

I trained Qwen3.5 to jailbreak itself with RL, then used the failures to improve its defenses

The author built an RL-based automated red-teaming loop for Qwen3.5, raising defense rate from 64% to 92% while benign accuracy fell from 92% to 88%, and the attacker found 7 tactic families.

Why it matters: HKR-H/K/R all pass: a named first-person RL red-team loop with concrete rates and failure modes. Source is a single Reddit post without paper/code validation, so it stays below P1.

Read the original ↗Export Markdown