GPT safety training launders gender bias instead of removing it
Harm Laundering in GPT Models: Gender Discrimination Transformed Rather Than
This EMNLP 2026 paper examines 450K gender-directed completions across 15 models from GPT-2 to GPT-5. Toxicity scores keep dropping, but discrimination changes shape: sexual violence clusters in GPT-2's women-directed output vanish by GPT-4, while men-directed completions gain positive framing—caregiving, emotional range, ally identity—that women-directed ones don't. At GPT-5, a 1,997-document topic cluster frames breast cancer as a men's rights debate; zero equivalent clusters appear for women. Three independent classifiers score this content as non-toxic. Topic diversity for women drops 36% relative to men at the GPT-4 alignment boundary. REGARD representational harm correlates with release date (ρ=+0.55), while Detoxify does not (ρ=−0.23). The authors call this 'harm laundering' and provide a three-stage detection protocol.
Why it matters: EMNLP 2026 paper with strong empirical backbone (450k completions, full GPT lineage) and a quotable new concept. HKR all hit, but as a single paper rather than a product launch, capped at 82.