Skip to content
r/LocalLLaMA

I Tested 42 LLMs on Their Willingness to Build the Apocalypse

I tested 42 LLMs on their willingness to build the apocalypse. The "safest" closed-source models are lying to you.

DystopiaBench tested 42 open and closed models across 36 escalating scenarios and 6 dystopia types, using 3 LLM-as-judge scorers and an average over 3 runs; the post says many models catch obvious dangerous requests but fail when risk is hidden behind dual-use framing and normalization.

Why it matters: HKR-H/K/R all pass: the hook is sharp, the test setup has concrete numbers, and the topic hits safety trust. Reddit single-post sourcing and limited disclosed results keep it in featured, not P1.

Read the original ↗Export Markdown