Skip to content
Hacker News front page

A $99 MUD benchmark for LLM agent behavior, with 13 models tested

Can a MUD evaluate LLMs? A $99 proof of concept

CrucibleBench drops models into a persistent MUD with NPC memory and trust mechanics, scoring them over 50 turns on hidden social objectives. 13 models, 650 runs, $99.59 total. The main finding isn't rankings: a single LLM-judge component shifted leaderboard positions by up to 6 spots while aggregate reliability stats stayed silent. GPT-5.4 ranked #1 under full scoring but fell to #5 with judge-dependent dimensions removed; Claude Sonnet 4.6 rose from #4 to #1. Dialogue looping was the top failure mode across all models—14% to 66% of frontier runs repeated 8+ talk commands at one NPC. The authors stress this is a proof-of-concept, not a validated social-intelligence measure or a predictor of real-world deployment.

Why it matters: A MUD-based social benchmark for LLMs, fully run for $99—the cost transparency alone is a hook. The real signal isn't the ranking but the finding that the judge LLM can swing positions by 6 slots, a concrete warning for anyone relying on leaderboards. Score held at 78 because ...

Read the original ↗Export Markdown