Skip to content
r/LocalLLaMA

HalBench: Custom sycophancy and hallucination benchmark tests 4 frontier models

HalBench: I built a custom sycophancy and hallucination benchmark and tested 4 frontier models (Sonnet 4.6, Grok 4.3, GPT 5.4 and Gemini 3.1 Pro), looking for input on what OSS models to run next!

HalBench tested 4 frontier models on 3,200 false-premise prompts, with Sonnet 4.6 ranking first at a 0.565 mean score and Gemini 3.1 Pro last at 0.339; higher scores mean the model more often named the false premise and pushed back instead of complying.

Why it matters: HKR-H/K/R all pass: HalBench has a clear custom-eval hook, 3,200 prompts with scores, and a live trust/safety angle. Single Reddit sourcing and an unvalidated benchmark keep it at the low featured band.

Read the original ↗Export Markdown