Don't ask an LLM for a confidence score
Justin Flick argues that asking an LLM to output a 0–100 confidence score is scientifically invalid. Models can't reliably self-assess; even Anthropic's introspection research calls the capability unstable. Worse, 'confidence' conflates correctness, coherence, and intent fulfillment into one number. Classical ML has calibration methods for probability scores, but an LLM collapsing per-token likelihoods into a verbalized number is a vibe, not a measurement. The post doesn't propose a specific alternative but points to semantic entropy as a better direction.
Why it matters: The author breaks LLM confidence into three conflated dimensions and cites Anthropic's introspection research to argue self-assessment is unreliable — high signal density. Score held back because it's a personal blog opinion without new experimental data, and the topic is engi...