LLM Judges Verify Presence, Not Absence: Omission Blindness in AI Clinical Notes
LLM judges fail to detect omissions in AI-generated clinical notes. A new benchmark of 500 note pairs shows detection accuracy for added/altered content at 0.79-0.94, but for omissions only 0.50-0.63—barely above chance. Restructuring the task helps: first list all facts from the transcript, then check each against the note. A two-step pipeline achieves 2.7% false alarms; a single-prompt method catches 12% more omissions at 6.2% false alarms and one-tenth the cost. Two physicians validated the pipeline as more reliable. Both methods miss omissions when the fact is restated elsewhere in the note. Dataset and code are open-sourced.