Google AI team shares how to write reliable rubrics for LLM-as-a-judge evaluations
Google AI 团队分享如何为 LLM-as-a-Judge 评测编写可靠的评分标准
This is part two of Google AI's series on LLM-as-a-judge. The core idea: write rubrics as strict, objective true/false questions to cut down on judge hallucinations and noisy scores. Four rules: keep each question atomic, avoid overlapping checks, use boolean judgments instead of subjective ratings, and treat rubrics like formal specs. The post doesn't name which model they use as the judge or provide quantitative comparison data.