Skip to content
Hacker News front pagetomncooper

Decision models like Jev don't beat LLM-as-a-judge or traditional classifiers

Red Hat 的护栏评测发现,Jev 等决策模型在速度或准确率上未稳定胜过 LLM-as-a-judge、预训练分类器及开源决策模型。提示词注入检测中,Qwen3.6-35B 准确率为 89.31%,Jev 为 86.35%;内容安全检测中,Jev 以 86.20% 领先。评测的远程调用包含英美之间的网络延迟,风险策略调优对不同模型的效果也不一致。

Read the original ↗Export Markdown