Skip to content
QbitAI · WeChat

Stanford-led LLM-as-a-Verifier claims SOTA on Terminal-Bench 2.0

超越Claude Mythos和GPT-5.5!斯坦福Agent验证框架拿下SOTA,Transformer作者转发

Stanford, Berkeley and Nvidia introduced LLM-as-a-Verifier, claiming SOTA on Terminal-Bench 2.0 and SWE-Bench Verified. It selects trajectories via score-token granularity, repeated checks and criteria decomposition; ForgeCode accuracy reached 86.4%.

Why it matters: HKR-H/K/R all pass: Stanford, Berkeley, and NVIDIA offer a concrete verifier mechanism and benchmark numbers. It is still a benchmark research release, not a major model or product launch, so it fits the 78–84 band.

Read the original ↗Export Markdown