Four report cards in one week: Step 5 Preview benchmarks, Jev's confidence blind spot, and what California's AI order actually requires
What happened
阶跃星辰的 Step 5 Preview 公布了四类成绩,但自报的 Terminal-Bench 分数还没上公开排行榜,权重许可条款也没披露。独立开发者用 111 道公开难题测了 TypeSafe 的判断接口 Jev,发现它答错时平均把握 0.690,答对时 0.684,把握高低分不出对错。加州州长签了 AI 行政令,但只是让政府部门去起草修法建议,目...
Coverage
Follow the reports to see the story from different sides.
- Computing Life · Share · YageFour report cards in one week: Step 5 Preview benchmarks, Jev's confidence blind spot, and what California's AI order actually requires
StepFun's Step 5 Preview posted four types of scores, but its self-reported Terminal-Bench result isn't on the public leaderboard yet, and weight license terms remain undisclosed. An independent dev tested TypeSafe's Jev classifier on 111 public hard questions: Jev's average confidence was 0.690 when wrong and 0.684 when right—confidence doesn't separate correct from incorrect. California's AI executive order only directs state agencies to draft legislative proposals; the only active law binding companies is 2025's SB 53.