Skip to content
Trending storyPast story

Four report cards in one week: Step 5 Preview benchmarks, Jev's confidence blind spot, and what California's AI order actually requires

1 report1 sourceupdated 4 days ago

What happened

Summary

阶跃星辰的 Step 5 Preview 公布了四类成绩,但自报的 Terminal-Bench 分数还没上公开排行榜,权重许可条款也没披露。独立开发者用 111 道公开难题测了 TypeSafe 的判断接口 Jev,发现它答错时平均把握 0.690,答对时 0.684,把握高低分不出对错。加州州长签了 AI 行政令,但只是让政府部门去起草修法建议,目...

Coverage

Follow the reports to see the story from different sides.

Sep 25
  1. Computing Life · Share · Yage
    Four report cards in one week: Step 5 Preview benchmarks, Jev's confidence blind spot, and what California's AI order actually requires

    StepFun's Step 5 Preview posted four types of scores, but its self-reported Terminal-Bench result isn't on the public leaderboard yet, and weight license terms remain undisclosed. An independent dev tested TypeSafe's Jev classifier on 111 public hard questions: Jev's average confidence was 0.690 when wrong and 0.684 when right—confidence doesn't separate correct from incorrect. California's AI executive order only directs state agencies to draft legislative proposals; the only active law binding companies is 2025's SB 53.