Jev probability calibration test drifts from theory
What happened
On October 2, Hacker News featured a hands-on test of probability calibration in Jev, a TypeSafe classification model. The experiment covered 10 label distributions, each with 5 prompt templates and 20 variations per template, for 1,000 settings total, all for under $4. Jev's output probabilities drifted from the theoretical distribution, with an average total variation distance of 0.518, only slightly better than the 0.546 of uniform guessing. The model usually picked the right distribution type but tended to concentrate probability in a single bin. Follow-up math tests found Jev struggles with multi-step arithmetic and tracking powers of ten.
Written by AI from the coverage · updated 1 hour ago
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageHow accurately calibrated is Jev?
TypeSafe 的分类模型 Jev 在这项实测中输出的概率分布偏离理论分布,平均总变差距离为 0.518,仅略低于均匀猜测的 0.546。实验覆盖 10 类分布、每类 5 个提示词模板及每模板 20 种变化,共 1,000 个设置,全部实验成本低于 $4.00。Jev 通常能选对分布类型,却倾向于把概率集中在单个区间;后续数学测试还发现,它在多步算术和十的幂次追踪上存在困难。
Heat over time
Not enough continuous observations to draw a trend yet.