Jev model's calibrated probabilities fail across different data distributions
What happened
Jev 是 TypeSafe 推出的一个分类模型,不用自己准备训练数据就能用,直接给非结构化输入打出带概率的标签。但作者指出,它标榜的“校准概率”站不住脚:模型在训练数据上校准过,不代表在你的生产数据上还能校准。不同公司的数据分布不一样,Jev 却会对相同输入吐出相同的概率,这本身就矛盾。更离谱的是,有人测出 Jev 认为抛一枚公平硬币正面朝上的概率是...
From Hacker News 首页
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageJev and System One Models: Calibration Beats Accuracy
TypeSafe AI released Jev, a non-autoregressive 'System One' model that answers structured questions with probabilities in a single forward pass. The author argues calibration, not accuracy, is the real bottleneck for production classifiers, and Jev's training objective (RLCD) directly optimizes for honest probabilities. Jev achieves 70–500 ms latency and costs $0.042 per million input tokens. The author lacks API access, so performance claims are from TypeSafe's launch post; the post does not disclose independent benchmarks.
- Hacker News front pageJev's calibrated probabilities break when your data distribution differs
TypeSafe's Jev model promises calibrated probabilities for structured decisions, but the author argues calibration depends on the data distribution. A model calibrated on training data won't stay calibrated on your production data. Worse, Jev reportedly assigns 0.92 probability to a fair coin landing heads, and probability semantics shift across different primitives. Treat Jev's outputs as ranking scores, not true probabilities. If you need real calibration, fit Platt scaling on a few hundred of your own labeled examples.