What Is RLCD? The Secret Behind Jev
What happened
传统奖励模型输出一个绝对分数,但这个分数在不同问题或模型间没有稳定含义,真正有用的是相对比较。PPRM 把比较显式化为偏好概率,RLCD 进一步扩展到多个候选,用 Plackett-Luce 目标做多选偏好建模,再加概率校准。Jev 把这个奖励模型做成了产品:带类型输出、可并行推理,不再藏在生成器后面。正文没披露 Jev 的具体性能数字和部署成本。
Coverage
Follow the reports to see the story from different sides.
- Hacker News front pageWhat Is RLCD? The Secret Behind Jev
RLCD turns reward modeling from a scalar score into multiway preference plus probability calibration. Traditional reward models output an absolute number, but that number has no stable meaning—only relative comparisons matter. PPRM makes the comparison explicit as a preference probability, and RLCD extends it to multiple candidates using a Plackett-Luce objective. Jev turns that reward model into a product with typed outputs and parallel inference, no longer hidden behind a generator. The post does not disclose Jev's specific performance numbers or deployment cost.