Skip to content
Hacker News front page

What Is RLCD? The Secret Behind Jev

RLCD turns reward modeling from a scalar score into multiway preference plus probability calibration. Traditional reward models output an absolute number, but that number has no stable meaning—only relative comparisons matter. PPRM makes the comparison explicit as a preference probability, and RLCD extends it to multiple candidates using a Plackett-Luce objective. Jev turns that reward model into a product with typed outputs and parallel inference, no longer hidden behind a generator. The post does not disclose Jev's specific performance numbers or deployment cost.

Read the original ↗Export Markdown