What Is RLCD? The Secret Behind Jev

From pairwise reward modeling to calibrated, multiway decisions
Jev looks mysterious when viewed as an alternative to a language model. It becomes much simpler when viewed as the next step in reward modeling.
More specifically, RLCD is a schema-conditioned Plackett–Luce objective. Jev turns that objective into a product by adding typed outputs and parallel inference.
That is the secret: the reward model is no longer hidden behind a generator. The reward model becomes the model.
A conventional reward model receives a context \(x\) and a candidate answer \(a\) , then produces a scalar:
Outcome reward models score the final answer. Process reward models score individual reasoning steps. In both cases, the learned object is an absolute-looking number.
The problem is that this number is not actually absolute.
Get smarter about the news
Sign up free for a feed built around what you actually care about, Dive Deeper research on any story, and the full text of every article.
Create free accountAlready have an account? Sign in