From pairwise reward modeling to calibrated, multiway decisions

Jev looks mysterious when viewed as an alternative to a language model. It becomes much simpler when viewed as the next step in reward modeling.

The core idea is:

More specifically, RLCD is a schema-conditioned Plackett–Luce objective. Jev turns that objective into a product by adding typed outputs and parallel inference.

That is the secret: the reward model is no longer hidden behind a generator. The reward model becomes the model.

Reward Modeling Started with a Scalar

A conventional reward model receives a context $x$ and a candidate answer $a$, then produces a scalar:

Outcome reward models score the final answer. Process reward models score individual reasoning steps. In both cases, the learned object is an absolute-looking number.