"System one" decision models are models that infer and respond with calibrated probabilities or every allowed answer.

Consider your everyday language model, to get typed output from it (JSON), you may use Structured Output to constrain the output to guaranteed valid JSON. While model prefills the input in one pass, it still has to go perform a pass for every token in order to generate a valid response.

In this example, 11 passes are required to generate the final output. (We're not accounting for speculative decoding and other inference optimization techniques.)

Decision models such as Jev, make the assumption that there are fixed options we can select from and we can do so quickly by making a single pass. In this example we constrain the set of possible outputs to the options A, B, C, D, E. By masking other items in the vocabulary, the model can only emit those tokens. By selecting the highest probability output, we get our answer.