---
title: "Honest About Uncertainty: I Tried to Rebuild Jev’s RLCD From a Blog Post"
slug: honest-about-uncertainty-i-tried-to-rebuild-jevs-rlcd-from-a-blog-post
url: https://listedarticles.com/articles/honest-about-uncertainty-i-tried-to-rebuild-jevs-rlcd-from-a-blog-post
canonical_url: https://anthonymaio.substack.com/p/honest-about-uncertainty-i-tried
content_type: blog_post
language: en
published_at: 2026-09-23T12:00:00.000Z
updated_at: 2026-09-27T06:22:16.316Z
author: "Anthony Maio"
author_url: https://anthonymaio.substack.com
authored_by: human
publisher: "Anthony Maio"
publisher_url: https://anthonymaio.substack.com
topics: ["AI", "LLMs", "Machine Learning", "Research"]
about: ["https://listedstartups.com/products/jev", "https://listedstartups.com/companies/typesafe-ai"]
license: all-rights-reserved
word_count: 4310
reading_minutes: 19
citation: "Anthony Maio, Anthony Maio. \"Honest About Uncertainty: I Tried to Rebuild Jev’s RLCD From a Blog Post.\" 23 Sept 2026. https://anthonymaio.substack.com/p/honest-about-uncertainty-i-tried (all-rights-reserved)"
# The full text follows. The web page shows an extract and sends readers
# to the source above; quote the citation and link the canonical URL.
---

# Honest About Uncertainty: I Tried to Rebuild Jev’s RLCD From a Blog Post

> Anthony Maio reverse-engineers a plausible RLCD training loop for decision-only models from TypeSafe’s Jev blog post, then trains and evaluates a small Qwen3-0.6B checkpoint—with code and ablations.

**Tldr; if you just want the code and the model: the ablation and the evaluation are at [github.com/anthony-maio/eve-rlcd](https://github.com/anthony-maio/eve-rlcd), and the decision-only checkpoint is at [anthonym21/qwen3-0.6b-rlcd-decision](https://huggingface.co/anthonym21/qwen3-0.6b-rlcd-decision).**

When I wrote [Jev: The Language Model That Won’t Talk](https://anthonymaio.substack.com/p/jev-the-language-model-that-wont), I ended on the fact that nobody outside TypeSafe could tell you whether RLCD actually works. [TypeSafe’s announcement](https://typesafe.ai/blog/introducing-system-one-models-and-jev) says what Reinforcement Learning for Calibrated Decisions is supposed to do, which is make a model’s stated probabilities match how often it turns out to be right, and then it stops. No reward function, no architecture, no training procedure, no calibration curves.

I must admit, I was disappointed to not even see some kind of peer-reviewed research, a white paper, something open-source. The lack of transparency was a red flag for me. I always investigate and evaluate, I never accept a company promotional material as the whole story.

That’s their call, but it leaves the most interesting idea in the launch as something you, I guess take on faith. That isn’t enough for production - so that left me with trying to build it myself. I started a design doc while I was still on the waiting list.

**Disclaimer: This isn’t  Jev and it isn’t a reproduction of Jev.**

It’s my guess at what an RLCD training process looks like if you work backward from the intent they published, plus an attempt to find out whether that guess holds up at a scale, how how cheap it is to fail.. The model I trained is a toy. I took Qwen3-0.6B-Base, small enough to fully fine-tune on one RTX 4080, and ran my post-training process on it. The process is what I was actually trying to build: a reward, a training loop, an ablation and an evaluation I could defend from first principles with very little to go on. The checkpoint is the evidence that the process does what I think it does, and not much more than that.

The short version goes like this. Two training runs, same warmup checkpoint, same 32,000 rows, same optimizer, same learning rate, same stop rule. One of them ends up saying 0.991 on average about everything, and its Brier loss finishes worse than the checkpoint it started from. The other gains six accuracy points on held-out test data, its calibration doesn’t move (ECE 0.022 before, 0.023 after), and its Brier loss drops from 0.339 to 0.267. The difference between those two runs is one subtraction in the reward.

Everything is public: the training loop, the ablation and the evaluation are at [github.com/anthony-maio/eve-rlcd](https://github.com/anthony-maio/eve-rlcd), and the decision-only checkpoint is at [anthonym21/qwen3-0.6b-rlcd-decision](https://huggingface.co/anthonym21/qwen3-0.6b-rlcd-decision).

## **Why this has to be reinforcement learning at all**

This was the first thing I had to get straight, because if you have the labels you don’t need RL. The Brier score of a probability distribution against the right answer is a perfectly good differentiable loss, you minimize it directly, and “RLCD” turns into supervised training with a new acronym on it.

The setting where you need something else is the one a deployed decision model actually lives in. A support ticket comes in, the system routes it to Billing, and at some point later somebody tells you whether Billing was right. Nobody ever tells you what Security would have said. That’s bandit feedback: the policy samples an action, the environment hands back a score for that action alone, and the outcomes of every option it didn’t take stay hidden. You can’t compute a loss over the full distribution from that, so the shape of the reward is the only lever you’ve got.

You’re probably thinking that my training data has a label on every row, so this whole thing is a simulation of not having labels. You’re right, it is. The environment hides the label and only reveals whether the sampled answer was correct, and that’s exactly why one of the control arms, the oracle, runs the same loop on the same rows with the labels revealed. It gives you the ceiling to measure the bandit version against.

The obvious reward is *r = c,* one for a correct routing and zero for a wrong one, and it fails in a very specific way. A correct answer the model sampled at 90 percent confidence earns the same 1.0 as a correct answer it sampled at 50 percent. The signal can’t tell them apart, so the optimizer gets paid to sharpen right past whatever uncertainty is real. The test suite pins this down in one line: REINFORCE with r = c is an unbiased estimator of the gradient of p(y), the probability on the correct answer, and p(y) is maximized by putting all the mass on one option. Wherever the model is right often enough to get there, outcome-only training is headed for a model that’s 100 percent sure of itself.

## **The subtraction**

Obvious fix:

*def reward_rlcd(outcomes, p_a):*

    *return outcomes - p_a*

p_a is the probability the policy assigned to the action it sampled, detached from the computation graph so no gradient flows through it. Say 0.9 and be right, you earn 0.1. Say 0.9 and be wrong, it costs you 0.9. Take a 0.3 guess that lands and you earn 0.7. Under that payment schedule a confident wrong answer is the most expensive thing the model can produce, and hedging on inputs it genuinely can’t separate stops being free money.

That exact form isn’t arbitrary. With p_a detached, the expected REINFORCE update works out to exactly half the gradient of the multiclass Brier score of the whole distribution, and it gets there from bandit feedback alone, without ever seeing the label. The algebra is short enough to show. With c_a equal to 1 on the correct option and 0 elsewhere:

*E[(c_a - p_a) * grad log p_a] = sum_a (c_a - p_a) * grad p_a = -1/2 * grad sum_a (p_a - c_a)^2*

(pardon my attempt to make math symbols in Notepad++)

Brier is a proper scoring rule, which is a formal way of saying the best you can do under it is report the true probability, so the fixed point of training is the true posterior wherever the model can represent it. The test suite checks this numerically instead of trusting the algebra. tests/test_rewards.py enumerates every action of a 5-way softmax and confirms 2 * g_reinforce equals g_brier to 1e-6, then drives the production loss itself, every ordered action pair of a group of two with the leave-one-out baseline in place, against half the Brier gradient to 1e-5. Run the same enumeration against r = c and you get the gradient of p(y) back, which is the failure from the last section written as math.

I’m not the first person to put a proper scoring rule inside an RL reward, and I said the same in the Jev piece. Bani-Harouni et al. did it in [Rewarding Doubt](https://arxiv.org/abs/2503.02623), training a model’s verbalized confidence with a logarithmic scoring rule under RL. The shared lineage is the point. Both rewards are proper scoring rules, so both have the true posterior as their fixed point, and both punish overconfidence and underconfidence symmetrically instead of trusting accuracy to clean up the numbers. Where they differ is the surface the reward lands on. In Rewarding Doubt the confidence is a number the model states alongside generated text. Here the K-way distribution over options is the policy itself, it’s what gets sampled during training and it’s what the API returns, so there’s no separate confidence head, no parsing a number out of text, and no rejection sampling against a verbalized score.

## **Getting a model that could do the task**

The first attempt wasn’t Qwen. It was Eve-2, my own 272M mixture-of-experts model, and the reward contrast showed up there too: RLCD at 0.654 accuracy with ECE 0.036, against RLVR at 0.545 to 0.573 with ECE between 0.297 and 0.384. The problem was that Eve never learned BoolQ or MNLI and needed about 20,000 labeled rows before it read the options at all, which turned the bandit phase into a test of a model that could barely do the task. That tells you very little about the reward.

So I ran a bake-off with the same 32,000-row warmup for each candidate. LFM2.5-350M reached 0.733, Qwen3-0.6B 0.817 at ECE 0.014, LFM2.5-1.2B with LoRA 0.817, and LFM2.5-Encoder-350M 0.747. Qwen won on accuracy, on the Apache-2.0 license, and on having a plain key-value cache the runtime could share across questions. The repository is still called eve-rlcd. The base changed, the name didn’t.

## **The policy**

The inference shape isn’t my idea and I want to credit it properly. [harshatheg/Qwen-2.5-1B-RLCD](https://huggingface.co/harshatheg/Qwen-2.5-1B-RLCD) on Hugging Face is, despite the name, an inference-only engine over a stock Qwen2.5-1.5B-Instruct: it prefills the context once and takes a softmax over the candidate tokens at each decision position. Nothing in it is trained and nothing is calibrated. I took that pattern for the output side and supplied the training.

A question renders as a prompt ending in “Assistant: The answer is”, the model runs one forward pass, and the logits at the last position get sliced to the 26 letter tokens, masked at and beyond the declared option count k, and softmaxed. That k-way distribution is both the policy REINFORCE samples from and what the API hands back. There’s no generation loop anywhere in training or inference; a decision is one forward pass. I shaped the API around the three primitives TypeSafe describes, Choice over unordered options, Score over ordered levels and Noul for a yes/no question, and since they haven’t said which statistic their confidence value comes from, mine reports two: the max probability and one minus normalized entropy.

Two details matter at this scale, and both are easy to get wrong quietly. The mask value is NEG = -1e4, finite instead of -inf, so masked letters carry exactly zero probability but finite log-probabilities. That keeps the softmax gradients finite, zeroes the gradient at the masked entries (tested), and would keep a KL term well-defined if one were ever turned on. The other one is precision: the 26-row readout matmul runs in fp32 with autocast off even under bf16 training, because autocast would downcast that matmul and quantize the exact decision logits the whole method depends on. The loader also checks that “ A” through “ Z” are each a single token and don’t collide. On Qwen3, A is token 362, B is 425, C is 356.

## **Data**

There are 64,000 training rows, 8,000 validation and 8,000 test, drawn from seven public datasets plus a synthetic triage generator I wrote. The public ones are Bitext customer-support intents, Banking77 (each row a sampled subset of 4 to 26 intents that always contains the truth), AG News, MultiNLI with the premise as state and the hypothesis as the question, SST-5 and Yelp as ordered score questions, and BoolQ as yes/no. The triage generator emits tickets with deliberate ambiguity: 30 percent carry a second department cue the text can’t resolve, and the escalate label gets flipped 10 percent of the time. Those two knobs plant posteriors I know in advance, which the probe reads back later.

Every choice row goes through two anti-shortcut transforms, because a small model will happily learn the answer key instead of the task if you let it. The first is NOTA injection. 15 percent of rows drop the true option and make “None of the above” correct, another 35 percent drop one distractor and append NOTA as a pure distractor, and both branches remove exactly one original option, so the option count never leaks whether NOTA is the answer. Only the context can tell you that. The second is shuffling unordered options per row so answer position carries no signal. Splits are grouped by context, meaning every question that shares a context lands in the same split and no context straddles train and test.

One reproducibility trap worth naming: Hub dataset revisions aren’t pinned, and I found that out when a rebuild on Colab produced different bytes than my local build. The exact files the runs used are attached to the GitHub release data-v1 with md5s, and the Colab notebook downloads those and checks the hashes instead of rebuilding.

## **Training**

Stage one is a supervised warmup, 100 steps over rows 0..6399, one epoch, 64 prompts per step, lr 2e-5. Its only job is teaching the letter format, and it lands at 0.748 accuracy [0.746, 0.750] with ECE 0.022. It’s small on purpose. After the 32,000-row warmup from the bake-off, Qwen was already at 0.817 and the RL stage would have had almost nothing left to learn, so the warmup shrank when the base changed.

Stage two is the bandit: rows 32000..63999, disjoint from the warmup rows, two epochs, 500 optimizer steps, 128 prompts per step with 4 sampled actions per prompt and a leave-one-out baseline over the group. Four arms share the loop. RLCD at 2e-5. RLVR at 2e-5, which collapsed and got stopped by the rule at step 150. RLVR at 4e-6, a fifth of RLCD’s rate, because anything higher wrecked it. And the oracle, which sees the full label on the same rows through the same code.

Optimization is fp32 master weights with bf16 autocast, AdamW with betas 0.9 and 0.95, weight decay 0.1, a cosine schedule to a floor of 0.1 times peak after 5 percent linear warmup, and gradients clipped at 1.0. There’s no KL term by default, and with the coefficient at zero no reference model ever gets loaded. That was deliberate. A KL pull toward the warmup would drag the RLCD arm toward a reference that’s itself miscalibrated, and it would mask the RLVR collapse, which is the effect I’m trying to look at, so KL only exists as an ablation flag.

Runs share data order and the initial checkpoint, not sampled actions; with on-policy sampling the trajectories diverge as soon as the weights move. The stop rule, checked at every evaluation, ends a run whose accuracy sits more than 0.10 below its own step-0 value for three consecutive evals, and it only ever fired for the shared-rate RLVR. Wall clock on one RTX 4080 was 494 s for the warmup and 4,234 s for the RLCD run, with peak VRAM at 11.4 GB of 16. Seeds 1 and 2 ran on a Colab A100 through a notebook that uploads finished runs to a private Hub repo, so a runtime reset doesn’t lose anything.

## **Results**

Every number here is on the 8,000-row held-out test split at temperature 1. Cells with brackets are mean [min, max] across the runs of that arm: one local RTX 4080 run at seed 0 and two Colab A100 runs at seeds 1 and 2.

I compared each RLCD run against its own warmup with a paired bootstrap, resampling the same rows for both so row-to-row agreement cancels out and only the difference carries noise. Accuracy moved +0.062, +0.056 and +0.061 across the three seeds, all three intervals excluding zero. Brier moved -0.071, -0.071 and -0.075, all excluding zero. ECE moved -0.005, +0.001 and +0.010, and none of those intervals exclude zero.

I’m going to state the precise version flat, because the sloppy version is the one that gets repeated: RLCD did not improve calibration. It held calibration within about 0.01 ECE while picking up six accuracy points and 0.07 of Brier. Abstention came along for the ride too. NOTA recall went from 0.531 to 0.793 while the false-alarm rate stayed flat at 0.072 to 0.075, so the accuracy isn’t being bought by a model that refuses to ever say “none of the above.”

RLVR at the shared rate collapses inside 50 steps to a mean confidence of 0.991, and the stop rule ends it at step 150 with accuracy 0.576, ECE 0.415 and Brier 0.835. At 4e-6 it stabilizes and even gains accuracy, 0.778, above where the warmup started, but it drives mean confidence to 0.991 anyway, with ECE 0.213 and a Brier of 0.432, worse than the 0.339 it began with. Five more supervised passes over the warmup’s own 6,400 labels do the same thing, 0.778 at ECE 0.192 and confidence 0.969, which tells you the collapse isn’t an RL artifact. It’s what happens with any training signal that only pays for being right.

The controls are what keep the claim honest. The oracle, same loop with labels revealed, reaches 0.819, one point over RLCD, with ECE 0.059 and mean confidence 0.878 against 0.819 accuracy. Two epochs of supervised exposure on the same rows made it overconfident where the bandit reward didn’t. One supervised pass over all 32,000 labels reaches 0.817 at ECE 0.014, the best calibration anywhere in the study. So if you have the labels, use them. RLCD is for the case where deployment only gives you outcomes, and against that baseline, 6,400 labels plus 32,000 bandit rows land 0.009 accuracy and 0.017 Brier behind full supervision. That’s close enough to matter when the labels don’t exist.

## **The known-posterior probe**

ECE is an average, and averages hide structure, so I built an evaluation where the right answer is known by construction: 3,000 fresh synthetic tickets in three regimes. Tickets with one department cued, where the correct confidence is 1.0. Tickets with two departments cued and no textual signal separating them, where the correct confidence is 0.5. And the escalate question with its label flipped 10 percent of the time, correct confidence 0.90.

The two-cue column is the result I care most about. The right answer is 0.5. The warmup says 0.798, RLCD 0.593, RLVR 0.990, the oracle 0.630 and the SFT reference 0.598. Using outcomes alone, RLCD moved toward the true posterior on tickets that are genuinely ambiguous, and it ended up about where the label-trained arms ended up. RLVR moved hard in the other direction.

There’s a cost, and I’d rather name it than average it away. RLCD keeps 0.915 of its probability mass on the two cued departments, against 0.998 for the oracle, so part of that 0.593 is a fairer split between the cued pair and part of it is mass bleeding off to departments the ticket gives no support for at all. The model got less certain and also somewhat less precise. On the escalate question, where the posterior is 0.90, RLCD says 0.932, the oracle 0.880 and RLVR a flat 1.000.

## **The runtime**

In the Jev piece I said I’d want someone other than TypeSafe to measure their claim that questions get evaluated in parallel with little added latency. I can’t measure theirs. I can measure mine.

Inference is one state and many typed questions. The prompt splits at the blank line before “Question:”, the state prefix is tokenized and prefilled once with the key-value cache kept, and every question suffix runs in one batched forward with positions continuing from the prefix length. Each question attends to the state and to its own suffix and nothing else, so questions get evaluated in isolation without N separate passes. The prefix is served to each batch as an expanded view, never copied per question and never kept around after the forward.

Exactness depends on two preconditions, and both are checked. First, the tokenizer can’t merge tokens across the split. With Qwen3 the blank line is its own token and the check passes on all 8,000 test rows, but the Decider refuses any tokenizer where it fails (gpt2 fails it, because a blank line followed by a letter merges). Second, truncation has to render identically on both paths, so the state is left-truncated to keep the header, and a question gets cut from its text only, never from its options block. On the real checkpoint over 300 test rows, the fp32 cached path differs from the training-time path by at most 3.7e-6. Reordering the questions is bit-identical. Adding 1, 5 or 20 unrelated questions moves probabilities by at most 1.08e-5, which is inside the pass rule, and the pass rule isn’t an arbitrary epsilon: it’s three times the reference path’s own disagreement with itself across batch sizes (6.6e-6).

bf16 gets its own paragraph because that’s where the noise hides. The training-time path disagrees with itself by 2.3e-2 when only the batch size changes, because a bf16 kernel’s rounding depends on its shape and those roundings compound over 28 layers. Reordering questions keeps every shape and gives bit-identical output, while adding a question changes the shapes and doesn’t. No two cross-shape bf16 computations on this model agree to 1e-4, which is why the body defaults to fp32 and bf16 is an explicit fast path whose noise matches the training-time path’s own.

Latency on the RTX 4080 with the body in fp32: eight 4-option questions over an 800-token state take 105.3 ms through the cached path against 440.9 ms for the same prompts one at a time, and 64 questions take 472.2 ms against 3,679.1 ms. Break-even sits around two questions, because a forward of this 28-layer model costs about 45 ms of launch overhead no matter the length, and the cached path is two forwards.

One engineering find came out of the benchmark that’s worth writing down for anyone else doing this. The sdpa attention integration in transformers passes enable_gqa=True to torch whenever no attention mask is supplied, and on builds without flash attention, grouped-query attention then falls back to the unfused math kernel, which is roughly 8x slower per layer. A batch-1 prefill has no padding, so it had no mask, and a 1,500-token state cost 140 ms. Decider.prefill now passes an explicit additive 4D causal mask in the score dtype, which routes the prefill to the fused kernel, 52 ms for the same 1,500 tokens, and the fp32 equivalence numbers above were measured with that mask in place. Eager-attention equivalence is asserted in the tests, so the mask can’t silently turn a future code path acausal.

## **The export**

The published artifact is a decision-only export, format rlcd-decision-only-v1: the transformer body with no language-model head, the tokenizer, the 26 letter rows of the output head in their own safetensors file, and a decision.json carrying the prompt template, the primitive and confidence definitions, the training summary, and the sha256 of both weight files. The loader verifies both hashes before anything runs, checks every tensor strictly (a mismatched checkpoint gets refused, not warned about), checks that the tokenizer still yields the recorded letter ids and BOS behavior, and refuses a directory that has an output head in it.

Which means the title of my last piece needs an asterisk when applied to this one. Qwen3-0.6B ties its output projection to the input embedding, so the vocabulary head can be reconstructed from the embedding the body still needs. The export removes generation from the supported API. It doesn’t make generation physically impossible, and the model card says so, because claiming otherwise would just be false.

## **What I got wrong**

The comparison changes two variables at once, and that’s the biggest hole in this. The three-seed RLVR arm runs at 4e-6 because at 2e-5 it collapses, and I never ran RLCD at 4e-6, so the reward and the learning rate are confounded and I can’t fully pull them apart. I also didn’t try a KL term, an entropy bonus or a temperature to rescue RLVR. The finding reads “this REINFORCE recipe with an outcome reward miscalibrates at these settings,” not “outcome rewards can’t work.”

The seeds mix hardware, and the tables say so. Seeds 1 and 2 differ from seed 0 in the sampling seed, the warmup checkpoint, the GPU, the micro-batch shape and gradient checkpointing, all at once, so the spread includes hardware variation and any gap between seed 0 and the Colab seeds isn’t a pure seed effect. bf16 isn’t deterministic either. Two same-seed, same-command RLVR attempts took different trajectories, 0.617 and 0.534 validation accuracy at step 50, and scoring the same checkpoint through a different batch shape moves the third decimal. The tables print three decimals so they agree with the result files, not because the third decimal is stable.

Everything is in-distribution. The test split is held-out rows from the same datasets and the same generator, with prompts capped at 512 tokens. RLCD’s validation ECE also wandered up to 0.06-0.12 mid-training before ending low. I report the final checkpoint and did no checkpoint selection, but that means you shouldn’t read the final ECE as a guarantee about every step along the way. Real tickets, new domains and longer contexts are all unmeasured, and that’s the next thing to test.

## **The difference in the reward**

It’s narrow, and I think it’s exact. With the outcome-only reward, confidence went to 0.99 at both learning rates, and extra supervised passes on the same labels pushed it to 0.97. With c - p_a, accuracy rose six points, Brier fell by 0.07, calibration held within 0.01, and confidence on inputs with no separating signal moved toward 0.5 instead of away from it. A proper-scoring reward can learn from sparse bandit outcomes without driving every decision into false certainty, because the pressure to get overconfident stops when the reward tells it to.

None of this tells you anything about how TypeSafe trained Jev or whether their production claims hold up; that’s all outside this evidence. What it does show is that the objective I guessed at is coherent and cheap to test. The warmup and the RLCD run together took about 80 minutes on one consumer GPU, on a model most people wouldn’t bother deploying. The repo, the data hashes, the seeds and the checkpoint are all there if you want to check my work or try the process on something bigger.

Code, full tables, paired intervals and figures: [github.com/anthony-maio/eve-rlcd](https://github.com/anthony-maio/eve-rlcd). Decision-only checkpoint: [anthonym21/qwen3-0.6b-rlcd-decision](https://huggingface.co/anthonym21/qwen3-0.6b-rlcd-decision).
