Tldr; if you just want the code and the model: the ablation and the evaluation are at github.com/anthony-maio/eve-rlcd, and the decision-only checkpoint is at anthonym21/qwen3-0.6b-rlcd-decision.

When I wrote Jev: The Language Model That Won’t Talk, I ended on the fact that nobody outside TypeSafe could tell you whether RLCD actually works. TypeSafe’s announcement says what Reinforcement Learning for Calibrated Decisions is supposed to do, which is make a model’s stated probabilities match how often it turns out to be right, and then it stops. No reward function, no architecture, no training procedure, no calibration curves.

I must admit, I was disappointed to not even see some kind of peer-reviewed research, a white paper, something open-source. The lack of transparency was a red flag for me. I always investigate and evaluate, I never accept a company promotional material as the whole story.