Ask an LLM how sure it is about something and when it answers you, it gives you a number based on vibes. “I’m about 90% confident” comes out of the same next-token machinery as everything else it says. The 90 is a word choice, not a measurement, and it will say it just as warmly about something it made up. Watch it happen. Give an LLM the last few days of weather and ask it to forecast tomorrow’s: { "09/27/2026": 59, // date and temp in Fahrenheit "09/26/2026": 57, "09/25/2026": 64, "09/24/2026": 63 } The obvious thing is to ask for a single number, and while we’re at it, ask how sure it is: llm.generate( prompt: "Given the last 4 days of temps, forecast tomorrow: " + weather, schema: { "date": string, "temperature": number, "confidence": string } ) // => { "date": "09/28/2026", "temperature": 60, "confidence": "about 90%" } Sixty degrees, about 90% confident. Sounds great. Now ask where the 90 came from. Nothing in that call gave the model a way to measure how sure it is; it produced “about 90%” the same way it produced “60”, by predicting what a confident forecaster would say next. Its real uncertainty comes from two places: how much data and prior knowledge it has, called epistemic uncertainty, and the plain unpredictability of the future, called aleatoric uncertainty. The first shrinks as we collect more history; the second never goes away. Suppose we had only one day of history, and it was a freak hot day. The model should be far less sure about tomorrow, and “about 90%” would come out just the same. The language of uncertainty is the probability distribution. Instead of asking for one temperature, take the full range ever recorded in the area, split it into buckets, and ask the model to put a probability on each. Those probabilities sum to 1 and together form a distribution over tomorrow’s temperature. A confident model piles most of its probability into one or two buckets; an uncertain one spreads it out. // one bucket per 10°F, covering every temp on record llm.generate( prompt: "Yesterday was 95°F. Forecast tomorrow's high.", schema: { "30-39": probability, ... "100-109": probability } ) // => { "50-59": 0.20, "60-69": 0.30, // "70-79": 0.22, "80-89": 0.15, ... } A forecast distribution is only a claim about the world. To check it we need the real distribution: across all the days the model made this same forecast, how often tomorrow’s high actually landed in each bucket. We aren’t grading whether the model got tomorrow right. It can’t be perfectly right, because part of the future is simply unpredictable. We’re grading whether it’s honest about how sure it is. A calibrated model’s claims line up with the outcomes: of all the days it gave 60-69°F a 30% chance, about 30% landed there. An overconfident model piles probability into a couple of buckets while outcomes scatter wider. An underconfident model spreads its bets while reality clusters tightly. A calibrated model will still be wrong plenty of the time, but when it says it’s more sure of one outcome than another, you can trust that number. So how does RLCD teach a model to do this? The same way a forecaster learns: by checking forecasts against what happened. The model makes its forecast, then tries a handful of slightly different versions of it: one a little more sure of the 60s, one leaning a little warmer, and so on. The next day the real high comes in the 70s. Each version is scored by how much chance it gave the 70s; more chance on what actually happened means a higher score. The model shifts toward the versions that beat the average and away from the ones that fell short, then repeats this over thousands of questions. The clever part is the score itself. A model that always acts sure gets punished hard whenever it’s wrong. A model that always hedges never scores well, even when it’s right. The only way to maximize the score over many questions is to say exactly how sure it should be: if something happens 60% of the time, the best score comes from saying 60%, not 40% and not 90%. So the model isn’t just learning to be right; it’s learning to be honest about how likely it is to be right. Laya, trained this way, reports that its stated confidence is off by about 6 points on average. Once a probability is calibrated, you can do arithmetic with it instead of eyeballing it. An uncalibrated “90% confident” from a model is a word choice. A calibrated 90% is a number you can wire straight into expected-value math. To see this with a real decision instead of a temperature bucket, imagine JEV as the forecaster. Give it a stock’s recent price action, options flow, and news, and ask for a calibrated distribution over where tomorrow’s close lands. from typesafe_sdk import Choice, TypeSafeClient client = TypeSafeClient() response = client.system_one( state="NVDA closed today at $180. Price action, options flow, and news over the last 5 trading days.", questions={ "closing_range": Choice( instructions="What range will NVDA's closing price fall in tomorrow?", criteria={ "165-170": "Closes between $165 and $170", "170-175": "Closes between $170 and $175", "175-180": "Closes between $175 and $180", "180-185": "Closes between $180 and $185", "185-190": "Closes between $185 and $190", "190-195": "Closes between $190 and $195", "195-200": "Closes between $195 and $200", }, ), }, ) probabilities = response.answers["closing_range"].probabilities
Essay · AI, LLMs, Machine Learning, Research, AI Safety
How to Solve Hallucination
First published at www.robw.fyi.
Robert W argues LLM confidence scores are vibe-words, not calibrated probabilities—and outlines RLCD-style approaches that treat hallucination as a solvable measurement and training problem.