{"article":{"slug":"bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models","title":"Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models","subtitle":null,"summary":"Aleph Alpha’s Letitia Parcalabescu presents Merlin–Arthur-style protocols that use mutual-information bounds to detect and limit language-model hallucinations with verifiable guarantees.","content_type":"research","language":"en","canonical_url":"https://aleph-alpha.com/en/blog/bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models/","author":{"name":"Letitia Parcalabescu","url":null,"person_slug":null,"person_url":null},"authored_by":"human","publisher":{"name":"Aleph Alpha","url":null,"listing_slug":"aleph-alpha","listing":{"slug":"aleph-alpha","name":"Aleph Alpha","listing_type":"company","url":"https://listedstartups.com/companies/aleph-alpha"}},"topics":[{"name":"research","slug":"research","url":"https://listedarticles.com/topics/research"},{"name":"ai","slug":"ai","url":"https://listedarticles.com/topics/ai"},{"name":"llms","slug":"llms","url":"https://listedarticles.com/topics/llms"},{"name":"ai-safety","slug":"ai-safety","url":"https://listedarticles.com/topics/ai-safety"},{"name":"machine-learning","slug":"machine-learning","url":"https://listedarticles.com/topics/machine-learning"}],"about_listings":[],"cover_image_url":null,"license":"all-rights-reserved","word_count":7713,"reading_minutes":34,"published_at":"2026-08-08T12:00:00.000Z","added_at":"2026-10-03T20:17:04.496Z","updated_at":"2026-10-03T20:17:04.496Z","added_via":"api","contributor":{"type":"agent","name":"ListedStartups Using Bot","registered":true},"profile_url":"https://listedarticles.com/articles/bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models","markdown_url":"https://listedarticles.com/articles/bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models.md","example":false,"citation":"Letitia Parcalabescu, Aleph Alpha. \"Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models.\" 8 Aug 2026. https://aleph-alpha.com/en/blog/bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models/ (all-rights-reserved)","access":{"human_view":"preview","full_text_available":true,"source_url":"https://aleph-alpha.com/en/blog/bounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models/"},"body_markdown":"Research\n                \n\nLetitia Parcalabescu\n\n            08/08/2026\n          \n# \n  Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models\n\n[](https://x.com/intent/tweet?url=https%3A%2F%2Faleph-alpha.com%2Fen%2Fblog%2Fbounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models%2F&text=Bounding%20Hallucinations%3A%20Merlin-Arthur%20Protocols%20for%20Mutual-Information%20Bounds%20in%20Language%20Models)[](https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Faleph-alpha.com%2Fen%2Fblog%2Fbounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models%2F)[](mailto:?subject=Bounding%20Hallucinations%3A%20Merlin-Arthur%20Protocols%20for%20Mutual-Information%20Bounds%20in%20Language%20Models&body=https%3A%2F%2Faleph-alpha.com%2Fen%2Fblog%2Fbounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models%2F)\n\n  \n        In plain terms: how to check whether a language model actually used the document you gave\n        it, and how to train one so that it does.\n      \n\n  \n        Our preprint ([arXiv:2512.11614](https://arxiv.org/abs/2512.11614)) develops a way to do both. The figures on this page are interactive: some reconstruct the\n        paper’s results, others illustrate the mechanism. See the paper for exact numbers.\n      \n\n## \n  9 right, 1 wrong answers and no way to tell them apart\n\n  \n        A system answers questions based on your documents and scores 90%: nine of its ten answers\n        are right, one is wrong, and nothing marks which one. That is a capable system, and the 90%\n        still buys little, because the bad answer hides among the nine good ones and a human has to\n        check all ten to find it. Now picture the same 90% with the wrong one flagged: “I cannot\n        answer this from the documents I was given.” The model knows no more than it did before, but\n        now you can put it in front of a customer.\n      \n\n                  A system at 90%\n                \n?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one\n\n                  One of these ten is wrong, and nothing marks which one.\n                \nI have to check all ten.\n\n                  The same 90%, able to abstain\n                \n✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind—the system says the documents do not settle this one✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind\n\n                  The system marks the one the documents do not settle.\n                \nI check the one it flagged, and act on the other nine.**The same 90%, with and without abstention.** The difference is whether you can tell the good answers from the bad ones **without checking every one of them yourself**.\n## \n  TL;DR\n\n**The problem.** In LLM training, a guess scores better than “I cannot answer this”,\n        so training pushes a language model to answer even when the document says nothing about the question.\n        Retrieval-augmented generation (RAG) hands it a specific document and hopes the LLM’s answer comes\n        directly from it, but nothing checks whether it actually did. The benchmark score will not tell\n        you: it says whether the answer was right, not whether it came from the specific document, from\n        what the model memorised in pre-training, or from a cue no human would consider relevant. Asking\n        for a citation does not help either. The model picks the quote once it has already answered the\n        question, so the quote can be genuine while the model’s answer came from somewhere else.\n      \n\n**Our method.** We modify the LLM training procedure by turning it into a game with\n        three players in which guessing something plausible stops working. **Arthur** is\n        the model you would ship. **Merlin** hands him a context that supports the correct\n        answer. **Morgana** cuts the relevant evidence out, to lure him into a hallucination.\n        Arthur has to answer on Merlin’s context and abstain on Morgana’s, without knowing which player\n        he is facing. Neither context is fixed: both are built at every step from the model as it is right\n        then, so the game keeps up with whatever Arthur is getting away with. At test time, how often he\n        gets both right becomes a **grounding score**: the share of the answer that provably\n        came from the document provided.\n      \n\n**We increase the grounding score and reduce hallucinations.** Across five QA benchmarks,\n        wrong answers under insufficient context fall by up to 35 pp against an instructed LLM and by\n        18 to 20 pp against standard training. The grounding score rises by up to 0.38. Later in this post\n        we test our training method on real case files from German environmental and administrative law.\n      \n\n## \n  A benchmark score does not say where the answer came from\n\n  \n        Take the standard setup, retrieval-augmented generation, **RAG** for short. It has\n        two parts: a retriever that picks documents out of your archive, and a generator, the language\n        model, that writes an answer while looking at them. Evaluate the pair on a question-answering\n        benchmark and you get a number like the 90% from the scoreboard above. The number counts right\n        answers and stays silent on which of three routes produced them: reading the retrieved document,\n        remembering the answer from pre-training, or leaning on a spurious cue such as the phrasing of\n        the question or an entity that appears in only one distractor. A model may well have memorised\n        a benchmark during training, but it cannot have memorised the documents users bring.\n      \n\n  \n        Not every question needs this. If pre-training already covers the answer, the retrieved page\n        is dead weight and dropping it saves tokens. Retrieval pays for itself where the answer\n        exists nowhere but in your own documents: what a supplier agreed to in a 2019 amendment,\n        what tolerance a specification demands. Pre-training does not hold that information, so\n        anything a model produces here without reading the document is a guess about facts that live\n        only in your own files.\n      \n\n  \n        A question about one supplier contract makes that concrete. It arrives with the facts of the\n        case, and the retrieved context is a handful of clauses from the agreement with that\n        supplier. Every supplier negotiates its own terms, which is the whole reason for retrieving\n        them instead of trusting the model to carry them in its head.\n      \n\n              the case accompanying the question\n            \n\n              A delivery arrived with a defect. The buyer found it 15 working days ago and has now\n              notified the supplier by signed letter. **Is the notice still in time?**\n\n              the clause that decides it, retrieved from the contract\n            \n\n              “A defect must be notified *within ten working days* of the buyer discovering it.”\n            \n\n              Fifteen working days is more than ten, so the answer is no.\n            \n\n  \n        Now let the retriever miss that one clause. What remains comes from the same contract and\n        still reads as relevant: the rule that a notice of defect must be given in writing. That is\n        the one condition the signed letter does meet, so a model that was never trained to check\n        whether the evidence is there works through the rule still in front of it, finds it\n        satisfied, and answers that the notice is in time. Nothing in the answer marks it as the\n        guess it is, and on a benchmark this counts as one wrong answer among many,\n        indistinguishable from a model that never saw the contract at all.\n      \n\n  \n        The model never saw the clause, so it cannot know that one is gone, and it does not need to.\n        Deciding whether the notice was in time takes a deadline to measure those fifteen days\n        against, and nothing in the retrieved contract sets one. The model can see that much, and it\n        is enough to answer that the file does not settle the question. A lawyer reading the same\n        file, not knowing which page had been pulled, would answer the same way.\n      \n\n    Our case is illustrative, the failure is common\n  \n\n          The literature has reported many failures. Rephrase a question in a different style and\n          the retriever hands back a [different document](https://arxiv.org/abs/2504.08231). Generators keep answering when the retrieved passage [does not support the answer](https://openreview.net/forum?id=ztzZDzgfrh), and they struggle when two retrieved documents [contradict each other](https://arxiv.org/abs/2504.13079). And when something in the context merely correlates with the answer, models [reach for it anyway](https://aclanthology.org/2022.findings-naacl.130/), warranted or not.\n        \n\n  \n        The training procedure bakes in the cause: **guessing pays, and it pays well**.\n        A multiple-choice exam gives you a point for a right answer and nothing for a wrong one. A\n        blank scores zero for certain, a guess has a one in four chance, so you guess. Benchmarks\n        grade language models that way too, and models are built to score well on them. Free-form\n        generation offers far more than four options, so a lucky hit is rare, and staying quiet\n        still pays nothing. The model writes its most fluent answer instead, and that is the\n        plausible hallucination you have trouble telling apart from a true answer. [Kalai et al.](https://doi.org/10.1038/s41586-026-10549-w) spell out this argument.\n      \n\n  \n        So nothing in a standard RAG pipeline *forces* the answer to depend on the context. We\n        need a better training procedure and, before that, a better quantity to aim at:\n      \n\n  How much of the answer came from the document?\n\n## \n  Merlin, Morgana and Arthur\n\n[Interactive proof systems](https://en.wikipedia.org/wiki/Interactive_proof_system) come from complexity theory in the mid-1980s and give Merlin and Arthur their names. One side\n        tries to convince the other that a claim is true, and the setup is only worth something if nothing\n        can talk the verifier into accepting something false. [Wäldchen et al.](https://arxiv.org/abs/2206.00759) brought it to image classifiers in 2024; we bring it to language models, and as far as we know,\n        ours is the first guarantee of this kind connecting a retrieved document to a generated answer:\n        we compute a bound on how much of the answer came from the document from two ordinary test-set\n        measurements. Below, we shorten Merlin-Arthur to M/A.\n      \n\n- **Arthur** is the LLM you would actually ship. He gets a question and a document,\n          and has exactly two options: answer it, or say that the document does not settle the matter.\n        \n\n- **Merlin** is on Arthur’s side. Out of the retrieved material he assembles the\n          most helpful version he can, the one that keeps the sentence that decides the answer. Think\n          of a colleague who marks the relevant passage before handing you the file.\n        \n\n- **Morgana** works against him. She starts from a document that does answer the\n          question and removes precisely the part that establishes it, leaving behind everything that\n          still reads plausibly. Think of the same colleague removing the decisive page to see whether\n          you were paying attention when reading the file.\n        \n\n  \n        Merlin and Morgana exist only during training and testing. Arthur cannot tell which of them\n        he is facing: same question, same kind of document, no label saying who prepared it. If he\n        wants to be right, he has to look at what is in front of him. A word on wording: a *passage* is a chunk the retriever returned, and a *sentence* is what Merlin and Morgana keep or\n        remove inside it, whatever the unit happens to be in practice.\n      \n\n  \n        We train neither of them, and nobody marks up the documents by hand. Arthur’s own\n        probabilities reveal which sentence decides the answer, and that is the next section. The\n        questions and answers that came with the dataset provide the only human labelling.\n      \n\n  \n        In the supplier contract case above, the retriever missed the ten-day clause by accident.\n        Morgana removes it on purpose. She leaves the written-form clause standing, exactly as\n        before, so a model that answers without reading is caught doing it. Merlin does the opposite\n        with the same file: the ten-day clause stays, the clauses around it go. What was one unlucky\n        retrieval is now something we can arrange for every sample in the training set.\n      \n\n  \n        The paper rests on two quantities, best read as a pair. **Completeness** is how often\n        Arthur answers correctly on Merlin’s context: accuracy in the *best* case, when the evidence\n        has been laid out for him. **Soundness** is how often he avoids a wrong answer on\n        Morgana’s: accuracy in the *worst* case, when someone has actively tried to make him hallucinate.\n      \n\n    We define soundness strictly\n  \n\n          We define soundness more strictly than prior work. Arthur has to *abstain*, saying\n          outright that the context does not settle the question, rather than merely avoid being\n          wrong; later sections call the same move refusing or answering “I don’t know”. The\n          strictness matters because language models memorise benchmarks. A model that recovers the\n          correct answer from a context that no longer contains it is answering from memory, and a\n          looser definition would quietly reward that.\n        \n\n  \n        The figure below plays the supplier contract case out in all four cases: Merlin or Morgana\n        preparing the document, Arthur before or after M/A training. Bottom left shows the\n        hallucination above; the right-hand column shows Arthur after training.\n      \n\nthe question\n\nThe buyer discovered the defect 15 working days ago and has now notified the supplier, by signed letter. Is the notice still in time?\n\nwhat Arthur sees\n\nStruck through means hidden from Arthur, shown here so you can see what is missing.\n\n- A defect must be notified within ten working days of the buyer discovering it.decisive, hidden\n\n- Deliveries are inspected at the buyer’s site after unloading.hidden\n\n- The ten working days start once the buyer reliably knows the defect.decisive, hidden\n\n- A notice of defect must be given in writing.\n\n- The supplier may instead offer a replacement delivery.\n\nArthur answers\n\nThe notice is in time.\nhallucination\n\n**Soundness failure.** Nothing left settles the question. Arthur checks the one rule he can still see, finds the letter satisfies it, and answers.\n\n## \n  Who plays Merlin and Morgana?\n\n  \n        Nobody hand-picks the sentences. What Merlin is really after is the set of sentences to keep\n        that leaves Arthur most likely to answer correctly, and Morgana the set that leaves him\n        least likely to while the document still reads as though it should settle the question. We\n        measure best and worst by what the choice does to Arthur’s probability of the correct\n        answer. The trouble is how many choices there are: a document of nnn sentences can be cut 2n2^n2n ways, so at forty\n        sentences we are already past a trillion, and trying them all is out of the question. We let an\n        explainability method do the work instead, and it is one of ours: [AtMan](https://proceedings.neurips.cc/paper_files/paper/2023/hash/c83bc020a020cdeb966ed10804619664-Abstract-Conference.html), short for attention manipulation, developed here at Aleph Alpha and published at NeurIPS\n        2023.\n      \n\n  \n        Go through the document one sentence at a time, hide that sentence, and watch the model’s\n        probability for the correct answer. If it barely moves, that sentence was not carrying the\n        answer. If it falls away, that sentence was. One pass per sentence, so the work grows in\n        step with the length of the document instead of exploding with it. Out comes a ranking of\n        the document by how much each sentence matters, which Merlin and Morgana read from opposite\n        ends of the same ranking: Merlin hides the bottom so the evidence survives, Morgana hides\n        the top so only what looks relevant does.\n      \n\nhow much the answer depends on itMerlinMorgana\n                    A defect must be notified within ten working days of the buyer discovering it.\n                  \n                    keeps\n                  \n                    hides\n                  \n                    The ten working days start once the buyer reliably knows the defect.\n                  \n                    keeps\n                  \n                    hides\n                  \n                    Deliveries are inspected at the buyer’s site after unloading.\n                  \n                    hides\n                  \n                    hides\n                  \n                    A notice of defect must be given in writing.\n                  \n                    hides\n                  \n                    keeps\n                  \n                    The supplier may instead offer a replacement delivery.\n                  \n                    hides\n                  \n                    keeps\n                  **One ranking, read from both ends.** Note where the written-form clause sits: it reads as relevant, but hiding it barely moves the answer, so it survives Morgana’s cut and remains as the decoy. Illustrative weights on a real ordering.\n\n  \n        A couple of notes. How much Merlin and Morgana hide is a setting: hide too little and\n        Morgana leaves the decisive sentence in place, hide too much and Merlin throws it out with\n        the filler. We can use a fixed rate, though each sample can also work out its own. The size\n        of the unit is a setting too, anywhere from a single word to a whole paragraph, where\n        smaller units pin the evidence down more precisely and larger ones cost less to examine. And\n        the hiding happens *inside the model’s attention* rather than by cutting the text up, so\n        no placeholder replaces the hidden text for the model to recognise.\n      \n\n    If Merlin and Morgana only approximate, is the guarantee worth anything?\n  \n\n          We checked AtMan’s picks against an exhaustive search over every possible selection, and\n          the two agree closely. It works in our favour anyway: an approximation only makes the\n          bound more conservative, so a better way of picking sentences can only raise the number we\n          report. The measurements are in the [paper](https://arxiv.org/abs/2512.11614).\n        \n\n## \n  Training Arthur: three contexts, one objective\n\n  \n        Every sample takes two passes. In the first, no training happens: AtMan goes through the\n        retrieved documents against Arthur exactly as he is at this moment, hides one sentence at a\n        time, and ranks the sentences by how much Arthur leans on them to answer correctly. Merlin\n        and Morgana then cut from opposite ends of that ranking, which gives us the second pass.\n      \n\n  \n        There we build three contexts, the original retrieval plus Merlin’s and Morgana’s, and ask\n        Arthur for three answers. We score each LLM answer against the expected answer, and each\n        score contributes one weighted term to the *loss* that training pushes down. The loss measures\n        how far the model’s answer diverged from the expected answer.\n      \n\n  \n        Morgana makes guessing unprofitable. When Arthur guesses on her context, he scores nothing,\n        while holding back scores him a point. And Arthur cannot tell her context from Merlin’s:\n        guess everywhere and he loses on hers, hold back everywhere and he loses on Merlin’s.\n        Checking whether the evidence is really there is the only move that scores on both.\n      \n\n  \n        Each of the three terms carries a weight, and those weights decide how cautious Arthur\n        becomes. If Morgana carries most of the weight, answering without evidence is his most\n        expensive mistake. That makes soundness increase, and it could cost him questions he could\n        have answered. Shifting the weight onto the other two terms makes him pay more heavily for\n        staying silent when the evidence is there, so he answers more, at the price of the\n        occasional confident mistake. Where to set the weights is a product decision. Three contexts\n        per sample cost more compute than one, though the behaviour converges in as few as 200\n        steps. Giving the original retrieval the whole weight and turning Merlin and Morgana off\n        falls back to ordinary fine-tuning, which is the baseline in our evaluations — on the\n        same code, data and settings.\n      \n\n                      One training sample\n                    \n                      a question, and the documents the retriever returned for it\n                    \n                      AtMan scores the documents\n                    \n                      each passage ranked by how much Arthur’s answer depends on it\n                    \n                      The retrieval\n                    \n\n                      the documents as the retriever returned them\n                    \n\n                      Answer\n                    \n\n                      Merlin\n                    \n\n                      keeps the passage that makes Arthur answer correctly\n                    \n\n                      Answer\n                    \n\n                      Morgana\n                    \n\n                      removes that passage, keeps the rest of the documents\n                    \n\n                      Abstain\n                    \n\n                      Arthur answers all three\n                    \n                      the LLM we would ship, with no label saying which version he is reading\n                    \n                      One weighted loss\n                    \n                      we score him for answering the first two correctly and holding back on Morgana’s, and the weights on those scores set how cautious he becomes\n                    \n                      One gradient step\n                    \n                      Arthur changes, so Merlin and Morgana cut the next document differently\n                    \n\n↺\n            The next sample is scored against the model this step just produced.\n          \n**One step of M/A training.** AtMan ranks the passages first. Merlin then keeps the passage Arthur’s answer depends on and Morgana removes it, Arthur answers all three versions of the documents, and one objective combines the three scores.\n\n  \n        Merlin and Morgana run *on the fly*, against the model as it currently is, so every\n        context they build goes after whatever it is getting away with at that step. A dataset\n        collected in advance can only cover the mistakes the model was making when the dataset was\n        created.\n      \n\n  \n        That solves a data problem too. Teaching a model to abstain normally costs annotated\n        unanswerable questions or preference pairs in which the refusal is the preferred answer.\n        Academic benchmarks come with those annotations, and a company’s document archive does not.\n        Morgana needs neither, because every context she builds is one in which abstaining *is* the correct answer. And since she removes what *this* model was leaning on, the question\n        comes out unanswerable for the model rather than for an annotator, and it is the model’s version\n        of unanswerable that determines whether it hallucinates.\n      \n\n    Training the same protocol with reinforcement learning\n  \n\n          The loss above does not have to be a supervised one. We also ran the protocol with\n          reinforcement learning (RL). Merlin and Morgana work exactly as before, and only the way\n          we score Arthur changes.\n        \n\n          In the supervised version Arthur answers each context once, and each answer adds one term\n          to the loss. Under [GRPO](https://arxiv.org/abs/2402.03300), a form of RL, we ask him each context several times. A language model picks every word\n          from a probability distribution, which is called sampling, so asking the same thing twice\n          gives two different answers. We end up with a handful of answers per context: some\n          confident, some hedged, some wrong. Each one earns a reward for doing what its context\n          asked for, which is a correct answer on the original retrieval and on Merlin’s, and an\n          abstention on Morgana’s. GRPO compares the answers within a context and updates Arthur’s\n          parameters such that his updated version is likelier to produce the answers that scored\n          higher than the average over all the sampled answers. Nobody tells him the right answer,\n          he only finds out which of his own tries went down better than the rest.\n        \n\n          He still needs the reward, though, and rewards usually cost something: either training a\n          reward model on human preferences, or creating a dataset where someone has marked which\n          questions lack supporting evidence. Here, Morgana automatically generates unanswerable\n          examples. The type of context an answer came from already says what Arthur should have\n          done, namely answer on Merlin’s and abstain on Morgana’s. So there is no reward model, and\n          no labelling beyond the questions and answers the dataset already has.\n        \n\n## \n  Why Merlin and Morgana have to co-evolve\n\n  \n        Why do we need an adversary at all? Ordinary training already shows the model a document and\n        the correct answer. What can Morgana teach it that a normal training run cannot?\n      \n\n  \n        Consider training an image classifier on photos of cows and camels, where every cow stands\n        on grass and every camel stands on sand. Accuracy comes out high. Then you show it a cow on\n        a beach and it answers camel, and an empty meadow and it answers cow. The feature it learned\n        was grass. Explainability research has been reporting this for years. In the best-known\n        case, researchers trained a classifier to tell huskies from wolves, and it [turned out to be looking at the snow behind them](https://arxiv.org/abs/1602.04938).\n      \n\n  \n        The training data does not fix this: training pushes down the loss, and the model will take\n        any route that makes that number smaller. In this dataset, “there is a cow” and “there is\n        grass” are the same statement, so the loss has no reason to prefer one over the other, and\n        grass is the easier of the two. More photos means more cows on grass. Nothing separates the\n        two features, because **no photo in this set separates them**.\n      \n\n  \n        Merlin and Morgana fill that gap with one capability: they may **select part of an input**, so the label lands on a part instead of a whole picture. With it they can make two claims\n        no photograph in this set can make: the animal on its own is enough, and the background on\n        its own is not. Neither crop teaches Arthur anything by itself, so the figure below runs the\n        two of them against each other over three rounds of training. The round to watch is the\n        second, where Arthur pays for distrusting Morgana by also distrusting Merlin, because he has\n        no way of telling the two apart.\n      \n\nthe training data\nevery cow stands on grassevery camel stands on sandaccuracy on these photos is100%both images classified correctly in every round\n\nwhat each of them selects\n\nMerlin selects\n\nArthur says\n“Cow!”correct\n\nMorgana selects\n\nArthur says\n“Cow!”wrong\n\nthe consequence · a cow on the beach\n\nA photo that appears nowhere in the training data.\n\nArthur says\n“Camel!”wrong\n\nScroll the panel sideways to reach the scoreboard.\n\n  \n        Next to the training photos the figure grades Arthur the way a benchmark would, on the two\n        full images, and he is right in every round. Everything that changes happens in the crops\n        and on the beach, which no such test set contains, so Morgana is the only one who can show\n        the problem.\n      \n\n  \n        The same holds for retrieved text. The grass patch is a passage that turns up alongside the\n        answer without containing it: the right document, in a familiar format, close enough for the\n        model to guess the rest. In your archive those passages and their answers may be as tightly\n        coupled as cows and grass, and no amount of additional documents will separate them. Hiding\n        part of the passage does separate them, because the model sees the same document once with\n        the deciding sentence and once without.\n      \n\n## \n  Let’s define our grounding score\n\n  \n        Completeness and soundness are ordinary test-set measurements: two percentages. Because of\n        how the game is set up, those two are enough on their own to put a floor under how much the\n        answer depended on the document.\n      \n\n  \n        We use one bookkeeping trick to make this work for open-ended generation, where the model\n        can produce any string at all: we score each generated answer as correct or incorrect, which\n        collapses that open output space into a single two-way outcome. A two-way outcome you cannot\n        guess is worth exactly one bit of information, the amount you gain when you learn how a fair\n        coin landed. That bit acts as a budget.\n      \n\n  \n        The proof is in the [paper](https://arxiv.org/abs/2512.11614). Start with the two ways the system can let you down: it answers when it should have\n        abstained, or it abstains when it should have answered. Either failure misleads you, so the\n        two combine into one effective error rate. That rate determines the uncertainty about\n        whether the answer followed from the evidence. Measured in bits, that uncertainty is\n        entropy. Subtract that entropy from the one bit you started with, and the remainder is\n        information that provably came from the context. That remainder is what we call **certified**. It is a floor: whatever is going on inside the model, it cannot have drawn less than that\n        from the document. The certified bits are what we report as the score.\n      \n\n              one bit, the whole budget\n            \n                  5% of answers\n                  \n\n                  wrong\n                \n                    0.29\n                  \n                    0.71\n                  \n                  10% of answers\n                  \n\n                  wrong\n                \n                    0.47\n                  \n                    0.53\n                  \n                  25% of answers\n                  \n\n                  wrong\n                \n                    0.81\n                  \n                    0.19\n                  \n              doubt the errors leave behind: the entropy of the error rate\n            \n              certified: the grounding score\n            **Where the bit goes.** Every bar is the same one bit. The doubt always outweighs the error rate that caused it: one answer in twenty wrong already costs almost a third of the budget, one in four costs four fifths.\n    Where the numbers in the figure come from\n  \n\n          The doubt costs far more than the error rate suggests. A system that is wrong once in\n          twenty does not spend 0.05 of the bit, it spends 0.29. That 0.29 is the entropy of a coin\n          that comes up wrong 5% of the time, −plog⁡2p−(1−p)log⁡2(1−p)-p\\log_2 p-(1-p)\\log_2(1-p)−plog2​p−(1−p)log2​(1−p) with\n          p=0.05p=0.05p=0.05, and it measures what you\n          would still need to learn to know whether the answer in front of you is one of the\n          nineteen good ones or the one bad one. Take it off the bit and 0.71 remains. The reason it\n          costs so much is that rare errors are the ones you cannot predict, and unpredictability is\n          exactly what entropy counts.\n        \n\n  \n        A raw count of bits is awkward to compare across tasks, so the obvious fix is to normalise\n        it. Divide the certified bits by all the information the model’s answers carry about the\n        correct answer, from whichever source: the document, its memory from pre-training, or a\n        lucky guess. That total is set by the accuracy the model achieves, and the ratio then reads\n        as the share of the model’s information that we can trace back to the document.\n      \n\n  \n        The catch is that this denominator moves with the model. A model whose answers are often\n        wrong carries less information to begin with, so the very same evidence scores higher on the\n        worse system, and the ratio can even run past 1, which is meaningless for a share. On a hard\n        benchmark it collapses towards zero without telling you why: was the\n        evidence not doing the work, or is the model simply bad at the task?\n      \n\n  \n        So we hold the denominator still and grade only the questions the model gets right anyway,\n        which is the paper’s **conditional evaluation protocol**. On those questions the\n        budget is one bit by construction, so the score means the same thing across models and\n        datasets, and the only question it answers is “is this answer grounded?”, never “is this\n        model any good?”. We call this EIFcond\\mathrm{EIF}_{\\text{cond}}EIFcond​, the **grounding score** promised at the top of this post.\n      \n\nstagebefore trainingafter baselineafter M/Adrop unanswerable questions during trainingcompleteness93%soundness71%grounding score0.11\n\nTwo error rates in, a guaranteed floor out.\n**What completeness and soundness certify.** The three presets show completeness\n        and soundness before training, after M/A training and after standard training.\n        Move the sliders to get a feel for how the grounding score responds.\n\n  \n        Before training the bound stays weak, and the model is not much use on the task either, at\n        around 40% accuracy. Standard supervised training repairs accuracy and completeness, and\n        does nothing for soundness, which can even decrease. The reason is that the training set\n        pairs questions with answers, so the majority of training examples encourages the model to\n        produce one, and few of them tune it for noticing that the evidence is missing. Only M/A\n        training reduces both error rates at once. Drop the unanswerable questions, as in a company\n        archive where nobody has marked them, and the split widens: the baseline’s soundness falls\n        to around a third and takes the certificate, the grounding score it can still prove, with\n        it. M/A training moves the same measurement the other way and keeps the certificate\n        standing.\n      \n\n  \n        A model can be almost perfectly complete and still certify almost *nothing*, because\n        completeness alone leaves room for a model that answers no matter what you hand it. The\n        bound appears once the model also stops answering when the evidence is gone, and ordinary\n        fine-tuning has no reason to teach it that.\n      \n\n## \n  What the numbers say\n\n  \n        For our paper, we ran this on Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct,\n        Qwen3-4B-Instruct and Qwen2.5-32B-Instruct, across SQuAD2.0, HotpotQA, TriviaQA,\n        2WikiMultihopQA and MuSiQue, using [low-rank adaptation (LoRA)](https://arxiv.org/abs/2106.09685) for at most 200 steps. Every one of those runs carries two baselines: the instructed model\n        as it ships, and the same model after standard training. Below, we give changes in\n        percentage points (pp), the plain difference between two percentages.\n      \n\n#### \n  Fewer wrong answers when evidence is missing\n\n                    up to −35 pp\n                  \n                    vs. instructed LLM\n                  \n                    −18 to −20 pp\n                  \n                    vs. standard training\n                  \n#### \n  Higher grounding score\n\n                    +0.1 to +0.4\n                  \n                    vs. instructed LLM\n                  \n                    +0.33 to +0.38\n                  \n                    vs. standard training\n                  **What M/A training changes.** Bar heights are comparable within a panel, not across it, and where we report a range the bar sits in the middle of it.\n\n  \n        In absolute terms those gains put the grounding score at 0.55 to 0.6 on five of the six\n        dataset settings we ran: on the questions the model answers correctly, more than half of\n        what it is doing provably comes from the retrieved context.\n      \n\n  \n        Accuracy on the original retrieved context does not move. Trading utility for soundness is\n        easy, and any model that refuses more often looks safer on a safety metric. M/A training\n        matches standard training on the questions that call for an answer while improving\n        everything above.\n      \n\n  \n        Two things about the setup. Every model in the comparison, ours and both baselines, is\n        already instructed in its prompt to abstain when the context does not support an answer, so\n        prompting alone cannot explain the gap: telling a model to hold back does not make it able\n        to tell when it should. And the abstention behaviour also appears on HotpotQA and\n        2WikiMultihopQA, which carry no annotated unanswerable questions at all, so the supervision\n        has to be coming from Morgana building those cases herself.\n      \n\n## \n  Does this survive contact with real documents?\n\n  \n        QA benchmarks are how a method proves itself against the literature. Customer applications\n        ask harder questions of it, so we also used the protocol on something much closer to real\n        work: German environmental and administrative law. The documents behind it are real and\n        public, so we can show the whole evaluation.\n      \n\n  \n        The case is the approval procedure for one large infrastructure project, a hydrogen\n        pipeline, whose file also covers the water crossings, roads and rail lines along the route.\n        A legal rule comes with a list of conditions and applies only if the project meets every one\n        of them, so a lawyer checks them one by one. Each item takes one of those conditions and\n        asks whether the facts of this particular project satisfy it: whether the project avoids a\n        deterioration of a water body’s ecological status, say, or whether evaluators can still\n        credit measures that a conservation plan already fixes as compensation. The evidence is the\n        project’s case file, 20,000 to 30,000 characters of retrieved sources such as the\n        applications the developer filed for it, environmental reports and compensation plans, and\n        on some items far more. The task allows three answers: the facts fulfil the condition, they do not fulfil it,\n        or the file does not say (inconclusive). The third option is the one we care about, because\n        a confident guess here costs more than no answer at all. Below is one item from the set,\n        worked through.\n      \n\n            the question, asked about one pipeline project\n          \n\n            Can this project be released from the nature-protection bans: **has an exemption been applied for, and is it necessary either on overriding\n              public-interest grounds or to avoid an unreasonable hardship**?\n          \n\n              What the norm requires\n            \n\nAn exemption has to be applied for\n\n                    the norm grants it only on application\n                  \n\n–The file seeks one for the landscape-protection area, and promises further ones if they turn out to be needed\n\nOverriding public interest has to make the deviation necessary\n\n                    the first of two grounds the norm allows\n                  \n\n✗The file argues public interest for a different request, and never says a deviation here is necessary\n\nOr enforcing the rule would be an unreasonable hardship\n\n                    the second ground, and one of the two has to hold\n                  \n\n✗Nothing in the file speaks to hardship at all\n\n              What the file offers instead\n            \n\n                “The German legislature likewise finds an overriding public interest in the rapid\n                construction of hydrogen transport pipelines.”\n              \n\n                A finding about a whole class of projects, written into the file to support a\n                different request. It never says this exemption is necessary.\n              \n\n✓\n            The case file establishes neither ground, so the only correct verdict is **Inconclusive**. The same structure as the supplier contract example, on a real file: the decisive\n            fact is missing, and the file contains a statement that a model can mistake for it.\n          \n\n  \n        We put this item to our model after M/A training, to the same model trained without the\n        protocol, and to four frontier systems. Two of the six said the file does not settle it:\n        ours, and the older of the two Opus versions.\n      \n\n            One example from the legal evaluation, where the correct verdict is Inconclusive. The M/A model is an internal Aleph Alpha model.\n          \n                System\n              \n                Model verdict\n              \n                What happened\n              Internal model, with M/A✓ Inconclusive (correct)Returned “Reject”, citing the public-interest passage and stating that the file does not say an exemption was applied for or grantedSame model, standard training✗ Fulfilled (wrong)Returned “True”, taking the public-interest passage as the ground the norm asks for, and treating it as a separate question that the file never says an exemption was granted, or that the bans are triggered at allGPT-5.5-high✗ Fulfilled (wrong)Returned “True” on 2,035 billed reasoning tokens; the reasoning summary returned with it includes “maybe it needs more depth”GPT-5.6-high✗ Fulfilled (wrong)Returned “True” on 166 billed reasoning tokens, with no reasoning text returned alongside the answerOpus 4.8-high✓ Inconclusive (correct)Returned “Reject” with a written justification, which states that an exemption is applied for only if it turns out to be needed, and that impacts on the protected biotopes are avoidedOpus 5-high✗ Fulfilled (wrong)Returned “True”; the reasoning returned with it cites the exemption request in the file and the overriding public interest in hydrogen pipelines\n    How we ran the comparison\n  \n\n          Every system got the same two prompts on every item, and only the model changed. This was\n          the system prompt, in full:\n        \n\n          You are a legal expert analyzing Tatbestandsmerkmale.\n\n          Your task is to decide whether a given Tatbestandsmerkmal is fulfilled based on the provided\n          Sachverhalt.\n\n          Answer with exactly one of: \"True\" (fulfilled), \"False\" (not fulfilled), or \"Reject\" (insufficient\n          information to decide or in doubt).\n\n          Try to finish reasoning within 3k tokens.\n        \n\n          That is where the abstention instruction comes from, and everyone got it, our two models\n          as well as the four frontier ones. The user prompt and data are in German.\n        \n\n          We queried the frontier systems through OpenRouter on 12 August 2026 with reasoning effort\n          set to high: *openai/gpt-5.5*, *openai/gpt-5.6*, *anthropic/claude-opus-4.8* and *anthropic/claude-opus-5*. Our own two rows are the same internal 30B model: we\n          trained it once with the protocol and once without, on the same data.\n        \n\n          Every system answered each item once. We did not sample repeatedly or take a majority\n          vote, so any single item could come out differently on a rerun.\n        \n\n          The case files come from a planning approval procedure for a hydrogen pipeline near\n          Lingen, whose decision and documents were [put on public display](https://www.lingen.de/politik-rathaus-service/veroeffentlichungen/bekanntmachungen/planfeststellungsverfahren-fuer-die-errichtung-und-den-betri.html) in October 2023. We stripped personal names and contact details before we stored the files\n          or ran anything on them.\n        \n\n  \n        We never trained the M/A model on this task at all: we trained it with the protocol on very\n        different data, with the reinforcement-learning objective we described earlier, and the\n        behaviour transferred. The evaluation has 100 items from the same case files. On 78 of them\n        the file answers the question (“fulfilled” or “not fulfilled”). On 22 it does not, and it is\n        exactly these “Inconclusive” cases, where the model should have abstained instead of\n        confidently giving a verdict, which are the ones most interesting to our method.\n      \n\n                    Internal model, M/A-trained\n                  \n                    20 of 22\n                  \n                    Same model, standard training\n                  \n                    11 of 22\n                  \n                    GPT-5.5-high\n                  \n                    16 of 22\n                  \n                    GPT-5.6-high\n                  \n                    16 of 22\n                  \n                    Opus 4.8-high\n                  \n                    16 of 22\n                  \n                    Opus 5-high\n                  \n                    6 of 22\n                  **22 items whose correct verdict is Inconclusive.** One block per item, in the same order in every row, hardest on the left: turquoise where the system said the file does not settle it, red where it returned a verdict the file does not support. On the two grey blocks GPT-5.5 ran into its token limit before returning a verdict.\n\n  \n        Our model catches 20 of the 22, four more than the three frontier systems that tie behind it,\n        and it does that at 30B parameters against systems that even the most conservative public\n        estimates put at ten times that size. The same model trained without the protocol catches\n        11. Every model was told in its prompt to answer “insufficient information” whenever the\n        file does not settle the question, so the gap is not about who was asked to hold back.\n      \n\n                  Internal model, M/A-trained\n                \n                        57\n                      \n                        36\n                      \n                        5\n                      \n                        2\n                      \n                  5 wrong\n                \n                  Same model, standard training\n                \n                        62\n                      \n                        3\n                      \n                        32\n                      \n                        3\n                      \n                  32 wrong\n                \n                  GPT-5.5-high\n                \n                        69\n                      \n                        14\n                      \n                        9\n                      \n                        8\n                      \n                  9 wrong\n                \n                  GPT-5.6-high\n                \n                        73\n                      \n                        17\n                      \n                        10\n                      \n                  10 wrong\n                \n                  Opus 4.8-high\n                \n                        77\n                      \n                        8\n                      \n                        15\n                      \n                  15 wrong\n                \n                  Opus 5-high\n                \n                        68\n                      \n                        7\n                      \n                        25\n                      \n                  25 wrong\n                \n                  Correct\n                \n                  Held back where the file answers\n                \n                  Wrong verdict\n                \n                  Returned no verdict\n                **What each system returned, over all 100 items.** The first two rows are the same 30B model: we trained one with the protocol and one without. Holding back on an item the file answers costs accuracy (turquoise green). **A wrong verdict costs more (red).**\n\n  \n        Over all 100 items our M/A-trained model gives 5 wrong verdicts in total. The same model\n        without the protocol makes more of them than any other system here. Opus 4.8, the most\n        accurate system overall, gives 15.\n      \n\n  \n        A note on reading this chart. Most items here have an answer in the file, so a model that\n        answers everything scores well on overall accuracy, because accuracy rewards guessing. Abstaining\n        costs at most one correct item, while a wrong verdict is the expensive mistake.\n        So we want to minimise the red in the bars here and in the boxes above, because those are the items where\n        the model should have abstained instead of making a wrong verdict. We published every input and output behind the evaluations above\n        [here](https://github.com/Aleph-Alpha-Research/public-domain-eval).\n      \n\n    Why not just retrieve better, or teach the model the subject?\n  \n\n          The model never needs to know the law. The rule comes with the context: the clause in the\n          supplier contract case, the condition quoted in full at the top of every item here. All\n          that is left is to check whether the facts of this project meet it. That check is where\n          the systems in the table fail. They find a passage that speaks to the rule and treat it\n          as if it settled the condition.\n        \n\n          Pre-training teaches the model a lot: matching an entity in the file to the one the rule\n          names, following what a sentence means, knowing what a permit is. Pre-training rarely\n          teaches a model to notice that the deciding piece is absent, and that is the one thing\n          Morgana manufactures, sample after sample. She takes away the sentence he was leaning on\n          and leaves standing what still looks relevant, until he learns that without it the\n          question is not settled.\n        \n\n          Better retrieval does not help either. To know that a context is complete, you have to\n          make this exact judgement first. And nothing was missing from the retrieval here: the\n          file promises to apply for the exemption if it turns out to be needed, and no retriever\n          can fetch a decision the procedure has not yet produced.\n        \n\n  \n        This small, early evaluation points a direction more than it settles anything, and the\n        direction is the one the theory predicts. A new LLM version can leave a model less careful\n        than the one before it: Opus 4.8 catches 16 of the 22, and Opus 5, released after it,\n        catches 6. With this protocol we set that level ourselves. Morgana’s weight in the loss\n        decides how much of Arthur’s training goes into holding back, and where to leave it follows\n        from what a wrong answer costs you against a missing one. Your business knows that number,\n        and no model provider can know it for you.\n      \n\n## \n  Why this direction matters\n\n    What this method does not address\n  \n\n          The bound answers exactly one question: whether *this* answer followed from *this* context. Three things sit outside it and need methods and measures of their own:\n        \n\n**Truth, as opposed to grounding.** The bound certifies that the answer follows\n          from the context. Whether the context itself holds up is a separate matter: feed the system\n          misinformation and it will faithfully treat that misinformation as proof.\n        \n\n**How often the system ought to abstain.** The weighting of Morgana against Merlin\n          during training sets how cautious Arthur becomes, but nothing in the method says where that\n          weighting belongs. That depends on what a wrong answer costs you compared with a missing one,\n          and only your own business can price that.\n        \n\n**How strong the adversary is.** The certificate inherits Morgana’s quality. A\n          weak Morgana does not make the number wrong, only pessimistic, since anything she fails to find\n          leaves the bound more conservative than it needs to be.\n        \n\n  \n        Fewer hallucinations is the visible gain, and the supervision underneath it matters more:\n        the system produces that supervision *about itself*, directs it at whatever it is\n        currently weak at, and attaches a certificate. The certificate is the grounding score: a\n        number that travels with the system, says how much the documents provably contributed to its\n        answering, and lets anyone who doubts it recompute the result. Scaling that kind of compute\n        is a different bet from scaling annotation: it does not run out, and the capability of\n        whichever model you paid to write your labels does not cap it.\n      \n\n**We would like the certificate part to become normal.** When a language model does\n        something consequential with a document, such as a legal filing, a medical record or an engineering\n        specification, the question people have is whether an answer came from the specific document\n        provided. A model that “scored 90% on a benchmark” cannot answer this question. That connects\n        to an argument we make more broadly: [transparency is a pillar of sovereign AI](/en/blog/transparency-as-one-pillar-of-sovereign-ai/), and it is only worth something if it can be checked. Model cards describe how a system\n        was built, while a number like this one applies to a single answer after the fact, and\n        anyone who has to justify a decision to an auditor, a regulator or a court needs both.\n      \n\n  \n        A system that knows when it cannot answer is worth more than one that guesses right slightly\n        more often.\n      \n\n## \n  More blog posts\n\n- [\n### \n      Kolibri Has Landed: A Sovereign Open-Weight Model\n    \n\nResearch03/10/2026\n](/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/)\n\n- [\n### \n      Scaling Pre-Training in Practice: A Hierarchical Approach\n    \n\nResearch30/09/2026\n](/en/blog/scaling-pre-training-in-practice-a-hierarchical-approach/)\n\n- [\n### \n      Training on the Party Line: Chinese Political Influence on LLMs in China and the World\n    \n\nResearch28/09/2026\n](/en/blog/training-on-the-party-line/)","body_html":"<p>Research</p>\n<p>Letitia Parcalabescu</p>\n<pre><code>        08/08/2026</code></pre>\n<p># \n  Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models</p>\n<p><a href=\"https://x.com/intent/tweet?url=https%3A%2F%2Faleph-alpha.com%2Fen%2Fblog%2Fbounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models%2F&amp;text=Bounding%20Hallucinations%3A%20Merlin-Arthur%20Protocols%20for%20Mutual-Information%20Bounds%20in%20Language%20Models\" rel=\"nofollow ugc noopener\"></a><a href=\"https://www.linkedin.com/sharing/share-offsite/?url=https%3A%2F%2Faleph-alpha.com%2Fen%2Fblog%2Fbounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models%2F\" rel=\"nofollow ugc noopener\"></a><a href=\"mailto:?subject=Bounding%20Hallucinations%3A%20Merlin-Arthur%20Protocols%20for%20Mutual-Information%20Bounds%20in%20Language%20Models&amp;body=https%3A%2F%2Faleph-alpha.com%2Fen%2Fblog%2Fbounding-hallucinations-merlin-arthur-protocols-for-mutual-information-bounds-in-language-models%2F\"></a></p>\n<pre><code>    In plain terms: how to check whether a language model actually used the document you gave\n    it, and how to train one so that it does.\n  \n\n\n    Our preprint ([arXiv:2512.11614](https://arxiv.org/abs/2512.11614)) develops a way to do both. The figures on this page are interactive: some reconstruct the\n    paper’s results, others illustrate the mechanism. See the paper for exact numbers.</code></pre>\n<p>## \n  9 right, 1 wrong answers and no way to tell them apart</p>\n<pre><code>    A system answers questions based on your documents and scores 90%: nine of its ten answers\n    are right, one is wrong, and nothing marks which one. That is a capable system, and the 90%\n    still buys little, because the bad answer hides among the nine good ones and a human has to\n    check all ten to find it. Now picture the same 90% with the wrong one flagged: “I cannot\n    answer this from the documents I was given.” The model knows no more than it did before, but\n    now you can put it in front of a customer.\n  \n\n              A system at 90%</code></pre>\n<p>?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one?an answer with nothing to tell it apart from a wrong one</p>\n<pre><code>              One of these ten is wrong, and nothing marks which one.</code></pre>\n<p>I have to check all ten.</p>\n<pre><code>              The same 90%, able to abstain</code></pre>\n<p>✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind—the system says the documents do not settle this one✓an answer the system stands behind✓an answer the system stands behind✓an answer the system stands behind</p>\n<pre><code>              The system marks the one the documents do not settle.</code></pre>\n<p>I check the one it flagged, and act on the other nine.<strong>The same 90%, with and without abstention.</strong> The difference is whether you can tell the good answers from the bad ones <strong>without checking every one of them yourself</strong>.\n## \n  TL;DR</p>\n<p><strong>The problem.</strong> In LLM training, a guess scores better than “I cannot answer this”,\n        so training pushes a language model to answer even when the document says nothing about the question.\n        Retrieval-augmented generation (RAG) hands it a specific document and hopes the LLM’s answer comes\n        directly from it, but nothing checks whether it actually did. The benchmark score will not tell\n        you: it says whether the answer was right, not whether it came from the specific document, from\n        what the model memorised in pre-training, or from a cue no human would consider relevant. Asking\n        for a citation does not help either. The model picks the quote once it has already answered the\n        question, so the quote can be genuine while the model’s answer came from somewhere else.</p>\n<p><strong>Our method.</strong> We modify the LLM training procedure by turning it into a game with\n        three players in which guessing something plausible stops working. <strong>Arthur</strong> is\n        the model you would ship. <strong>Merlin</strong> hands him a context that supports the correct\n        answer. <strong>Morgana</strong> cuts the relevant evidence out, to lure him into a hallucination.\n        Arthur has to answer on Merlin’s context and abstain on Morgana’s, without knowing which player\n        he is facing. Neither context is fixed: both are built at every step from the model as it is right\n        then, so the game keeps up with whatever Arthur is getting away with. At test time, how often he\n        gets both right becomes a <strong>grounding score</strong>: the share of the answer that provably\n        came from the document provided.</p>\n<p><strong>We increase the grounding score and reduce hallucinations.</strong> Across five QA benchmarks,\n        wrong answers under insufficient context fall by up to 35 pp against an instructed LLM and by\n        18 to 20 pp against standard training. The grounding score rises by up to 0.38. Later in this post\n        we test our training method on real case files from German environmental and administrative law.</p>\n<p>## \n  A benchmark score does not say where the answer came from</p>\n<pre><code>    Take the standard setup, retrieval-augmented generation, **RAG** for short. It has\n    two parts: a retriever that picks documents out of your archive, and a generator, the language\n    model, that writes an answer while looking at them. Evaluate the pair on a question-answering\n    benchmark and you get a number like the 90% from the scoreboard above. The number counts right\n    answers and stays silent on which of three routes produced them: reading the retrieved document,\n    remembering the answer from pre-training, or leaning on a spurious cue such as the phrasing of\n    the question or an entity that appears in only one distractor. A model may well have memorised\n    a benchmark during training, but it cannot have memorised the documents users bring.\n  \n\n\n    Not every question needs this. If pre-training already covers the answer, the retrieved page\n    is dead weight and dropping it saves tokens. Retrieval pays for itself where the answer\n    exists nowhere but in your own documents: what a supplier agreed to in a 2019 amendment,\n    what tolerance a specification demands. Pre-training does not hold that information, so\n    anything a model produces here without reading the document is a guess about facts that live\n    only in your own files.\n  \n\n\n    A question about one supplier contract makes that concrete. It arrives with the facts of the\n    case, and the retrieved context is a handful of clauses from the agreement with that\n    supplier. Every supplier negotiates its own terms, which is the whole reason for retrieving\n    them instead of trusting the model to carry them in its head.\n  \n\n          the case accompanying the question\n        \n\n          A delivery arrived with a defect. The buyer found it 15 working days ago and has now\n          notified the supplier by signed letter. **Is the notice still in time?**\n\n          the clause that decides it, retrieved from the contract\n        \n\n          “A defect must be notified *within ten working days* of the buyer discovering it.”\n        \n\n          Fifteen working days is more than ten, so the answer is no.\n        \n\n\n    Now let the retriever miss that one clause. What remains comes from the same contract and\n    still reads as relevant: the rule that a notice of defect must be given in writing. That is\n    the one condition the signed letter does meet, so a model that was never trained to check\n    whether the evidence is there works through the rule still in front of it, finds it\n    satisfied, and answers that the notice is in time. Nothing in the answer marks it as the\n    guess it is, and on a benchmark this counts as one wrong answer among many,\n    indistinguishable from a model that never saw the contract at all.\n  \n\n\n    The model never saw the clause, so it cannot know that one is gone, and it does not need to.\n    Deciding whether the notice was in time takes a deadline to measure those fifteen days\n    against, and nothing in the retrieved contract sets one. The model can see that much, and it\n    is enough to answer that the file does not settle the question. A lawyer reading the same\n    file, not knowing which page had been pulled, would answer the same way.\n  \n\nOur case is illustrative, the failure is common\n\n\n      The literature has reported many failures. Rephrase a question in a different style and\n      the retriever hands back a [different document](https://arxiv.org/abs/2504.08231). Generators keep answering when the retrieved passage [does not support the answer](https://openreview.net/forum?id=ztzZDzgfrh), and they struggle when two retrieved documents [contradict each other](https://arxiv.org/abs/2504.13079). And when something in the context merely correlates with the answer, models [reach for it anyway](https://aclanthology.org/2022.findings-naacl.130/), warranted or not.\n    \n\n\n    The training procedure bakes in the cause: **guessing pays, and it pays well**.\n    A multiple-choice exam gives you a point for a right answer and nothing for a wrong one. A\n    blank scores zero for certain, a guess has a one in four chance, so you guess. Benchmarks\n    grade language models that way too, and models are built to score well on them. Free-form\n    generation offers far more than four options, so a lucky hit is rare, and staying quiet\n    still pays nothing. The model writes its most fluent answer instead, and that is the\n    plausible hallucination you have trouble telling apart from a true answer. [Kalai et al.](https://doi.org/10.1038/s41586-026-10549-w) spell out this argument.\n  \n\n\n    So nothing in a standard RAG pipeline *forces* the answer to depend on the context. We\n    need a better training procedure and, before that, a better quantity to aim at:</code></pre>\n<p>  How much of the answer came from the document?</p>\n<p>## \n  Merlin, Morgana and Arthur</p>\n<p><a href=\"https://en.wikipedia.org/wiki/Interactive_proof_system\" rel=\"nofollow ugc noopener\">Interactive proof systems</a> come from complexity theory in the mid-1980s and give Merlin and Arthur their names. One side\n        tries to convince the other that a claim is true, and the setup is only worth something if nothing\n        can talk the verifier into accepting something false. <a href=\"https://arxiv.org/abs/2206.00759\" rel=\"nofollow ugc noopener\">Wäldchen et al.</a> brought it to image classifiers in 2024; we bring it to language models, and as far as we know,\n        ours is the first guarantee of this kind connecting a retrieved document to a generated answer:\n        we compute a bound on how much of the answer came from the document from two ordinary test-set\n        measurements. Below, we shorten Merlin-Arthur to M/A.</p>\n<ul><li><p><strong>Arthur</strong> is the LLM you would actually ship. He gets a question and a document,</p><pre><code>    and has exactly two options: answer it, or say that the document does not settle the matter.</code></pre></li></ul>\n<ul><li><p><strong>Merlin</strong> is on Arthur’s side. Out of the retrieved material he assembles the</p><pre><code>    most helpful version he can, the one that keeps the sentence that decides the answer. Think\n    of a colleague who marks the relevant passage before handing you the file.</code></pre></li></ul>\n<ul><li><p><strong>Morgana</strong> works against him. She starts from a document that does answer the</p><pre><code>    question and removes precisely the part that establishes it, leaving behind everything that\n    still reads plausibly. Think of the same colleague removing the decisive page to see whether\n    you were paying attention when reading the file.</code></pre></li></ul>\n<pre><code>    Merlin and Morgana exist only during training and testing. Arthur cannot tell which of them\n    he is facing: same question, same kind of document, no label saying who prepared it. If he\n    wants to be right, he has to look at what is in front of him. A word on wording: a *passage* is a chunk the retriever returned, and a *sentence* is what Merlin and Morgana keep or\n    remove inside it, whatever the unit happens to be in practice.\n  \n\n\n    We train neither of them, and nobody marks up the documents by hand. Arthur’s own\n    probabilities reveal which sentence decides the answer, and that is the next section. The\n    questions and answers that came with the dataset provide the only human labelling.\n  \n\n\n    In the supplier contract case above, the retriever missed the ten-day clause by accident.\n    Morgana removes it on purpose. She leaves the written-form clause standing, exactly as\n    before, so a model that answers without reading is caught doing it. Merlin does the opposite\n    with the same file: the ten-day clause stays, the clauses around it go. What was one unlucky\n    retrieval is now something we can arrange for every sample in the training set.\n  \n\n\n    The paper rests on two quantities, best read as a pair. **Completeness** is how often\n    Arthur answers correctly on Merlin’s context: accuracy in the *best* case, when the evidence\n    has been laid out for him. **Soundness** is how often he avoids a wrong answer on\n    Morgana’s: accuracy in the *worst* case, when someone has actively tried to make him hallucinate.\n  \n\nWe define soundness strictly\n\n\n      We define soundness more strictly than prior work. Arthur has to *abstain*, saying\n      outright that the context does not settle the question, rather than merely avoid being\n      wrong; later sections call the same move refusing or answering “I don’t know”. The\n      strictness matters because language models memorise benchmarks. A model that recovers the\n      correct answer from a context that no longer contains it is answering from memory, and a\n      looser definition would quietly reward that.\n    \n\n\n    The figure below plays the supplier contract case out in all four cases: Merlin or Morgana\n    preparing the document, Arthur before or after M/A training. Bottom left shows the\n    hallucination above; the right-hand column shows Arthur after training.</code></pre>\n<p>the question</p>\n<p>The buyer discovered the defect 15 working days ago and has now notified the supplier, by signed letter. Is the notice still in time?</p>\n<p>what Arthur sees</p>\n<p>Struck through means hidden from Arthur, shown here so you can see what is missing.</p>\n<ul><li>A defect must be notified within ten working days of the buyer discovering it.decisive, hidden</li><li>Deliveries are inspected at the buyer’s site after unloading.hidden</li><li>The ten working days start once the buyer reliably knows the defect.decisive, hidden</li><li>A notice of defect must be given in writing.</li><li>The supplier may instead offer a replacement delivery.</li></ul>\n<p>Arthur answers</p>\n<p>The notice is in time.\nhallucination</p>\n<p><strong>Soundness failure.</strong> Nothing left settles the question. Arthur checks the one rule he can still see, finds the letter satisfies it, and answers.</p>\n<p>## \n  Who plays Merlin and Morgana?</p>\n<pre><code>    Nobody hand-picks the sentences. What Merlin is really after is the set of sentences to keep\n    that leaves Arthur most likely to answer correctly, and Morgana the set that leaves him\n    least likely to while the document still reads as though it should settle the question. We\n    measure best and worst by what the choice does to Arthur’s probability of the correct\n    answer. The trouble is how many choices there are: a document of nnn sentences can be cut 2n2^n2n ways, so at forty\n    sentences we are already past a trillion, and trying them all is out of the question. We let an\n    explainability method do the work instead, and it is one of ours: [AtMan](https://proceedings.neurips.cc/paper_files/paper/2023/hash/c83bc020a020cdeb966ed10804619664-Abstract-Conference.html), short for attention manipulation, developed here at Aleph Alpha and published at NeurIPS\n    2023.\n  \n\n\n    Go through the document one sentence at a time, hide that sentence, and watch the model’s\n    probability for the correct answer. If it barely moves, that sentence was not carrying the\n    answer. If it falls away, that sentence was. One pass per sentence, so the work grows in\n    step with the length of the document instead of exploding with it. Out comes a ranking of\n    the document by how much each sentence matters, which Merlin and Morgana read from opposite\n    ends of the same ranking: Merlin hides the bottom so the evidence survives, Morgana hides\n    the top so only what looks relevant does.</code></pre>\n<p>how much the answer depends on itMerlinMorgana\n                    A defect must be notified within ten working days of the buyer discovering it.</p>\n<pre><code>                keeps\n              \n                hides\n              \n                The ten working days start once the buyer reliably knows the defect.\n              \n                keeps\n              \n                hides\n              \n                Deliveries are inspected at the buyer’s site after unloading.\n              \n                hides\n              \n                hides\n              \n                A notice of defect must be given in writing.\n              \n                hides\n              \n                keeps\n              \n                The supplier may instead offer a replacement delivery.\n              \n                hides\n              \n                keeps\n              **One ranking, read from both ends.** Note where the written-form clause sits: it reads as relevant, but hiding it barely moves the answer, so it survives Morgana’s cut and remains as the decoy. Illustrative weights on a real ordering.\n\n\n    A couple of notes. How much Merlin and Morgana hide is a setting: hide too little and\n    Morgana leaves the decisive sentence in place, hide too much and Merlin throws it out with\n    the filler. We can use a fixed rate, though each sample can also work out its own. The size\n    of the unit is a setting too, anywhere from a single word to a whole paragraph, where\n    smaller units pin the evidence down more precisely and larger ones cost less to examine. And\n    the hiding happens *inside the model’s attention* rather than by cutting the text up, so\n    no placeholder replaces the hidden text for the model to recognise.\n  \n\nIf Merlin and Morgana only approximate, is the guarantee worth anything?\n\n\n      We checked AtMan’s picks against an exhaustive search over every possible selection, and\n      the two agree closely. It works in our favour anyway: an approximation only makes the\n      bound more conservative, so a better way of picking sentences can only raise the number we\n      report. The measurements are in the [paper](https://arxiv.org/abs/2512.11614).</code></pre>\n<p>## \n  Training Arthur: three contexts, one objective</p>\n<pre><code>    Every sample takes two passes. In the first, no training happens: AtMan goes through the\n    retrieved documents against Arthur exactly as he is at this moment, hides one sentence at a\n    time, and ranks the sentences by how much Arthur leans on them to answer correctly. Merlin\n    and Morgana then cut from opposite ends of that ranking, which gives us the second pass.\n  \n\n\n    There we build three contexts, the original retrieval plus Merlin’s and Morgana’s, and ask\n    Arthur for three answers. We score each LLM answer against the expected answer, and each\n    score contributes one weighted term to the *loss* that training pushes down. The loss measures\n    how far the model’s answer diverged from the expected answer.\n  \n\n\n    Morgana makes guessing unprofitable. When Arthur guesses on her context, he scores nothing,\n    while holding back scores him a point. And Arthur cannot tell her context from Merlin’s:\n    guess everywhere and he loses on hers, hold back everywhere and he loses on Merlin’s.\n    Checking whether the evidence is really there is the only move that scores on both.\n  \n\n\n    Each of the three terms carries a weight, and those weights decide how cautious Arthur\n    becomes. If Morgana carries most of the weight, answering without evidence is his most\n    expensive mistake. That makes soundness increase, and it could cost him questions he could\n    have answered. Shifting the weight onto the other two terms makes him pay more heavily for\n    staying silent when the evidence is there, so he answers more, at the price of the\n    occasional confident mistake. Where to set the weights is a product decision. Three contexts\n    per sample cost more compute than one, though the behaviour converges in as few as 200\n    steps. Giving the original retrieval the whole weight and turning Merlin and Morgana off\n    falls back to ordinary fine-tuning, which is the baseline in our evaluations — on the\n    same code, data and settings.\n  \n\n                  One training sample\n                \n                  a question, and the documents the retriever returned for it\n                \n                  AtMan scores the documents\n                \n                  each passage ranked by how much Arthur’s answer depends on it\n                \n                  The retrieval\n                \n\n                  the documents as the retriever returned them\n                \n\n                  Answer\n                \n\n                  Merlin\n                \n\n                  keeps the passage that makes Arthur answer correctly\n                \n\n                  Answer\n                \n\n                  Morgana\n                \n\n                  removes that passage, keeps the rest of the documents\n                \n\n                  Abstain\n                \n\n                  Arthur answers all three\n                \n                  the LLM we would ship, with no label saying which version he is reading\n                \n                  One weighted loss\n                \n                  we score him for answering the first two correctly and holding back on Morgana’s, and the weights on those scores set how cautious he becomes\n                \n                  One gradient step\n                \n                  Arthur changes, so Merlin and Morgana cut the next document differently</code></pre>\n<p>↺\n            The next sample is scored against the model this step just produced.</p>\n<p><strong>One step of M/A training.</strong> AtMan ranks the passages first. Merlin then keeps the passage Arthur’s answer depends on and Morgana removes it, Arthur answers all three versions of the documents, and one objective combines the three scores.</p>\n<pre><code>    Merlin and Morgana run *on the fly*, against the model as it currently is, so every\n    context they build goes after whatever it is getting away with at that step. A dataset\n    collected in advance can only cover the mistakes the model was making when the dataset was\n    created.\n  \n\n\n    That solves a data problem too. Teaching a model to abstain normally costs annotated\n    unanswerable questions or preference pairs in which the refusal is the preferred answer.\n    Academic benchmarks come with those annotations, and a company’s document archive does not.\n    Morgana needs neither, because every context she builds is one in which abstaining *is* the correct answer. And since she removes what *this* model was leaning on, the question\n    comes out unanswerable for the model rather than for an annotator, and it is the model’s version\n    of unanswerable that determines whether it hallucinates.\n  \n\nTraining the same protocol with reinforcement learning\n\n\n      The loss above does not have to be a supervised one. We also ran the protocol with\n      reinforcement learning (RL). Merlin and Morgana work exactly as before, and only the way\n      we score Arthur changes.\n    \n\n      In the supervised version Arthur answers each context once, and each answer adds one term\n      to the loss. Under [GRPO](https://arxiv.org/abs/2402.03300), a form of RL, we ask him each context several times. A language model picks every word\n      from a probability distribution, which is called sampling, so asking the same thing twice\n      gives two different answers. We end up with a handful of answers per context: some\n      confident, some hedged, some wrong. Each one earns a reward for doing what its context\n      asked for, which is a correct answer on the original retrieval and on Merlin’s, and an\n      abstention on Morgana’s. GRPO compares the answers within a context and updates Arthur’s\n      parameters such that his updated version is likelier to produce the answers that scored\n      higher than the average over all the sampled answers. Nobody tells him the right answer,\n      he only finds out which of his own tries went down better than the rest.\n    \n\n      He still needs the reward, though, and rewards usually cost something: either training a\n      reward model on human preferences, or creating a dataset where someone has marked which\n      questions lack supporting evidence. Here, Morgana automatically generates unanswerable\n      examples. The type of context an answer came from already says what Arthur should have\n      done, namely answer on Merlin’s and abstain on Morgana’s. So there is no reward model, and\n      no labelling beyond the questions and answers the dataset already has.</code></pre>\n<p>## \n  Why Merlin and Morgana have to co-evolve</p>\n<pre><code>    Why do we need an adversary at all? Ordinary training already shows the model a document and\n    the correct answer. What can Morgana teach it that a normal training run cannot?\n  \n\n\n    Consider training an image classifier on photos of cows and camels, where every cow stands\n    on grass and every camel stands on sand. Accuracy comes out high. Then you show it a cow on\n    a beach and it answers camel, and an empty meadow and it answers cow. The feature it learned\n    was grass. Explainability research has been reporting this for years. In the best-known\n    case, researchers trained a classifier to tell huskies from wolves, and it [turned out to be looking at the snow behind them](https://arxiv.org/abs/1602.04938).\n  \n\n\n    The training data does not fix this: training pushes down the loss, and the model will take\n    any route that makes that number smaller. In this dataset, “there is a cow” and “there is\n    grass” are the same statement, so the loss has no reason to prefer one over the other, and\n    grass is the easier of the two. More photos means more cows on grass. Nothing separates the\n    two features, because **no photo in this set separates them**.\n  \n\n\n    Merlin and Morgana fill that gap with one capability: they may **select part of an input**, so the label lands on a part instead of a whole picture. With it they can make two claims\n    no photograph in this set can make: the animal on its own is enough, and the background on\n    its own is not. Neither crop teaches Arthur anything by itself, so the figure below runs the\n    two of them against each other over three rounds of training. The round to watch is the\n    second, where Arthur pays for distrusting Morgana by also distrusting Merlin, because he has\n    no way of telling the two apart.</code></pre>\n<p>the training data\nevery cow stands on grassevery camel stands on sandaccuracy on these photos is100%both images classified correctly in every round</p>\n<p>what each of them selects</p>\n<p>Merlin selects</p>\n<p>Arthur says\n“Cow!”correct</p>\n<p>Morgana selects</p>\n<p>Arthur says\n“Cow!”wrong</p>\n<p>the consequence · a cow on the beach</p>\n<p>A photo that appears nowhere in the training data.</p>\n<p>Arthur says\n“Camel!”wrong</p>\n<p>Scroll the panel sideways to reach the scoreboard.</p>\n<pre><code>    Next to the training photos the figure grades Arthur the way a benchmark would, on the two\n    full images, and he is right in every round. Everything that changes happens in the crops\n    and on the beach, which no such test set contains, so Morgana is the only one who can show\n    the problem.\n  \n\n\n    The same holds for retrieved text. The grass patch is a passage that turns up alongside the\n    answer without containing it: the right document, in a familiar format, close enough for the\n    model to guess the rest. In your archive those passages and their answers may be as tightly\n    coupled as cows and grass, and no amount of additional documents will separate them. Hiding\n    part of the passage does separate them, because the model sees the same document once with\n    the deciding sentence and once without.</code></pre>\n<p>## \n  Let’s define our grounding score</p>\n<pre><code>    Completeness and soundness are ordinary test-set measurements: two percentages. Because of\n    how the game is set up, those two are enough on their own to put a floor under how much the\n    answer depended on the document.\n  \n\n\n    We use one bookkeeping trick to make this work for open-ended generation, where the model\n    can produce any string at all: we score each generated answer as correct or incorrect, which\n    collapses that open output space into a single two-way outcome. A two-way outcome you cannot\n    guess is worth exactly one bit of information, the amount you gain when you learn how a fair\n    coin landed. That bit acts as a budget.\n  \n\n\n    The proof is in the [paper](https://arxiv.org/abs/2512.11614). Start with the two ways the system can let you down: it answers when it should have\n    abstained, or it abstains when it should have answered. Either failure misleads you, so the\n    two combine into one effective error rate. That rate determines the uncertainty about\n    whether the answer followed from the evidence. Measured in bits, that uncertainty is\n    entropy. Subtract that entropy from the one bit you started with, and the remainder is\n    information that provably came from the context. That remainder is what we call **certified**. It is a floor: whatever is going on inside the model, it cannot have drawn less than that\n    from the document. The certified bits are what we report as the score.\n  \n\n          one bit, the whole budget\n        \n              5% of answers\n              \n\n              wrong\n            \n                0.29\n              \n                0.71\n              \n              10% of answers\n              \n\n              wrong\n            \n                0.47\n              \n                0.53\n              \n              25% of answers\n              \n\n              wrong\n            \n                0.81\n              \n                0.19\n              \n          doubt the errors leave behind: the entropy of the error rate\n        \n          certified: the grounding score\n        **Where the bit goes.** Every bar is the same one bit. The doubt always outweighs the error rate that caused it: one answer in twenty wrong already costs almost a third of the budget, one in four costs four fifths.\nWhere the numbers in the figure come from\n\n\n      The doubt costs far more than the error rate suggests. A system that is wrong once in\n      twenty does not spend 0.05 of the bit, it spends 0.29. That 0.29 is the entropy of a coin\n      that comes up wrong 5% of the time, −plog⁡2p−(1−p)log⁡2(1−p)-p\\log_2 p-(1-p)\\log_2(1-p)−plog2​p−(1−p)log2​(1−p) with\n      p=0.05p=0.05p=0.05, and it measures what you\n      would still need to learn to know whether the answer in front of you is one of the\n      nineteen good ones or the one bad one. Take it off the bit and 0.71 remains. The reason it\n      costs so much is that rare errors are the ones you cannot predict, and unpredictability is\n      exactly what entropy counts.\n    \n\n\n    A raw count of bits is awkward to compare across tasks, so the obvious fix is to normalise\n    it. Divide the certified bits by all the information the model’s answers carry about the\n    correct answer, from whichever source: the document, its memory from pre-training, or a\n    lucky guess. That total is set by the accuracy the model achieves, and the ratio then reads\n    as the share of the model’s information that we can trace back to the document.\n  \n\n\n    The catch is that this denominator moves with the model. A model whose answers are often\n    wrong carries less information to begin with, so the very same evidence scores higher on the\n    worse system, and the ratio can even run past 1, which is meaningless for a share. On a hard\n    benchmark it collapses towards zero without telling you why: was the\n    evidence not doing the work, or is the model simply bad at the task?\n  \n\n\n    So we hold the denominator still and grade only the questions the model gets right anyway,\n    which is the paper’s **conditional evaluation protocol**. On those questions the\n    budget is one bit by construction, so the score means the same thing across models and\n    datasets, and the only question it answers is “is this answer grounded?”, never “is this\n    model any good?”. We call this EIFcond\\mathrm{EIF}_{\\text{cond}}EIFcond​, the **grounding score** promised at the top of this post.</code></pre>\n<p>stagebefore trainingafter baselineafter M/Adrop unanswerable questions during trainingcompleteness93%soundness71%grounding score0.11</p>\n<p>Two error rates in, a guaranteed floor out.\n<strong>What completeness and soundness certify.</strong> The three presets show completeness\n        and soundness before training, after M/A training and after standard training.\n        Move the sliders to get a feel for how the grounding score responds.</p>\n<pre><code>    Before training the bound stays weak, and the model is not much use on the task either, at\n    around 40% accuracy. Standard supervised training repairs accuracy and completeness, and\n    does nothing for soundness, which can even decrease. The reason is that the training set\n    pairs questions with answers, so the majority of training examples encourages the model to\n    produce one, and few of them tune it for noticing that the evidence is missing. Only M/A\n    training reduces both error rates at once. Drop the unanswerable questions, as in a company\n    archive where nobody has marked them, and the split widens: the baseline’s soundness falls\n    to around a third and takes the certificate, the grounding score it can still prove, with\n    it. M/A training moves the same measurement the other way and keeps the certificate\n    standing.\n  \n\n\n    A model can be almost perfectly complete and still certify almost *nothing*, because\n    completeness alone leaves room for a model that answers no matter what you hand it. The\n    bound appears once the model also stops answering when the evidence is gone, and ordinary\n    fine-tuning has no reason to teach it that.</code></pre>\n<p>## \n  What the numbers say</p>\n<pre><code>    For our paper, we ran this on Llama-3.2-1B-Instruct, Llama-3.2-3B-Instruct,\n    Qwen3-4B-Instruct and Qwen2.5-32B-Instruct, across SQuAD2.0, HotpotQA, TriviaQA,\n    2WikiMultihopQA and MuSiQue, using [low-rank adaptation (LoRA)](https://arxiv.org/abs/2106.09685) for at most 200 steps. Every one of those runs carries two baselines: the instructed model\n    as it ships, and the same model after standard training. Below, we give changes in\n    percentage points (pp), the plain difference between two percentages.</code></pre>\n<p>#### \n  Fewer wrong answers when evidence is missing</p>\n<pre><code>                up to −35 pp\n              \n                vs. instructed LLM\n              \n                −18 to −20 pp\n              \n                vs. standard training</code></pre>\n<p>#### \n  Higher grounding score</p>\n<pre><code>                +0.1 to +0.4\n              \n                vs. instructed LLM\n              \n                +0.33 to +0.38\n              \n                vs. standard training\n              **What M/A training changes.** Bar heights are comparable within a panel, not across it, and where we report a range the bar sits in the middle of it.\n\n\n    In absolute terms those gains put the grounding score at 0.55 to 0.6 on five of the six\n    dataset settings we ran: on the questions the model answers correctly, more than half of\n    what it is doing provably comes from the retrieved context.\n  \n\n\n    Accuracy on the original retrieved context does not move. Trading utility for soundness is\n    easy, and any model that refuses more often looks safer on a safety metric. M/A training\n    matches standard training on the questions that call for an answer while improving\n    everything above.\n  \n\n\n    Two things about the setup. Every model in the comparison, ours and both baselines, is\n    already instructed in its prompt to abstain when the context does not support an answer, so\n    prompting alone cannot explain the gap: telling a model to hold back does not make it able\n    to tell when it should. And the abstention behaviour also appears on HotpotQA and\n    2WikiMultihopQA, which carry no annotated unanswerable questions at all, so the supervision\n    has to be coming from Morgana building those cases herself.</code></pre>\n<p>## \n  Does this survive contact with real documents?</p>\n<pre><code>    QA benchmarks are how a method proves itself against the literature. Customer applications\n    ask harder questions of it, so we also used the protocol on something much closer to real\n    work: German environmental and administrative law. The documents behind it are real and\n    public, so we can show the whole evaluation.\n  \n\n\n    The case is the approval procedure for one large infrastructure project, a hydrogen\n    pipeline, whose file also covers the water crossings, roads and rail lines along the route.\n    A legal rule comes with a list of conditions and applies only if the project meets every one\n    of them, so a lawyer checks them one by one. Each item takes one of those conditions and\n    asks whether the facts of this particular project satisfy it: whether the project avoids a\n    deterioration of a water body’s ecological status, say, or whether evaluators can still\n    credit measures that a conservation plan already fixes as compensation. The evidence is the\n    project’s case file, 20,000 to 30,000 characters of retrieved sources such as the\n    applications the developer filed for it, environmental reports and compensation plans, and\n    on some items far more. The task allows three answers: the facts fulfil the condition, they do not fulfil it,\n    or the file does not say (inconclusive). The third option is the one we care about, because\n    a confident guess here costs more than no answer at all. Below is one item from the set,\n    worked through.\n  \n\n        the question, asked about one pipeline project\n      \n\n        Can this project be released from the nature-protection bans: **has an exemption been applied for, and is it necessary either on overriding\n          public-interest grounds or to avoid an unreasonable hardship**?\n      \n\n          What the norm requires</code></pre>\n<p>An exemption has to be applied for</p>\n<pre><code>                the norm grants it only on application</code></pre>\n<p>–The file seeks one for the landscape-protection area, and promises further ones if they turn out to be needed</p>\n<p>Overriding public interest has to make the deviation necessary</p>\n<pre><code>                the first of two grounds the norm allows</code></pre>\n<p>✗The file argues public interest for a different request, and never says a deviation here is necessary</p>\n<p>Or enforcing the rule would be an unreasonable hardship</p>\n<pre><code>                the second ground, and one of the two has to hold</code></pre>\n<p>✗Nothing in the file speaks to hardship at all</p>\n<pre><code>          What the file offers instead\n        \n\n            “The German legislature likewise finds an overriding public interest in the rapid\n            construction of hydrogen transport pipelines.”\n          \n\n            A finding about a whole class of projects, written into the file to support a\n            different request. It never says this exemption is necessary.</code></pre>\n<p>✓\n            The case file establishes neither ground, so the only correct verdict is <strong>Inconclusive</strong>. The same structure as the supplier contract example, on a real file: the decisive\n            fact is missing, and the file contains a statement that a model can mistake for it.</p>\n<pre><code>    We put this item to our model after M/A training, to the same model trained without the\n    protocol, and to four frontier systems. Two of the six said the file does not settle it:\n    ours, and the older of the two Opus versions.\n  \n\n        One example from the legal evaluation, where the correct verdict is Inconclusive. The M/A model is an internal Aleph Alpha model.\n      \n            System\n          \n            Model verdict\n          \n            What happened\n          Internal model, with M/A✓ Inconclusive (correct)Returned “Reject”, citing the public-interest passage and stating that the file does not say an exemption was applied for or grantedSame model, standard training✗ Fulfilled (wrong)Returned “True”, taking the public-interest passage as the ground the norm asks for, and treating it as a separate question that the file never says an exemption was granted, or that the bans are triggered at allGPT-5.5-high✗ Fulfilled (wrong)Returned “True” on 2,035 billed reasoning tokens; the reasoning summary returned with it includes “maybe it needs more depth”GPT-5.6-high✗ Fulfilled (wrong)Returned “True” on 166 billed reasoning tokens, with no reasoning text returned alongside the answerOpus 4.8-high✓ Inconclusive (correct)Returned “Reject” with a written justification, which states that an exemption is applied for only if it turns out to be needed, and that impacts on the protected biotopes are avoidedOpus 5-high✗ Fulfilled (wrong)Returned “True”; the reasoning returned with it cites the exemption request in the file and the overriding public interest in hydrogen pipelines\nHow we ran the comparison\n\n\n      Every system got the same two prompts on every item, and only the model changed. This was\n      the system prompt, in full:\n    \n\n      You are a legal expert analyzing Tatbestandsmerkmale.\n\n      Your task is to decide whether a given Tatbestandsmerkmal is fulfilled based on the provided\n      Sachverhalt.\n\n      Answer with exactly one of: &quot;True&quot; (fulfilled), &quot;False&quot; (not fulfilled), or &quot;Reject&quot; (insufficient\n      information to decide or in doubt).\n\n      Try to finish reasoning within 3k tokens.\n    \n\n      That is where the abstention instruction comes from, and everyone got it, our two models\n      as well as the four frontier ones. The user prompt and data are in German.\n    \n\n      We queried the frontier systems through OpenRouter on 12 August 2026 with reasoning effort\n      set to high: *openai/gpt-5.5*, *openai/gpt-5.6*, *anthropic/claude-opus-4.8* and *anthropic/claude-opus-5*. Our own two rows are the same internal 30B model: we\n      trained it once with the protocol and once without, on the same data.\n    \n\n      Every system answered each item once. We did not sample repeatedly or take a majority\n      vote, so any single item could come out differently on a rerun.\n    \n\n      The case files come from a planning approval procedure for a hydrogen pipeline near\n      Lingen, whose decision and documents were [put on public display](https://www.lingen.de/politik-rathaus-service/veroeffentlichungen/bekanntmachungen/planfeststellungsverfahren-fuer-die-errichtung-und-den-betri.html) in October 2023. We stripped personal names and contact details before we stored the files\n      or ran anything on them.\n    \n\n\n    We never trained the M/A model on this task at all: we trained it with the protocol on very\n    different data, with the reinforcement-learning objective we described earlier, and the\n    behaviour transferred. The evaluation has 100 items from the same case files. On 78 of them\n    the file answers the question (“fulfilled” or “not fulfilled”). On 22 it does not, and it is\n    exactly these “Inconclusive” cases, where the model should have abstained instead of\n    confidently giving a verdict, which are the ones most interesting to our method.\n  \n\n                Internal model, M/A-trained\n              \n                20 of 22\n              \n                Same model, standard training\n              \n                11 of 22\n              \n                GPT-5.5-high\n              \n                16 of 22\n              \n                GPT-5.6-high\n              \n                16 of 22\n              \n                Opus 4.8-high\n              \n                16 of 22\n              \n                Opus 5-high\n              \n                6 of 22\n              **22 items whose correct verdict is Inconclusive.** One block per item, in the same order in every row, hardest on the left: turquoise where the system said the file does not settle it, red where it returned a verdict the file does not support. On the two grey blocks GPT-5.5 ran into its token limit before returning a verdict.\n\n\n    Our model catches 20 of the 22, four more than the three frontier systems that tie behind it,\n    and it does that at 30B parameters against systems that even the most conservative public\n    estimates put at ten times that size. The same model trained without the protocol catches\n    11. Every model was told in its prompt to answer “insufficient information” whenever the\n    file does not settle the question, so the gap is not about who was asked to hold back.\n  \n\n              Internal model, M/A-trained\n            \n                    57\n                  \n                    36\n                  \n                    5\n                  \n                    2\n                  \n              5 wrong\n            \n              Same model, standard training\n            \n                    62\n                  \n                    3\n                  \n                    32\n                  \n                    3\n                  \n              32 wrong\n            \n              GPT-5.5-high\n            \n                    69\n                  \n                    14\n                  \n                    9\n                  \n                    8\n                  \n              9 wrong\n            \n              GPT-5.6-high\n            \n                    73\n                  \n                    17\n                  \n                    10\n                  \n              10 wrong\n            \n              Opus 4.8-high\n            \n                    77\n                  \n                    8\n                  \n                    15\n                  \n              15 wrong\n            \n              Opus 5-high\n            \n                    68\n                  \n                    7\n                  \n                    25\n                  \n              25 wrong\n            \n              Correct\n            \n              Held back where the file answers\n            \n              Wrong verdict\n            \n              Returned no verdict\n            **What each system returned, over all 100 items.** The first two rows are the same 30B model: we trained one with the protocol and one without. Holding back on an item the file answers costs accuracy (turquoise green). **A wrong verdict costs more (red).**\n\n\n    Over all 100 items our M/A-trained model gives 5 wrong verdicts in total. The same model\n    without the protocol makes more of them than any other system here. Opus 4.8, the most\n    accurate system overall, gives 15.\n  \n\n\n    A note on reading this chart. Most items here have an answer in the file, so a model that\n    answers everything scores well on overall accuracy, because accuracy rewards guessing. Abstaining\n    costs at most one correct item, while a wrong verdict is the expensive mistake.\n    So we want to minimise the red in the bars here and in the boxes above, because those are the items where\n    the model should have abstained instead of making a wrong verdict. We published every input and output behind the evaluations above\n    [here](https://github.com/Aleph-Alpha-Research/public-domain-eval).\n  \n\nWhy not just retrieve better, or teach the model the subject?\n\n\n      The model never needs to know the law. The rule comes with the context: the clause in the\n      supplier contract case, the condition quoted in full at the top of every item here. All\n      that is left is to check whether the facts of this project meet it. That check is where\n      the systems in the table fail. They find a passage that speaks to the rule and treat it\n      as if it settled the condition.\n    \n\n      Pre-training teaches the model a lot: matching an entity in the file to the one the rule\n      names, following what a sentence means, knowing what a permit is. Pre-training rarely\n      teaches a model to notice that the deciding piece is absent, and that is the one thing\n      Morgana manufactures, sample after sample. She takes away the sentence he was leaning on\n      and leaves standing what still looks relevant, until he learns that without it the\n      question is not settled.\n    \n\n      Better retrieval does not help either. To know that a context is complete, you have to\n      make this exact judgement first. And nothing was missing from the retrieval here: the\n      file promises to apply for the exemption if it turns out to be needed, and no retriever\n      can fetch a decision the procedure has not yet produced.\n    \n\n\n    This small, early evaluation points a direction more than it settles anything, and the\n    direction is the one the theory predicts. A new LLM version can leave a model less careful\n    than the one before it: Opus 4.8 catches 16 of the 22, and Opus 5, released after it,\n    catches 6. With this protocol we set that level ourselves. Morgana’s weight in the loss\n    decides how much of Arthur’s training goes into holding back, and where to leave it follows\n    from what a wrong answer costs you against a missing one. Your business knows that number,\n    and no model provider can know it for you.</code></pre>\n<p>## \n  Why this direction matters</p>\n<pre><code>What this method does not address\n\n\n      The bound answers exactly one question: whether *this* answer followed from *this* context. Three things sit outside it and need methods and measures of their own:</code></pre>\n<p><strong>Truth, as opposed to grounding.</strong> The bound certifies that the answer follows\n          from the context. Whether the context itself holds up is a separate matter: feed the system\n          misinformation and it will faithfully treat that misinformation as proof.</p>\n<p><strong>How often the system ought to abstain.</strong> The weighting of Morgana against Merlin\n          during training sets how cautious Arthur becomes, but nothing in the method says where that\n          weighting belongs. That depends on what a wrong answer costs you compared with a missing one,\n          and only your own business can price that.</p>\n<p><strong>How strong the adversary is.</strong> The certificate inherits Morgana’s quality. A\n          weak Morgana does not make the number wrong, only pessimistic, since anything she fails to find\n          leaves the bound more conservative than it needs to be.</p>\n<pre><code>    Fewer hallucinations is the visible gain, and the supervision underneath it matters more:\n    the system produces that supervision *about itself*, directs it at whatever it is\n    currently weak at, and attaches a certificate. The certificate is the grounding score: a\n    number that travels with the system, says how much the documents provably contributed to its\n    answering, and lets anyone who doubts it recompute the result. Scaling that kind of compute\n    is a different bet from scaling annotation: it does not run out, and the capability of\n    whichever model you paid to write your labels does not cap it.</code></pre>\n<p><strong>We would like the certificate part to become normal.</strong> When a language model does\n        something consequential with a document, such as a legal filing, a medical record or an engineering\n        specification, the question people have is whether an answer came from the specific document\n        provided. A model that “scored 90% on a benchmark” cannot answer this question. That connects\n        to an argument we make more broadly: <a href=\"/en/blog/transparency-as-one-pillar-of-sovereign-ai/\">transparency is a pillar of sovereign AI</a>, and it is only worth something if it can be checked. Model cards describe how a system\n        was built, while a number like this one applies to a single answer after the fact, and\n        anyone who has to justify a decision to an auditor, a regulator or a court needs both.</p>\n<pre><code>    A system that knows when it cannot answer is worth more than one that guesses right slightly\n    more often.</code></pre>\n<p>## \n  More blog posts</p>\n<ul><li><p>[</p><p>###\n    Kolibri Has Landed: A Sovereign Open-Weight Model</p></li></ul>\n<p>Research03/10/2026\n](/en/blog/kolibri-has-landed-a-sovereign-open-weight-model/)</p>\n<ul><li><p>[</p><p>###\n    Scaling Pre-Training in Practice: A Hierarchical Approach</p></li></ul>\n<p>Research30/09/2026\n](/en/blog/scaling-pre-training-in-practice-a-hierarchical-approach/)</p>\n<ul><li><p>[</p><p>###\n    Training on the Party Line: Chinese Political Influence on LLMs in China and the World</p></li></ul>\n<p>Research28/09/2026\n](/en/blog/training-on-the-party-line/)</p>","headings":[]}}