Research

Letitia Parcalabescu

        08/08/2026

# Bounding Hallucinations: Merlin-Arthur Protocols for Mutual-Information Bounds in Language Models

    In plain terms: how to check whether a language model actually used the document you gave
    it, and how to train one so that it does.
  


    Our preprint ([arXiv:2512.11614](https://arxiv.org/abs/2512.11614)) develops a way to do both. The figures on this page are interactive: some reconstruct the
    paper’s results, others illustrate the mechanism. See the paper for exact numbers.

## 9 right, 1 wrong answers and no way to tell them apart

    A system answers questions based on your documents and scores 90%: nine of its ten answers
    are right, one is wrong, and nothing marks which one. That is a capable system, and the 90%
    still buys little, because the bad answer hides among the nine good ones and a human has to
    check all ten to find it. Now picture the same 90% with the wrong one flagged: “I cannot
    answer this from the documents I was given.” The model knows no more than it did before, but
    now you can put it in front of a customer.
  

              A system at 90%