Superhuman AI could produce endless mathematics and still make the field worse
Notes from Daniel Litt's thread, based on a talk he gave at an OpenAI summit on the future of mathematics.
The thread starts from a strong assumption: AI becomes robustly superhuman at mathematics. Litt then asks a less familiar question. Could mathematical progress still stall?
His scenario is deliberately pessimistic, not a prediction. It imagines today's publication incentives surviving after proofs, papers, conjectures, and eventually whole theories become cheap to generate. The result is a strange failure mode: vastly more correct mathematics, but fewer people reading it, understanding it, discussing it, or learning how to create it themselves.
TL;DR
- Mathematical output is already rising while shared discussion spaces such as MathOverflow appear to be shrinking.
- AI systems often converge on the same visible problems, producing duplicated proofs whose marginal value may be little more than the cost of the tokens.
- If careers still reward papers and solved conjectures after those outputs become cheap, the rational strategy is to run a high-volume "conjecture slot machine."
- Models that can reconstruct a paper from a few key ideas make it dangerous to discuss unfinished work, weakening the open exchange on which mathematics depends.
- Formal verification may preserve correctness without preserving understanding. Humans can trust statements while losing contact with the ideas behind them.
- The long-term risk is not simply that AI replaces mathematicians. It is that mathematical institutions select for people who produce the most output while engaging least with the mathematics.
- The institutions need new incentives for human understanding, judgment, mentorship, trust, and genuinely useful science.
2026: output explodes, attention does not
Litt begins with an increase in the number of combinatorics papers posted to arXiv. Other areas show a similar, if less dramatic, rise. More mathematics is being produced, but it is unclear how much of it is interesting, correct, or meaningfully read.
At the same time, parts of the mathematical community are weakening. A chart in the thread shows a long decline in MathOverflow activity, with a sharper drop since early 2025. Some of that activity has moved to Discord and other private spaces, but Litt points out something harder to explain away: MathOverflow has fewer questions and fewer answers. He could not find a compensating increase in answers to older questions or another positive interpretation of the data.
The contrast matters. The field is getting more output and less visible conversation.
Even significant AI results are beginning to duplicate one another. Three groups produced similar proofs of Feige's 1/e conjecture at almost the same time; two said AI found the result. OpenAI and Anthropic models have independently replicated other recently announced results. Models and the people operating them tend to attack the same salient problems.
Once several systems can solve a problem, the next independent solution adds little. Its mathematical value may be only the compute spent producing it and a small amount of information confirming that current models can handle the task.
2027: the conjecture slot machine
Academic mathematics rewards papers, theorems, solved conjectures, citations, and priority. These are useful proxies when producing a good result requires expertise, sustained effort, and contact with the underlying ideas.
AI changes the price of the proxy.
If institutions keep rewarding papers after models can generate them cheaply, the dominant career strategy becomes what Litt calls playing the slot machine for conjectures. A researcher can ask an agent to choose problems, solve them, check the proofs, and turn the results into papers. Someone who cares about correctness might publish several short papers a day. Someone who does not can publish far more.
The person named as author may contribute little expertise and may not read much of the work. Nor will anyone else: human attention does not scale with machine output.
This is a Goodhart-style failure. The profession uses papers and proofs to encourage good science, human understanding, and the development of expertise. Once machines can maximize the measured outputs directly, those outputs stop reliably representing the values behind them.
Open work becomes dangerous
One of the most immediate risks is the loss of informal exchange.
Models are approaching the point where they can reconstruct a paper from a few important ideas. Some autonomous results already have a "last mile" quality: a model finishes a problem after deep recent work by humans.
In that environment, telling a colleague about unfinished work creates a priority risk. Even saying that a certain problem appears solvable may be enough for someone else to point an agent at it and publish first. Litt says several colleagues have already told him that they are unwilling to discuss work in progress for this reason.
Mathematics depends on more than final papers. Seminars, correspondence, half-formed conjectures, failed approaches, and conversations between researchers help ideas develop. A race for machine-assisted priority could push all of that into smaller trusted circles, or stop it altogether.
2028: correctness survives, understanding frays
There are clear benefits. Cheap autoformalization could expose gaps and errors in the literature and repair them. Formal proof systems can provide a reliable check when humans no longer have time to inspect every generated argument.
But formal correctness does not guarantee that readers understand what was proved.
Litt points to cases where a formalized statement differs from the English statement it is meant to represent. The difference may not be obvious to readers. Models will probably become good at checking whether formal definitions match their intended meaning, but that still leaves humans relying on a machine-mediated chain they do not personally understand.
The field may retain confidence in statements while losing contact with ideas. A theorem is verified, but few people know why its definitions are natural, how its proof connects to other work, or what conceptual compression it offers. Translating formal mathematics back into good human explanations helps, but it also consumes time and attention.
2029 and beyond: mathematicians as lab operators
Litt imagines models eventually performing the whole research cycle: building theories, proposing conjectures, proving or refuting them, and iterating. Human mathematicians become something like laboratory scientists who allocate agents and compute to questions they find interesting.
That arrangement might produce extraordinary discoveries. It also leaves several institutional problems unanswered:
- Who reads and evaluates the resulting work?
- How do new mathematicians develop expertise?
- How does the field preserve different styles of thought when many researchers use the same few models?
- Why should institutions keep funding mathematical agents if the output is barely understood or used?
- What happens when career incentives favor people who are good at generating papers rather than people who care about mathematics?
A profession can survive automation of a task. It has a harder time surviving the loss of the practices that reproduce the profession itself.
The risks in one list
Litt's summary slide names five:
- Mathematicians lose contact with the mathematics.
- The field loses diversity of thought because researchers depend on the same models and tools.
- Institutions fail to train young mathematicians, both in expertise and professional values.
- Trust and the social structure of mathematics deteriorate.
- Obsolete incentives change who enters and succeeds in the profession.
These failures would not come from AI alone. The current system already overvalues publication counts, priority, and prestige. Highly capable AI acts as a stress test: institutions crack first where their incentives were already weak.
What mathematics is actually trying to preserve
The useful distinction in Litt's argument is between institutional values and the mechanisms currently used to support them.
The values include good science, human understanding, expertise, trust, mentorship, and a healthy intellectual community. Papers, theorems, jobs, grants, and prestige are mechanisms. They are not the values themselves.
That framing connects directly to Grant Sanderson's argument about AI and mathematical progress. Sanderson separates proof from understanding and taste: a system may generate correct proofs without knowing which ideas are fertile or which explanation gives humans a better mental model. Litt adds the institutional consequence. If proof becomes abundant but the reward system still treats proof production as scarce, the field can optimize itself away from understanding.
His final questions are therefore practical:
- How should mathematical institutions change so they reward good science and the creation of human expertise?
- Which functions and values of the mathematical community remain robust when AI becomes highly capable?
Litt ends optimistically. Mathematics can survive and flourish, and humans may gain access to ideas that are currently unreachable. But adaptation means more than adding agents to the existing publication pipeline. The field has to decide what it wants to keep scarce, rewarded, and human.
The unsettling possibility is not a future with no mathematics. It is a future full of mathematics that nobody inhabits.