Daniel Litt has recently written about the near future of mathematics. I think it’s a great essay and you should go read it now. However, despite making very good points, I believe it falls into traps that plague much of the writing on this topic by prominent figures in math academia. Rather than attacking any particular point being made I will argue that the framing is misleading in a sense that will hopefully be clarified below.
The essay begins by addressing the disagreement within the academic community as to the goal of mathematics. The suggested goal is recursively defined as:
Producing and understanding high quality mathematics.
Producing high quality mathematicians.
It then argues that, under the assumption of robustly superhuman AI capabilities, the institutional structures through which these goals have been pursued are becoming ill-suited for the purpose and proposes some reorientations to address this misalignment. The resulting conclusions are highly agreeable but the reader is left with very few operationally meaningful judgements on the genuinely contentious issues. This is a general pattern with many essays on the topic. By now I have conceded that if I want this topic treated properly I have to write about it myself. This is my attempt to do so.
What is mathematics?
Irrespective of the views of any particular mathematician, however decorated, as to the output of mathematics, the input to mathematics is, and has always been, discourse. By mathematics I mean pure mathematics throughout: the study of mathematics for the purpose of better understanding mathematics itself, which may or may not unlock previously inaccessible domains for applied mathematics. Applied mathematics, broadly construed, is the development of mathematics for application to other domains of science, and it borrows its standards from those domains: a method is good if it works on the problem it was built for. While definitions in mathematics can be, and often are, motivated by applications to other domains, mathematics has no external domain from which to borrow its standards of quality. Indeed, unlike in other sciences, the objects of study in mathematics are entirely contingent on past work. The very definitions of the objects studied in mathematics were put forth by past mathematicians who found them useful for studying questions about objects defined by a previous generation, and so on recursively.
Accordingly, the standards of quality and impact are downstream of a wide and vibrant community of mathematicians, with correctness merely serving as a filter on mathematical work. In this sense the output of mathematics too is not simply theorems or proofs but, more generally, mathematical discourse. There is no proprietary mathematics in roughly the same sense as there is no proprietary philosophy; a large collection of correct mathematical propositions locked inside a company is bound to be forgotten if no community of mathematicians can discuss, reinterpret and breathe life into them. Math academia relies on the same discourse to decide who are, in Litt’s term, “high quality mathematicians”. A novel result can be discovered independently, copied from a draft or obtained in a private conversation, and assigning credit requires knowing which happened.
A word about credit. It is fundamentally cringe to talk about credit. Vying for credit in the open is undignified. As Tywin Lannister wisely notes: “Any man who must say ‘I am the king’ is no true king”. But is this really a useful lesson for us? Do we want credit to be assigned only by absolute coercive political force? Of course not. We would like to live in a high trust society in which credit is judged by individuals naturally and intelligently according to merit and is never appropriated in bad faith. Unfortunately we live in a Malthusian age, in which the population of claimants and the intensity of their pursuit of credit has outgrown its supply. The resulting surplus then appears to us as grift. But I digress.
Mathematicians, relative to the difficulty of their occupation and the opportunity cost of pursuing it, are paid very modestly. A central component of their compensation is therefore the joy of making new beautiful discoveries, sharing them with a community who can appreciate them and receiving recognition for it. This is not just a utilitarian requirement for allocating grants. The desire to be recognized for one’s creative output is deeply human, and arguably a major driving force of human achievement across cultures. Without it our heritage would be poor and lifeless. Credit is in other words part of the pay, and in mathematics it has so far been paid out through the same discourse that sets the standards.
Mathematics has held up better in the Malthusian age than most fields: correctness filters out the false results, and small communities usually know where a correct one came from. Whether credit continues to disperse according to merit or is settled by politics is therefore dependent on whether one can still tell where a result came from. That is a question about what AI does to this discourse, the input to mathematics, and it comes before any particular aspiration for the terminal goal of mathematics. To answer it we must agree on some scoped definition for what AI is. This is perhaps the first trap in discussions of the topic. Namely, the failure to disambiguate two importantly distinct aspects of AI:
What AI does
What AI is
Let’s begin with the first of these.
What AI does
Studying AI systems from outside the walls of the labs developing them is a filthy, unbecoming, and fundamentally cucked activity. The labs control everything from the architecture of the model, to the training data, algorithms and all the way to the context management and tool calls during inference which might contain any number of undisclosed references or heavy computations behind the scenes. Much of this information is central for properly evaluating and interpreting the results achieved by these systems. Unfortunately this means we are informed on AI capabilities mostly through the following information channels:
Self reports by labs with untold amounts of resources employing talented mathematicians with unrestricted access to frontier internal models and tooling deployed at enormous cost to crack prestigious problems for publicity purposes.
Benchmark results whose relevance to practical math is questionable, which are vulnerable to train-on-test syndrome and whose maintaining orgs often have questionable incentives downstream of funding.
A horde of button pushers who use AI as a black-box solver, motivated either by applications to other domains or by sheer clout chasing.
A smaller collection of mathematicians who use the models, mostly silently, for their own learning and research purposes.
What follows is from the fourth channel and based on my anecdotal experience hence should be taken with a grain of salt.
It is not easy, nor wise, to place the capabilities of AI systems on any one-dimensional axis. Only mildly less foolish would be to model capabilities on a two-dimensional plane. This is what we will do. Specifically we can conceptualize AI as alternating between two tasks:
Semantically searching the literature, whether explicitly by fetching material from journals or arXiv or implicitly by retrieving information stored in the weights.
Stochastically proof-searching, with literature-informed heuristics stored in the weights, by composing techniques from 1 in a feedback loop against an anchoring oracle, either numerical (e.g. NumPy, SciPy), symbolic (Sage, Lean etc.) or another judge LLM (possibly even the very same model checking its own steps).
I think there is a strong argument to be made that results relying on external verifiers have scaled significantly further than those that do not. Even superficially, at the level of the domains with the most AI results, both in quality and in volume, one can see a theme in the simplicity of verification rather than in the simplicity of the problems themselves. The most impressive problems solved seem to be accessible either through a shallow combination of already existing techniques or through an inscrutable, highly technical execution of existing techniques, often with heavy numerical and/or symbolic feedback. We have yet to see much evidence, if any, of novel abstractions developed as intermediate steps to solve a particularly difficult problem with no human guidance. (Anecdotally, my intuition here relies on trying problems from my own field of specialty which I have solutions for but will not reveal as they are largely unpublished.)
My current guess is that the human ability to synthesize higher level abstractions can emerge with scale but that under the current paradigm the improvement will be logarithmic rather than exponential. I think one can learn something here by analogy from certain aspects of code generation, such as choosing good architectural abstractions or properties in software, which have failed to improve much in the last couple of years. I suspect these capabilities are importantly bottlenecked by pretraining (architecture, size, etc) with RL being much more effective at boosting the search capabilities with abstractions held fixed. If things continue this way I expect impressive symbolic proofs which are increasingly incomprehensible to both humans and AI, with improvements in exposition failing to keep up with the complexity of the proofs being generated.
I do not believe we can comfortably assume models will soon take over proof generation entirely. My intuition is less tied to whether the models can come up with a new idea and more to how long the proof is and how narrow the keyhole; disproofs are often a narrower search target than proofs. A proof of Hodge, BSD, or RH that does not rely on soundness bugs, would be a significant update for me in believing the models can scale from here to autonomous solvers of all of our hardest problems.
I will consider the case where the models become robustly superhuman provers while their ability to form higher level abstractions improves only moderately, roughly in step with their exposition. I do so for two reasons. First, because that is what I believe will happen for a while and second, since the alternative would be the assumption that AI models can take on all parts of the profession in which case I believe the problems become entirely political.
Human mathematics could exist in such a world as a perfectly healthy hermeneutic discourse on top of an empirical background consisting of a growing formal math library. The library is then to human mathematics roughly what experimental data is to physics. Math in this world is split between two frontiers, an empirical frontier with more of an engineering flavor and with credit still given for proving impressive things, and a theoretical frontier which looks a bit more like humanities in its nature. The abstractions of the theoretical frontier are then fed back into the library to compress and refactor the existing proofs into nicer conceptual ones while the empirical frontier advances far ahead with slop.
For this type of mathematical discourse to exist there must still be some way of assigning credit for contributions to it (more on why later), that is, of knowing where they came from, which brings us to the second aspect of AI.
What AI is
For our purpose an AI service can be treated as an input/output box. A mathematician puts information in and receives information out; the relevant question is whether information supplied by one mathematician can affect what another later receives. Without getting into an analysis of the TOS of the leading US labs I will take the following as a postulate:
AI labs periodically distill user data to improve their models on observed tasks.
Evidence of this can be found in a couple tweets by (past and current) high ranking officials at OAI.
Whether this happens by training on chat logs, their distilled summaries, or synthetically constructed corpora from iterating on derived rubrics is unimportant. What matters is only that the box has memory across users, so that inputs of one user affect the outputs seen by another user with no record of the information exchange. An immediate consequence is that AI labs are, whether willing or not, participants in mathematical discourse. The models are trained on the existing discourse, mathematicians produce new discourse in conversation with the models, the next generation of models is distilled from that, and so on recursively in exactly the manner described in the first section, except that one of the participants reproduces what it has been told without any record of who told it. The easiest way to see the problem this creates is through the following hypothetical scenario.
Imagine a group of researchers (group A) working tightly with AI on some mathematical problem. After many months of laboring on it the lab serving their models distills some of their interactions in a way that improves the model in the relevant domain. Another research group (group B) using the updated model subsequently makes a breakthrough precisely on the problem studied by group A and using substantially similar methods. Group B has no indication that the relevant outputs originated in the work of group A and publishes before group A manages to complete their work. Group B has, unbeknownst to them, committed what would once have been considered academic fraud.
The scenario is not entirely hypothetical and in fact, depending on where one’s trust falls in the Alpöge and Buckmaster versus OAI dispute discussed below, has arguably occurred already. Note that under Litt’s assumption of “robustly superhuman capabilities” this applies to expository work just as well as to raw results.
The scenario shows that under the postulate the entirely ordinary operation of current AI companies is fundamentally incompatible with any faithful system of credit assignment in the open academic sciences. When intellectual work cannot be traced back to its origins the entire hierarchy of modern academia collapses, and what remains is organized by connections, inherited status and politics. The only distinguishing feature of mathematics here is that, being fundamentally a discourse, it can only exist in an open academic tradition, whereas the other sciences can at least in principle retreat into proprietary labs. Mathematics with no public exchange of ideas is definitionally impossible. As such, under the postulate, the continued operation of AI companies is fundamentally incompatible with any objectively meritocratic conceptualization of academic mathematics.
The arguments above do not depend on the conduct of any particular company. There is however a separate issue with the conduct of the companies: they are now themselves producing mathematics, using internal models and resources unavailable outside the walls of the labs. Access to these tools carries an implicit responsibility which we have unfortunately not seen taken seriously by either lab; both still seem to prioritize short term PR over academic norms and conventions.
Credit is one half of authorship, and we dealt with it above; responsibility is the other half, and what it consists of has been roughly the same convention across math academia for at least a century, possibly more. A journal consists of editors and referees who evaluate mathematical work for publication but neither are responsible for the correctness of the results they publish, nor for supplying additional details when an omission is later found to warrant them. These functions belong to the author, whose job is to vouch for the correctness of the results and to bear the responsibility for correcting any mistakes and filling any omissions. A named author which cannot supply these functions, be it a human, a corporation, a robot or an alien for that matter, is definitionally incapable of being an author, irrespective of any ethical controversy over whether it should be. A model cannot vouch for anything and cannot be held to fill a gap next month, so a proof credited to e.g. Claude has no author unless whoever posts it takes the function on, and taking it on means being able to say why one believes the proof. Many of the incidents of the past few months are clarified by this alone.
The Navier--Stokes dispute has already received a great deal of attention; I prefer to focus instead on an anecdotal example closer to my own experience, namely the complex six sphere. On Aug 24, 2026 Levent posted a thread containing a brief two page description of a construction of a holomorphic six sphere, credited to Claude (model unspecified), together with 100+ pages of essentially unreadable slop presented as proof and credited to Opus.
Fortunately the two page writeup was sufficient to specify the construction uniquely, so I spent a few days trying to understand it. When I finally got a good idea of why it works1 I decided to skim the Opus writeup again, at which point I became highly suspicious that the entire document was in fact a deformalization of a Lean proof Levent did not disclose. I asked him publicly on X but he did not respond.
Later, in a different context, he was complaining that the anonymous accounts attacking lab mathematicians have implicitly become the representatives of a mathematical community which has largely migrated off X, and invoked the six sphere:
@alpoge(Levent Alpöge): …i somehow don’t think active number theorists are going to angirly reply to me ten times every day over posting a correct argument for S^6
This time I confronted him more aggressively, inquiring into which evidence led him to declare correctness with such confidence: he did not write the argument, he probably did not read it either, so how did he know?
This is a question to ask of AI generated mathematics in general, namely what exactly made the person presenting the result believe it. I think honesty is a central pillar of mathematics and as such a proof should be presented as closely as possible to the actual reasons its authors believe it. If the reason is that mathematicians have read and understood the TeX then the authors should say so, and if it is a Lean formalization then the Lean formalization is the proof being relied on and the TeX merely an exposition of it. The same goes for some internal harness or collection of models, which if constitute part of the evidence for correctness are then necessarily part of the proof. The question also applies to the labs: what was the reason in their case?
I believe that both OAI and Anthropic run proof-search harnesses (which they also use for RL training) in which TeX is interleaved with Lean so that agent swarms can coordinate and iterate given a prompt from the user, and that the AI generated manuscripts we have seen are this interleaved TeX and Lean written up and cleaned after the fact. This is based partly on having used the models in exactly this way myself (at a proprietary capacity for my employer, not in any academic capacity) and partly on how the manuscripts look. I do not claim to have evidence sufficient to indict any particular lab, and nothing in the TeX would distinguish such a process from either of the ones the labs disclose in any case.
However, if any of these manuscripts was indeed produced as a deformalization without proper disclosure we should view this as a form of proof laundering. Mathematics is critically about methods. Obscuring how a result was obtained removes important information. (It moreover hinders our ability to assess capabilities but that is not a mathematical concern.) A Lean proof found by trial-and-error search which is then deformalized into a 100+ manuscript of highly unreadable informal proof is evidence for an entirely different bundle of capabilities than producing an informal proof directly which would be closer to the process of human mathematicians, most of whom are not even familiar with Lean’s syntax let alone its semantics.
Let’s take the example of the complex six sphere. It turns out that, upon a careful spelling out of the argument, the same construction using the same representation but pulled back to a family of groups gives infinitely many non-deformation-equivalent holomorphic six spheres of algebraic dimension one.2 Any human who stumbled by this construction would almost surely have recognized this. I am accordingly suspicious that the construction wasn’t understood by anyone involved in producing these documents, and the actual reason for believing the manuscript was correct was a hidden Lean proof.
The responsibilities of an author when it comes to correctness apply recursively: an author is responsible for the correctness of their results, not their arguments, and is therefore further implicitly responsible for the correctness of any results on which they depend. This includes results cited from works by other mathematicians, then further on the results on which those in turn depend, and so on recursively. In particular, if the author is relying on a formal object as the proof, then the correctness of the checker used to validate is then itself part of what the author is vouching for. I’m not worried about unfixable soundness bugs. What I care about is preserving a standard of mathematical proof dating back at least to Hilbert, where a proof should include a complete account of the dependencies it rests upon. If Lean code is to replace traditional proofs the author must account for the metamathematical step that makes such a replacement legitimate. Finally, there is the separate issue of resource disparities.
According to OAI, they began working on the Navier--Stokes millennium problem after hearing a rumor that Alpöge and Buckmaster were cracking forced Euler using an internal Anthropic model. Using an internal model and double digit millions of dollars’ worth of compute, an amount no academic mathematician could ever match, they found a counterexample which they published soon thereafter. Alpöge and Buckmaster later claimed that OAI stole ideas from Buckmaster’s codex chat logs.
Irrespective of which party you believe in the story, it provides a particularly extreme demonstration of the hypothetical scenario described earlier. Indeed, anyone with access to sufficiently capable models can take any mathematical idea developed in the open, throw tokens at it until it yields, then sweep the reward in prestige. It need not even be the labs who engage in this activity to greatly damage the social fabric of the mathematical community.
The labs are therefore simultaneously competing for mathematical results, selling mathematical assistance, and controlling most of the information required to evaluate either. This adversarial situation, where information about proofs and methods by which proofs were obtained are withheld whether deliberately or out of neglect, is one the mathematical community has never had to address. Most senior mathematicians are likely not sufficiently familiar with the technology to even know which questions to ask. As such, if we wish to reason seriously about the institutional effects of AI on mathematics we cannot treat labs as merely neutral suppliers of an abstract service.
What can be done?
The natural conclusion of all this is that credit for mathematical work is dead. But that is also a lazy conclusion. Indeed, it does not immediately follow merely from the coexistence of math and AI. A fixed model can search the literature, proof-search in Lean and even write TeX, without any of the user’s inputs ever making it into any future model subsequently served to anyone else. Prohibiting such reuse may well slow down the progress in model development by removing valuable data from the training set but that is not a problem for mathematics.
There is in other words no technical contradiction between extremely powerful AI and mathematics in which provenance remains faithful. The technical fix already exists. In the trustless limit one can run open-weight models with properly documented training histories locally. Unfortunately that is financially unfeasible as of today. However, allowing a bit of trust one can consider services offering credible ZDR commitments. Unfortunately, even this may not be enough if people choose to trade their ZDR premium for extra tokens. As for opt-out buttons, I will simply assert here that they are operationally meaningless given a careful reading of the TOS.
The provenance problem can therefore only be solved by setting some academic standard, possibly in addition to a regulatory one. Note that this is a much narrower problem than solving copyright on the internet in general: the only new issue created by AI is the transmission of seemingly private communication. The problem of omitting citations to unofficial drafts or preprints found on the internet predates AI and is handled about as well as it ever will be by prestige and reputation costs. The relevant regulation is accordingly a prohibition on using private consumer data to improve models, extended to the data markets through which such data would otherwise be traded, and enforcing it requires:
A technical guarantee that the prohibition is respected.
Enough transparency to enforce it against the lab.
Academic rules against using services which do not provide it.
Enforcement is a much smaller problem here than for copyright in general. Math and academic science in general consists of mostly small communities with very strong reputation costs for mishaps, so once the service used and its guarantees are themselves observable, reputation can do essentially all of the remaining work. I do not think such regulation is likely to be passed, let alone practically enforced, but I do think it is possible in much the same sense as a robustly superhuman AI mathematician is possible, and the claim of possibility has to be made before evaluating plausibility.
This proposal doesn’t address the disparaging inequality in access to compute, and even still in this hypothetical world I expect a heavy price to be paid in the culture of how we do mathematics longer term. Openly sharing an unfinished idea gives anybody with adequate compute resources the opportunity to finish it first without violating any rules. Once proof production becomes cheap relative to idea formation there is a strong incentive to keep ideas private until the expensive part of the work is no longer scoopable. In the two-frontier picture this is the empirical frontier being scooped and the theoretical frontier emptying out: the standards of high quality mathematics are created by the community, and once easy scooping pushes the community into a dark forest of private work the standards go down with it.
There is another, easier to achieve, survival mode, in which mathematics becomes a humanities discipline outright and its hierarchy is organized openly by taste and connections. One may well prefer it, at the cost of whatever benefits objectivity confers on an academic institution. Personally I have always viewed mathematics as more an art than a science and so I find this fate entirely acceptable but I am also sympathetic to many mathematicians who find the scientific aspects of the field important and who would like to preserve them. I furthermore suspect that the required change is far more radical than what I expect academic institutions are capable of.
For all the reasons above I believe that under some narrow chain of circumstances in government policy, corporate leadership and academic institutional reform, academic mathematics can survive extremely powerful AI, even if a more secretive culture is impossible to avoid. I do not expect these circumstances to occur though. The capabilities themselves are not the issue. Between the credit ambiguity and resource asymmetries it would take a herculean stand of our leaders and institutions to hold the tide of short term corporate incentives and save the fragile values of our community. This leads me to believe academic mathematics is effectively dead.
1
See my thread of September 3, 2026.2
I do not claim this result and invite people to publish it if it’s correct or otherwise dunk on me here or on X.
