Mathematics and AI
Leslie Ann Goldberg, 28 Sept 2026 (this page gets updated often)
This web page is about the impact of AI-generated proofs on pure mathematics (including the impact on my research area, which is at the intersection of combinatorics, probability, and theoretical computer science). I had originally created the webpage as an annotated list of links. The links are still here, towards the bottom of the page, and they may well be the most valuable part of this page! However, I have been asked to give a more clear account of my own opinion, so here it is.
- It is undeniable that current generative AI models can solve (and have solved) major mathematical problems, and that they are already better at solving some of these problems than most of us humans. It is surprising (and kind of amazing) that they can do this, but a key ingredient is that they have ingested the entire corpus of all past work in mathematics, and no human mathematician can do this.
- I did not predict this development, and I did not believe that it had occurred until I saw it myself, at the end of July 2026. During the month or two since then, it has become possible to solve some difficult problems (including some that I had previously tried to solve, without success) with "one-shot" prompts of the form "Solve problem X, perhaps building on Y and Z. Don't give up. Check your results carefully. Keep trying. Build a Lean certificate. Give me LaTeX source for a paper with the result." Of course, not every problem can be solved this way, but some real problems can be, and have been.
- The announcement of the solution to Navier Stokes on 8 Sept 2026 was a turning point. I'm not going to delve here into the controversy about this solution, or about the work of Alpöge and Buckmaster – there is plenty on the web about that. The point that I want to make is simply that the power of AI in this domain is now undeniable.
- There has been very substantial response in the maths community, much of which is available from the links below. There has also been backlash to this response. In my opinion, this is a very critical turning point for mathematics – and it is, in fact, an existential crisis. It is the biggest crisis in mathematics since the end of Hilbert's program.
-
In case there are any non-mathematicians reading this, I will attempt to explain why it is a crisis (this explanation extends through most of the rest of
these bullet points).
The context for the crisis is that we, as mathematicians,
had not clearly elucidated, for ourselves, our fundamental purpose. We are being forced to do that now, which is a good thing.
The vast introspection that this has caused (which I will explain below) leaves most of us (at least those of us that identify as "pure" mathematicians)
steadfast in our belief that mathematics, and the understanding of it, by humans, is a good thing for the world, and will continue to be so.
At the same time, we are worried
that the combination of
- the culture that has developed in mathematics (particularly our mechanism for evaluating the quality of mathematicians), and
- the existence of these very powerful AI tools
- I don't want to keep emphasising the difference between "pure" and "applied" mathematics, because of course it is a continuum. Mathematics, from across this continuum, leads to valuable practical contributions. Nevertheless, I want to identify here the sort of mathematics that I am discussing, and this is what is commonly known as "pure mathematics". (The terminology is odd – for example, the work that I do contributes to Computer Science, but it is very much in the tradition of pure maths.)
- We need to understand the goal of the discipline in order to understand the current crisis. In short, mathematics exists, independently of our appreciation of it, and the role of the mathematician is to study, to understand, and to discover the beautiful structure that is out there. Mathematics of this sort creates many practical benefits for mankind. These benefits cannot be predicted in advance. Rather, the discoveries which come from pure, curiosity-driven, enquiry turn out to be useful after the fact (think of calculus, which underlies all of engineering and technology or of number theory, which underlies cryptography). Despite this great usefulness, our motivation, is not, and cannot be, solving immediate practical problems – that is too short-sighted. To make the most significant progress on practical problems, the world needs long-term foundational understanding, and that is what mathematicians provide.
- Any tool that aids long-term understanding should be welcomed. Also, we should be open-minded about accepting change. So why are people so upset about recent developments? The reason is that the "AI solutions" that are coming out are solving major open problems without adding any understanding. A key example here is the disproof of the Jacobian conjecture. This conjecture stood from 1884 until a couple of months ago. Its solution - a counter-example of a few lines announced in July 2026 - can readily be checked by anyone. But it doesn't really teach us anything.
- Many AI-generated proofs come with Lean certificates. For those that don't know, let me briefly say what that is. Lean is a proof system that allows us to rigorously check proofs. The proofs need to be written in its language, but AI is good at doing this automatically. Once it is done, a human need only check the translation - after that, the code guarantees correctness (though note this caveat). However, formal mathematics is very hard for humans to understand. Here is an analogy: If I gave you directions, telling you how to walk from A to B, I might say "Go down this street. Turn left at that street..." You would learn the route from A to B. But a Lean proof is more like a highly-detailed instruction: "Pick up your left foot. Move it six inches. Now pick up your right foot. Move it six inches forward and two to the right. Avoid the pebble". A detailed set of instructions like that may get you from A to B (if you followed it precisely), but you wouldn't learn anything.
- There are essentially three reasons why these AI outputs are currently unsatisfying. Sometimes they come without explanation - a single-bit answer, with nothing to learn. Sometimes they come as a formal proof, and it is difficult to digest, and to find the meaning. The third reason is just the scale. The discipline is being flooded with new "results" – and it is hard to for humans to catch up and to make sense of it all.
- None of this is problematic, in itself. So what is the problem?
- The problem, and a prediction: I believe that mathematical solutions generated by AI models (and one-shot solutions with essentially no meaningful human input) are going to dominate mathematics for a while. Probably for about 20 years. For those of us that already have tenure (or the equivalent status in the relevant country), this is no problem at all. We can continue to work towards understanding. Or those that prefer can instead become typers-of-prompts. However, unless the community responds carefully at this stage, this interim period may be damaging for the future of mathematics as a discipline. Here is the reason: During this current interim period, it has become very difficult, and probably impossible, to identify suitable problems for PhD students and young mathematicians. Mathematics is difficult, and it takes time to learn enough to solve problems. The way that our culture has evolved, PhD students and young mathematicians (and even older mathematicians!) are judged according to their track record in solving open problems. But there is little incentive for young mathematicians now – whatever a PhD student can do, over four years of hard thought, can probably be done by next month's new model, in a matter of minutes. Young mathematicians have to train by solving problems – but what if all of these problems can more-easily be solved by AI? Where does this leave them? Really deep problems, of the sort that AI cannot yet approach, are likely to take decades to solve, not years. Based on these considerations, many young people are leaving the field, and who can blame them? But mathematics itself continues to be important. Once AI finishes its big "mop up", the world is going to need humans to introduce the key ideas that are going to be needed for further progress. Curiosity-driven research is still going to be important. But introducing these key ideas takes a long time. If we don't nurture the younger generation of mathematicians, there will be nobody to produce these ideas. (I don't believe that AI can do that now, or that it will ever be able to. Even if it could, the purpose of mathematical research is surely advancing human understanding, so we need to retain humans with the relevant skills.)
- The solution: Whatever the solution is, it lies in the hands of the mathematics community. Terence Tao says that we have made the mistake of using open problems as a "proxy" for generating human mathematical understanding and that we must now de-emphasise the proxy and re-emphasise human understanding. I mostly agree, but I'm not sure that we were consciously thinking of these problems as a proxy. Probably we had not thought hard enough about our fundamental purpose. So it is good that we are doing that now. I agree with Tao, and others, that the fundamental purpose should be to achieve human understanding of mathematics. Here are a few thoughts about how we might move forward. Undoubtedly, many of them are wrong, or sub-optimal, and I'll end up editing this list.
- Solving previously-stated open problems is fine, but we need to stop thinking of this as the goal (and we need to reduce the rewards for solved problems).
- I think we should have a special section on ArXiv for AI-generated proofs. Nothing should appear there without a Lean certificate. There should be a facility for human authors to add commentary (and even linked papers) that shed more light on these AI-generated solutions or offer alternative proofs or clearer explanations.
- I don't think journals should even consider papers where it is stated that major contributions came from generative AI. This is my view, despite the obvious worry that it causes the incentive for authors to lie. Journals are at breaking point, and there is no capacity for refereeing the large dump of AI-generated papers. In the future, journals will probably be about excellent exposition, and about the explanation of general high-level ideas. Correctness-checking can and should be left to theorem provers such as Lean.
- We will have to change the way that we select postdocs and faculty members, and the way that we make tenure decisions for pure mathematicians. Right now there is too much incentive to use AI to solve problems, and then to lie about it, claiming credit where no credit is due. Let's instead focus evaluation on the quality of exposition/understanding as demonstrated through lectures, interviews, and recommendations. There is no easy way to get there – in my department (and many others) we have many hundreds of applications for every faculty position. Obviously, we cannot interview all of them. But at least for now, at least for more pure areas, we are going to have to find some way to evaluate the more high-level contribution. Asking candidates to describe this contribution in application materials is perhaps a start. Limiting the number of papers that will be considered may also be useful.
- Why don't the young mathematicians who are affected by this just work on more applied problems? This is a question that I have been asked many times. Yes, some of them will, and some of them should. However, I think, for the good of the world, in the future, we need to nurture at least a small set of pure mathematicians who keep the skill set. It will be needed again, once this transition period is over. (By then, AI will merely be a useful tool, amongst others...)
- What should the AI industry do about this? On 21 Sept, the Advisory Group on Maths and AI was announced. They confirmed what many of us had been hearing as gossip for weeks - that OpenAI had solved many more significant results in mathematics and want advice from the community about how to release these results. My opinion is that it doesn't matter - they should simply put the statements, along with Lean certificates, on their webpage. The maths community doesn't need to be coddled, and there is no advantage in holding back truth. We must examine the results, and take them in.
- Why is this web page about maths when there is so much more going on? The current disruption to the mathematics community will surely also arise in other academic disciplines. In fact, many of the issues that we are discussing, particularly, how to ensure a pipeline of humans with the right skill set, has already come up, for example, in software engineering. There are other issues that are worth discussing, but that I have omitted here. For example, the way that AI solutions can obscure proper credit for ideas, and the way that educational assessment must change, in light of AI. People are also right to worry about societal risks. I think Amodei's essay (12 Sept) about the risk that the internet could be disrupted is worth taking very seriously. This web page is about mathematics merely because this is what I know best.
- Terry Tao's blog is a continuing source of articles about this topic. Some articles are written by him. Others have guest authors.
- Proofs and Prompts is a communal blog on this topic. There is a lot here and it is frequently updated.
- Leiden declaration on AI and Maths. This is a thoughtful declaration, released on 2 June 2026, which includes positive suggestions about how to address these new challenges. Details are here. They have also created a news page with updates and announcements.
- A Severe Misalignment of AI in Mathematics Written by 25 Fields Medallists and released on 11 Sept 2026. Available here. This letter explains the problem, and it is wonderful that so many Fields medallists got behind it. The letter does not offer a solution, so that is up to us, now.
- Mathematics in the age of AI. Public lecture by Terence Tao at the 2026 ICM on 24 July 2026. I agree with his analysis. Here is an ArXiv paper based on the talk and here are the slides. I think all mathematicians should read this. You can also read more of Tao's writings on mathstodon (without joining any social media).
- Report of the Summit on PhD Math Education in the Age of AI, Sept 2026.
- Steven Kelk makes lots of interesting and thoughtful posts on this topic – see his page here. (This page is accessible without joining social media.)
- Alberto Romero's blog The Algorithmic Bridge has some thought-provoking articles, including The Month AI Conquered Math: The Full Story and Millennium Pastimes: Has math been solved by AI?
- Po-Ling Loh from Cambridge Statistics has written an interesting article about this topic on pages 12-14 here.
- I don't agree with the conclusions of this paper by Max Weinreich. It doesn't seem plausible to me that the maths community could (or should) resolve not to use an available (and useful) technology. Nevertheless, I put this here because I respect his bravery and because I think his views should be heard.
AI use in my own current projects
AI is a useful tool, and ultimately we will need to learn how to use it constructively in mathematics. My own view is that this will come after the "mop-up phase" (the phase that we are in now, where AI models can solve many open problems that are beyond the grasp of humans, merely because it has ingested the whole of the internet, and humans can't do that). Eventually, the key ingredient will be new human ideas, but we are not there now.
I'm not a typist, so I don't see any point in spending my time posing problems to AI and reporting the conclusions. Other people are doing that. That's fine.
Meanwhile, I think it is valuable for mathematicians to contribute to this discourse. We are the ones who need to work together to move the discipline past this crisis. We'll also need to contribute to digesting and explaining the many results that AI has given us. Checking correctness is boring (I certainly don't plan to be a line-checker for AI proofs) but extracting ideas is a good human activity.
In terms of old-fashioned, problem-solving collaborations, here is the approach that I take now. This is not meant to be morally prescriptive. It is just what I am doing. My personal approach to AI in current collaborations:
- Obviously, any collaborator may use LLMs as search engines, to learn things. Learning is good. Generative AI is excellent for search (it enables looking things up without even knowing what they are called!) It is also obviously very sensible to use generative AI for boring routine tasks like making latex pictures and diagrams.
- I don't join problem-solving projects without checking first that co-authors don't plan to send the actual research question to an LLM, or to ask a machine to produce the main proof or proof idea. (This is not a moral statement - I just don't want to be involved in such a project.)
- There are many grey areas: What if you just want to speed up the proof of a routine easy lemma by using an LLM? In my view this is OK, and is likely to be standard in the future (despite the risk that we get a little less skillful!), but I prefer to work in situations where co-authors would check with each other before doing this. Also, I don't really want to do it myself. That is just personal preference.