# What Today’s Best Models Still Can’t Do in Math

The a16z Show · 2026-09-01 · 63 min · https://www.youtube.com/watch?v=tQI35CSNB08

## Transcript

Daniel Litt

The goal of mathematics is not to produce mathematics papers. It’s to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that’s pretty unsatisfying.

Lisha Li

Comparing Anthropic with OpenAI, do you detect any differences in how that is similar to human reasoning?

Daniel Litt

They definitely are not good at it autonomously, but with some hints, you can get them to do something interesting. A lot of progress in mathematics comes from letting a thousand different flowers bloom and people pursuing their own curiosity, and then the boundaries of knowledge expand in some fairly uniform way.

Lisha Li

What has been the most impressive result so far?

Daniel Litt

My favorite fully autonomous result by an AI so far is the solution to the Irish unit-distance problem. There was some lemma I wanted to prove, and none of the frontier models could do it. So I worked out a ton of examples on my own, and I realized, well, maybe here’s some reason why it could be true. Once I had that statement, the models were able to very quickly prove that sort of better statement.

Lisha Li

How should the mathematics community best adapt and benefit from this?

### Meet Daniel Litt: A Practicing Mathematician's Evolving Views on AI

I am so excited to have you on, Daniel. Daniel is a professor of mathematics at the University of Toronto. Toronto is my hometown, so that’s also very exciting. But what is most special here is that Daniel is an actual practicing mathematician, and in addition, he’s been incredibly vocal about his evolving views of AI in math. I feel like every time I check in with you—if I just don’t check in with you for 2 weeks, something different has been revealed, and then you’re very—what do you call it?

Daniel Litt

I have a lot of opinions.

Lisha Li

You have a lot of opinions. Exactly. So I want to get into that. One of the things I’m most interested in is not just a discussion of how the capabilities advance. I feel like in math, that’s definitely the headline, and so on, but you’ve also been very thoughtful about how practicing mathematicians should respond. That gives us a chance to talk about what is special about math. It’s not just, “Hey, AI has been making a lot of progress here,” but an opportunity to delve into what mathematicians actually do.

So maybe, with that arc in mind, we can start with what has been the most impressive result so far, given all the recent progress, for you. Then maybe we’ll take it from there.

### The Most Impressive Result: The Erdős Unit Distance Problem

Daniel Litt

Yeah. So, okay, there have now been a lot of results. Some were produced autonomously, some were produced semi-autonomously, and in some, the AI contribution is just not at all clear. They’re in a lot of different areas, so anything I say is—I can only really comment on things that I have some expertise in. It’s quite possible that if you talk to a different mathematician, you’ll get different answers here.

My favorite fully autonomous result by an AI so far is the solution to the Irish unit-distance problem, which I think was announced in mid-May. At least what I liked about that is that it seemed to me to be, in some ways, a little bit creative. I think some of the results we’ve seen have had the flavor of taking some known techniques and applying them in maybe a clever way, or they’ve been results I would characterize as last-mile results, where some recent, quite deep work was done by a group of human mathematicians and then the AI took the final step.

Lisha Li

Yeah. But with the distance problem, I think it was something where the result was a little unexpected. First of all, my sense was that people working in the area thought it was true, and then a counterexample was found. But it also brought in some techniques from another area. I think those techniques were not especially deep or new; they were sort of classical ideas from the ’60s, but they were new to this area of studying point configurations in the plane, and so that was pretty cool. Afterward, we got to see that it was fruitful.

Daniel Litt

A bunch of mathematicians took those ideas and used them to find counterexamples to a bunch of other interesting open questions. For example, there’s the sum-product conjecture over the real numbers.

That’s at least the way I like to think about how cool a result is: you look at it post hoc and see, well, whatever new ideas were introduced, if any, were kind of useful to do other things. Did they improve our understanding of something? I think that’s maybe, so far, the main example I know of a result of that form.

Lisha Li

Yeah, I think it’s really meaningful that you’re commenting on this because that result came out, to your point, in May, and there have been so many headlines so far. It’s probably hard for somebody who’s not a practicing mathematician to appreciate the differences among these headlines, and you’ve already started laying out a taxonomy of what’s different in that proof. It would be interesting to use that as both an excuse to talk about where you sense the model differences are and what mathematicians actually do.

So, in this case, the most naive understanding of what mathematicians do is that we’re pushing around symbols in a logical manner. This is why RL is so successful at this, because you can both verify it somewhat cheaply compared to other domains, and also because the rules are quite amenable. If you’re superhuman at that, you might be good at math.

But I think that, of course, betrays most of what is interesting about mathematics, which is perhaps—I think you said this as well—but anybody who’s tried to do math knows it’s about the understanding, getting at truth, remaining confused, and developing intuitions. Very little of the tool for doing that stuff is anything other than having really strong abilities to push out logical implications.

### What AI Is Actually Doing for Working Mathematicians Today

So maybe you can speak to this: when you say it’s most impressive and creative, decoupling the inhuman feats of logical implication from where it’s being creative, what is it helping engender in terms of mathematical activity as well?

Daniel Litt

Yeah. Okay. So first of all, you characterized it as inhuman in some way. I actually think the argument was very human-like. I think OpenAI released a chain of thought, and it was very recognizable. If I tried to imagine my chain of thought in trying to solve a problem, it might look kind of like that.

Lisha Li

Yeah. We haven’t seen the raw chain of thought.

Daniel Litt

Maybe they cleansed it a little bit.

Lisha Li

Yeah.

Daniel Litt

Or maybe the model likes to swear a lot in the middle of its chain of thought, and they clean that up or something. We don’t know. But at least the summary seemed pretty readable. I would say that’s actually typical of most of the results that I’ve studied. They don’t seem inhuman at all; they seem absolutely like something a human mathematician could produce.

They’re typically understandable, if not so well written, if you just look at the raw model output. It’s not like there’s some “move 37” or whatever. It’s like a human mathematician doing math. It’s like a human mathematician doing certain types of math.

The models are still not very good at some mathematical activities. Here, I don’t mean by field, but rather certain things you do when you try to solve a problem. The models don’t seem to be doing some of them, but they’re very good at certain other things. The ways they might be a little bit inhuman are that they don’t get tired and they know a lot. But if you have to just read the final output, it doesn’t seem inhuman.

Lisha Li

It’s interesting because, obviously, we can’t get too much information from the labs producing these models about why or how the training recipes work or how they’re advancing in reasoning. But at least one of the things we do know, and I think OpenAI spearheaded this, is that reasoning in natural language is actually what they happen to scale up, and it’s not actually pushing a lot of Lean-verified proofs as the training corpus. That’s kind of amazing.

On another side, what’s interesting is that it’s not clear that a lot of the mathematical training data, if they use it to a large extent at all, is reflective of how mathematicians think, if that’s fair, because a lot of it is not legible as traces, right? Most papers are crisp and polished. Textbooks certainly just show very little motivation for how something is developed, which is why it’s usually easier to follow a research direction by actually talking to the researchers and hearing how they’re thinking about it.

So I’m curious, before maybe even going to the taxonomy as an excuse, when you’re examining these models and their results, comparing Anthropic with OpenAI, do you detect any differences in how that is similar to human reasoning? And then also, if you have any comments or insights on perhaps why natural language scales so well that way, even though it’s—

Daniel Litt

Okay.

So, first of all, I really like your point, by the way, that they’re mostly doing natural-language reasoning rather than Lean. I think that suggests to me that these—you hear a lot of people say math is a verifiable domain, like that explains the progress, whatever—but my sense is that because they’re primarily scaling in formal reasoning, probably the techniques are going to generalize to other domains pretty well. That’s just my guess.

Lisha Li

Okay.

Daniel Litt

You asked a little bit about Claude versus ChatGPT. My sense is that they’re pretty similar in terms of capabilities. I’ve played around a lot more with ChatGPT than with Fable, but it seems like there are a lot of cases where OpenAI will drop a solution to some problem and Anthropic will say, “Oh, we know, too.” Exactly.

So it actually seems like they’re solving a very similar collection of problems, and it’s a relatively small portion of what human mathematicians do. One thing that’s interesting is that we see the solutions have a certain flavor, right? There are things that rely on the models’ strengths, like their ability to grind out a long computation, pull together technical ideas from many areas, or draw on many papers that a human mathematician might not have read.

But they seem weaker in things like intuition or having some big-picture point of view. A lot of what I do as a mathematician is have some kind of philosophy that’s very nonrigorous. Maybe I think one thing is kind of analogous to another thing, and then a lot of what I’m working out is figuring out how to make that precise and trying to measure the extent to which I’ve succeeded in understanding it. Can I solve a problem? Can I find an interesting phenomenon that I don’t understand, and then come to understand it?

So far, even in the best results the models are producing, you don’t see that much of this kind of reasoning. It’s more like they’re very, very good at applying some known techniques.

Lisha Li

To be clear, that’s a very powerful thing to do—to be very good at applying all known techniques.

Daniel Litt

There are mathematicians who have had great careers doing very high-quality work of that flavor, and I think a lot of what the models are producing is high quality in that way. But it’s some kind of fairly narrow band of what mathematicians care about so far.

I do think there are signs of all the frontier models starting to be able to do more fuzzy things. I’ve tried to get both Fable and ChatGPT 5.6, I guess, to do some kind of theory building. They’re not good. They’re definitely not good at doing it autonomously, at least with whatever scaffolding I’ve set up. But with some hints, you can kind of get them to do something interesting.

When you give the models hints, it’s always a little hard to tell what part is from the model and what part is from you. But my experience is that if they can do it with a hundred bits of hints or whatever, in 6 months maybe they can do it without hints. I do think there are signs that they’re also somehow picking up some of this implicit and unwritten mathematical knowledge.

### Intuition, Taste & Why Math Isn't Just About Proofs

Lisha Li

Okay. I would love to go into the intuition part and where it sucks, to put it in a very basic way. But you actually mentioned a small detail, which is that from the public information, you’ve gleaned that Anthropic and OpenAI are probably neck and neck, but you personally are using a lot more ChatGPT. Why is that? Or 5.6?

Daniel Litt

Well, why is that? I don’t know. I mean, I just think it’s inertia. I’m pretty sure of it. One thing is that ChatGPT got better at math earlier. For a long time, the Claude models were just not useful for research math, and then I think maybe around Opus 4.5 or Opus 4.6 they more or less caught up.

Some experimentation suggests to me that they’re pretty neck and neck, so for my own work, except when I’m just experimenting, I mostly just stick with one.

Lisha Li

As people outside of labs like us, it’s really interesting just to compare how they differ at the frontier. And to your point, it might be a little bit of momentum. I do think that, at least from my anecdotal experience, 5.6 has been a lot clearer in exposition. Obviously, the models are just incredibly jagged at the frontier, and this might not be true, but in the explanations of results, I always find that 5.2 is giving a more accurate theory of mind of what it assumes I know and don’t know.

Whereas Fable might be explaining something very trivial, but then just jump to, “Well, you know, obviously you should know these things.”

Daniel Litt

I’m not sure. I find they’re both pretty bad at theory of mind.

Lisha Li

Okay, great. So, from your perspective, you’re probably asking much deeper questions. The thread that I really wanted to pull on was when you were talking about the models maybe starting to get better intuitions or even theory building. Before we even dive into that, it would be useful to talk through what your primary activity as a mathematician is, especially in your area of algebraic geometry, which probably has a very different flavor than a combinatorialist or someone in some other area.

If you can give a brief lay of the land, and then explain what your mathematical activity was pre-AI and maybe how it’s changing with AI.

Daniel Litt

I think there are a lot of different kinds of mathematicians. There are a lot of different spectra on which one can put a mathematician. Definitely, a lot of mathematicians like solving open problems.

Lisha Li

Mm-hmm.

Daniel Litt

I’m one of those. I like to solve an open problem. I think of myself as a problem solver. Another taxonomy you could have is a problem solver versus a theory builder.

Lisha Li

Mm-hmm.

Daniel Litt

At least for me, the point of an open problem is that it’s supposed to measure your failure to understand something. It’s kind of like a benchmark. One problem I really like is the Grothendieck p-curvature conjecture. It measures something about our failure to understand differential equations. There’s some very basic object we would like to understand, and if we can’t answer this conjecture, we know we don’t understand it.

Lisha Li

Mm-hmm. Okay.

Daniel Litt

So in practice, how do you get at a problem that’s supposed to be measuring something you don’t understand? Of course, you try to understand the thing better. In practice, that means you try to find the smallest situation where you can’t understand something, fiddle around with it, and then, once you win, stare at what you developed to win and try to turn that into some theory.

That’s one thing you might do. You might try to solve a problem and, in so doing, develop some kind of new understanding of the situation. You might also just have some feeling that this thing is related to this other thing. You might start building a table: property A is related to property A′, property B is related to property B′, and so on and so forth.

For example, in my work, a lot of it is motivated by an analogy between the homology of algebraic varieties and representations of fundamental groups.

Lisha Li

Okay, so that’s some fancy stuff.

Daniel Litt

This analogy is very fruitful. Really, any phenomenon that appears on one side, you can find an analog on the other side. Trying to realize that dream has led to a lot of beautiful mathematics over the last 30 or 40 years, by people like Carlos Simpson and Takuro Mochizuki and others.

Here, someone noticed that there was an analogy, and then that analogy led to a huge amount of development. There’s not really an open problem at the end, although, of course, as you develop this, you come up with lots of open problems. There’s just a philosophy that you’re trying to realize, and that philosophy is super nonrigorous, actually. It’s not symbol pushing at all.

That’s another kind of activity that I like. Beyond that, a lot of what you do when you try to understand something is actually trying to figure out what the right question is. You have some object you feel like you don’t understand, and figuring out what you don’t know is actually a very challenging thing to do.

You can go online and find a list of open conjectures or whatever, but this doesn’t really capture, in a lot of ways, what we don’t know. Often, finding the conjecture is really, really hard. A good example of this is the Birch and Swinnerton-Dyer conjecture, which is one of the Millennium Prize Problems. It’s a beautiful relationship between L-functions of elliptic curves and the set of solutions to the corresponding equations—the rank of the group of solutions.

That’s sort of what that means. It was discovered by Birch and Swinnerton-Dyer, and it was the first big-data conjecture. Birch and Swinnerton-Dyer had found all these statistics on elliptic curves in the 1960s. This was one of the first ever computer-aided bits of mathematics: they graphed these statistics and noticed that the slope of some line on a graph was related to some other algebraic invariant they knew. That was the source of this conjecture.

A lot of the time, you’re just working out examples and trying to figure out an explanation for that experiment. Going back to AI, I think these are things where AI seems, so far, to help a lot more with some things than with others. The more vague a phenomenon is, or the less you have a precise question in mind, the less useful it happens to be.

You were asking how I use it in my daily life. What I’ve found is that the projects I have that predate AI—the projects I’ve been thinking about for 3, 4, or 5 years—it’s just not that useful. It’s primarily a substitute for Google or something. I might use it to learn about some related topic, or something where I would have earlier Googled something and then read a paper; maybe I’ll discuss it with AI instead. It saves some time, for sure, but it’s not really doing deep intellectual work for me.

But I’m now—well, I suck at coding, and so now I have my good friend who’s really good at coding. Now I have all these coding projects, because suddenly, if I had a question where coding would have been really useful, I would have procrastinated on it for 6 months until I—

Lisha Li

A tireless PhD student.

Daniel Litt

Exactly. So, yeah, now I’ve picked up all these projects where coding is really useful. The models are very good for massively parallel things. If you want to find an example of something, you can ask it to work through 1,000 examples in parallel—or maybe 10 examples at a time, with 10 different subagents—and that’s really useful. But these are different activities, which are kind of on top of what I was doing before.

Lisha Li

I’d love to dig in.

Daniel Litt

Yeah—yeah, go ahead.

### Deep Thinking vs Pattern Matching: What Models Are Missing

Lisha Li

Yeah, sorry, because you were mentioning that there are projects you’ve been working on for 3, 4, or 5 years, and I don’t know if it’s correct to say those are more of the theory-building aspect of it, because you did characterize yourself as an open-problem solver. What is the thing that drives you? Because the deep-thinking part, for this audience, might also be useful to explain. Most of pure math isn’t motivated by anything external, whereas applied math has at least some external motivation for why a certain formal structure might be interesting to study.

Whereas pure math seems almost sociological. To some extent, I think Thurston made some comment or point about this in the 1970s: that it is a sociological phenomenon. As more and more mathematicians start examining something, you’ll maybe converge on some interesting structures. But it’s not—there’s some reason why people would prefer to study something, or think it’s beautiful. What is it that drives you in particular, and then maybe you can make a more general comment about the profession?

Daniel Litt

Yeah. I mean, definitely some people are motivated by beauty or aesthetic considerations. I try not to be motivated by that. I think that’s a bad way.

One thing is that it sort of limits you, right? One failure mode I see among young mathematicians sometimes is that you have something and you think you know how to prove it, and then the proof feels really ugly, and you decide—well, what if you’re wrong and it’s not ugly? Why limit yourself? You should—

Lisha Li

For this audience, what is “ugly”? I have an intuition of what’s ugly, but what does that mean? Spell it out.

Daniel Litt

I don’t really know. People sometimes feel this way. Maybe it involves a lot of grinding calculation that’s not illuminating.

But you should win by any means necessary, in my opinion. I like to think of what I’m doing as doing kind of physics, except with concepts. Instead of beauty, I try to think about what is fundamental and what’s going to open up further understanding most. Am I introducing a new idea that will be broadly useful for understanding this object?

I guess in some sense there’s an aesthetic consideration there, but I try to think that the orientation of trying to do good science rather than trying to do art is what I prefer. There’s a huge variety of opinions here, and lots of mathematicians think of themselves as being closer to poets or something.

What I do think is broadly true is that progress in mathematics seems to come from people pursuing their personal curiosity. It’s been crucial historically that there are a lot of different people with different views of what’s interesting. Then the frontier of knowledge expands and expands, and suddenly you have these opportunistic situations where a new idea has been introduced and you can suddenly cascade through a bunch of other things that we didn’t understand before.

Lisha Li

Yeah. I mean, going back to the 3-, 4-, or 5-year problems and where the models are not useful, let me ask it another way: when you’re doing the deep thinking, is it just that it isn’t clear that you formulate it as a problem, and more that you’re thinking about—what are the fundamental physics of—

Daniel Litt

Yeah, so in some cases there is a well-stated problem here. I’ve been thinking about things for maybe 10 years now at this point. For certain problems, sometimes there is just a conjecture that I would like to prove.

I think one reason the models might not be useful for some of these things is that the conjectures are true. For example, the general belief in the community was that the unit distance conjecture was true, and then it turned out to be false. What that means is that there’s a specific construction you can do to refute it.

On the other hand, a lot of the things I think about—I know, maybe someone will come up with a counterexample to the curvature conjecture tomorrow and I’ll look like a fool—but in general, these conjectures fit into some very broad theoretical framework, which means that we actually have a lot of evidence that they’re true. There’s not a construction you can do to refute them. You need to somehow—there’s this giant framework where certain pieces of it are only conjectural, and you probably need to resolve some of those conjectures to win.

We also have a pretty good sense that very serious new ideas are needed to resolve those conjectures. Of course, you can’t be sure. Maybe there’s some very clever construction that will let you avoid having a big new idea. It’s quite possible we’ll find this out.

My sense is that, for at least a lot of the things I’ve been thinking about, they’re just not accessible to applying known techniques in a very technically strong way. You need to develop a new technique. To be clear, I’m not saying the models won’t be able to do this; it’s just that so far they seem not to.

Lisha Li

Actually, that’s exactly the point I wanted to delve into. To your point, we can try to calibrate and forecast why they would get better at this, but it is true that, when it comes to providing a construction for a counterexample, they seem to be strong. If you have to start developing either new theory or, to your point, techniques or machinery to prove why a conjecture is true, they struggle more.

### Why the Unit Distance Result Was Actually Creative

It’s probably because a lot of what they’re drawing on is just techniques that have happened in other areas, and they’re porting them over. To your point, that’s why maybe the unit distance problem was such a more creative result, because it was maybe doing more of that on its own. It was drawing in something from an unexpected area.

Daniel Litt

Yeah. Exactly.

Lisha Li

And so, maybe to ask the question: it’s not that, because AI is getting so good so fast, we can count it out. But why? What do you think it has to— I guess you’re spelling out what it has to get better at, but maybe you have some more feelings on why it’s particularly hard to develop that new theory and technology.

Daniel Litt

Yeah, it’s a good question. I think you just need a different—my guess, actually, is that it’s probably totally doable. It just hasn’t been done yet. Maybe you just need a different RL environment. I don’t know. At this point, my expectation is simply that the trajectory will continue upward. I’m not a skeptic of continued capabilities growth.

Lisha Li

I think what is definitely true is that the skill of developing a theory, or building your understanding of some poorly understood object, is a fuzzier one. Mhm.

Daniel Litt

So it might be harder. You can tell it, “Develop your understanding of zeta functions,” and then, once it proves your hypothesis, you give it a reward. But it’s harder, I think, to come up with intermediate things that you can reward.

Lisha Li

That said, mathematics as a whole provides a lot of conjectures of varying levels of difficulty. Maybe this explains why there seems to be a little bit of progress in these areas. Presumably, they’re trying to get it to solve lots of problems, and some of those problems develop at least some of the skills of theory building.

Humans are able to develop these skills. Sometimes they get rewards from their PhD advisers, and their adviser says, “Oh, that’s a good idea,” based on some element of human taste or whatever.

Daniel Litt

That might be something one can do too. My expectation is that, as part of continued capabilities growth, we’ll see growth in these areas too.

Lisha Li

Yeah, yeah. I like the framing where you’re casting these increasingly difficult conjectures as a form of curriculum for both humans and, of course, AI. Maybe this is getting a little too philosophical for some people’s tastes, but it gets at the question of why we’re particularly good—or maybe particularly bad—at math, too.

What is it that we’re either struggling to do, or that some people are particularly good at, when they develop new theory? It’s not really—or maybe it’s related to—why we can formulate good structures for physics as well. It’s not obvious that all parts of the world are understandable and legible in that way, but some parts are, and therefore we try to do it because there are maybe compressive pressures on our minds. We need to—we can’t understand anything without compressing stuff more finely. I don’t know if that’s also the correct interpretation of—

Daniel Litt

I think that’s—I mean, I’m a little bit skeptical of compression as a metric of interest, but it’s definitely an anthropological reason for why we do a certain kind of theory building. I just think it’s true. In fact, I think our inability to just grind is kind of important to our ability to make discoveries.

As an example, I have 1 paper out so far where the models were kind of useful. They proved a couple of lemmas. This was a situation where we had kind of proved the main result, and then there were some lemmas I was unhappy with. They seemed not to be optimal, so I worked with Gemini Deep Think—which at the time was also on the frontier, though it no longer is—to prove the lemmas.

There was a lemma I wanted to prove, and the models couldn’t do it. None of the frontier models could do it. I worked out a ton of examples on my own and realized, “Oh, well, maybe here’s some reason why it could be true.” I found a better statement of the lemma. Once I had that statement, I probably could have done it pretty fast, but the models were able to very quickly prove that better statement.

Our inability to prove it led to an improvement in the result. Now you can take the original lemma I had that the models weren’t able to prove, put it into ChatGPT 5.6 Pro, and it will output the worst proof you’ve ever seen: 10 pages of brutal calculation with no insight whatsoever.

This would have been a perfectly fine proof, but it would not have led to the discovery of what I think is a beautiful conceptual explanation for why this thing we discovered was true. We found a better proof because we couldn’t do—I mean, I say we couldn’t do the calculation, but what actually happened is that I realized this horrible grind proof would work, and I could not bring myself to do it. I looked for another argument. Now the models can do very long technical calculations pretty reliably.

Lisha Li

Yeah, yeah. Well, I mean, to your credit—I don’t know—I think I understand why you’re saying: don’t shy away from trying to prove something just because it seems ugly. You have to take the first step, and eventually you’ll try to work toward insight.

You might be shy about admitting it, but it is still maybe driven by aesthetics or just pure—I want to understand. If understanding just means it’s a little bit simpler or more compressed, then I’ve gotten understanding. It’s a deep philosophical question, too: What is it?

Daniel Litt

Right. I mean, of course, there’s a reason we do informal mathematics rather than writing out long, formal strings of symbols in ZFC or whatever. Besides length, we’re somehow trying to put it in a way where we’re actually getting some non-rigorous understanding out.

Lisha Li

Yeah. I mean, and I think somehow—

Daniel Litt

Yeah. And even how that relates to why it helps. Again, is it because of our inability to grind, or is it because if we’re pushing toward something that’s more compressive, it also hopefully helps us understand other fields as well?

It’s somewhat magical that when we try to optimize for both, they tend to coincide. I don’t know if that’s a fair quantitative statement, but why that is is kind of magical.

Lisha Li

It’s at least sometimes true, yeah.

Daniel Litt

Yeah, exactly. We certainly bias toward the cases in which it is. That’s why we call that a good theory to build.

Lisha Li

But it’s very much, I think, what we can study and understand, and therefore it’s also very convenient that it was rich in mathematical results there.

### How Should the Math Community Adapt to AI?

Okay, so I think there are 2 things that we can pull on. Given that you’ve stated—or not stated, but believe—that the mathematical ability of AI is going to continue advancing and can do some of the things you’re doing nowadays, how should the mathematics community best adapt and benefit from this?

I’m just saying, as somebody who doesn’t have the time to practice mathematics anymore, this is very great because I can maybe dabble more. There are a lot of results that can come out. But I can also see where you’ve made the more precise point that we cannot motivate the right kind of behavior around understanding and development, so I’d love to hear more about your views there.

Daniel Litt

Yeah. So, first of all, it is clearly really exciting that there are increasingly capable models producing high-quality results. At least some high-quality results, and also a lot of slop.

Lisha Li

Yeah, some good stuff.

Daniel Litt

As the models get really capable, my hope is that they will answer a lot of the questions that have kept me up at night, and I’ll get to learn the answers. I think that’s really exciting. A lot of people got into math largely because they enjoyed learning math.

Lisha Li

Yeah.

Daniel Litt

The first thing you do as a math student is learn stuff that other people did, and you do that for 20 years before you start doing—maybe not quite 20 years, but—

Lisha Li

15 years.

Daniel Litt

If you’re lucky, 20 years. If you started early. Yeah.

Lisha Li

Yeah. Yeah.

Daniel Litt

That said, the goal of mathematics is not to produce mathematics papers; it’s to produce some kind of understanding. Maybe some of that understanding resides in model weights or something, but to me, that’s pretty unsatisfying.

My own motivation for doing mathematics is that I would like to satisfy my personal curiosity. I think people should have the capability to do that, and that requires a pretty substantial apparatus. It’s simply not the case that you can study the questions I think are fundamental unless you’ve invested a huge amount of time and effort getting to the point where you can meaningfully do so.

Moreover, even the people—you have the small group of people doing fancy research mathematics on the frontier or whatever—rely on a huge apparatus of thousands, millions, or billions of people who are trying to learn to think mathematically. You need an entire mathematical community to support a small group of people who are on the frontier. Without the pipeline, the pipeline doesn’t exist.

So if you think that’s important—the development of human capital that can meaningfully engage with frontier mathematics—I think you still have to incentivize those people to invest their time and effort in getting to the point where they can engage, and then do so in a high-quality, meaningful way. Right now, I think the existing incentive structures for math research do not encourage people to do that.

Right now, if you’re a postdoc on the market and want to get a job, maybe for the next couple of years before the community adapts, the best way to do that is to produce a lot of papers that perhaps prove old conjectures or whatever. You can do that by playing the slot machine until the model produces a hopefully correct proof of such a result. You don’t even have to pick the theorem in advance.

Here’s an experiment you can do: You can take Codex and say, “Go online, find 5 recent conjectures in algebraic geometry, and prove them.” I’ve run this experiment, and with some back-and-forth, I was able to get 3 quite bad but correct papers in an hour. They’re now sitting on my hard drive, waiting for me to email the relevant people, but this is not a good use of my time to investigate these results.

You definitely see people doing this. There’s been a huge uptick in arXiv papers, mostly not very interesting. Some of them are interesting.

Lisha Li

But a lot of it is clearly low quality, even if the result is something that would have been highly rewarded a year ago. There’s just no evidence that a human being is actually engaged with it. There’s no development of human capital or understanding.

Sometimes we’ve seen examples where 3, 4, or 5 papers with the exact same proof of the exact same theorem have come out within a couple of days of each other, which is clearly a situation where someone is playing the slot machine.

Daniel Litt

ChatGPT is consistently finding the same thing.

Lisha Li

But that’s also interesting. It’s kind of mode-collapsed on certain paths of reasoning.

Daniel Litt

Yeah. And I think it’s not obvious that this problem goes away as the models get better. Maybe it does, but maybe it doesn’t.

A lot of what I was saying earlier is that a lot of progress in mathematics comes from letting a thousand different flowers bloom and allowing people to pursue their own curiosity. Then the boundaries of knowledge expand in a hopefully fairly uniform way in this really high-dimensional space of mathematics. Opportunistically, you suddenly get some applications or answers to old questions.

It’s not clear to me that if we subordinate mathematical exploration to what the models want to pursue, what we’re getting is one mathematician duplicated a thousand times, or actually a million different mathematicians doing a million different things.

Lisha Li

Aside from math having to adapt by changing its incentive structures, I think this is perhaps dangerous, or at least something the labs have to pay attention to. On the one hand, what’s been successful for them is that this emergent reasoning capability is obviously incredibly powerful, and it’s creating a lot of great PR headlines.

But to your point, and also where my interests tend to lie, human mathematicians are coming from all sorts of weird intuitions. That is why you end up developing a frontier that is so diverse that you can actually draw from it, make connections, and produce a lot more.

If it’s true that most of these proofs being pushed out by the labs are converging on very similar things because they’re technically drawing on the same body of literature—and that’s where they’re strong right now—it’s not clear where they’re developing their intuitions. They’re developing them from practicing mathematicians, but it’s not clear where that other, emergent, stronger, diverse intuition might be coming from. It might come, but I don’t have a good theory for where it comes from.

It’s also not clear that, with test-time compute and various post-training scaling, you can even introduce new capabilities. There’s a huge debate about whether you could introduce new capabilities, and math is precisely the context in which I think this should be studied: where it’s good, where it fails, and where human mathematicians are good. That’s exactly a precise question to study.

So this is a long way of saying that it’s unclear to me whether we’ll produce a superior result without human aid if we don’t incentivize enough people to interact with the models. Unlike other domains, it might not actually continue producing a superior result without human aid.

Daniel Litt

Yeah, let me push a bit further. Let’s suppose the models become really robustly superhuman, even without adding meaningful cognitive capacity. Okay.

Lisha Li

I claim that we still want human mathematicians.

Daniel Litt

Okay. Great.

Daniel Litt

Human mathematicians. So why? It’s because there’s a question about how we’re designing society, right? Maybe the optimal situation is that you have the models doing all sorts of math research, and that leads, in this highly nonuniform and diverse way, to lots of applications and improved understanding. Maybe you don’t need humans to add that kind of diversity.

But just because something is optimal doesn’t mean you do it. There’s no reason to think that if we hand over control of math research to the models and just let them do their own thing, they’ll do the optimal thing. In fact, if we instrumentalize what we want them to do—if we say, “Make our lives better,” or whatever—it might not be the case that what they decide to do is pursue a wide variety of interesting research. They might just try to take the direct path. We don’t know what’s going to happen.

I don’t know if you believe there’s any value in this broad-based fundamental research, which I do. I think that’s one of the most valuable things that humans—or whatever the models can be doing—can pursue. I think the easiest way to guarantee it happens is to keep a community with broad interests that is pushing the models to do it and helping us design a society where that’s what we’re pushing for.

Daniel Litt

Yeah. I think at least my hopeful vision for the future is that humans are not totally disempowered. We have some control over where we’re going, and if that’s the case, what we end up doing is going to be driven by human interests.

You want to have people who have lots of different interests and also the capabilities to pursue them. You want people who are smart and engaged, who can do not just mathematical thinking but all sorts of thinking.

Lisha Li

Yeah. A great fear is that, as AI advances, we don’t develop the right ergonomic interfaces to encourage us to continue to be good thinkers. It’s so easy to relinquish that because you just hand it off, and the models aren’t even good at that level of thinking where it’s the top-level structure. Despite that, it’s so easy to do so.

### Where AI Will Impact Applied Math First

Especially for math, another selfish reason is to math-max. I think it’s actually a great pedagogical excuse to get really rigorous about thinking through various things. This is not why mathematicians do it, but as somebody who is part of what we do, I think it’s a great framework for thinking about many things, not just mathematics.

Daniel Litt

As somebody who is now also a parent of a 2-year-old, I think about this a lot, too. It’s not about grinding. Grinding is still good, but I don’t want to undersell it too much.

I used to collaborate with some Hungarian mathematicians, Béla among them, and I heard that in Budapest they would just teach group theory in primary school. We should definitely do that. We should continue doing that.

Now that AI is somewhat good at explaining and is far more accessible, we should probably proliferate that even more. Maybe that helps bring more people to the frontier rather than just—

Yeah, I mean, this is something I’m concerned about. Of course, I mostly talk about math because that’s where I live, but one nice thing about thinking about this is that we’re one of the first professions to have a significant impact from high-quality models.

Although I think maybe we're one of the first professions for it to happen so publicly. My sense is that there are plenty of other professions that are—

Lisha Li

100%.

Daniel Litt

Coding. But also, anything you do at a computer—probably a huge amount is being done by the models at this point—and there isn't a public reckoning about it.

Lisha Li

Your math capabilities are useful for the company and the labs to talk about, I think, probably a bit more publicly than everyone else. But, yeah, I think one reason to try to maintain human capital in this area is that it's a model for all professions. Presumably, we still want people who are meaningfully engaging with the world, who are experts, and who have talents and trained skills.

It's actually very convenient that the math profession is so entwined with education here, because I think we're also seeing some challenges among college and secondary education coming from AI, too.

Daniel Litt

So, as you said, it's also an amazing tool to learn.

Lisha Li

People have started talking about a bimodal distribution in their classes. There are some people who are really figuring out how to take advantage of new tools, and other people who are just letting them do their homework and then bombing everything else.

Daniel Litt

Unfortunately, I don't think that adapts fast enough. We definitely want to be living in a world where we're producing better thinkers. I think it's just—

Lisha Li

Even if we talk about the pure optimization game, I think that's better for us. And so, as a human being, I'm like, that would be pretty—

### Taking Advantage of AI Without Losing the Craft

Daniel Litt

It would be inconvenient if we became worse thinkers just as AI ascends. It's too easy to let that happen. So we should be thinking hard about how to take advantage of this and harness it for our own improvement, right? Well—

Lisha Li

Yeah. I think it's interesting to observe that, at their current level of capabilities, the models let you do a lot more. They let you do things you wouldn't have done before, cheaply enough to do them now. But it's not clear to me that, in many cases, they're actually improving the quality of outputs.

Daniel Litt

Yeah. I think this is common: you have a new technology that's doing something a little bit worse than what was previously done, but much cheaper, and so you suddenly get a lot of low-quality outputs that are displacing previous high-quality outputs. But I think it's possible to use the tools in a way that actually improves quality along every dimension. It just requires some thoughtfulness and some redesign of institutions to incentivize that.

Lisha Li

Yeah. Well, hopefully capitalism works there. I do feel like the most high-value things require people to use AI effectively, and right now the models are not good enough without human experts participating. But to your point, there's a vast majority of perhaps more junior and entry-level roles, and it's harder for those roles to adapt as well. The thing that would be a mistake is to use the AI models in a way that doesn't—

Daniel Litt

Basically, you need to be ascending and using the models to deepen your understanding. It's so easy for human nature to make us lazy, and you have to resist that, because that is the moment that you will lose, basically. Forgive the very competitive language, but it really is that it's so easy to relinquish the thinking to the models.

The models can't really think. As things are ascending so fast, it's critical that you continue developing those faculties and actually leverage AI to improve those faculties rather than relinquishing them.

Lisha Li

Yeah. Yeah. I totally agree. I know I get to talk—

Daniel Litt

And not even kind of a positive, obviously.

### Comparing Anthropic vs OpenAI in Math

Lisha Li

Yeah. No, exactly. Suddenly there's a spotlight on math, and I can nerd out about it more. I was actually kind of curious: have you got any comments on the—

Daniel Litt

I guess this is another constructive result: the elliptic curve of rank 30 that just came out yesterday. So, if that—

Lisha Li

That's right.

Daniel Litt

Yeah. Yeah.

Lisha Li

Well, we don't have any details about it yet.

Daniel Litt

I know. There are rumors.

Lisha Li

Yeah, we don't know how it happened.

Daniel Litt

And a collaborator whose name I unfortunately forget. Maybe you can add it in post.

Lisha Li

I just know the Twitter handle.

Daniel Litt

Yeah, we can.

Lisha Li

A lot of these nice recent results have come from Levent Poge, with unclear amounts of autonomy. My sense is that some of them are semiautonomous rather than fully autonomous.

Daniel Litt

With this one, we don't know anything about the methods. Without knowing about the methods, it's very hard to say how significant it is.

Lisha Li

Yeah.

Daniel Litt

What I would say is that, historically, such results have been understood by the community as cool, but I wouldn't say they're a big deal. The typical place a result like this might go is a website of records—

Lisha Li

Like, not—

Daniel Litt

It's not an Annals-level result.

Lisha Li

Yeah, it's not an Annals-level result. But that said, it's cool, and there definitely were a few very, very talented mathematicians who like these kinds of questions, with Noam Elkies being maybe the most famous example. Elkies and Klagsbrun are the ones who have been pushing this record for a while, and they recently found a rank-29 example—

Daniel Litt

Yeah, you know, so those are mathematicians. I think people like it, and it's cool that the models can do this sort of thing. I think Ava Howell may be the third collaborator.

Lisha Li

Okay.

### Why Some Labs Have Gone More Secretive

Daniel Litt

They have not yet told us how they did it.

Lisha Li

Have they been mostly more secretive? I think they've released some traces for stuff, but—

Daniel Litt

Yeah, so for this one, I think they haven't yet, unless I missed it. Levent likes to tweet out his results, but he has been slowly releasing some kind of PDF write-ups, too. I think he's just having fun on the internet.

Lisha Li

Yeah. Yeah.

Lisha Li

It reminds me of when you were saying that you have to evaluate how the results came about. I had Mark Sellke and Mehtab Sawhney from OpenAI on recently, and they were saying that what's been most charming or delightful is that the proofs have been relatively short. They're not 200-page proofs, which maybe corresponds to your grinding point as well.

But maybe this is optimized for—in retrospect, it was picked because it was short—or maybe, on average, the stuff that you throw at GPT, Saul, or Fable tends to be shorter and more legible to humans, rather than just going haywire and grinding it out.

Lisha Li

Yeah. So, I think it is nice that they'll sometimes produce short, clever proofs. Of course, everyone likes a short, clever proof. I think my sense is that the reason they're not producing long, complicated proofs is that they cannot—

Daniel Litt

The ability to check correctness is not yet there. So, even if you ask the models to produce—

Lisha Li

For a short proof, you can then ask them, “Is that correct?” They will often say no.

Daniel Litt

They're much more reliable than they were 6 months ago, for sure, but they will still sometimes produce things that are just wrong—and they know they're wrong.

Lisha Li

Yes.

Daniel Litt

I think the problem with producing a very long thing is that they might not know they're wrong. And so what I wonder is, presumably internally, OpenAI and Anthropic have probably solved a lot more problems than they've released.

Lisha Li

I imagine quite a few of them are just ones they're not sure are true.

Daniel Litt

So, for example, with this recent list of 10 problems released by OpenAI, those were all formalized in Lean, which is, of course, very good evidence that they're true. I have no doubt that there were a lot more that they could not formalize in Lean because the prerequisite results have not been put into Mathlib yet, for example. There were probably more that were longer, which also makes them challenging to check.

Lisha Li

Yeah.

Daniel Litt

Yeah, so this is my guess. We do actually see very long generated proofs on the arXiv. For example, someone recently posted a claimed proof of resolution of singularities in positive characteristic that was 800 AI-generated pages. It's definitely—I mean, I'm sorry, I haven't read it; I haven't gotten far—but there's no way it's correct. This would be a major result. It's just not within the capacities of the current models, if you're reasonably well calibrated.

Lisha Li

Yeah. Yeah.

Daniel Litt

It's definitely the case that no human has read it. The models are definitely not able to check this kind of thing. So, well, I mean, yeah, this is my expectation: the reason it's producing short, clever things is just that that's what we can check. You can get it to produce long things that are grindy or hard to check, but then...

Lisha Li

Yeah, yeah, we'll have to get there. This is actually more what it says about frontier capabilities rather than, in a perhaps more negative way, "Hey, it's just so good at these clever things." No, I think that's totally fair. Especially if you look at an adjacent domain like code, right? I remember reading something that Cursor put out about testing their long-horizon harness.

In this case, they're trying to reproduce SQLite in Rust, and it was just so telling how far we are from that. It seems like a very comparable task: it's very long, and you have to make sure there's something to verify that it's a correct implementation by testing a suite of cases. You actually need a harness there; it's not just the raw models. It takes a while, and you can actually compare different frontier versus non-frontier models, who's the planner, and so on—the differences in their capabilities as well.

Daniel Litt

Yeah. I also think in practice that, to elicit a long proof, you kind of need a harness.

Lisha Li

Yeah.

Daniel Litt

When you make a harness whose goal is to elicit a proof, I think it often decreases reliability because you're just trying to produce output. ChatGPT 5.6 Pro is very careful—it really tries not to say wrong stuff, for example, although it happens, and then you'll ask it, "Was that correct?" and it'll say, "No."

But when you're trying to get it to elicit something, you try to get it to be creative. You try to get it out of this very rigorous rut in order to actually get it somewhere. If your interest is in—there are a lot of people who are trying to arbitrage the prestige mechanics of academic mathematics and elicit a lot of proofs which are not necessarily being checked—in order to do that, I think you just decrease the reliability to get a lot of stuff. Any harness that can elicit a 250-page paper is probably not being very careful about what it's producing.

Lisha Li

And you're saying that it's decreasing reliability—or it's not reliable—just because the capability isn't there yet, so it's forcing a longer-horizon task on it.

Daniel Litt

Yeah. In practice, how does a human check it? A human can't check a 250-page paper either. You can't reliably check it line by line. What you try to do to understand it is understand the overall global structure of the argument and stress-test it in various ways: Would this argument apply to something else that I know to be false? Does it work in this special case?

The models seem not to be able to do that kind of more—I don't know—fuzzy unit testing of a proof very well yet. Actually, one of my favorite tests for the models, which they haven't succeeded at yet, is this: There's a paper—I won't name it—that came out a couple of years ago that was wrong. It was in a close enough area to myself that I immediately downloaded it and started reading it, and it was very hard to find the actual specific error, but it was also very clear from the structure of the argument that it couldn't work. It was, you know, me and a bunch of other experts arguing something that was too strong to be true, if you took that a little bit further.

Lisha Li

Yeah.

Daniel Litt

A bunch of other experts and I immediately realized it was wrong, and we emailed the author. We went back and forth until someone figured out what the precise, specific error was.

Lisha Li

Okay.

Daniel Litt

So far, the models seem not to have been able to do this. The specific error is quite subtle, but they're also not able to do this kind of overall, big-picture sanity-checking, which is how people check papers in practice.

Lisha Li

Yeah. Yeah. It mirrors why, in code, it's so clear that it's good at the syntax, but higher-level architectural stuff is still very weak. Maybe it'll get there, but it probably needs some harness help. Who knows? People have evolving opinions on how much the harness and the model have to co-evolve, and which one is necessary, but the next model requires less. I feel like, in math, it would be very interesting to see how you—if you do any experiments with the harness there—how that improves, because it is a mark of general reasoning. Yeah. Yeah.

Daniel Litt

Yeah. I have my own sort of bad little harness in Codex and Claude Code, but I personally do not enjoy autonomous mathematics very much, so I mostly do not use the harness. I mostly try to use it to help me understand stuff.

### Raising a Mathematician: Teaching Math to a Toddler

Lisha Li

Okay. No, it's totally fair. Exactly. You don't want to automate your job away, because that does involve you being in the loop to understand it, which is necessary to participate. Maybe to finish off, I'd love to ask—and this could be just something you haven't thought about or have actually thought a lot about—I think you also have a toddler, right?

Daniel Litt

Yeah.

Lisha Li

Yeah. And so you have a three-year-old?

Daniel Litt

Yeah, three-year-old.

Lisha Li

Great. Right. So you're one year more advanced and probably have more thoughts on this. How are you thinking about her education in math—not to grind, but really just how to pass on the love of it—and how to react to AI?

Daniel Litt

Yeah. So she's three. She's never used AI. She is starting to add; that's about as far as we are in math.

Lisha Li

That's more than she can do: add single-digit numbers by counting on her fingers and count up to maybe 30 reliably and 50 semireliably. So I'm very proud of that.

Daniel Litt

That's good. Yeah.

Lisha Li

Yeah, I definitely encourage that. We talk about shapes and stuff. A couple of days ago, I woke her up and she was hiding under the blankets, and I was like, "What are you up to under there?" She said, "Oh, I'm doing some math."

Daniel Litt

Oh, I did. I think you tweeted about that. That was adorable.

Lisha Li

Yeah, it was great. So I think she has some sense that I like math, and she's into it because of that.

I don't know. I think the world is probably going to look pretty different in 20 years, or whenever she's fully adult and doing her own thing. But a lot of what we educate people for is pretty robust to changes in the nature of the world. I think the reason to learn math has always been to think clearly and better understand the world, and presumably that's something you want to do even if there are extremely capable AIs.

This is also true of reading a lot of books and doing the humanities and so on. I think the actual values of the math profession and education more broadly are things we definitely want to try to preserve. I hope to instill that in my daughter. How much institutions have to change to make sure that happens is maybe an open question I think about a lot. But, at least at a personal level, I'm definitely trying to convince my three-year-old that math is super cool.

Lisha Li

Oh, yeah, 100%.

Daniel Litt

One of her first words was "icosahedron."

Lisha Li

Was what?

Daniel Litt

She has a little icosahedron toy. My parents gave it to her when she was one.

Not literally a first word, but she learned the Platonic solids quite early.

Lisha Li

Oh, very good. Very good. Very fun.

Daniel Litt

Next up, group theory. I mean, it's very natural. That's right. Yeah.

Lisha Li

Actually, when I teach her addition and subtraction, we'll do it in the context of a general group.

Daniel Litt

Oh, very good.

Lisha Li

Well, at least you give the motivation. I think a lot of people probably skip that part. Maybe this is how math grad students can also focus on training the next, much younger generation to use AI in service of actually getting better at math rather than—

Daniel Litt

Just lacking understanding.

Lisha Li

Wonderful. Okay. Well, thank you so much, Daniel.

Daniel Litt

Thank you. It was a lot of fun.

Lisha Li

Yeah, it was a lot of fun. I think there's going to be a lot more progress very soon. I'd love to maybe catch up and chat again.

Daniel Litt

Sounds great.
