当今最优秀的模型在数学方面仍然做不到什么
- Daniel Litt 最喜欢的完全自主成果,是5月中旬解决爱尔兰单位距离问题的结果;他认为这“在某种程度上有点创造性”,而不是对人类工作的最后一步收尾。 Lisha指出,相关领域的研究者原本认为命题成立,但有人借助20世纪60年代的经典技术找到了反例,而这些技术此前并未用于研究平面上的点配置。数学家随后用这套思路为其他开放问题找到反例,包括实数域上的和积猜想。Litt 的评判标准是事后检验:新想法是否能被用于其他问题,或是否加深了理解?
- 能力前沿高度失衡:模型“非常、非常擅长运用已知技术”——能够埋头做计算,也能整合许多论文中的技术性想法,但在直觉、大局观哲学和理论构建方面明显薄弱。 Fable 和 ChatGPT 5.6 都不擅长自主构建理论,不过在提示下可以做出有意思的东西。Litt 的保守判断是,如果模型只需“100比特的提示”就能做到,也许6个月后就能在没有提示的情况下做到。
- AI 生成的数学证明读起来像人写的,而不是来自异形智能——“并不存在什么‘第37步’”。 Lisha指出,各实验室似乎在扩展自然语言推理,而不是依靠大规模 Lean 验证证明语料。Litt认为,这意味着尽管人们常用可验证性解释数学进步,这类技术或许也能泛化到其他领域;但他强调,这“只是我的猜测”。
- 短而巧的证明构成的是验证能力的上限,而不只是风格选择:“它们之所以没有产出冗长复杂的证明,是因为它们做不到”;模型生成长篇输出时“可能根本不知道自己错了”。 OpenAI 最近列出的10个问题已经用 Lean 形式化,而 Litt 认为,一份声称解决正特征下奇点消解问题、长达800页的 AI 生成证明几乎肯定是错的。能够诱导模型写出250页论文的编排框架,往往是以牺牲可靠性为代价。
- 激励机制已经面临风险:博士后可以“反复拉老虎机”,直到模型吐出一个希望正确的证明,然后靠此发表论文。 Litt 曾让 Codex 找出5个近期的代数几何猜想并加以证明;经过一些来回沟通,Codex 在1小时内生成了3篇“很糟糕、但正确的论文”。他和 Lisha 还讨论了这样的案例:几天之内,3–5篇论文使用了“完全相同的证明证明完全相同的定理”。系统性风险不是100万人追逐不同的好奇心,而是“1个数学家被复制1,000遍”。
- Anthropic 与 OpenAI“旗鼓相当”,Claude 在 Opus 4.5 或 Opus 4.6 左右追了上来,此前很长时间对研究数学都没什么用。 一个新的秩30椭圆曲线结果颇有意思,但在不知道其方法的情况下无法评估;近期一些成果涉及 Levent Poge,但自主程度并不清楚。从历史上看,这类纪录很酷,却“不是 Annals 级别的成果”。
- 即便模型变得“真正稳健地超越人类”,Lisha 仍说“我们还是需要人类数学家”。 Litt认为,最优并不意味着会开展广泛、多元的基础研究:自主系统可能会选择一条直接路径。由兴趣广泛的群体组成的数学共同体,才能保留人类的控制权,把模型引向多样化研究,并维持培养数学思维者的路径。
- Litt 的核心原则是:“数学的目标不是产出数学论文,而是产出某种理解。” 如果理解只是储存在模型权重里,对他而言并不令人满意。人类数学家应继续发展自身能力,利用 AI 加深理解,而不是把思考交给 AI。
1. 唯一达到 Litt 标准的自主成果:爱尔兰单位距离反例
- Litt 对 AI 数学成果的分类是:有些完全自主,有些半自主,还有一些连 AI 到底贡献了什么都“不太清楚”。许多成果只是“最后一步”——深层次工作由人类完成,AI 负责迈出最后一步。他最喜欢的完全自主成果,是5月中旬公布的爱尔兰单位距离问题解答,因为它“在某种程度上有点创造性”。
- 它的创造性来自哪里?Lisha说,相关领域的研究者原本认为命题成立,但最终找到了反例。这个结果引入了“另一个领域”的技术——“60年代的经典想法”,本身既不深也不新,却是首次用于研究平面上的点配置。它随后产生了连锁效果:数学家用这些想法为其他开放问题找到反例,“比如实数域上的和积猜想”。
- Litt 的评分函数是事后检验:观察成果引入了哪些新想法,再追问这些想法是否“有助于做其他事情”,或是否“改善了我们对某件事的理解”。
2. 证明是人类式的,自然语言推理或许可以泛化
- 针对 Lisha 所说的“非人类壮举”,Litt 的修正是:公开的思维链摘要“非常容易辨认”,大致就是人类数学家的推理。就他研究过的各项成果而言,“并不存在什么‘第37步’”;它们“像是一个人类数学家在做某种类型的数学”。真正更像非人类的地方在于,模型不会疲惫,而且知道很多东西。
- Lisha观察到,各实验室似乎在扩展自然语言中的推理能力,而不是依靠大规模 Lean 验证证明语料。Litt说,人们常常认为数学是一个可验证领域,但“因为它们主要在扩展形式化推理”,他猜测这些技术可能泛化到其他领域。他强调,这“只是我的猜测”。
3. 实验室战报:5.6 对 Fable,在锯齿状前沿上难分伯仲
- Litt认为,两家实验室解决的是“一组非常相似的问题”——OpenAI 发布一个解法,Anthropic 就说自己也知道怎么解;而这组问题只占人类数学家工作中“相对很小的一部分”。他主要使用 ChatGPT,部分原因只是惯性:很长一段时间里,Claude 模型“对研究数学根本没用”,直到 Opus 4.5 或 Opus 4.6 左右才基本追上。
- Lisha根据个人体验做出的区分是:5.6 对自己知道什么、不知道什么有更清晰的心智模型,而 Fable 会先解释一个简单问题,然后突然跳到很远的地方。Litt 的结论是,两者“理论构建能力都很差”。他曾尝试让 Fable 和 ChatGPT 5.6 构建理论;“它们显然不擅长自主完成”,但在提示下能产出有意思的东西。他补充说,很难判断其中多少来自模型,多少来自提供提示的人。
4. 数学家实际在做什么:开放问题是基准,类比是引擎
- Litt认为自己首先是解题者,而不是理论构建者,但开放问题的作用是衡量理解上的缺口:“它有点像一个基准测试。”他举的例子是 Grothendieck p-曲率猜想,它衡量的是人们对微分方程理解上的不足。
- 另一种工作模式由哲学驱动:Litt 的研究源于一个类比——代数簇的同调与基本群的表示之间存在对应。“一边出现的任何现象,都能在另一边找到类似物。”沿着这条类比,Carlos Simpson、Takuro Mochizuki 等人开展了数十年的数学研究。“实际上,这种哲学非常不严谨,完全不是符号推演。”
- 找到正确的问题往往才是最难的部分。Birch and Swinnerton-Dyer 猜想是“第一个大数据猜想”:Birch 和 Swinnerton-Dyer 在1960年代收集椭圆曲线的统计数据,将其绘图后,发现一条斜率与某个代数不变量之间存在关系。
- AI并不适合这些活动:“一个现象越模糊,或者你脑中越没有一个精确的问题,它就越没用。”对于他已经推进3–5年的项目,AI“主要是 Google 的替代品”。但由于 Litt 不擅长编程,AI也解锁了一批涉及大规模并行寻找例子的项目——让 AI 并行处理1,000个例子,或一次处理10个、由10个子代理分工;没有 AI,这些项目可能会被他推迟几个月。
5. 反对美学至上:“不择手段地取胜”
- Litt努力不让审美成为动力,并指出年轻数学家常见的一种失误:因为证明“感觉非常丑陋”就放弃它。他的回应是:“如果你错了,而它其实并不丑呢?为什么要限制自己?”他说:“你应该不择手段地取胜。”
- 他更愿意用“做某种以概念为对象的物理学”作为指南:追问什么是基础性的,什么能进一步打开理解,而不是试图创作数学艺术。他也承认,很多数学家认为自己更接近诗人。
- 他更广泛的社会学判断,也贯穿后面对 AI 的讨论:进步来自人们追逐个人好奇心。许多人对“什么有趣”有不同理解,才能让“一千种不同的花”绽放;前沿由此扩张,新想法也有机会层层传递,最终解开旧问题。
6. 模型为何难以构建理论——而无法蛮力推进反而导出了更好的定理
- Litt认为,模型难以处理他的一些困难猜想,一个原因是这些猜想通常被认为是真的,而且嵌在宽泛的理论框架中。错误的猜想可以通过一个具体构造直接反驳;但这类猜想可能需要先解决其他猜想,并发展出“非常严肃的新想法”。他强调:“我不是说模型做不到;只是到目前为止,它们似乎还做不到。”
- 他并不怀疑能力会继续增长。他猜测这件事“可能完全可行”,也许需要一个不同的强化学习环境。难点在于理论构建更难设计奖励:证明一个猜想很容易奖励,但对理解的中间性改进则不然。
- 他最有力的案例来自一篇论文。当时仍处于前沿的 Gemini Deep Think 在证明引理方面很有用,但所有前沿模型都无法证明其中一个引理。Litt 逐一研究了大量例子,意识到一个更好的命题可能成立,随后模型很快就证明了这个改进后的命题。
- 如今,ChatGPT 5.6 Pro 能证明原来的引理,但给出的是“你见过的最糟糕的证明”——10页惨烈计算,没有任何洞见。Litt意识到,蛮力证明确实能奏效,却实在不愿亲手去做,于是转而寻找概念性论证。“我们无法直接蛮力推进,这一点对我们发现新东西的能力其实很重要。”
7. 垃圾论文经济学:老虎机与 arXiv 上的模式坍缩
- 激励问题在于,学界适应之前的未来几年里,求职市场上的博士后可能通过“反复拉老虎机,直到模型产出一个希望正确的证明”来发表论文。Litt 做过一次实验:让 Codex“找出5个近期的代数几何猜想,并证明它们”。经过一些来回沟通,Codex 在1小时内生成了3篇“很糟糕、但正确的论文”;这些论文就放在他的硬盘里,等着他联系相关人士。
- 低质量的标志未必是错误,而是缺乏人类参与的痕迹:“看不出任何证据表明确实有一个人在参与其中”,也没有人力资本或理解的积累。Lisha提到,有时几天内会出现“3篇、4篇或5篇论文,使用完全相同的证明证明完全相同的定理”。她称之为一种模式坍缩;Litt说,ChatGPT 总是在找到同一个东西。
- Litt不认为随着模型改进,问题就会自然消失。如果数学探索被模型所追逐的目标支配,整个领域可能得到的是“1个数学家被复制1,000遍”,而不是“100万名不同的数学家在做100万件不同的事”。
8. 短证明是校验上限,模型还无法给论证做单元测试
- Lisha转述 Mark Sellke 和 Mehtab Sawhney 的观察:AI证明往往短得讨喜。Litt的解释是,“它们之所以没有产出冗长复杂的证明,是因为它们做不到”:模型还没有足够的能力检查正确性,生成长篇论证时“可能根本不知道自己错了”。
- 他猜测,OpenAI 和 Anthropic 解决的问题可能多于公开发布的数量,但其中一些无法形式化,因为前置结果尚未进入 Mathlib;而更长的证明也更难检查。OpenAI 最近列出的10个问题已经用 Lean 形式化,这为其正确性提供了“非常有力的证据”。
- 一个典型案例是一份800页的 AI 生成证明,声称解决正特征下的奇点消解问题。Litt没有读过它,但说“它不可能正确”,也没有任何人类读过它,而当前模型无法检查这样的论证。
- 为诱导模型写出长证明而搭建的框架,可能反而降低可靠性。试图让模型生成一篇250页的论文,可能是在把它推离谨慎模式;“任何能够诱导模型写出250页论文的框架,可能都没有认真检查它产出的内容”。
- Litt最喜欢的一项未通过测试,涉及一篇未具名的错误论文:具体错误非常微妙,但整体结构立即让专家意识到,它证明了一个过强的结论。人类会从全局对长篇论证做压力测试——看它是否会推出已知为假的结论,或检查它在特殊情形下是否成立。模型仍无法可靠地完成这种对证明的“模糊单元测试”。
9. 即便 AI 稳健地超越人类,只要培养链条还在,人类仍有任务
- Litt的核心主张是:“数学的目标不是产出数学论文,而是产出某种理解。”某些理解或许存在于模型权重中,但“对我来说,这相当令人不满意”。
- 数学前沿研究依赖“成千上万、数百万乃至数十亿人”学会数学式思考。现有激励机制却可能转而奖励博士后发表大量论文,包括由模型生成的论文,而不奖励人力资本或理解的积累。
- Lisha说,即便模型变得稳健地超越人类,“我们还是需要人类数学家”。Litt回应,即使自主研究在效率上达到最优,也不能保证模型会追求广泛的基础研究,而不是沿着通往工具性目标的直接路径前进。确保研究多样性的最简单方式,是维持一个兴趣广泛的共同体,让它推动模型,也帮助塑造社会。
- Lisha 和 Litt 还讨论了便宜但略逊一筹的产出挤压高质量工作的风险。Litt认为,AI可以在每个维度上提升质量,但这需要审慎思考和制度重构。Lisha描述了一种课堂上的“双峰分布”:一部分学生学会使用工具,另一部分学生让 AI 替自己完成作业,随后在其他地方遭遇失败。
- 谈到育儿,Litt说,他3岁的女儿已经开始学习加法;他开玩笑说,“icosahedron”是她最早说出的词之一,同时解释说她很早就学会了柏拉图立体。Lisha说,自己的女儿能可靠地数到大约30,数到50则算得不太稳定;她还描述过一次,发现女儿钻在毯子下面说:“哦,我在做数学。”Lisha说,她把加减法放在一般群的语境里教;Litt说,下一步就是群论。两人的共同教育判断是:即使世界上存在能力极强的 AI,数学仍能帮助人清晰思考、理解世界。
完整逐字稿
The goal of mathematics is not to produce mathematics papers. It’s to produce some kind of understanding. Maybe some of that understanding resides in model weights. To me, that’s pretty unsatisfying.
Comparing Anthropic with OpenAI, do you detect any differences in how that is similar to human reasoning?
They definitely are not good at it autonomously, but with some hints, you can get them to do something interesting. A lot of progress in mathematics comes from letting a thousand different flowers bloom and people pursuing their own curiosity, and then the boundaries of knowledge expand in some fairly uniform way.
What has been the most impressive result so far?
My favorite fully autonomous result by an AI so far is the solution to the Irish unit-distance problem. There was some lemma I wanted to prove, and none of the frontier models could do it. So I worked out a ton of examples on my own, and I realized, well, maybe here’s some reason why it could be true. Once I had that statement, the models were able to very quickly prove that sort of better statement.
How should the mathematics community best adapt and benefit from this?
1. Meet Daniel Litt: A Practicing Mathematician's Evolving Views on AI
I am so excited to have you on, Daniel. Daniel is a professor of mathematics at the University of Toronto. Toronto is my hometown, so that’s also very exciting. But what is most special here is that Daniel is an actual practicing mathematician, and in addition, he’s been incredibly vocal about his evolving views of AI in math. I feel like every time I check in with you—if I just don’t check in with you for 2 weeks, something different has been revealed, and then you’re very—what do you call it?
I have a lot of opinions.
You have a lot of opinions. Exactly. So I want to get into that. One of the things I’m most interested in is not just a discussion of how the capabilities advance. I feel like in math, that’s definitely the headline, and so on, but you’ve also been very thoughtful about how practicing mathematicians should respond. That gives us a chance to talk about what is special about math. It’s not just, “Hey, AI has been making a lot of progress here,” but an opportunity to delve into what mathematicians actually do.
So maybe, with that arc in mind, we can start with what has been the most impressive result so far, given all the recent progress, for you. Then maybe we’ll take it from there.
2. The Most Impressive Result: The Erdős Unit Distance Problem
Yeah. So, okay, there have now been a lot of results. Some were produced autonomously, some were produced semi-autonomously, and in some, the AI contribution is just not at all clear. They’re in a lot of different areas, so anything I say is—I can only really comment on things that I have some expertise in. It’s quite possible that if you talk to a different mathematician, you’ll get different answers here.
My favorite fully autonomous result by an AI so far is the solution to the Irish unit-distance problem, which I think was announced in mid-May. At least what I liked about that is that it seemed to me to be, in some ways, a little bit creative. I think some of the results we’ve seen have had the flavor of taking some known techniques and applying them in maybe a clever way, or they’ve been results I would characterize as last-mile results, where some recent, quite deep work was done by a group of human mathematicians and then the AI took the final step.
Yeah. But with the distance problem, I think it was something where the result was a little unexpected. First of all, my sense was that people working in the area thought it was true, and then a counterexample was found. But it also brought in some techniques from another area. I think those techniques were not especially deep or new; they were sort of classical ideas from the ’60s, but they were new to this area of studying point configurations in the plane, and so that was pretty cool. Afterward, we got to see that it was fruitful.
A bunch of mathematicians took those ideas and used them to find counterexamples to a bunch of other interesting open questions. For example, there’s the sum-product conjecture over the real numbers.
That’s at least the way I like to think about how cool a result is: you look at it post hoc and see, well, whatever new ideas were introduced, if any, were kind of useful to do other things. Did they improve our understanding of something? I think that’s maybe, so far, the main example I know of a result of that form.
Yeah, I think it’s really meaningful that you’re commenting on this because that result came out, to your point, in May, and there have been so many headlines so far. It’s probably hard for somebody who’s not a practicing mathematician to appreciate the differences among these headlines, and you’ve already started laying out a taxonomy of what’s different in that proof. It would be interesting to use that as both an excuse to talk about where you sense the model differences are and what mathematicians actually do.
So, in this case, the most naive understanding of what mathematicians do is that we’re pushing around symbols in a logical manner. This is why RL is so successful at this, because you can both verify it somewhat cheaply compared to other domains, and also because the rules are quite amenable. If you’re superhuman at that, you might be good at math.
But I think that, of course, betrays most of what is interesting about mathematics, which is perhaps—I think you said this as well—but anybody who’s tried to do math knows it’s about the understanding, getting at truth, remaining confused, and developing intuitions. Very little of the tool for doing that stuff is anything other than having really strong abilities to push out logical implications.
3. What AI Is Actually Doing for Working Mathematicians Today
So maybe you can speak to this: when you say it’s most impressive and creative, decoupling the inhuman feats of logical implication from where it’s being creative, what is it helping engender in terms of mathematical activity as well?
Yeah. Okay. So first of all, you characterized it as inhuman in some way. I actually think the argument was very human-like. I think OpenAI released a chain of thought, and it was very recognizable. If I tried to imagine my chain of thought in trying to solve a problem, it might look kind of like that.
Yeah. We haven’t seen the raw chain of thought.
Maybe they cleansed it a little bit.
Yeah.
Or maybe the model likes to swear a lot in the middle of its chain of thought, and they clean that up or something. We don’t know. But at least the summary seemed pretty readable. I would say that’s actually typical of most of the results that I’ve studied. They don’t seem inhuman at all; they seem absolutely like something a human mathematician could produce.
They’re typically understandable, if not so well written, if you just look at the raw model output. It’s not like there’s some “move 37” or whatever. It’s like a human mathematician doing math. It’s like a human mathematician doing certain types of math.
The models are still not very good at some mathematical activities. Here, I don’t mean by field, but rather certain things you do when you try to solve a problem. The models don’t seem to be doing some of them, but they’re very good at certain other things. The ways they might be a little bit inhuman are that they don’t get tired and they know a lot. But if you have to just read the final output, it doesn’t seem inhuman.
It’s interesting because, obviously, we can’t get too much information from the labs producing these models about why or how the training recipes work or how they’re advancing in reasoning. But at least one of the things we do know, and I think OpenAI spearheaded this, is that reasoning in natural language is actually what they happen to scale up, and it’s not actually pushing a lot of Lean-verified proofs as the training corpus. That’s kind of amazing.
On another side, what’s interesting is that it’s not clear that a lot of the mathematical training data, if they use it to a large extent at all, is reflective of how mathematicians think, if that’s fair, because a lot of it is not legible as traces, right? Most papers are crisp and polished. Textbooks certainly just show very little motivation for how something is developed, which is why it’s usually easier to follow a research direction by actually talking to the researchers and hearing how they’re thinking about it.
So I’m curious, before maybe even going to the taxonomy as an excuse, when you’re examining these models and their results, comparing Anthropic with OpenAI, do you detect any differences in how that is similar to human reasoning? And then also, if you have any comments or insights on perhaps why natural language scales so well that way, even though it’s—
Okay.
So, first of all, I really like your point, by the way, that they’re mostly doing natural-language reasoning rather than Lean. I think that suggests to me that these—you hear a lot of people say math is a verifiable domain, like that explains the progress, whatever—but my sense is that because they’re primarily scaling in formal reasoning, probably the techniques are going to generalize to other domains pretty well. That’s just my guess.
Okay.
You asked a little bit about Claude versus ChatGPT. My sense is that they’re pretty similar in terms of capabilities. I’ve played around a lot more with ChatGPT than with Fable, but it seems like there are a lot of cases where OpenAI will drop a solution to some problem and Anthropic will say, “Oh, we know, too.” Exactly.
So it actually seems like they’re solving a very similar collection of problems, and it’s a relatively small portion of what human mathematicians do. One thing that’s interesting is that we see the solutions have a certain flavor, right? There are things that rely on the models’ strengths, like their ability to grind out a long computation, pull together technical ideas from many areas, or draw on many papers that a human mathematician might not have read.
But they seem weaker in things like intuition or having some big-picture point of view. A lot of what I do as a mathematician is have some kind of philosophy that’s very nonrigorous. Maybe I think one thing is kind of analogous to another thing, and then a lot of what I’m working out is figuring out how to make that precise and trying to measure the extent to which I’ve succeeded in understanding it. Can I solve a problem? Can I find an interesting phenomenon that I don’t understand, and then come to understand it?
So far, even in the best results the models are producing, you don’t see that much of this kind of reasoning. It’s more like they’re very, very good at applying some known techniques.
To be clear, that’s a very powerful thing to do—to be very good at applying all known techniques.
There are mathematicians who have had great careers doing very high-quality work of that flavor, and I think a lot of what the models are producing is high quality in that way. But it’s some kind of fairly narrow band of what mathematicians care about so far.
I do think there are signs of all the frontier models starting to be able to do more fuzzy things. I’ve tried to get both Fable and ChatGPT 5.6, I guess, to do some kind of theory building. They’re not good. They’re definitely not good at doing it autonomously, at least with whatever scaffolding I’ve set up. But with some hints, you can kind of get them to do something interesting.
When you give the models hints, it’s always a little hard to tell what part is from the model and what part is from you. But my experience is that if they can do it with a hundred bits of hints or whatever, in 6 months maybe they can do it without hints. I do think there are signs that they’re also somehow picking up some of this implicit and unwritten mathematical knowledge.
4. Intuition, Taste & Why Math Isn't Just About Proofs
Okay. I would love to go into the intuition part and where it sucks, to put it in a very basic way. But you actually mentioned a small detail, which is that from the public information, you’ve gleaned that Anthropic and OpenAI are probably neck and neck, but you personally are using a lot more ChatGPT. Why is that? Or 5.6?
Well, why is that? I don’t know. I mean, I just think it’s inertia. I’m pretty sure of it. One thing is that ChatGPT got better at math earlier. For a long time, the Claude models were just not useful for research math, and then I think maybe around Opus 4.5 or Opus 4.6 they more or less caught up.
Some experimentation suggests to me that they’re pretty neck and neck, so for my own work, except when I’m just experimenting, I mostly just stick with one.
As people outside of labs like us, it’s really interesting just to compare how they differ at the frontier. And to your point, it might be a little bit of momentum. I do think that, at least from my anecdotal experience, 5.6 has been a lot clearer in exposition. Obviously, the models are just incredibly jagged at the frontier, and this might not be true, but in the explanations of results, I always find that 5.2 is giving a more accurate theory of mind of what it assumes I know and don’t know.
Whereas Fable might be explaining something very trivial, but then just jump to, “Well, you know, obviously you should know these things.”
I’m not sure. I find they’re both pretty bad at theory of mind.
Okay, great. So, from your perspective, you’re probably asking much deeper questions. The thread that I really wanted to pull on was when you were talking about the models maybe starting to get better intuitions or even theory building. Before we even dive into that, it would be useful to talk through what your primary activity as a mathematician is, especially in your area of algebraic geometry, which probably has a very different flavor than a combinatorialist or someone in some other area.
If you can give a brief lay of the land, and then explain what your mathematical activity was pre-AI and maybe how it’s changing with AI.
I think there are a lot of different kinds of mathematicians. There are a lot of different spectra on which one can put a mathematician. Definitely, a lot of mathematicians like solving open problems.
Mm-hmm.
I’m one of those. I like to solve an open problem. I think of myself as a problem solver. Another taxonomy you could have is a problem solver versus a theory builder.
Mm-hmm.
At least for me, the point of an open problem is that it’s supposed to measure your failure to understand something. It’s kind of like a benchmark. One problem I really like is the Grothendieck p-curvature conjecture. It measures something about our failure to understand differential equations. There’s some very basic object we would like to understand, and if we can’t answer this conjecture, we know we don’t understand it.
Mm-hmm. Okay.
So in practice, how do you get at a problem that’s supposed to be measuring something you don’t understand? Of course, you try to understand the thing better. In practice, that means you try to find the smallest situation where you can’t understand something, fiddle around with it, and then, once you win, stare at what you developed to win and try to turn that into some theory.
That’s one thing you might do. You might try to solve a problem and, in so doing, develop some kind of new understanding of the situation. You might also just have some feeling that this thing is related to this other thing. You might start building a table: property A is related to property A′, property B is related to property B′, and so on and so forth.
For example, in my work, a lot of it is motivated by an analogy between the homology of algebraic varieties and representations of fundamental groups.
Okay, so that’s some fancy stuff.
This analogy is very fruitful. Really, any phenomenon that appears on one side, you can find an analog on the other side. Trying to realize that dream has led to a lot of beautiful mathematics over the last 30 or 40 years, by people like Carlos Simpson and Takuro Mochizuki and others.
Here, someone noticed that there was an analogy, and then that analogy led to a huge amount of development. There’s not really an open problem at the end, although, of course, as you develop this, you come up with lots of open problems. There’s just a philosophy that you’re trying to realize, and that philosophy is super nonrigorous, actually. It’s not symbol pushing at all.
That’s another kind of activity that I like. Beyond that, a lot of what you do when you try to understand something is actually trying to figure out what the right question is. You have some object you feel like you don’t understand, and figuring out what you don’t know is actually a very challenging thing to do.
You can go online and find a list of open conjectures or whatever, but this doesn’t really capture, in a lot of ways, what we don’t know. Often, finding the conjecture is really, really hard. A good example of this is the Birch and Swinnerton-Dyer conjecture, which is one of the Millennium Prize Problems. It’s a beautiful relationship between L-functions of elliptic curves and the set of solutions to the corresponding equations—the rank of the group of solutions.
That’s sort of what that means. It was discovered by Birch and Swinnerton-Dyer, and it was the first big-data conjecture. Birch and Swinnerton-Dyer had found all these statistics on elliptic curves in the 1960s. This was one of the first ever computer-aided bits of mathematics: they graphed these statistics and noticed that the slope of some line on a graph was related to some other algebraic invariant they knew. That was the source of this conjecture.
A lot of the time, you’re just working out examples and trying to figure out an explanation for that experiment. Going back to AI, I think these are things where AI seems, so far, to help a lot more with some things than with others. The more vague a phenomenon is, or the less you have a precise question in mind, the less useful it happens to be.
You were asking how I use it in my daily life. What I’ve found is that the projects I have that predate AI—the projects I’ve been thinking about for 3, 4, or 5 years—it’s just not that useful. It’s primarily a substitute for Google or something. I might use it to learn about some related topic, or something where I would have earlier Googled something and then read a paper; maybe I’ll discuss it with AI instead. It saves some time, for sure, but it’s not really doing deep intellectual work for me.
But I’m now—well, I suck at coding, and so now I have my good friend who’s really good at coding. Now I have all these coding projects, because suddenly, if I had a question where coding would have been really useful, I would have procrastinated on it for 6 months until I—
A tireless PhD student.
Exactly. So, yeah, now I’ve picked up all these projects where coding is really useful. The models are very good for massively parallel things. If you want to find an example of something, you can ask it to work through 1,000 examples in parallel—or maybe 10 examples at a time, with 10 different subagents—and that’s really useful. But these are different activities, which are kind of on top of what I was doing before.
I’d love to dig in.
Yeah—yeah, go ahead.
5. Deep Thinking vs Pattern Matching: What Models Are Missing
Yeah, sorry, because you were mentioning that there are projects you’ve been working on for 3, 4, or 5 years, and I don’t know if it’s correct to say those are more of the theory-building aspect of it, because you did characterize yourself as an open-problem solver. What is the thing that drives you? Because the deep-thinking part, for this audience, might also be useful to explain. Most of pure math isn’t motivated by anything external, whereas applied math has at least some external motivation for why a certain formal structure might be interesting to study.
Whereas pure math seems almost sociological. To some extent, I think Thurston made some comment or point about this in the 1970s: that it is a sociological phenomenon. As more and more mathematicians start examining something, you’ll maybe converge on some interesting structures. But it’s not—there’s some reason why people would prefer to study something, or think it’s beautiful. What is it that drives you in particular, and then maybe you can make a more general comment about the profession?
Yeah. I mean, definitely some people are motivated by beauty or aesthetic considerations. I try not to be motivated by that. I think that’s a bad way.
One thing is that it sort of limits you, right? One failure mode I see among young mathematicians sometimes is that you have something and you think you know how to prove it, and then the proof feels really ugly, and you decide—well, what if you’re wrong and it’s not ugly? Why limit yourself? You should—
For this audience, what is “ugly”? I have an intuition of what’s ugly, but what does that mean? Spell it out.
I don’t really know. People sometimes feel this way. Maybe it involves a lot of grinding calculation that’s not illuminating.
But you should win by any means necessary, in my opinion. I like to think of what I’m doing as doing kind of physics, except with concepts. Instead of beauty, I try to think about what is fundamental and what’s going to open up further understanding most. Am I introducing a new idea that will be broadly useful for understanding this object?
I guess in some sense there’s an aesthetic consideration there, but I try to think that the orientation of trying to do good science rather than trying to do art is what I prefer. There’s a huge variety of opinions here, and lots of mathematicians think of themselves as being closer to poets or something.
What I do think is broadly true is that progress in mathematics seems to come from people pursuing their personal curiosity. It’s been crucial historically that there are a lot of different people with different views of what’s interesting. Then the frontier of knowledge expands and expands, and suddenly you have these opportunistic situations where a new idea has been introduced and you can suddenly cascade through a bunch of other things that we didn’t understand before.
Yeah. I mean, going back to the 3-, 4-, or 5-year problems and where the models are not useful, let me ask it another way: when you’re doing the deep thinking, is it just that it isn’t clear that you formulate it as a problem, and more that you’re thinking about—what are the fundamental physics of—
Yeah, so in some cases there is a well-stated problem here. I’ve been thinking about things for maybe 10 years now at this point. For certain problems, sometimes there is just a conjecture that I would like to prove.
I think one reason the models might not be useful for some of these things is that the conjectures are true. For example, the general belief in the community was that the unit distance conjecture was true, and then it turned out to be false. What that means is that there’s a specific construction you can do to refute it.
On the other hand, a lot of the things I think about—I know, maybe someone will come up with a counterexample to the curvature conjecture tomorrow and I’ll look like a fool—but in general, these conjectures fit into some very broad theoretical framework, which means that we actually have a lot of evidence that they’re true. There’s not a construction you can do to refute them. You need to somehow—there’s this giant framework where certain pieces of it are only conjectural, and you probably need to resolve some of those conjectures to win.
We also have a pretty good sense that very serious new ideas are needed to resolve those conjectures. Of course, you can’t be sure. Maybe there’s some very clever construction that will let you avoid having a big new idea. It’s quite possible we’ll find this out.
My sense is that, for at least a lot of the things I’ve been thinking about, they’re just not accessible to applying known techniques in a very technically strong way. You need to develop a new technique. To be clear, I’m not saying the models won’t be able to do this; it’s just that so far they seem not to.
Actually, that’s exactly the point I wanted to delve into. To your point, we can try to calibrate and forecast why they would get better at this, but it is true that, when it comes to providing a construction for a counterexample, they seem to be strong. If you have to start developing either new theory or, to your point, techniques or machinery to prove why a conjecture is true, they struggle more.
6. Why the Unit Distance Result Was Actually Creative
It’s probably because a lot of what they’re drawing on is just techniques that have happened in other areas, and they’re porting them over. To your point, that’s why maybe the unit distance problem was such a more creative result, because it was maybe doing more of that on its own. It was drawing in something from an unexpected area.
Yeah. Exactly.
And so, maybe to ask the question: it’s not that, because AI is getting so good so fast, we can count it out. But why? What do you think it has to— I guess you’re spelling out what it has to get better at, but maybe you have some more feelings on why it’s particularly hard to develop that new theory and technology.
Yeah, it’s a good question. I think you just need a different—my guess, actually, is that it’s probably totally doable. It just hasn’t been done yet. Maybe you just need a different RL environment. I don’t know. At this point, my expectation is simply that the trajectory will continue upward. I’m not a skeptic of continued capabilities growth.
I think what is definitely true is that the skill of developing a theory, or building your understanding of some poorly understood object, is a fuzzier one. Mhm.
So it might be harder. You can tell it, “Develop your understanding of zeta functions,” and then, once it proves your hypothesis, you give it a reward. But it’s harder, I think, to come up with intermediate things that you can reward.
That said, mathematics as a whole provides a lot of conjectures of varying levels of difficulty. Maybe this explains why there seems to be a little bit of progress in these areas. Presumably, they’re trying to get it to solve lots of problems, and some of those problems develop at least some of the skills of theory building.
Humans are able to develop these skills. Sometimes they get rewards from their PhD advisers, and their adviser says, “Oh, that’s a good idea,” based on some element of human taste or whatever.
That might be something one can do too. My expectation is that, as part of continued capabilities growth, we’ll see growth in these areas too.
Yeah, yeah. I like the framing where you’re casting these increasingly difficult conjectures as a form of curriculum for both humans and, of course, AI. Maybe this is getting a little too philosophical for some people’s tastes, but it gets at the question of why we’re particularly good—or maybe particularly bad—at math, too.
What is it that we’re either struggling to do, or that some people are particularly good at, when they develop new theory? It’s not really—or maybe it’s related to—why we can formulate good structures for physics as well. It’s not obvious that all parts of the world are understandable and legible in that way, but some parts are, and therefore we try to do it because there are maybe compressive pressures on our minds. We need to—we can’t understand anything without compressing stuff more finely. I don’t know if that’s also the correct interpretation of—
I think that’s—I mean, I’m a little bit skeptical of compression as a metric of interest, but it’s definitely an anthropological reason for why we do a certain kind of theory building. I just think it’s true. In fact, I think our inability to just grind is kind of important to our ability to make discoveries.
As an example, I have 1 paper out so far where the models were kind of useful. They proved a couple of lemmas. This was a situation where we had kind of proved the main result, and then there were some lemmas I was unhappy with. They seemed not to be optimal, so I worked with Gemini Deep Think—which at the time was also on the frontier, though it no longer is—to prove the lemmas.
There was a lemma I wanted to prove, and the models couldn’t do it. None of the frontier models could do it. I worked out a ton of examples on my own and realized, “Oh, well, maybe here’s some reason why it could be true.” I found a better statement of the lemma. Once I had that statement, I probably could have done it pretty fast, but the models were able to very quickly prove that better statement.
Our inability to prove it led to an improvement in the result. Now you can take the original lemma I had that the models weren’t able to prove, put it into ChatGPT 5.6 Pro, and it will output the worst proof you’ve ever seen: 10 pages of brutal calculation with no insight whatsoever.
This would have been a perfectly fine proof, but it would not have led to the discovery of what I think is a beautiful conceptual explanation for why this thing we discovered was true. We found a better proof because we couldn’t do—I mean, I say we couldn’t do the calculation, but what actually happened is that I realized this horrible grind proof would work, and I could not bring myself to do it. I looked for another argument. Now the models can do very long technical calculations pretty reliably.
Yeah, yeah. Well, I mean, to your credit—I don’t know—I think I understand why you’re saying: don’t shy away from trying to prove something just because it seems ugly. You have to take the first step, and eventually you’ll try to work toward insight.
You might be shy about admitting it, but it is still maybe driven by aesthetics or just pure—I want to understand. If understanding just means it’s a little bit simpler or more compressed, then I’ve gotten understanding. It’s a deep philosophical question, too: What is it?
Right. I mean, of course, there’s a reason we do informal mathematics rather than writing out long, formal strings of symbols in ZFC or whatever. Besides length, we’re somehow trying to put it in a way where we’re actually getting some non-rigorous understanding out.
Yeah. I mean, and I think somehow—
Yeah. And even how that relates to why it helps. Again, is it because of our inability to grind, or is it because if we’re pushing toward something that’s more compressive, it also hopefully helps us understand other fields as well?
It’s somewhat magical that when we try to optimize for both, they tend to coincide. I don’t know if that’s a fair quantitative statement, but why that is is kind of magical.
It’s at least sometimes true, yeah.
Yeah, exactly. We certainly bias toward the cases in which it is. That’s why we call that a good theory to build.
But it’s very much, I think, what we can study and understand, and therefore it’s also very convenient that it was rich in mathematical results there.
7. How Should the Math Community Adapt to AI?
Okay, so I think there are 2 things that we can pull on. Given that you’ve stated—or not stated, but believe—that the mathematical ability of AI is going to continue advancing and can do some of the things you’re doing nowadays, how should the mathematics community best adapt and benefit from this?
I’m just saying, as somebody who doesn’t have the time to practice mathematics anymore, this is very great because I can maybe dabble more. There are a lot of results that can come out. But I can also see where you’ve made the more precise point that we cannot motivate the right kind of behavior around understanding and development, so I’d love to hear more about your views there.
Yeah. So, first of all, it is clearly really exciting that there are increasingly capable models producing high-quality results. At least some high-quality results, and also a lot of slop.
Yeah, some good stuff.
As the models get really capable, my hope is that they will answer a lot of the questions that have kept me up at night, and I’ll get to learn the answers. I think that’s really exciting. A lot of people got into math largely because they enjoyed learning math.
Yeah.
The first thing you do as a math student is learn stuff that other people did, and you do that for 20 years before you start doing—maybe not quite 20 years, but—
15 years.
If you’re lucky, 20 years. If you started early. Yeah.
Yeah. Yeah.
That said, the goal of mathematics is not to produce mathematics papers; it’s to produce some kind of understanding. Maybe some of that understanding resides in model weights or something, but to me, that’s pretty unsatisfying.
My own motivation for doing mathematics is that I would like to satisfy my personal curiosity. I think people should have the capability to do that, and that requires a pretty substantial apparatus. It’s simply not the case that you can study the questions I think are fundamental unless you’ve invested a huge amount of time and effort getting to the point where you can meaningfully do so.
Moreover, even the people—you have the small group of people doing fancy research mathematics on the frontier or whatever—rely on a huge apparatus of thousands, millions, or billions of people who are trying to learn to think mathematically. You need an entire mathematical community to support a small group of people who are on the frontier. Without the pipeline, the pipeline doesn’t exist.
So if you think that’s important—the development of human capital that can meaningfully engage with frontier mathematics—I think you still have to incentivize those people to invest their time and effort in getting to the point where they can engage, and then do so in a high-quality, meaningful way. Right now, I think the existing incentive structures for math research do not encourage people to do that.
Right now, if you’re a postdoc on the market and want to get a job, maybe for the next couple of years before the community adapts, the best way to do that is to produce a lot of papers that perhaps prove old conjectures or whatever. You can do that by playing the slot machine until the model produces a hopefully correct proof of such a result. You don’t even have to pick the theorem in advance.
Here’s an experiment you can do: You can take Codex and say, “Go online, find 5 recent conjectures in algebraic geometry, and prove them.” I’ve run this experiment, and with some back-and-forth, I was able to get 3 quite bad but correct papers in an hour. They’re now sitting on my hard drive, waiting for me to email the relevant people, but this is not a good use of my time to investigate these results.
You definitely see people doing this. There’s been a huge uptick in arXiv papers, mostly not very interesting. Some of them are interesting.
But a lot of it is clearly low quality, even if the result is something that would have been highly rewarded a year ago. There’s just no evidence that a human being is actually engaged with it. There’s no development of human capital or understanding.
Sometimes we’ve seen examples where 3, 4, or 5 papers with the exact same proof of the exact same theorem have come out within a couple of days of each other, which is clearly a situation where someone is playing the slot machine.
ChatGPT is consistently finding the same thing.
But that’s also interesting. It’s kind of mode-collapsed on certain paths of reasoning.
Yeah. And I think it’s not obvious that this problem goes away as the models get better. Maybe it does, but maybe it doesn’t.
A lot of what I was saying earlier is that a lot of progress in mathematics comes from letting a thousand different flowers bloom and allowing people to pursue their own curiosity. Then the boundaries of knowledge expand in a hopefully fairly uniform way in this really high-dimensional space of mathematics. Opportunistically, you suddenly get some applications or answers to old questions.
It’s not clear to me that if we subordinate mathematical exploration to what the models want to pursue, what we’re getting is one mathematician duplicated a thousand times, or actually a million different mathematicians doing a million different things.
Aside from math having to adapt by changing its incentive structures, I think this is perhaps dangerous, or at least something the labs have to pay attention to. On the one hand, what’s been successful for them is that this emergent reasoning capability is obviously incredibly powerful, and it’s creating a lot of great PR headlines.
But to your point, and also where my interests tend to lie, human mathematicians are coming from all sorts of weird intuitions. That is why you end up developing a frontier that is so diverse that you can actually draw from it, make connections, and produce a lot more.
If it’s true that most of these proofs being pushed out by the labs are converging on very similar things because they’re technically drawing on the same body of literature—and that’s where they’re strong right now—it’s not clear where they’re developing their intuitions. They’re developing them from practicing mathematicians, but it’s not clear where that other, emergent, stronger, diverse intuition might be coming from. It might come, but I don’t have a good theory for where it comes from.
It’s also not clear that, with test-time compute and various post-training scaling, you can even introduce new capabilities. There’s a huge debate about whether you could introduce new capabilities, and math is precisely the context in which I think this should be studied: where it’s good, where it fails, and where human mathematicians are good. That’s exactly a precise question to study.
So this is a long way of saying that it’s unclear to me whether we’ll produce a superior result without human aid if we don’t incentivize enough people to interact with the models. Unlike other domains, it might not actually continue producing a superior result without human aid.
Yeah, let me push a bit further. Let’s suppose the models become really robustly superhuman, even without adding meaningful cognitive capacity. Okay.
I claim that we still want human mathematicians.
Okay. Great.
Human mathematicians. So why? It’s because there’s a question about how we’re designing society, right? Maybe the optimal situation is that you have the models doing all sorts of math research, and that leads, in this highly nonuniform and diverse way, to lots of applications and improved understanding. Maybe you don’t need humans to add that kind of diversity.
But just because something is optimal doesn’t mean you do it. There’s no reason to think that if we hand over control of math research to the models and just let them do their own thing, they’ll do the optimal thing. In fact, if we instrumentalize what we want them to do—if we say, “Make our lives better,” or whatever—it might not be the case that what they decide to do is pursue a wide variety of interesting research. They might just try to take the direct path. We don’t know what’s going to happen.
I don’t know if you believe there’s any value in this broad-based fundamental research, which I do. I think that’s one of the most valuable things that humans—or whatever the models can be doing—can pursue. I think the easiest way to guarantee it happens is to keep a community with broad interests that is pushing the models to do it and helping us design a society where that’s what we’re pushing for.
Yeah. I think at least my hopeful vision for the future is that humans are not totally disempowered. We have some control over where we’re going, and if that’s the case, what we end up doing is going to be driven by human interests.
You want to have people who have lots of different interests and also the capabilities to pursue them. You want people who are smart and engaged, who can do not just mathematical thinking but all sorts of thinking.
Yeah. A great fear is that, as AI advances, we don’t develop the right ergonomic interfaces to encourage us to continue to be good thinkers. It’s so easy to relinquish that because you just hand it off, and the models aren’t even good at that level of thinking where it’s the top-level structure. Despite that, it’s so easy to do so.
8. Where AI Will Impact Applied Math First
Especially for math, another selfish reason is to math-max. I think it’s actually a great pedagogical excuse to get really rigorous about thinking through various things. This is not why mathematicians do it, but as somebody who is part of what we do, I think it’s a great framework for thinking about many things, not just mathematics.
As somebody who is now also a parent of a 2-year-old, I think about this a lot, too. It’s not about grinding. Grinding is still good, but I don’t want to undersell it too much.
I used to collaborate with some Hungarian mathematicians, Béla among them, and I heard that in Budapest they would just teach group theory in primary school. We should definitely do that. We should continue doing that.
Now that AI is somewhat good at explaining and is far more accessible, we should probably proliferate that even more. Maybe that helps bring more people to the frontier rather than just—
Yeah, I mean, this is something I’m concerned about. Of course, I mostly talk about math because that’s where I live, but one nice thing about thinking about this is that we’re one of the first professions to have a significant impact from high-quality models.
Although I think maybe we're one of the first professions for it to happen so publicly. My sense is that there are plenty of other professions that are—
100%.
Coding. But also, anything you do at a computer—probably a huge amount is being done by the models at this point—and there isn't a public reckoning about it.
Your math capabilities are useful for the company and the labs to talk about, I think, probably a bit more publicly than everyone else. But, yeah, I think one reason to try to maintain human capital in this area is that it's a model for all professions. Presumably, we still want people who are meaningfully engaging with the world, who are experts, and who have talents and trained skills.
It's actually very convenient that the math profession is so entwined with education here, because I think we're also seeing some challenges among college and secondary education coming from AI, too.
So, as you said, it's also an amazing tool to learn.
People have started talking about a bimodal distribution in their classes. There are some people who are really figuring out how to take advantage of new tools, and other people who are just letting them do their homework and then bombing everything else.
Unfortunately, I don't think that adapts fast enough. We definitely want to be living in a world where we're producing better thinkers. I think it's just—
Even if we talk about the pure optimization game, I think that's better for us. And so, as a human being, I'm like, that would be pretty—
9. Taking Advantage of AI Without Losing the Craft
It would be inconvenient if we became worse thinkers just as AI ascends. It's too easy to let that happen. So we should be thinking hard about how to take advantage of this and harness it for our own improvement, right? Well—
Yeah. I think it's interesting to observe that, at their current level of capabilities, the models let you do a lot more. They let you do things you wouldn't have done before, cheaply enough to do them now. But it's not clear to me that, in many cases, they're actually improving the quality of outputs.
Yeah. I think this is common: you have a new technology that's doing something a little bit worse than what was previously done, but much cheaper, and so you suddenly get a lot of low-quality outputs that are displacing previous high-quality outputs. But I think it's possible to use the tools in a way that actually improves quality along every dimension. It just requires some thoughtfulness and some redesign of institutions to incentivize that.
Yeah. Well, hopefully capitalism works there. I do feel like the most high-value things require people to use AI effectively, and right now the models are not good enough without human experts participating. But to your point, there's a vast majority of perhaps more junior and entry-level roles, and it's harder for those roles to adapt as well. The thing that would be a mistake is to use the AI models in a way that doesn't—
Basically, you need to be ascending and using the models to deepen your understanding. It's so easy for human nature to make us lazy, and you have to resist that, because that is the moment that you will lose, basically. Forgive the very competitive language, but it really is that it's so easy to relinquish the thinking to the models.
The models can't really think. As things are ascending so fast, it's critical that you continue developing those faculties and actually leverage AI to improve those faculties rather than relinquishing them.
Yeah. Yeah. I totally agree. I know I get to talk—
And not even kind of a positive, obviously.
10. Comparing Anthropic vs OpenAI in Math
Yeah. No, exactly. Suddenly there's a spotlight on math, and I can nerd out about it more. I was actually kind of curious: have you got any comments on the—
I guess this is another constructive result: the elliptic curve of rank 30 that just came out yesterday. So, if that—
That's right.
Yeah. Yeah.
Well, we don't have any details about it yet.
I know. There are rumors.
Yeah, we don't know how it happened.
And a collaborator whose name I unfortunately forget. Maybe you can add it in post.
I just know the Twitter handle.
Yeah, we can.
A lot of these nice recent results have come from Levent Poge, with unclear amounts of autonomy. My sense is that some of them are semiautonomous rather than fully autonomous.
With this one, we don't know anything about the methods. Without knowing about the methods, it's very hard to say how significant it is.
Yeah.
What I would say is that, historically, such results have been understood by the community as cool, but I wouldn't say they're a big deal. The typical place a result like this might go is a website of records—
Like, not—
It's not an Annals-level result.
Yeah, it's not an Annals-level result. But that said, it's cool, and there definitely were a few very, very talented mathematicians who like these kinds of questions, with Noam Elkies being maybe the most famous example. Elkies and Klagsbrun are the ones who have been pushing this record for a while, and they recently found a rank-29 example—
Yeah, you know, so those are mathematicians. I think people like it, and it's cool that the models can do this sort of thing. I think Ava Howell may be the third collaborator.
Okay.
11. Why Some Labs Have Gone More Secretive
They have not yet told us how they did it.
Have they been mostly more secretive? I think they've released some traces for stuff, but—
Yeah, so for this one, I think they haven't yet, unless I missed it. Levent likes to tweet out his results, but he has been slowly releasing some kind of PDF write-ups, too. I think he's just having fun on the internet.
Yeah. Yeah.
It reminds me of when you were saying that you have to evaluate how the results came about. I had Mark Sellke and Mehtab Sawhney from OpenAI on recently, and they were saying that what's been most charming or delightful is that the proofs have been relatively short. They're not 200-page proofs, which maybe corresponds to your grinding point as well.
But maybe this is optimized for—in retrospect, it was picked because it was short—or maybe, on average, the stuff that you throw at GPT, Saul, or Fable tends to be shorter and more legible to humans, rather than just going haywire and grinding it out.
Yeah. So, I think it is nice that they'll sometimes produce short, clever proofs. Of course, everyone likes a short, clever proof. I think my sense is that the reason they're not producing long, complicated proofs is that they cannot—
The ability to check correctness is not yet there. So, even if you ask the models to produce—
For a short proof, you can then ask them, “Is that correct?” They will often say no.
They're much more reliable than they were 6 months ago, for sure, but they will still sometimes produce things that are just wrong—and they know they're wrong.
Yes.
I think the problem with producing a very long thing is that they might not know they're wrong. And so what I wonder is, presumably internally, OpenAI and Anthropic have probably solved a lot more problems than they've released.
I imagine quite a few of them are just ones they're not sure are true.
So, for example, with this recent list of 10 problems released by OpenAI, those were all formalized in Lean, which is, of course, very good evidence that they're true. I have no doubt that there were a lot more that they could not formalize in Lean because the prerequisite results have not been put into Mathlib yet, for example. There were probably more that were longer, which also makes them challenging to check.
Yeah.
Yeah, so this is my guess. We do actually see very long generated proofs on the arXiv. For example, someone recently posted a claimed proof of resolution of singularities in positive characteristic that was 800 AI-generated pages. It's definitely—I mean, I'm sorry, I haven't read it; I haven't gotten far—but there's no way it's correct. This would be a major result. It's just not within the capacities of the current models, if you're reasonably well calibrated.
Yeah. Yeah.
It's definitely the case that no human has read it. The models are definitely not able to check this kind of thing. So, well, I mean, yeah, this is my expectation: the reason it's producing short, clever things is just that that's what we can check. You can get it to produce long things that are grindy or hard to check, but then...
Yeah, yeah, we'll have to get there. This is actually more what it says about frontier capabilities rather than, in a perhaps more negative way, "Hey, it's just so good at these clever things." No, I think that's totally fair. Especially if you look at an adjacent domain like code, right? I remember reading something that Cursor put out about testing their long-horizon harness.
In this case, they're trying to reproduce SQLite in Rust, and it was just so telling how far we are from that. It seems like a very comparable task: it's very long, and you have to make sure there's something to verify that it's a correct implementation by testing a suite of cases. You actually need a harness there; it's not just the raw models. It takes a while, and you can actually compare different frontier versus non-frontier models, who's the planner, and so on—the differences in their capabilities as well.
Yeah. I also think in practice that, to elicit a long proof, you kind of need a harness.
Yeah.
When you make a harness whose goal is to elicit a proof, I think it often decreases reliability because you're just trying to produce output. ChatGPT 5.6 Pro is very careful—it really tries not to say wrong stuff, for example, although it happens, and then you'll ask it, "Was that correct?" and it'll say, "No."
But when you're trying to get it to elicit something, you try to get it to be creative. You try to get it out of this very rigorous rut in order to actually get it somewhere. If your interest is in—there are a lot of people who are trying to arbitrage the prestige mechanics of academic mathematics and elicit a lot of proofs which are not necessarily being checked—in order to do that, I think you just decrease the reliability to get a lot of stuff. Any harness that can elicit a 250-page paper is probably not being very careful about what it's producing.
And you're saying that it's decreasing reliability—or it's not reliable—just because the capability isn't there yet, so it's forcing a longer-horizon task on it.
Yeah. In practice, how does a human check it? A human can't check a 250-page paper either. You can't reliably check it line by line. What you try to do to understand it is understand the overall global structure of the argument and stress-test it in various ways: Would this argument apply to something else that I know to be false? Does it work in this special case?
The models seem not to be able to do that kind of more—I don't know—fuzzy unit testing of a proof very well yet. Actually, one of my favorite tests for the models, which they haven't succeeded at yet, is this: There's a paper—I won't name it—that came out a couple of years ago that was wrong. It was in a close enough area to myself that I immediately downloaded it and started reading it, and it was very hard to find the actual specific error, but it was also very clear from the structure of the argument that it couldn't work. It was, you know, me and a bunch of other experts arguing something that was too strong to be true, if you took that a little bit further.
Yeah.
A bunch of other experts and I immediately realized it was wrong, and we emailed the author. We went back and forth until someone figured out what the precise, specific error was.
Okay.
So far, the models seem not to have been able to do this. The specific error is quite subtle, but they're also not able to do this kind of overall, big-picture sanity-checking, which is how people check papers in practice.
Yeah. Yeah. It mirrors why, in code, it's so clear that it's good at the syntax, but higher-level architectural stuff is still very weak. Maybe it'll get there, but it probably needs some harness help. Who knows? People have evolving opinions on how much the harness and the model have to co-evolve, and which one is necessary, but the next model requires less. I feel like, in math, it would be very interesting to see how you—if you do any experiments with the harness there—how that improves, because it is a mark of general reasoning. Yeah. Yeah.
Yeah. I have my own sort of bad little harness in Codex and Claude Code, but I personally do not enjoy autonomous mathematics very much, so I mostly do not use the harness. I mostly try to use it to help me understand stuff.
12. Raising a Mathematician: Teaching Math to a Toddler
Okay. No, it's totally fair. Exactly. You don't want to automate your job away, because that does involve you being in the loop to understand it, which is necessary to participate. Maybe to finish off, I'd love to ask—and this could be just something you haven't thought about or have actually thought a lot about—I think you also have a toddler, right?
Yeah.
Yeah. And so you have a three-year-old?
Yeah, three-year-old.
Great. Right. So you're one year more advanced and probably have more thoughts on this. How are you thinking about her education in math—not to grind, but really just how to pass on the love of it—and how to react to AI?
Yeah. So she's three. She's never used AI. She is starting to add; that's about as far as we are in math.
That's more than she can do: add single-digit numbers by counting on her fingers and count up to maybe 30 reliably and 50 semireliably. So I'm very proud of that.
That's good. Yeah.
Yeah, I definitely encourage that. We talk about shapes and stuff. A couple of days ago, I woke her up and she was hiding under the blankets, and I was like, "What are you up to under there?" She said, "Oh, I'm doing some math."
Oh, I did. I think you tweeted about that. That was adorable.
Yeah, it was great. So I think she has some sense that I like math, and she's into it because of that.
I don't know. I think the world is probably going to look pretty different in 20 years, or whenever she's fully adult and doing her own thing. But a lot of what we educate people for is pretty robust to changes in the nature of the world. I think the reason to learn math has always been to think clearly and better understand the world, and presumably that's something you want to do even if there are extremely capable AIs.
This is also true of reading a lot of books and doing the humanities and so on. I think the actual values of the math profession and education more broadly are things we definitely want to try to preserve. I hope to instill that in my daughter. How much institutions have to change to make sure that happens is maybe an open question I think about a lot. But, at least at a personal level, I'm definitely trying to convince my three-year-old that math is super cool.
Oh, yeah, 100%.
One of her first words was "icosahedron."
Was what?
She has a little icosahedron toy. My parents gave it to her when she was one.
Not literally a first word, but she learned the Platonic solids quite early.
Oh, very good. Very good. Very fun.
Next up, group theory. I mean, it's very natural. That's right. Yeah.
Actually, when I teach her addition and subtraction, we'll do it in the context of a general group.
Oh, very good.
Well, at least you give the motivation. I think a lot of people probably skip that part. Maybe this is how math grad students can also focus on training the next, much younger generation to use AI in service of actually getting better at math rather than—
Just lacking understanding.
Wonderful. Okay. Well, thank you so much, Daniel.
Thank you. It was a lot of fun.
Yeah, it was a lot of fun. I think there's going to be a lot more progress very soon. I'd love to maybe catch up and chat again.
Sounds great.