[BidClub_]
Latent Space · · 92 分钟

🔬GPT‑5 如何推导理论物理与量子引力新结果——Alex Lupsasca,OpenAI

Alex Lupsasca

YouTube
TL;DR
  • Alex Lupsasca 认为,前沿物理已经从AI辅助跨入了在边界明确的研究任务上达到超人水平的阶段。 ChatGPT o3 将一项原本需要数天的计算压缩至11分钟;GPT-5 在30分钟内复现了他最重要的论文之一;近期模型还解决了专家追踪一年的问题。他对这一里程碑的判断是绝对的,但对其适用范围保持条件限定:“至少在某些方向上,AI已经达到超人水平”(“at least in some directions AI has become superhuman”)。

  • 这项旗舰成果把随粒子数阶乘式爆炸的胶子计算,替换成了规模只随粒子数线性增长的公式。 人类发现,在特殊时空设定下,当粒子变得共线时,单负螺旋度胶子树振幅并不一定为零;GPT-5.2 Pro 简化了5点和6点情形,猜出了通式;一个更强的内部模型又在12小时内独立重新发现并证明了它。“他甚至还没下飞机,我们就已经解决了这个问题”(“We solved the problem before he even got off the plane.”)。

  • 第2篇论文表明,第1项成果并非偶然:公开版 ChatGPT-5.2 Pro 将这一方法从胶子推广到了数学上不同的引力子振幅。 模型选中了有向矩阵树定理,完成了多轮检查,并通过一场110页的对话产出了论文的大部分内容。团队称,模型大约在第1篇论文发布后3天就能给出第2篇;最终耗时3周,是因为人类进行了细致核验。“大部分时间花在验证答案,而不是写作上”(“Most of the time was spent verifying the answer, not writing.”)。

  • 研究经济学正在从生产计算转向选择问题、验证大量产出。 Lupsasca 说,只要引导得当,对于难度不高或类似已知计算的问题,模型可能每天产出1篇论文;10个并行对话则能侦察不同路径,快速淘汰死路。稀缺的人类贡献将变成品味、专家引导和验证,而不是原始符号劳动。

  • AI目前更像一名能力惊人的研究生,而不是能自主设定科学议程的人。 给它一个尖锐、定义清楚的问题,它就能完成前沿计算;但“优秀物理学家与伟大物理学家的区别,在于知道该问什么问题”。这道护城河正在收窄:当被要求为胶子论文提出3个后续方向时,GPT给出的基本就是Lupsasca自己的3个首选。

  • 赋予研究者“AI超能力”的同一项能力,也在威胁训练体系和信息质量。 传统上需要6至12个月的研究问题如今可以被“碾碎”,学生借此建立技术与信心的成长仪式正在弱化;与此同时,措辞糟糕的提示词也能生成包装精美的“AI垃圾内容”,充斥 arXiv。Lupsasca 的应对是提高发表门槛,而不是花一年写出30篇增量论文。

  • 验证、形式化和科学传播正在成为新的基础设施机会。 自然语言推理已经强到让形式化证明显得没那么紧迫,但大规模并行生成又把负担推回给人类,使 Lean 等工具重新变得更有价值。Lupsasca 也怀疑静态论文还能否再撑20年:他设想的将是交互式研究对象,能解释全局、展示推导,并按需展开细节。

摘要 · 为研究而整理的核心内容

1. 前沿物理从邮件工具跃迁至研究级推理

  • 仅仅1年多前,Lupsasca 还认为AI最多适合处理邮件,与那些构成其研究核心的“重要理论物理计算”无关。ChatGPT o3 只用11分钟就完成了一项他原本需要数天的计算;据他所知,普通科学软件也无法完成这项任务。

  • GPT-5 带来了更彻底的转变。只给出一个热身提示,它就在30分钟内复现了Lupsasca最重要、最困难的论文之一——他认为只有少数几个人能够完成的计算。“天啊,这改变了一切。”他回忆自己当时这样想;在休假期间加入 OpenAI,成了顺理成章的选择。

  • Lupsasca 说,GPT-5.4 又实现了一次大幅跃升,即便消费者只看邮件质量,几乎也察觉不到变化。他对 GPT-5 最初反响冷淡的抱怨很直接:“GPT-3就能写邮件。它写邮件还能好多少?重点不在这里。”前沿科学能力的进步速度,远远超过了消费者看得见的体验。

  • 研究团队如今会主动传来证据。一位通信对象称,Codex 在10分钟内生成了技术要求很高的 SYK 模型模拟,此前物理学家为搭建它苦苦挣扎。RJ 指出,许多物理学家也具备编程能力;Brandon 则强调,有能力的研究者本来就会尝试这件事,最后得出的结论是:“Codex 现在真的很强了。”

2. 散射振幅压缩了量子场论中的信息

  • Lupsasca 从量子场论必须调和的一组张力讲起:相对论绝对禁止信息超光速传播,而量子力学则认为物理量本质上带有模糊性。量子场论是20世纪用来同时容纳这两项原则、并通过概率结果描述力的理论框架。

  • 这些概率来自对称为量子振幅的复数量求平方。散射振幅规定了具有给定能量、动量和偏振的粒子相撞后,以另一种构型出现时可能发生什么;如果知道任意 n 的 n点振幅,就“大致上”等于知道整个理论。

  • 偏振引入了螺旋度:光子的横向箭头在传播时可以右旋或左旋,形成正或负螺旋度。传递束缚原子核的强力的胶子,也带有类似标签。树振幅提供领先阶贡献,圈图则增加相互作用,因此带来耦合常数更高次幂的、逐渐变小的修正。

3. 共线漏洞重新打开了教材认定为零的振幅

  • 如果所有胶子都具有正螺旋度,树振幅就会为零:这种相互作用被禁止,无法发生。教材将基本相同的论证延伸到单负构型,即其中1个胶子具有相反螺旋度,并同样将该振幅视为零。

  • 下一个情形是2个负螺旋度胶子,它 famously 不会消失。Parke 和 Taylor 在1980年代的计算汇总了大量几乎完全相互抵消的项,最终只剩下一条简洁公式。这些振幅后来被称为“最大螺旋度破缺”(MHV)振幅;但Lupsasca如今更偏好理论色彩较弱的“双色负”描述。

  • Alfredo Guevara、David Skinner 和 Andrew Strominger 发现了一个漏洞:零值论证假定粒子从一般方向抵达。当部分粒子严格共线,并在所需的特殊时空签名下考察问题——通俗地说,就是“2个空间维度和2个时间维度”——单负振幅就可能非零。

  • 人类可以表示出答案,却无法揭示它的简洁形式。3个粒子给出1项,4个给出2项,5个给出8项,6个则产生32项复杂表达式;对一般 n 而言,费曼图数量按阶乘增长。整整1年,他们都在寻找 Parke–Taylor 式奇迹简化的单负对应物。

4. GPT将阶乘展开压缩成线性公式

  • 在 Strominger 按计划到 OpenAI 合作之前,团队把5点、8项的表达式输入公开版 ChatGPT Pro。模型找到了一个相空间区域:其中1个粒子的频率符号与其他粒子不同;在该区域里,表达式收缩成3项的乘积。据称模型写了 Python,检查约5,000种可能,并自行选出了这种简化方式。

  • 6点测试更具冲击力:32项、每项都包含嵌套乘积的表达式,收缩成了4项的乘积。“哇,好吧。这真的很漂亮。”研究人员这样回应。他们没有7点表达式,因为手工展开1个这样的表达式成本高到无法承受。

  • 当被要求推断所有 n 的模式时,GPT-5.2 Pro 提出了一条复杂度按线性增长、而非阶乘增长的公式:粒子数翻倍,项数也只需翻倍。Lupsasca 认为,这就是 Parke–Taylor 公式的单负版本;但公开模型能猜出公式,却无法证明它。

  • 一个更强的内部模型接到的是经过清晰表述的问题,没有候选公式,也没有低点数答案。12小时后,它独立重新发现了同一个表达式,并给出3步推导。论文在陈述结果之后的其余部分,本质上就是该模型生成的证明。

5. 科学主张独立于不寻常的发现手段

  • Lupsasca 对功劳划分得很清楚:人类发现了共线漏洞,并证明振幅不必为零;AI找到了简洁的通式及其证明。论文标题《单负胶子树振幅非零》把重点放在物理学上,而不是把工作包装成一次AI演示。

  • 作者没有在摘要中提及AI,只在一段文字中交代 GPT-5.2 Pro 的猜想和内部模型的证明。Lupsasca 打了个比方:读者查阅几十年前的结果时,并不关心当年运行关键计算的是哪个 MS-DOS 版本,或用了几摞软盘。方法在历史上很重要,但不应遮蔽持久的物理学成果。

  • 主持人提出了一个值得保留的质疑:从展示出的5点、6点公式走向通式,对研究生来说可能看起来很自然。Lupsasca 的回答是,证明来自一场没有任何极限情形提示的新会话,构成了独立检验;但他仍拒绝过早给论文排名,因为真正的影响力要由未来几十年的后续工作来衡量。

6. 公开版 GPT 将发现从胶子推广至引力子

  • 引力子是引力波假设中的不可分量子,但没有任何实验直接测量到它们,正确的量子引力理论也仍未知。用场论语言说,胶子是自旋1粒子,引力子是自旋2粒子,偏振数据的某些方面翻倍,使数学问题实质上不同。

  • 胶子论文发布3周后,团队发布了《单负引力子树振幅非零》。他们称,论文其实可以在第1篇发布约3天后就推出,因为 ChatGPT 就是这么快给出了答案;拖延来自仔细检查、撰写完善论文和补充背景。“大部分时间花在验证答案,而不是写作上。”Lupsasca 认为,1年前这还是荒谬的倒置。

  • 这一次不需要内部模型。团队把胶子论文交给公开版 ChatGPT-5.2 Pro,强调其附录,说明引力问题需要的两处改动,然后说:“祝你好运。你是一名杰出的理论物理学家。”模型意识到,可以用有向矩阵树定理来组织新的计算——这是一套已知数学工具,但人类专家此前没有想到在这里使用它。

  • 在一场110页的对话中,模型提出下一步、完成合理性检查、处理树求和与约化公式,并反复请求继续的许可。人类的回复往往只是“继续”,或者要求进行一项明确验证。Lupsasca 将这种工作模式称为“氛围物理”(“vibe physics”)。

7. 结果何时成为有意义的科学,仍取决于人类引导

  • 人类判断依然清晰地体现在模型没有主动产生的部分。Strominger 撰写了更宽泛的引言;通用AI草稿没有解释结果为何重要。研究人员还单独发起了另一项研究,考察振幅在与天体全息和量子引力相关的对称性下如何变换。

  • 但从第3节开始,Lupsasca 说,发表的引力子论文基本保持了模型草稿的面貌,数学部分由公开版 ChatGPT Pro 推导完成。因此,他的判断相当明确:这是“量子引力领域一项真正扎实的成果”,主要由AI产出,但问题由专家选择,工作由专家引导,每一步由专家核验,科学视角也由专家提供。

  • 边界已经移动,但并未消失。AI如今解决了几位领域专家追踪一年的问题,但Lupsasca明确表示,它还没有解决一个让整个学界苦战几十年的问题。这是他想看到的下一个门槛。

8. AI压缩困惑,并把并行对话变成研究侦察兵

  • 物理教育传统上把艰苦计算当作必经训练:学生学习技术,也证明自己能够熬过困难工作。教授们会把可处理的6至12个月问题留在手边,用来连接课程学习与研究。Lupsasca 怀疑,当前模型可以“碾碎”其中许多作业,使旧有训练结构变得不稳定。

  • 但AI也能帮助学生穿过他所称的研究生课程与研究前沿之间的“沙漠”。它可以按所需层级拆解任何事实,连接陌生概念,避免数月无效困惑。尚未解决的问题是:当最强的导师也有能力替学生完成作业时,学生如何建立独立信心。

  • 在Lupsasca自己的工作中,困惑耗费的时间已经大幅缩短。完成计算后,他可以询问这与另一个已知事实如何相容,立即听到自己遗漏了什么,或对问题的框架哪里有误;过去需要散步、转做其他项目、在思维堵塞中消耗数日的过程,如今变成快速纠偏。

  • AI也改变了探索方式。他不再需要把稀缺精力押在从A经过B走到C的一条路径上,而是可以同时开启10个对话,让它们沿不同方向成为高速“侦察兵”。即便不完美的侦察兵,也能指出有希望的地形并留下路标;沿着一条部分绘制出的路线前进,远比第1个闯入未知区域容易。

9. 技术能力的扩张速度快于科学品味

  • 研究生进入物理学时,常问的问题包括:为什么空间有3个维度?大爆炸发生了什么?黑洞内部是什么?专业成熟意味着意识到,有些问题距离当前知识边界太远,无法有效切入;进步来自选择一个刚刚超出边界的问题。

  • Lupsasca 因此区分了能力与伟大。一个合格的物理学家可以学会问题所需的任何数学、代码或计算工具。“优秀物理学家与伟大物理学家的区别,在于知道该问什么问题”——这是最难的科学技能,通常也是最后才学会的技能。

  • 当前模型像能力异常强的研究生:一旦给出精确问题,它们可能在计算上达到超人水平。但它们还不能稳定判断哪个问题值得追踪,因此人类品味仍居核心;而把问题、细节层级和表述方式匹配给学生的学术能力,也可以直接迁移到提示AI协作者。

  • 这种优势可能只是暂时的。当Lupsasca把胶子论文的一页内容交给 ChatGPT Pro,要求它提出3个最佳后续方向时,模型给出的基本就是他的3个首选。他不愿透露内部正在进行的自主前沿发现工作,但表示,AI在选择问题方面已经强得出人意料。

10. 黑洞“Love”测试让这条轨迹变得私人化

  • 主持人质疑,表面上的创造力是否只是重新组合。Lupsasca 回答说,人类发明也可能只是“一台栈式机器的重新组合”,而 GPT 给人的感觉像一名有创造力的协作者。他保留了一个异议基准:据他了解,Terry Tao 曾将看似有创造力的AI证明追溯至晦涩参考文献,尚未见过真正令人印象深刻的新颖数学步骤。

  • Lupsasca 对 GPT-5 的决定性测试,使用了他当时尚未公开的论文,主题是黑洞为何没有“Love”——Love 数衡量潮汐响应,黑洞响应为零意味着存在某种保护对称性。由于模型的训练数据截止时间早于论文发表,他只提供了控制方程,并要求模型找出其对称性。模型最初错误地回答说不存在任何对称性。

  • 随后,他给出了一个显而易见的平直时空热身问题。GPT-5 Pro 在9分钟内正确解决了这个有着200年历史的问题,找到了3个生成元;在同一场对话中得到提示后,它又回到黑洞方程,并在18分钟内发现了新的对称性。“那就是我的第37步棋。”他说,那一刻,模型在不到半小时内复现了他最珍视的计算。

11. 低成本论文生成抬高门槛,也让验证变得稀缺

  • Lupsasca 认为,在专家引导下,对于接近已知工作的计算,模型已经能够每天产出1篇人类水平的论文。糟糕的问题同样能生成包装精美的胡言乱语,助长“AI垃圾内容”,并淹没 arXiv。学界再也不能把专业化呈现当成底层问题或推导可靠的证据。

  • 他的应对不是发表30个单负论文变体。随着生产成本下降,重要性的标准必须提高。他希望以这些新振幅为切入点,逐步攻向更困难的量子引力问题,最终挑战那些让整个学界苦战几十年的难题。

  • 静态论文本身也显得低效:AI完成计算,人类把它压缩成简短符号,结果又被重新输入AI展开。Lupsasca设想一种交互式论文:可以解释全局、放大某个推导,或展示数学家目前通常省略的直觉图景;同时保留写作迫使人澄清思路的有用纪律。

  • 验证可能成为今年的主导瓶颈。Lupsasca 曾认为形式化证明系统不可或缺,后来自然语言推理进步到足以接近人类数学讨论;如今模型可以同时攻克数千个问题,返回多到人类无法检查的证明。因此,Lean 等系统中的形式化、自动检查,以及更清晰的模型置信度信号,又重新变得有价值。

Brandon

Okay, so I think we're at a special time now where, at least in some directions, AI has become superhuman, at least on certain tasks. That's what led to these recent papers, which resolve a problem that had puzzled expert physicists in the field for over a year. They were unable to resolve it, and AI was able to do so very quickly. I think that's a certain milestone we've passed. You guys are bringing attention to this because, for the average person on the street who doesn't care about theoretical physics, this is not very noticeable, but I think it's a very profound change and we've really passed some kind of threshold.

Welcome to the AI for Science podcast, part of Lean Space Network. I'm Brandon. I develop RNA therapeutics using AI at Atomic AI. I'm joined by my co-host, RJ Honicky, CTO and founder of Miro Omics. It's a pleasure to introduce Alex Lupsasca, a professor at Vanderbilt University and a fellow at OpenAI. For a young researcher, he has quite a storied background. Among other things, he's the winner of the 2024 New Horizons Breakthrough Prize. It's the Oscars for science. I asked ChatGPT whether this is the most prestigious award someone at his career stage could win, and it recommended a second one called the IUPAP award, which turns out to be one he also won. Anyway, right now he's having fun at OpenAI, doing some really cool research pushing the foundations of theoretical physics using GPT models.

Alex Lubyansky

A pleasure to be here. The one message I wanted to convey is that I think we're on a trajectory which I personally find very surprising and surreal, but also amazing. I would say that a little over a year ago, AI was very useful for email, but not for the kind of work that I do, which I consider important theoretical physics calculations. I thought, “That's special. It's much harder than email, and AI is not going to be able to do that.”

Then a series of developments came in rapid succession that completely changed my mind. I can walk you through some of these examples. Specifically, ChatGPT o3 was the first really strong reasoning model that could do actual math that was useful for my research and could save me a lot of time. That's when I started to really pay attention and use it a lot more. I thought, “Wow, this is a great tool. I've got to get ahead of this and learn how to integrate it into my workflow.”

Then, when GPT-5 came out, it was able to reproduce one of my best papers, which took me a very long time to come up with, in about 30 minutes. That's when I really became AI-pilled. I thought, “Oh my God, this changes everything. It's the most important discovery in my lifetime. It's going to affect everything about how we do research.”

Frankly, I would go around telling a lot of my colleagues, “This crazy thing happened. Pay attention.” I was getting lots of different reactions, but I think people weren't quite getting it. I talked to OpenAI, and they were also really excited. I thought, “I don't know that much about AI, but I have to get in on this. To understand that this is happening and not be a part of it is a huge mistake, so I have to go to OpenAI.”

I was on sabbatical, so it was very easy to come here, and I joined the company. Then it just kept ramping up even beyond that. To the point where now I think most of my senior colleagues in physics are aware of where things are headed, and they're all getting on board.

RJ Honicky

Yeah, that's an awesome story. I find it really funny because that story almost reminds me of a lot of different people who had the same realization with Codex sometime last fall, especially. It just took off, and a bunch of people went from, “Oh man, this is 20% of my work. It's kind of a nice assistant,” to, “Oh crap, what just happened?” Even Andrej Karpathy went from that perspective to realizing what had happened.

Alex Lupsasca

Well, yeah, in August, actually, I remember when GPT-5 came out. At that point, I was really following AI pretty closely, and I think on Twitter the reception was lukewarm. A lot of people were saying, “We expected a lot more,” and, “It's not better at writing email.” I remember thinking, “Okay, GPT-3 could write email. How much better can it get at writing email? That's not the point.” But at the science frontier, the capabilities were really taking off.

RJ Honicky

Yeah, there was a lot of attention paid even to o3, but presumably GPT-5 was a huge jump.

Alex Lupsasca

And I think GPT-5.4 is also a huge jump. I don't know if it's noticeable from the outside, although I did see some chatter online. People are running these independent benchmarks, which do show this. I think people are realizing it, and in practice researchers are now all over AI and using it.

I'm getting inbound messages all the time because I'm the resident scientist doing physics at OpenAI, so everybody is sending me papers and chats saying, “Oh my God, this happened.” I got one just this week. Somebody said, “Codex just wrote up a simulation of the SYK model.” This is a very technical thing in quantum mechanics and gravity. A lot of research groups have been trying to run this simulation, and they couldn't do it, but Codex did it in 10 minutes, just because setting it up was so hard.

Brandon

Well, I think it's partly because of the Venn diagram: when you look at the people who have the physics knowledge and the people who have the top coding skills, maybe the overlap isn't that large, although I think it's been growing. But I think in this example there are a lot of really good people in physics with coding skills who would be trying to simulate these things. So I think Codex is just really good now.

Okay. Yeah, nice. I think we're at a special time now where, at least in some directions, AI has become superhuman, at least on certain tasks. That's what led to these recent papers, which resolve a problem that had puzzled expert physicists in the field for over a year. They were unable to resolve it, and AI was able to do so very quickly. I think that's a certain milestone we've passed, and I'm glad that you guys are bringing attention to this because, for the average person on the street who doesn't care about theoretical physics, this is not very noticeable. But I think it's a very profound change, and we've really passed some kind of threshold.

Let's specifically focus on the gluon paper in the physics part, and we can get to the AI part later.

Alex Lupsasca

Okay, so in physics there are 2 basic principles of nature that we think every law, or every theory, should respect. On the one hand, there's the principle of relativity, which, at a very high level, declares an absolute law that cannot be broken: you cannot transmit information faster than the speed of light. But then there's another principle, the uncertainty principle that underlies quantum mechanics, which says that everything is a little fuzzy. Your position and velocity have a little fuzziness to them.

You can see immediately, at this level of description, that there's a tension between these 2 principles, because one is an absolute law declaring that you cannot go faster than the speed of light, and the other is saying, “It's a little bit fuzzy.” This is just to give a sense of how, when you try to write down these principles mathematically, the equations don't really play nicely with each other. It's been a real struggle to come up with physical theories that can reconcile both principles simultaneously to describe the physical world around us.

I would say that the great achievement of 20th-century physics, which is really one of the greatest triumphs in human thought as far as I'm concerned, is the elaboration of this framework called quantum field theory. It's a general framework that can describe the physical forces of nature in a way that accommodates both of these principles.

In quantum field theory, which is our best theory to date, I will say it gets a little technical, but again, I'll try to keep it pretty high-level. What you're trying to compute or describe are the probabilities for certain events to occur. Because you're in this quantum-mechanical setting, you can't say with certainty what's going to happen when you have a certain experiment, but you want to predict probability distributions.

In quantum mechanics, probability distributions are obtained by squaring certain complex quantities. By complex, I don't mean complicated; I mean they're not real numbers. They're real plus imaginary numbers, which we call quantum amplitudes. So the goal of a theory is to predict quantum amplitudes, which are these complex objects whose squares give the quantum probabilities, and that's the most you can say about the outcome of an experiment.

These quantum amplitudes, in particular, include a variety called scattering amplitudes, which describe the following scenario. Suppose you have a bunch of particles that you throw at one another. This is what happens in particle colliders like the LHC at CERN in Geneva. You take a bunch of particles, smash them together, and stuff happens. They interact via the physical laws of nature. Various processes occur, and then other particles come out at the end of the interaction.

A scattering amplitude is the object that describes the probability for a particular type of interaction: we have some particles coming in with some energies and momenta, and some other particles coming out with other energies and momenta. These scattering amplitudes are functions of all the data describing the particles coming in and the particles going out.

In general, you can have arbitrarily many particles involved in an interaction. This is one of the hallmarks of quantum field theory: particles can be destroyed, so you don't necessarily have the same number of particles at the end as you had in the beginning. Particles can be created. Lots of things can happen.

In general, you want to describe all the possibilities, so you want to have an amplitude for an arbitrary number, n, of particles. That's called an n-point amplitude, because there are n particles coming in and out. It turns out in quantum field theory that if you have a particular force and you're able to compute the n-point amplitudes—these functions of the n parameters that square to the probabilities—then you know everything about the theory, more or less. There's always an asterisk, but it's basically the entire content of the theory.

Brandon

So, wait. If you have a theory that tells you any number of particles can come in and go out, then I can say, “I can declare anything about that system.”

Alex Lubyansky

Exactly. Then you know everything. Importantly, these amplitudes aren't just numbers; they're functions, because the probabilities that they compute depend on how much energy the particles have, what their momenta are, and also their polarizations.

A lot of particles, like the photon—the particle of light—have something called polarization. When you look at the surface of a lake and have polarized sunglasses, you can turn your head and see more or less sunlight reflected off the lake. That's because a photon, which you can think of as a little particle of light, carries, as it propagates, a little arrow perpendicular to the direction of propagation. That's called the polarization.

This polarization has a direction, and sunglasses can selectively let in light with one polarization and not the other. As light travels, its polarization can rotate. It can wind. It can do its own thing.

In general, if it winds in a right-handed way—if, as the particle travels, the polarization winds to the right—we call that positive helicity, or right-handed polarization. If it winds in the other direction, we call that left-handed helicity, or negative helicity. These amplitudes, which are the fundamental objects in quantum field theory and contain all the information there is to know about physical forces, depend not just on the energies and momenta, but also on the polarizations.

Now, I've told you about how there are 2 basic principles of nature: relativity and quantum mechanics. They come together in this framework of quantum field theory. I keep talking about forces, and there are 4 fundamental forces of nature.

There's electromagnetism, which is responsible for basically the properties of atomic elements in the periodic table, and therefore chemistry, biology, and everything that you see, touch, or feel. Pretty much all of it is due to electromagnetism: textures, colors, and so on. This force is mediated by the photon, which is the particle of light. That one is the most familiar to us.

Then there's gravity, which is another force that we feel very much because it keeps us to the ground. Then there are 2 nuclear forces: the weak and strong nuclear forces, which we don't really notice directly in our daily lives. The weak nuclear force is responsible for radioactive decay and other such processes, while the strong force, which is the strongest of them all, is what binds the nucleus together.

You learn in high school that like charges repel. If so, then why do protons stick together inside the nucleus of the atom? They should repel one another. Indeed, that's the case, but if you bring them really close, the strong force kicks in and overwhelms the relatively weaker electromagnetic force.

The strong force is mediated by the exchange of particles called gluons, because they're what glue together the nucleus of the atom. So, gluons are the particles of the strong force. Gravity is mediated by gravitons.

Brandon

I think the gluon paper was sort of the starting point for this, maybe—not. The gluon paper had a really specific result, right?

Yeah, absolutely. Maybe let me just flash the paper itself. We put this on the arXiv a little over a month ago. Here's the paper. Let me explain in a few sentences, now that I've given a lot of background, what the title means.

The title says, “Single-minus gluon tree amplitudes are nonzero.” This might sound forbidding, but I think we can unpack this for the audience. Gluons are the particles that carry the strong force, and gluon amplitudes are functions that describe the quantum probabilities for gluons to interact via the strong force.

The word “tree” here is a little bit of a technicality. It means we're only considering processes where no gluons are created or destroyed. If gluons are created or destroyed, then you get loops, which we can explain later, but this is just a technicality. We're considering special interactions where the same gluons that come in also come out.

Brandon

For anyone who's ever fit a polynomial, you can think of trees as being like a linear term, and loops can be higher-order terms.

Exactly. It's way more complicated than that, but conceptually, it's like the lowest order in a series.

Single-minus—now I have to explain that. Remember I told you earlier how particles have polarizations? When you try to study gluon amplitudes, this is a whole industry of physics. It's a very complicated field. People have written thousands of papers over the decades.

You always want to try to understand the simplest examples first. That's why you start with the tree amplitudes, or the leading effects, and then you worry about the loop corrections. You might think that the simplest example to start with is one in which all the particles have the same helicity. Say they're all right-handed, or, that is to say, they're all plus-helicity particles.

It's been known for a long time that, in that case, the amplitude is just 0, which means the interaction is forbidden and cannot happen. That's one way to think about it: it's just a symmetry that explicitly forbids this.

Brandon

So, you don't have to worry about calculating anything. You just know.

Yeah, just dimensional analysis. It's a very general argument.

Brandon

You don't need to do very much work.

It's true that it's the simplest example, but it's so simple that nothing happens. The answer is trivial. You might ask, what about the next level up? What if one of them has the opposite helicity, while all the others have plus helicity? That's what we would call a single-minus amplitude.

If you look at the lecture notes and textbooks that have been written on this, the same argument that rules out the all-plus amplitudes also appears to rule out the single-minus amplitudes. They're too simple. They can't really interact. There's nothing to see here; move on.

Then you might ask, okay, what about the next thing, where there are 2 particles with minus helicity and all the others have plus helicity? If there are n of them, there are n − 2 others that have positive helicity. These would be double-minus amplitudes.

People in the ’80s studied and computed these amplitudes. They're not 0. In particular, 2 physicists, Parke and Taylor, found this beautiful result. They did a lot of really hard work and computed these amplitudes—a very technical, difficult calculation. At the end, you get all these terms and you have to sum them all up. Almost all of them cancel, and at the end, you're left with this very simple formula that fits in half a line. It's now known as the Parke–Taylor formula for these amplitudes.

These amplitudes are now called MHV amplitudes, which stands for maximally helicity violating, because they have the largest—or so we thought possible—asymmetry between the plus- and minus-helicity particles. That's the most asymmetry.

Now, let's get to this paper, which came out last month. This is a paper written with Alfredo Guevara, who's a postdoc at the Institute for Advanced Study; David Skinner, a professor at Cambridge University; Andrew Strominger, a professor at Harvard who used to be my advisor; and also Kevin Weil, who studied as a particle physicist in a previous life.

How did this happen? Maybe we'll get into how I ended up at OpenAI a little bit later, but I ended up at OpenAI and started to improve the models' abilities to do physics. The models got really, really good at physics, and I thought, “Okay, it's so good now. We should try to solve some actual research problems at the frontier.”

I called up Andy, who used to be my advisor, and said, “Hey, Andy, do you want to come here to San Francisco, visit OpenAI, and we can try to solve one of your problems in physics?” I thought, “It's probably not going to work, but if it doesn't work, at least we'll figure out why it doesn't work. I can do this with a different physicist every month, and eventually something will work. In the meantime, we'll learn how to improve the models, so it's all fun and useful.”

Andy was the first one I invited to do this.

And he said, “Well, I have this perfect problem that I’ve been thinking about with Alfredo and David for the past year.” I’ll explain the problem, but the amazing thing is that we decided to start working on it using AI a little bit before Andy was scheduled to come, like the week before. In fact, using ChatGPT, we solved the problem before he even got off the plane.

swyx

Which was a huge surprise to him. [Laughter.]

To him. Yeah, I think to me, too, to be honest. We had not expected that. It’s a really cool story.

So, Andy, David, and Alfredo understood a year ago that this statement—that the single-minus amplitudes are zero—is not exactly correct, because the usual argument in the lecture notes and textbooks has a loophole. The loophole is that it assumes the particles are coming from generic directions. But in a certain regime where the particles are exactly aligned with one another—we say they’re collinear—the usual argument has a loophole, and it’s possible for the amplitudes not to be zero.

But then, if they’re not zero, what are they? Suddenly, these really simple amplitudes, previously thought to be zero, if they’re not zero, we should compute them, and they should be something really nice, simple, and special. Now, I’m sweeping a lot of details under the rug here. This has to work in some split-signature spacetime, and it connects to lots of other things they’ve been worrying about. We’re not going to worry about this.

I was actually hoping at the end that we could talk about what it means to be 2 dimensions in space and 2 dimensions in time, but I think part of this is doable.

swyx

Really mind-bending stuff.

The loophole is one about the alignment of the particles, but it’s also a loophole about the spacetime of the physics—the universe we’re living in.

So, they understood that they’re not zero, and they started to compute them. Alfredo is really, I think, the unsung hero of this story because he did a lot of really hard work to compute these things by hand. I’ll just show you an example.

In the paper, there’s a lot of formalism. Here is the beginning of the definition of the general answer. Then you have to define these vertex objects, V, and they’re complicated. They involve spinor-helicity functions of spinors, and then you have this recursive formula. It’s a whole mess.

Concretely, if you try to unpack this definition, remember these amplitudes are a function of the number of particles involved. There’s a 3-point amplitude where there are only 3 gluons in the interaction, and the answer is pretty simple. This is some function that we’ve defined here—not that complicated. Then this is the 4-point amplitude, where now there are 4 particles, and you can see that we go from 1 term to a sum of 2 terms here.

But then, once you get to 5 particles, you start to get a lot more terms. There are 8 of them being summed here. By the time you get to 6 particles, it explodes in your face. For those people not watching this on YouTube and listening, this equation takes up a quarter of the page. It’s 32 terms, each of which is a product of 4 terms, each of which is itself encapsulating a rather complicated formula.

swyx

Yeah, this is super nasty, and that’s as far as Alfredo got—or Andy, whatever. So Alfredo did it. Look, is this just an expansion of some sort? How hard is it to do this expansion?

Brandon

Very hard.

swyx

Okay.

David Skinner

Yeah. There’s a nice graphical way to understand this in terms of Feynman diagrams. I hadn’t planned to explain this, but it’s a visual subject.

The math is very complicated, and already back in the ’40s, Richard Feynman, who was one of the pioneers of quantum field theory, came up with this very visual way to organize our understanding of the subject. You can doodle these little cartoons that represent possible interactions.

The rules of quantum mechanics actually say that in these amplitudes, where you scatter a bunch of particles, you get to fix what comes in and what comes out because that’s the question you’re asking: What’s the probability for a certain interaction? But then everything that happens in between, you don’t get to choose, because the physical laws determine what happens.

In quantum mechanics, you’re supposed to consider all the possibilities—all the ways in which the incoming particles can interact and transform into the outgoing particles. You’re supposed to average or sum over all the possibilities to get the final amplitude for the process, as a sum over the amplitudes for each individual possibility for how you could get there.

swyx

So, just to be clear, there are incoming particles. They interact, and then there are all these different possibilities. They each have their own amplitudes, and then I select this one possibility and this one possibility, and I get one possible interaction. There’s an infinite number of those for each, and then I sum those infinite possibilities, I suppose, and I get the outcome?

Yeah. In principle, there are infinitely many pictures to sum over, but that’s why we organize them by how complex they are. It turns out that every time you get an interaction—every time there’s a vertex where 2 lines meet—that point interaction comes with a power of the coupling constant, which controls the strength of the interaction.

Every additional interaction makes the amplitude more suppressed, so it contributes less to the final answer. You want to first consider the diagrams with the fewest possible number of interactions because they will give you most of the final amplitude. Then, if you’re trying to get a more and more refined answer, you consider the more and more complicated cartoons with more and more interactions.

In fact, this is one of the ways in which the diagrams can get complicated: They can have loops. For instance, here you have a particle that decays into 2 particles, creating this loop because then they meet up again and disappear. In this interaction, you have intermediate particles being created and destroyed.

But whenever that happens, you get 2 extra vertices in your graph. These diagrams are suppressed because it’s less likely that you get these extra felicitous interactions, so you don’t need to worry about this as much. It’s like a small correction. Of course, in principle, you could keep going, but you’re never done except in very special circumstances.

swyx

Or higher-order powers in a polynomial or something. Or a Taylor series.

And so, to go back to the story back in the ’80s with the MHV amplitudes, which I think now is a bit of a misnomer, I would call them double-minus amplitudes because that’s where we’re going to get to in a second.

swyx

Right.

David Skinner

There was this heroic calculation where a lot of Feynman diagrams were summed. They were considering more and more interactions with more and more particles, and every time there were more and more terms, but they all canceled and in the end always gave a simple answer.

In fact, that’s what this PT term is. PT stands for Parke–Taylor. These formulas fit in a line, so it’s not that complicated, but it’s very surprising that such a messy calculation would clean up into such a simple result.

What Alfredo, Andy, and David did was understand that these single-minus amplitudes, in the special case where some of the particles are aligned, don’t have to be zero. You can do this very complicated Feynman diagram expansion to get the answer, which is not zero, but the problem is, if you do it this way, you can represent the answer in some horrendously messy, complicated way. If you unpack it, it’s extremely complicated.

When you consider the n-point amplitude—the probability of n particles interacting—the number of terms in your answer, which roughly corresponds to the number of diagrams you have to add up, grows factorially in n, the number of particles. Factorial growth is really bad. It’s superexponential; it goes faster than an exponential, so it blows up in your face. This is what you’re seeing here.

That’s because, roughly, you have to draw all the possible cartoons, and the possible combinations are a combinatorial problem. That’s where the factorial behavior comes from. But we know from the ’80s that in the actually more complicated double-minus case, Parke and Taylor found this miraculous simplification.

Andy, Alfredo, and David spent the last year chasing the analog of the Parke–Taylor formula—the very simple answer that was obtained in the ’80s for the double-minus amplitudes—but now for these single-minus amplitudes, which they understood are not zero. But then, what are they?

They were getting this really complicated answer. You never know in physics ahead of time if something will simplify. You have to believe in it to find the simplification. But because the double-minus one simplified, it felt like these should simplify, too.

We think they’re important for lots of things, and that these are somehow really important objects that are very fundamental. They should have a nice description. They spent a year looking for that.

swyx

There’s a funny—the next line, if you scroll down, is something like, “We need a simpler formula.”

Right. “A more concise formula is needed.” And this is where AI comes in. When I asked Andy, “Hey, do you have a problem in your pocket that we should use AI to target?” he said, “Well, I have just the perfect thing for you. We’ve been puzzling about this.”

swyx

It's really important. It's really interesting. It connects to all these things, and we don't know the answer.

When I was a grad student, if I had approached something like this, I probably would have plugged it into a computer algebra system, let it chug along, tried a few limiting cases, and seen if there were any magical simplifications. This type of thing is something where you oftentimes see, “We need a different approach.”

Exactly. Before Eddie even got here, we started to play with ChatGPT. Alfredo, Andy, and I were trying different things, with lots of different chats happening and going back and forth. David was involved as well.

The first thing that happened is that we fed the 5-point amplitude into ChatGPT and asked, “Can you simplify this?” It said, “There’s a special region, so there’s an extra assumption that you can make in which this answer simplifies to this one.”

So this assumption is equivalent to having 1 particle coming in and decaying into n − 1 of—

swyx

That’s one way to think about it, roughly. Okay, but we’re in 2 spacetime dimensions, so—

Yeah, it’s complicated. But basically, you can look at what we call phase space. It’s the entire space of possibilities for all the energies and momenta of incoming particles. There’s a special region in that phase space where 1 particle has a different sign of its frequency compared to the others.

In that region, there’s a big simplification that happens, which ChatGPT found. I should say this was the public model, but the Pro version that thinks really hard.

swyx

Was that a known fact that it was just able to relate to the problem, or was that something it put together?

As far as I know, it put that together. It said, “This 5-point function, which is a sum of 8 terms, each one of which is a product of 3 terms—they’re all pretty complicated.” It said, “Actually, this simplifies to this product of only 3 terms.”

We stared at this thing and thought, “Wow, that’s really nice. We didn’t know this.” In hindsight, once you know it, you can rederive it, but it takes a while to understand where this comes from. I think that was a leap of insight that the AI had.

At some point, it said, “I wrote Python code and ran through all 5,000 possibilities, and I deduced this.” It’s the equivalent of running a computer algebra system, but it just decided to do it on its own and came up with a huge simplification.

Brandon

Great. Yeah, awesome. Was this after making the assumption? Was this after the 1-particle decay assumption?

Alex Lubyansky

Yeah. It figured out there was some region in which things simplified. This was very experimental; we were talking about it a lot. It figured out there was a special region in which things simplified, and then ChatGPT came up with that simplification as well. All of them.

Then we were like, “Okay, let’s give it the 6-point function, which Alfredo heroically computed.” We didn’t have the 7-point function. I don’t think anybody could use the heat identity to expand it; it would be disgusting.

ChatGPT did its little thing, and then it was like, “Yep, simplifies to this.” We thought, “Whoa, okay, that is really nice.” Instead of 32 terms, it reduces to just 4 terms. It’s not a sum of 32 terms; it’s a product of only 4 terms.

Then we asked ChatGPT, “Can you guess the general formula for all n?” You could imagine using some programming language or symbolic manipulation software to do these reductions in specific examples. But to tackle the general case, I don’t know how to use a computer to do that. ChatGPT said, “Yeah, this is the answer in the general case.” Boom.

swyx

How long does that take?

Using Pro, it thinks for 20 minutes at a time. You go back and—

swyx

But it wasn’t like 6 days or something?

No, no, no. It was just over the course of several interactions.

The amazing thing is that the formula it proposed, instead of having this factorial growth—which is superexponential, where the number of terms blows up as you consider n, the number of particles, increasing—is actually linear. If you double the number of particles, you only double the number of terms. It’s the nicest possible behavior you could imagine.

This is the equivalent, I think, of the Parke–Taylor formula for the double-minus amplitudes that was known back in the ’80s, but now for the single-minus amplitudes. This was guessed by GPT-5.2 Pro, but it couldn’t quite derive it. So I said, “Hmm, it looks like this, but I don’t know how to prove that.”

Alex Lubyansky

Yeah, I think the model was not quite strong enough to prove it. Part of my work at OpenAI has been to develop stronger physics capabilities in the models. A lot of people have been adding lots of things; it’s not just my singular contribution. There’s a lot of great research happening, and it all comes together. It takes a village.

We had this internal model that could think for a very long time and was extra strong at physics. We gave it the whole problem from scratch without actually giving it this formula. We just formulated the problem in a very sharp way and asked the model to find the amplitude in the general case in this region, because we had identified that this was the special place to look.

It took 12 hours, which is a long time, but it came back with the same formula, which we had not given it. It rediscovered the correct formula. This time, it also found the proof that the formula is correct and derived it.

In fact, the remainder of the paper after we state the equation is devoted to the proof, which is basically what came out of the AI. We say, “The rest of this work is devoted to proving that the conjecture is correct.” There are 3 steps: first, you show this; second, you show blah; and third, you show blah. This is basically what the AI came up with.

Now I can finally summarize the paper. The title is *Single-minus gluon tree amplitudes are non-zero*. These are special interactions between gluons where only 1 of them has a different helicity from the others. They were previously thought never to occur, but these interactions can actually happen. The amplitudes are non-zero. That’s the main claim of the paper.

I think it’s quite surprising. I think it’s a really nice paper. The final result, I guess, has 2 parts. One is understanding that it’s not zero. That came from the humans a year ago. They were trying really hard to find a simple answer for what the amplitude is, and they were stumped for a year.

They were able to get this indirect representation, which is extremely complicated in terms of Feynman diagrams. But they were looking for the simple formula that is analogous to the Parke–Taylor formula from the ’80s for the more complicated amplitudes. That was done with the AI, and I think that’s a really interesting result.

Brandon

Yeah, it totally changes the way you should think about where we are in physics and how AI is going to change that. It’s a result that top researchers in this field were thinking about for a year, and then the AI solved it.

There are several things about the story that I think people didn’t understand on Twitter. Maybe scroll down to equations 35 to 38. I would say most people, even introductory grad students, would look at equations 35 to 38 and say that 39 is actually a very natural extension of this. I don’t think that’s that surprising. I think it’s interesting.

I didn’t know until just now that—[clears throat]—when you proved 39, that was a fresh session. That was without the limiting cases. You started from scratch.

Alex Lubyansky

Yes. I did it that way because it’s an extra way to be confident in the answer. If a different model independently comes up with it from scratch, then you’re not just spoon-feeding it the answer that you think is correct. That’s an extra confirmation.

We thought a lot about how to put this out into the world, and there’s no perfect way to do this. We could clearly have done a better job of communicating it. One thing that was important to us was not making this paper about AI, because I think this is a really interesting physics result. People will keep reading this paper, I hope, for a long time.

We didn’t put AI in the abstract because this is a physics result that stands on its own. There’s 1 paragraph about AI where we say, “The final formula was first conjectured by GPT-5.2 Pro and then proved by an internal OpenAI model.” That’s what happened. It’s true, but we didn’t really want to get into it because I don’t think that’s the point of the paper.

It’s really interesting how it happened, but the result stands on its own. If you read a paper today that was written 20 years ago and used a computer to do some critical step in the argument, and it had a whole discussion of how, “I loaded MS-DOS 3.1, it had 5 floppy disks, and I had to swap my floppy disk,” you wouldn’t care. That’s not why you’re reading the physics paper today.

We didn’t really want to go into that in the paper. We talked a little bit about it in the blog post that we released with OpenAI, which is this one. On Twitter, there were a lot of questions, and I wrote some tweets that I think clarified it. There was also a physicist who wrote a great blog post about actually understanding the story.

The Economist also put out a great article about it, and they really understood what happened. I thought it was great coverage. *Science* magazine also wrote about it. Harvard and the Institute for Advanced Study put out press releases. I think it got a lot of attention, but it’s kind of a subtle thing to explain.

It took us an hour to go through what happened and what was done, so it’s hard to explain. I think it would have been a distraction from the physics point of the paper to go into that.

swyx

Okay, let’s talk about the physics, then. Give us a sense, because my theoretical physics on the frontier comes from PBS Space Time, right? It’s a great channel.

Yeah, it’s a great channel, but it gives you a great high-level picture. It’s hard to know how this sits in the pantheon of papers that represent the cutting edge of theoretical physics.

swyx

Not exactly that. I want to understand: It seems like you’re comparing it with a previous result that is pretty significant, highly cited, and very important. How does this compare with that?

Okay, you’re putting me in a bit of a tough spot. I will say I think the result is surprising. That’s why the title is what it is: “Single-minus gluon tree amplitudes are non-zero.” If you’re somebody who works in this field, that should catch your attention.

Ultimately, it’s very hard to know in science, when you release something into the world, how it’s going to be received and how impactful it will be. I think the true value of a paper can only be assessed decades into the future, based on how much future work it leads to and what developments it opens up.

swyx

Maybe a better way of asking is: My understanding is that the previous paper opened up a whole line of thinking about—

Yeah, I think this is a great segue to the second paper that came out just 3 weeks later.

Brandon

Perfect. Then let’s talk about—

So, it got its own blog post. This is March 4, so I guess 2 weeks ago now. We were talking earlier about how there are 4 forces: the strong force, mediated by gluons, and gravity, which is mediated by gravitons. We can produce gluons at the LHC and measure their effects fairly directly. Gravitons, we think, are also around us, being produced all the time, even as I move my hands, but we’ve never done an experiment that directly measures gravitons.

They’re supposed to be the quantum of gravity, so they’re really interesting from a theoretical standpoint. Going back to RJ’s question earlier, what is a graviton? There are different answers we could give. Ultimately, the correct answer depends on what the theory of quantum gravity is, which we don’t know yet.

swyx

Yeah.

Alex Lubyansky

If you just naively try to take all of the tricks from field theory that we know from the Standard Model and apply them to gravity, things just break down. The theory is not self-consistent. There are various problems.

swyx

Yeah.

Guest

Just like in this room there’s light flowing around, there’s some indivisible bit of light that you eventually can’t break up into smaller bits. That’s the quantum of light. We call that the photon. The gravitational force is mediated by the exchange of gravitational force or gravitational waves.

If you try to take a gravitational wave and break it up into smaller and smaller pieces, at some point you get a quantum that you can’t break up anymore, and that would be the graviton. That’s how we understand them.

swyx

Okay, so the idea is that you can’t—you get to a certain point, and you can’t have less gravity than that. You either have some or none. Right? That’s one way to think about it?

Alex Lubyansky

Yeah. We wrote this paper, which is called “Single-minus graviton tree amplitudes are nonzero.” It’s almost the same title, except with “graviton” instead of “gluon.” That’s on purpose, because we wanted to extend the result.

It’s the same story in the sense that it was thought that all single-minus amplitudes are zero, but actually that’s not true for gravity either. Gravity is a lot more complicated, though, so if you want to compute the graviton amplitudes, it’s potentially a lot harder.

Brandon

Do gravitons have phase the same way that gluons do? Is that it?

They actually have spin 2 rather than spin 1. It’s getting into the weeds, so the numbers you have to use to describe them are a little bit different. They’re doubled in some sense.

Brandon

Okay, so their polarization is more complicated.

I see.

The special region in which the final answer simplifies has 2 labels because it’s a spin-2 particle, whereas in the gluon case there was only 1 label because it was a spin-1 particle.

Brandon

So this is like—

It’s not the same math. Gluons and gravitons do have some structural similarities compared to other types of particles.

Well, yes, in the sense that they’re particles of force.

Brandon

Yeah, yeah, but they’re sort of doubled.

Yeah, they’re sort of doubled. I guess the people watching this podcast probably like to geek out on this.

The modern definition of a particle in quantum field theory, which is our best-verified framework for nature, is that particles are irreducible representations of the Poincaré group.

swyx

We just lost 90% of our audience right now.

Yeah, okay. Maybe we cut this. There are mathematical representations, and they’ve all been classified. All the possibilities are known by Wigner, actually, a brilliant physicist.

It turns out that the representations of possible particles are completely labeled by the mass, spin, and charge of the particle. These are the 3 quantum numbers. Particles of long-range forces like gravity and electromagnetism have 0 mass. They have to have integer spin. Spin 1 is 3 of the 4 forces, and spin 2 is gravity. And then that’s it.

But let’s set that aside. The really cool thing about this paper is that, first of all, it came out 3 weeks after the first one, which is really fast. I think this is a great example of AI accelerating science.

In fact, we could have put this paper out 3 days after the first one, because that’s how fast we got the answer out of ChatGPT. But it took us 3 weeks because we wanted to check very carefully that it was correct. Most of the time was spent verifying the answer, not writing, which is insane, actually, if you take a step back.

If you told me a year ago, “You’re going to have this AI that just does really hard calculations for you, and then most of the human effort goes to verifying the answer,” I would have thought you were crazy. So, it’s very surreal.

We also had to write it up as a nice paper, which involved putting in the citations and references. That takes some time. I also had a baby in the meantime, so we lost some time there.

But we did this really fast. I think it’s an example of accelerating science. Another really cool thing is that, for this paper, we didn’t have to use an internal OpenAI model that had to think for hours. This was all done using the publicly available GPT Pro.

In fact, we shared 1 of the main prompts that we used. If you go to the blog post “Extending Single-minus Amplitudes to Gravitons” and scroll down to the text, there’s a link to 1 of the chats that we used. You can see we used GPT-5.2 Pro.

The amazing thing about this is that we gave it the gluon paper as a seed. We said, “Read and understand the paper. Make sure you understand the manipulations in the appendices, because that’s where most of the hard work goes.”

It comes back and says, “Yep, I understood the paper. Let me focus on the appendices. Here’s what happened.” Basically, the punchline is that GPT Pro, with the gluon paper as an anchor, was able to do the graviton calculation, which is really different mathematically, completely on its own—not from scratch, I guess, but from the gluon paper. It’s just a different thing, and it was strong enough to do it completely.

swyx

So it took the conceptual leap from the previous paper and just said, “Okay, what math do I need to make that same conceptual—”

Yeah, and it’s different math. That’s an important thing to emphasize. In particular, there’s a crucial application of something called the directed matrix-tree theorem.

Alfredo and David—we’ve been thinking about these things for a very long time—we were like, “Whoa, that’s really cool. That’s surprising. We hadn’t thought of that or seen that before.” That was known math, but maybe because it has such a broad understanding of math and physics, it’s able to say, “Oh, this is a good thing to apply in this case.”

Yeah, exactly. Here it understood the gluon paper, and then we said, “Okay, well, the task is to generalize this paper to the gravity case. Here are 2 key changes, but otherwise the manipulations should be similar.”

We tweaked some things at the get-go. Then we said, “Good luck. You’re a brilliant theoretical physicist.” We gave it 2 paragraphs. We gave it the gluon paper, a couple of paragraphs, and said, “Good luck.”

It thought for 20 minutes, and boom, it starts to think. It starts at the beginning and works through the implications. Really interesting stuff. Then it says, “Here’s what I would do next to turn this into the gravity paper. If you want, I can do blah.” And so we said, “Yeah, go ahead.”

Another thought for 31 minutes.

swyx

Thought for 31 minutes.

Yeah, this exchange is 110 pages. But I think it’s hilarious. I would describe this as vibe physics.

Because you can see its reasoning as it goes, it does a lot of hard work. It goes through lots of equations. It starts to do the—okay, now you have to use this different math; you have to use these tree calculations and loop-reduction formulas. There’s a lot happening: sums over trees and concrete checks.

One of the things I love is that it’s able to do the same things that a human would do: check some basic cases as a sanity check and to get intuition. It comes back every 3 minutes and says, “Here’s what remains to finish the full gravity paper.” Then there’s a list: “If you want, I can write the gravity analog.” “Yes, do that. This is the first step.”

It goes back and thinks for 34 minutes. Half-collinear support—it starts to do stuff. These formulas actually made it into the paper in some form. This is all correct. There’s a bunch of stuff.

At the end, it says, “If you want, the next most useful thing I can do is this.” And we’re like, “Yeah, verify this by performing the explicit check.” It goes on, and, just to cut to the end, finally we say, “Okay, write up the paper.” You can see the paper that it writes, and it’s very close to the final thing we actually put on the arXiv.

Brandon

Did it make suggestions that were not what you would have suggested as the next steps?

Alex Lubyansky

It’s very smart. It knows kind of where to go, and it’s useful to steer it. If you compare what it came up with with the actual paper that we put in, the intro—the abstract and introduction—was written by Andy, who’s an amazing writer. I think he gave this wider perspective on the problem, how it fits into physics, and how it connects to other things that the AI didn’t do. The intro it wrote was more generic.

But AI could write really well. We didn’t really try to make it. The other thing is that we added Section 2, which was not part of that initial exchange. It’s about how these graviton amplitudes transform under certain symmetries of physics.

That’s something that we’re really, really interested in because we eventually want to understand quantum gravity, as I mentioned earlier. Typically, the first step to uncovering a new theory is to understand what its symmetries are. That’s something that gives you some kind of ground to stand on.

In particular, Andy has been pushing this program of celestial holography, which is a whole thing we could get into, but it’s an exploration of the symmetries of quantum gravity. He really wanted to understand this. There’s a separate chat—we didn’t share that one—where we led the AI to explain how these answers fit into the symmetries that we know the theory should have. That’s something that went in there.

Actually, I think from Section 3 onward, it’s pretty much very close to what the AI wrote. I would say this is really remarkable. It’s a real, solid result in quantum gravity that was done pretty much completely by an AI, with humans steering it and asking the right questions.

All the math was derived by ChatGPT Pro, the public model you can access. Most of the time spent by us humans was checking everything and writing it up. That’s really wild. [laughter]

Brandon

As a physicist, you find yourself where a lot of coders have found themselves, where there’s a fundamental, maybe epistemological, question here. If, as a physicist, I could have done that—maybe I needed a little more background, but a lot of it was, “Yeah, go ahead.” Take this paper and give it some prompt.

You guys obviously prompted very well, but there wasn’t much more than that. Maybe an undergraduate in physics could have come up with a lot of it. How does the undergraduate in physics now learn when they don’t have to do the hard calculation themselves?

Alex Lubyansky

You’re opening up many different strands of conversation, which are all super interesting. Let’s try to unpack that a little bit. The most direct thing you asked is: How does the next generation learn?

Brandon

Yeah.

Alex Lubyansky

That’s a really good question. I think about this a lot. Now that a lot of senior physicists in the field are coming to grips with these new capabilities, one of the questions that comes up very quickly is, “How do we train the next generation?”

The way we were trained was by going through these difficult rites of passage, where you have to do these really arduous calculations. That’s how you build confidence in your own abilities and test your knowledge. It’s not just about what you’re capable of doing; it’s about knowing that you’re capable of doing it, proving it to yourself, and building that self-confidence. That is important. We don’t have a good answer. This is something that academia is going to have to grapple with.

One thing that is especially difficult is that, as a professor, I have graduate students. The gap between where classes take you—even graduate courses—and where research begins is actually huge, and it’s growing wider. Classes go very far, but only so far.

Usually, as a professor, when you take on new students, you keep in your pocket a few easy problems, in the sense that you know they’re going to work. There are some questions that you know, in principle, you could work out—not that they’re that difficult—but you give them to a student so that they go through the exercise of learning everything around the question and developing the technology.

You know enough about the problem that you’re sure there’s an answer, that the student can get there, and that you can advise the student in the process of discovering it. I think the issue is that many such problems now, I would say, these models can probably crush.

These are problems that we usually take—again, the time scale for a theoretical physics paper is 6 months to a year. That’s pretty typical. So if you tell a student, “Go away and think for 6 months about this one question. You have to work really hard, learn a lot of stuff around it, and do lots of calculations,” would even the most determined students go 6 months without asking ChatGPT for help? That’s a little bit weird.

It’s also an opportunity. I remember that time in my graduate school career. In my second year of grad school, I had taken all my graduate courses in my first year, and my second year was my first project. It was actually the hardest time for me in graduate school: traversing the desert from where classes take you to the research frontier.

It’s very hard, and there’s a lot of time spent banging your head against the wall. All the time, you’re confused and you don’t understand things, just because you need to absorb so much knowledge. AI can totally help you with that. It’s the best teacher and knows everything. It can unpack any complicated fact to any desired level of detail.

Actually, my experience as a trained professional physicist working on my own research using GPT now is that there are 2 key ways in which my research has completely changed. One is that I spend much less time being confused. I’ll do a calculation, get an answer, and think, “How does this fit in with this other fact that I know? How do I reconcile these things in my mind? I’m confused.”

Alessio Fanelli

Yeah, I do that all the time.

Nima Arkani-Hamed

In research, usually you take a step, and then you hit a roadblock or an obstacle. You’re confused, and you have to think for a few days. Maybe you go for a walk, work on another project, come back, and get a new idea. You spend a lot of time confused. That’s the nature of research.

With GPT, I’m like, “Hey, I just did this. I found this. How does this mesh with this other thing?” Then it’s like, “Oh, well, you forgot this thing,” or, “Oh, you didn’t quite think about it correctly,” or, “This is the standard fact.” The amount of time you spend confused dramatically shrinks, and you move so much faster. That’s one of the accelerating effects.

The other accelerating effect is that I only have so much free time and energy. Especially when you become a professor, you have to teach, you have students, and you have grants to administer. There are a lot of things you have to do. Your free time to think about research without distractions shrinks, and you only have so much energy to do hard calculations.

What you would usually do is, if you have a problem, you’re at point A and you want to get to point C, you think about the route. You think, “Oh, I have to go through point B first.” Actually, maybe there are multiple points, and you try to plot in your mind the course that you’re going to take before you start doing the hard work. You try to think really hard about where you’re going and chart a course.

With AI, you can launch 10 instances of a chat and have each one try a different route. You can send it as a scout that moves very fast into the unknown, pushing outwards, and very quickly get some feedback. You can see which approaches are not promising and which are much more promising.

If you follow them, there’s a huge difference between being the first to push into the unknown and following someone ahead of you. Even if ChatGPT doesn’t always get everything right, just having a scout that signposts some key steps along the way, which you can use to anchor your own movement, is extremely helpful.

Those are 2 concrete ways that AI has changed the way I work. If you’re entering research, having an assistant that can help you find your way to where you’re trying to go can be very good. It’s inevitably going to change how we work, how we operate, and how we train students.

Part of what’s exciting about my job is trying to figure out how all of this works. It’s not just a job for OpenAI; it’s actually a job for every researcher and professor more generally to think about this.

I think the future is very bright. We have some challenges to overcome, but on balance, this is such an amazing tool. I think it’s going to give human physicists AI superpowers because of what I just described. You can do so much more.

The kind of skill that is really useful to get great results out of AI is very similar to the kind of skill that you develop as an academic collaborating with other humans. This is like a collaborator. If you’re a professor who’s been advising students and postdocs, you know that a lot of what being a professor involves is knowing, for each student or postdoc you’re working with, exactly what question to give them.

It’s matching the problem to the person and knowing how to give them the question—with how much detail, and at what level of detail. Not too much and not too little. That’s actually what you have to think about when you interact with ChatGPT. It’s a transferable skill, and people who are good at this are about to get AI superpowers.

Alessio Fanelli

What you just described reminds me of several conversations we’ve had on the podcast so far, which keep coming back to this concept of taste. One thing that, especially in theoretical physics—high-energy physics—has maybe had a problem with, although I’m not sure if you want to describe it that way, is that it can be very trendy. Certain things become in fashion because maybe right now we’re in a world where we don’t have the data to define new directions, to really guide or constrain where we’re going.

I’m curious: how does something that is superhuman, in that it has basically all known physics, interact with a field where, at its core, what can often become popular—what people start working on—is based more on general aesthetics or what the community collectively thinks is cool at the time?

I can imagine it could vibe with so many different worlds. For example, just using Klein space, using this sort of 2+2-dimensional setup for this, was already an assumption that I think is actually kind of important in some ways and does provide feedback to our world.

But you could have asked ChatGPT to solve this problem in all sorts of ways, and maybe it could come up with all sorts of things that don’t really align with the useful taste of the community. How do you actually deal with a proliferation of really interesting results when it’s not clear where the field should go?

Nima Arkani-Hamed

You’re getting at the heart of what it means to make progress in theoretical physics and research. This is a hard question, and there is a simple answer: if there were, it would be research.

Let me say a couple of things. The first one is that when you go to graduate school in physics, it’s usually because you’re really interested in the big questions: Why are there 3 dimensions of space? What happened at the Big Bang? What’s inside a black hole?

These are the things I was thinking about because of science-fiction movies and books. What you realize is that, actually, even though these questions are really cool and exciting, they’re not really the most fruitful scientific questions. At any given time, there’s an edge of knowledge, and the role of scientists is to expand the edge of knowledge—to push into the unknown.

To do that, you want to find the questions that are right at the edge or just beyond the edge, but not so far that you can’t grapple with them. The question of why there are 3 dimensions of space is a really cool question, but I don’t know of anyone who has said anything really compelling about that. It’s just a question that’s beyond the edge.

As a professional physicist, I don’t spend my time thinking about this because I just don’t know of any pathway to solving the question. The process of training as a physicist involves coming to grips with what the edge of knowledge is, because that’s where the interesting, fruitful questions to make progress on as a scientist live.

Often, when you go through graduate studies, you worry, “Oh, my God, I have to learn about Feynman diagrams, all this math, and all these calculational methods.” It’s true; that’s a really hard thing to learn, and it takes a lot of work.

But in some sense, once you become a professional physicist, you should feel like you can learn any tool. You can pick up any tool that is needed for the task at hand. You should develop that confidence, and that’s what makes you a competent physicist.

A competent physicist is someone who can learn any new mathematical tool, piece of code, or whatever is needed to solve the problem at hand. In graduate school, it’s daunting—you have to learn a lot—but by the end, you should have a lot of skills in the toolkit and the confidence to pick up any new one as needed.

The difference between a good physicist and a great physicist is knowing what the right question is to ask. That’s actually the hardest part of being a scientist: knowing what the next fruitful question to tackle is. I think AI right now is a very good physicist—in fact, maybe superhuman—when it comes to certain computations.

It’s like an extremely technically skilled graduate student. You can give it a sharp, well-posed question, and it will do incredibly hard calculations correctly and come back to you with the answer. It’s super competent. But one of the things it doesn’t quite have yet is knowing what the right question to ask is.

Just like with humans, that’s actually the hardest skill to pick up. It’s the one that comes last.

Brandon

I know you’re not working directly on AI so much—I don’t know exactly how much you do—but do you get a sense that you can imagine a future where you just do better reinforcement learning, or change the architecture of the model completely so that it’s something other than a transformer, and the trajectory just keeps going like this?

It’s been a very rapid increase since o1 in terms of reasoning capabilities. Or do you get a sense that we’re getting near the edge of the frontier of knowledge now, so that the ability of the model to recombine knowledge in somewhat novel ways is kind of it?

I don’t want to play down any of these results, but it seems like a lot of what it did was recombination of known facts. Do you have any reason to believe that will continue, or are we going to say, “Okay, we know how to recombine stuff really well, and we can’t push beyond that,” without getting too philosophical?

Alex Lubyansky

I’m not sure that any of us are anything more than recombination of a stack machine working with GPT 4 on this problem. Me working with GPT-4 on this problem feels like working with a creative collaborator. It did things I didn’t know; I found them surprising. I’m not sure there’s a qualitative difference. I think it’s just a matter of degree.

As we continue scaling the capabilities, which is certainly happening, I don’t see why it’s going to stop. We definitely have a bunch of things in the pipeline that are going to keep coming this year. My horizon for seeing into the future is not that good beyond the year, but definitely we’re going to keep scaling up this year.

I don’t see any reason why it’s going to stop, and I think that’s going to make these models display feats of insight that look to us like real creativity. I would say this already happened in this project, at least. What is creative insight is a bit in the eye of the beholder, right?

I mean, AlphaGo, right? It came up with moves that were very surprising. I talked to Terry Tao a couple of weeks ago at UCLA.

We had an OpenAI event with IPAM, which is the Institute for Pure and Applied Mathematics. I talked to Terry Tao, and he said that, in his view, all of the proofs that he’s seen AI come up with in math—even the ones that at first seemed creative and surprising—were later tracked down and found to have really pulled facts out of some obscure reference.

I don’t want to put words in his mouth, but my understanding was that Terry Tao has not yet been impressed by a creative move in math. Terry Tao is a unique individual, though. I’ve been impressed. I consider myself—my bar is lower.

As we keep scaling this up, I can’t go into the details, but there’s a lot of effort at OpenAI. There are a lot of really smart, hardworking people who are pushing very hard to take this next step, and I think it’s going to come eventually. Just look at the trajectory that we’re on.

A year ago, I was a black hole physicist in academia, not really paying too much attention to AI. I thought, “Yeah, it’s cool for emails, but it’s not going to do what I do, which is special.” o3, which was really the first strong reasoning model, came out and was able to do a calculation for me that would have taken me days. It did it in 11 minutes.

I thought, “Wow.” That was shocking to me. We could go into the details if we have time. I could show the example because I saved it. It was really surprising to me.

Then I thought, “Okay, I’ve got to really start using this tool. There’s no other software that can do this kind of calculation, as far as I know. It’s really surprising and really cool.” Then GPT-5 came out 6 months later, and it was able to reproduce one of my hardest calculations, which I think the number of people in the world who could do that could be counted on your hands.

Brandon

When you say “reproduce,” do you mean this has been published or not published? Was it secret or internal?

Alex Lubyansky

Last summer, in June, I put out this paper, which I really like. It’s called “Why Is There No Love in Black Holes?”

Love is actually a technical term. It refers to Augustus Love, a British mathematician who studied the tides. When you have an object like the moon going around the Earth, it exerts tidal forces on the oceans. You can measure the tidal response of the Earth and its oceans to the moon via some coefficients that encode the strength of the tidal response, and these are called Love numbers in reference to Augustus Love.

Famously, black holes do not experience tides, so they have no love. There’s been a resurgence of interest in this fact in the last 5 years because people understood that it can be connected to a symmetry principle.

In physics, whenever something is zero—why should black holes never experience tides?—that’s surprising. Oftentimes, the answer is that there’s a symmetry principle at work that forbids the existence of tides and protects the structure of the black hole.

I found these new symmetries. These are differential operators that act on solutions to this equation, which describes perturbations of a black hole. These generators are symmetries because if you act on the solution to this equation, you get a new solution.

I thought this was very beautiful. I liked it very much, and it came out in June on the arXiv. In August, GPT-5 came out. The cutoff date for its training set precedes the release of this paper, so GPT did not see this paper during training.

When it came out, I thought, “Okay, I’m going to meet Mark Chen, who’s chief research officer at OpenAI.” He said, “Give GPT Pro a really hard problem. Let’s see how good it is.” I thought, “You want a hard problem? I’ll give you a hard problem.” I had just solved this problem and written a paper. I was very excited about it. I thought, “This is really deep and cool.”

I gave GPT the equation and said, “What are the symmetries?” I didn’t tell it that there were symmetries, because the default assumption should be that there aren’t any. It thought for 5 minutes and said, “Yeah, there are no symmetries,” which is what usually happens. That was wrong.

Mark Chen was visibly crestfallen. He said, “Oh. Well, okay, what if you give it an easier question?” Then I gave it the same question, but not for a black hole spacetime—for an empty, flat spacetime, which is a simpler problem. That’s actually how I approached this problem myself: you warm up on the easier question first.

I gave it the flat-space question, which is also in this paper. It’s this equation, which looks much simpler. This also has 3 symmetry generators, which are shown here. This is not new; these equations have been studied for 200 years. Everything in flat space has been known forever.

GPT-5 Pro thought for about 9 minutes, and it came up with the answer: a very beautiful, perfectly structured, perfectly correct answer. At the time, I also tried the other models from our competitors, and none of them could get this. GPT Pro was really ahead, and I think it continues to be the best model for this kind of mathematical physics work.

Mark Chen said, “Okay, this is great, but now that it’s done the warm-up problem, in the same chat instance, try the full problem again. Now that it’s been primed.” I thought, “Okay, why not?” I gave it the same question as before: “What are the symmetries of this equation?”—now the full black hole problem.

This time, it thought for 18 minutes, which I had never seen before, and it came up with the answer. Basically, in under 30 minutes, with 1 hint—which is the obvious warm-up problem to prime the model on first—it completely solved this problem. It was one of the nicest calculations that I’ve ever done, and that really blew my mind.

That was my Move 37, the Evan moment.

Brandon

Yeah, that’s how we call it in the AI world.

Alex Lubyansky

Once I saw that, I thought, “Okay, we’re on this crazy trajectory.” Eighteen months ago, it wasn’t useful. A year ago, it could do really hard calculations that would take me days. Eight months ago, it could reproduce some of my best work in under 30 minutes. In the last month, it solved these questions that we’ve discussed at length, which world experts had spent a year thinking about without being able to get to the answer.

I think it’s just going to keep getting better. Where are we going to be in 6 months or a year? I don’t see any reason why it would stop. I think we’re going to be having a very exciting year.

Brandon

Going back to these thoughts about scientific discovery and what these models can do versus just being superhuman at solving physics, people keep asking this question: hypothetically, could we train a version of ChatGPT where it’s never seen anything after 1904, and could it rediscover relativity?

I think there’s a very analogous question we could ask right here, which is a new conceptual result about single-minus gluon amplitudes that was sparked by human insight. There were some very specific assumptions that went into this, like understanding that working in Kerr spacetime is something that people have been thinking about and that has some useful, transferable insight. People have been thinking about maximally helicity-violating amplitudes for quite some time.

Have you ever tried using a model whose cutoff date was right before this paper and asked, given a Kerr metric, whether there’s anything interesting with regard to helicity violation? Or maybe turning it around and saying, “It’s long been thought—or long been known—that, with the exception of some set of measure zero due to Witten, there are no single-minus nonzero amplitudes.”

Have you tried either of these directions and asked it to discover a new insight, push the boundary as you were just talking about, and make a leap—in addition to not just solving a problem that you can give it, but actually getting that intuition?

Nima Arkani-Hamed

Yes.

Shawn Wang

You have tried this?

Alex Lubyansky

Not exactly the counterfactual version that you’re describing. I personally haven’t done that, but pushing the models at the frontier to try to make this type of leap is something that we’re very focused on.

I don’t want to talk about the internal research we’re doing, but I can say something publicly, I think. You can take this page from the paper and feed it to ChatGPT Pro—the best model we have right now—and ask it, “What should I do next? Give me the top 3 follow-up questions to ask based on this paper.”

I’ve done this experiment, and the top 3 questions it comes up with are my top 3 questions for what I should do next. I think the models are smart enough now and have enough background knowledge that, for this paper, GPT is about as good as me at finding the next thing to ask. That’s really interesting, and it opens up a lot of possibilities.

Brandon

Can you just do the agent loop where you say, “Okay, what’s the next question? Go ahead and solve that. What’s the next question?”

I guess this goes back to the question I was asking before: if you do that—and you probably have tried it, or something that OpenAI has tried—do you eventually get to some plateau where you’re not pushing the boundary of knowledge anymore? Or is the plateau just money, and if you had more money, you could go further?

Alex Lubyansky

Just to be very explicit, because I haven’t said this quite out loud: I think we now have models that can really turn out papers that are as good as human-written papers.

In fact, this is a bit of a problem because when a professional physicist uses this tool, steers the model, and checks the answer, they can get amazing results. But there are also people who feed it wrong questions that go off the deep end, and then submit that to arXiv. This is a problem that the academic community is trying to come to grips with now: AI slop in science. This is something we have to figure out.

With proper steering, you could probably turn out a paper a day now. If you give a question to ChatGPT, it’ll solve it if it’s not that hard of a question or if it’s a similar calculation to stuff that’s already been done. It can totally do it in 30 minutes, and then you could say, “Write it up as a paper,” and send it to arXiv.

I think we’re already in this moment. We’ve passed that threshold. This is the new reality, and more and more people are catching on to this all the time. Some of them are doing this, and this is why arXiv is now inundated with submissions.

So what’s the correct response to this? We put out these 2 papers in very fast succession. We could spend the rest of the year writing 30 more papers like this, but I don’t think that’s what we should be doing. Instead, now that we have this new tool that gives us AI superpowers, I think we should just raise the bar for what it means to write a good paper. We should aim higher, basically.

One thing that I’m excited about is that I think these single-minus amplitudes papers open the way to a whole direction of research, which I think is a line of attack on really interesting questions in quantum gravity. To go back to the start of the session, this is the missing piece of the puzzle of fundamental theoretical physics. I think we have a pretty clear line of attack through a series of questions, all of which I think will be amenable to solution with AI.

I’m excited to spend a good part of this year trying to follow this path and solve harder and harder problems. This paper gave an answer to a question that had stumped Andy, Alfredo, and David, who are experts in this, for a year. But we haven’t seen an AI solve a question that has stumped an entire community of physicists for decades. That hasn’t happened yet.

Given the trajectory that we’re on, at some point—hopefully not too far in the future—we should see that. I think that’s the exciting thing to try to move toward: pushing the envelope of what can be done.

Brandon

We wanted to ask a question: if you could remove 1 bottleneck for your domain—in this case, maybe it’s AI for physics, maybe it’s physics, or maybe it’s mostly AI—what would that be, and why?

Alex Lubyansky

Well, off the top of my head, I spend so much of my time writing papers. The way I think now is so far from papers that it just feels like they’re not the right way, somehow, to store and communicate knowledge.

I think an extreme version of this, which makes the problem more apparent, is math—especially certain parts of math where papers are very terse and take 4 pages. I had this experience when I was learning algebraic geometry in graduate school. I went to a mathematician and said, “What’s going on in this 4-page paper?” It was just very terse notation.

He said, “Forget what’s in the paper,” and took me to the blackboard and started to draw pictures. He said, “This is how you should think about it.” Then I was like, “Oh wow, this is amazing.” But none of that is in the paper.

Mathematicians, I think, have this cultural norm that they hide the messy work and write these beautiful, short, pristine papers. It depends on the subfield, but oftentimes that’s the case. The way they actually think about the subject as a living, breathing entity is very different from the way in which it’s recorded in papers.

Some of that is also true for physics. I love doing calculations, coming up with questions, and finding the answer. I would say the huge bottleneck is writing it up. Somehow, it feels like papers are not quite the way of the future—or at least the way we currently operate: I write it up, send it to a journal, and it takes 6 months. I don’t know. Why are we doing all of this? It feels like maybe there should be something better.

If you want to understand this paper, one thing you can do is upload it into ChatGPT and ask it to explain it to you. You can keep unfolding the complexity into more and more detailed explanations. If we move to a world where we use AI to do the calculation and get the result, then we have the step of condensing it into a paper. Then I send the paper to Brandon, and he puts it back into an AI. I mean, why are we doing this?

Brandon

Yes, right. That’s a little bit funny.

Alex Lubyansky

If you ask me whether I’d be confident that in 20 years we’ll have these sorts of static documents in which we publish our results as papers, I would think not. That doesn’t seem like the best thing we could be doing.

Maybe some kind of interactive paper that lives in an LLM. Maybe your whole paper is a ChatGPT page, and there’s a chatbot attached to the paper. You can say, “Explain the big picture,” or “Zoom into this fact.” I think we’re going to head in that direction. That would be a cool thing to see.

Writing a paper, though, is a useful exercise because it forces you to condense your thoughts and make them really clear. I’m not saying it’s a bad thing to do in general, but the way we do it is very slow.

Maybe another answer is that, in this project—the graviton paper—we got to a draft of a paper extremely fast, and then we spent most of our time checking the answer. I think that will effectively be a big bottleneck, maybe the next big bottleneck.

That’s one of the things the models are missing. If you ask me what we can really improve in the models for scientific research, I think we’ve touched on the 2 big things already, but just to spell them out: one is creativity, the spark of invention, and really taking the next step. I think that will come as we scale up the intelligence. We’ll see, but I don’t know that there’s something missing inherently. I think it’s just starting to make these leaps for me.

Maybe we should encourage the models to try to make bigger leaps, because large language models, after all, are trained to give you the middle-of-the-road answer. If you ask an AI like ChatGPT, “Write me an email about blah,” you want it to give you the expected answer, not sample from the tails—a wacky email. You want it to give you a reasonable thing.

For most tasks, you want that. But for scientific research, sometimes you want the idea that comes out of left field, the thinking outside the box, or really sampling far out of the distribution. That’s something we could do in principle, but that’s not how the models are set up. We’re not really favoring that, so we might have to make tweaks of this kind to enable the models to take bigger leaps.

The second thing is verification. We’re now in this new regime where the models are so capable that, for very hard computations at the frontier of knowledge, they can just do the whole thing. But is it correct? In this case, it was correct.

Sometimes I get emails from people saying, “I did this really long calculation, but there was a mistake somewhere.” Disappointing. The calculations are getting more and more complicated, longer and longer, but sometimes they mess up.

I think improving verification—or even just having the model indicate more directly how confident it is in its answer—is important. I think they’re smart enough to know whether they’re very confident in the answer versus when they’re just guessing in some step. Getting the AI to be more explicit about that is, I think, a way to improve it for research. That verification step is going to become maybe a bigger bottleneck this year.

Brandon

Yeah, Keren Hong from Axiom would agree with you emphatically. Formal verification is their thing, right?

Alex Lubyansky

Yeah. It’s interesting: a year ago, I would have said it was super important to have formal verification. Then the models got so smart that I thought, well, if Brandon and I talk about a mathematical proof and go over it, we’re not going to formalize it in set-theoretic notation, or reason the way Lean—which is this language for formal verification—does. We reason through the proof in natural language. We use words.

If a model is really smart enough, then it should be able to do the same thing. We’ve been seeing this huge increase in capability for mathematical reasoning and developing proofs using natural language. For a while, it looked like that wasn’t the thing to really focus on.

But now that we’re in this regime where you can just get ChatGPT to tackle thousands of questions at the same time, and it will return proofs for a significant fraction of them, the onus is back on humans to verify all the outputs. If that becomes a bottleneck, I think formalizing math and automating verification will become more valuable. That’s something we’re thinking a lot about as well.

Brandon

Thanks. What do you want the audience to take away from today? Is there 1 message that you want them to leave with?

Alex Lubyansky

Yeah, I think it’s important to get the word out that the models we’re developing at OpenAI are becoming really capable in scientific research.

I myself was a bit of an AI skeptic a year plus ago because I thought the models were very good at writing tasks but not mathematical tasks. That changed with o3, the first strong reasoning models. And then GPT-5 was able to do some of the hardest calculations that I can do and reproduce them correctly. Recently, in the past month, we’ve seen models solve open questions in theoretical physics. And now they’re solving problems in quantum gravity and quantum field theory.

So if you just extrapolate that into the future, imagine where we’re going to be in 6 months or a year. I think it’s kind of surreal to live through this time, but it’s really happening. It’s really amazing. And I think we’re going to see a lot of big changes happening in research.

So, yeah, pay attention to this space. Let’s stay tuned.

Brandon

That’s awesome. Thank you so much for taking the time. I learned a lot from our discussion, and I’m definitely going to keep up with what you’re up to.

Alex Lubyansky

Thank you. It’s been great to be here.

Brandon

Thank you. Thank you.