Brandon
Okay, so I think we're at a special time now where, at least in some directions, AI has become superhuman, at least on certain tasks. That's what led to these recent papers, which resolve a problem that had puzzled expert physicists in the field for over a year. They were unable to resolve it, and AI was able to do so very quickly. I think that's a certain milestone we've passed. You guys are bringing attention to this because, for the average person on the street who doesn't care about theoretical physics, this is not very noticeable, but I think it's a very profound change and we've really passed some kind of threshold.
Welcome to the AI for Science podcast, part of Lean Space Network. I'm Brandon. I develop RNA therapeutics using AI at Atomic AI. I'm joined by my co-host, RJ Honicky, CTO and founder of Miro Omics. It's a pleasure to introduce Alex Lupsasca, a professor at Vanderbilt University and a fellow at OpenAI. For a young researcher, he has quite a storied background. Among other things, he's the winner of the 2024 New Horizons Breakthrough Prize. It's the Oscars for science. I asked ChatGPT whether this is the most prestigious award someone at his career stage could win, and it recommended a second one called the IUPAP award, which turns out to be one he also won. Anyway, right now he's having fun at OpenAI, doing some really cool research pushing the foundations of theoretical physics using GPT models.
Alex Lubyansky
A pleasure to be here. The one message I wanted to convey is that I think we're on a trajectory which I personally find very surprising and surreal, but also amazing. I would say that a little over a year ago, AI was very useful for email, but not for the kind of work that I do, which I consider important theoretical physics calculations. I thought, “That's special. It's much harder than email, and AI is not going to be able to do that.”
Then a series of developments came in rapid succession that completely changed my mind. I can walk you through some of these examples. Specifically, ChatGPT o3 was the first really strong reasoning model that could do actual math that was useful for my research and could save me a lot of time. That's when I started to really pay attention and use it a lot more. I thought, “Wow, this is a great tool. I've got to get ahead of this and learn how to integrate it into my workflow.”
Then, when GPT-5 came out, it was able to reproduce one of my best papers, which took me a very long time to come up with, in about 30 minutes. That's when I really became AI-pilled. I thought, “Oh my God, this changes everything. It's the most important discovery in my lifetime. It's going to affect everything about how we do research.”
Frankly, I would go around telling a lot of my colleagues, “This crazy thing happened. Pay attention.” I was getting lots of different reactions, but I think people weren't quite getting it. I talked to OpenAI, and they were also really excited. I thought, “I don't know that much about AI, but I have to get in on this. To understand that this is happening and not be a part of it is a huge mistake, so I have to go to OpenAI.”
I was on sabbatical, so it was very easy to come here, and I joined the company. Then it just kept ramping up even beyond that. To the point where now I think most of my senior colleagues in physics are aware of where things are headed, and they're all getting on board.
RJ Honicky
Yeah, that's an awesome story. I find it really funny because that story almost reminds me of a lot of different people who had the same realization with Codex sometime last fall, especially. It just took off, and a bunch of people went from, “Oh man, this is 20% of my work. It's kind of a nice assistant,” to, “Oh crap, what just happened?” Even Andrej Karpathy went from that perspective to realizing what had happened.
Alex Lupsasca
Well, yeah, in August, actually, I remember when GPT-5 came out. At that point, I was really following AI pretty closely, and I think on Twitter the reception was lukewarm. A lot of people were saying, “We expected a lot more,” and, “It's not better at writing email.” I remember thinking, “Okay, GPT-3 could write email. How much better can it get at writing email? That's not the point.” But at the science frontier, the capabilities were really taking off.
RJ Honicky
Yeah, there was a lot of attention paid even to o3, but presumably GPT-5 was a huge jump.
Alex Lupsasca
And I think GPT-5.4 is also a huge jump. I don't know if it's noticeable from the outside, although I did see some chatter online. People are running these independent benchmarks, which do show this. I think people are realizing it, and in practice researchers are now all over AI and using it.
I'm getting inbound messages all the time because I'm the resident scientist doing physics at OpenAI, so everybody is sending me papers and chats saying, “Oh my God, this happened.” I got one just this week. Somebody said, “Codex just wrote up a simulation of the SYK model.” This is a very technical thing in quantum mechanics and gravity. A lot of research groups have been trying to run this simulation, and they couldn't do it, but Codex did it in 10 minutes, just because setting it up was so hard.
Brandon
Well, I think it's partly because of the Venn diagram: when you look at the people who have the physics knowledge and the people who have the top coding skills, maybe the overlap isn't that large, although I think it's been growing. But I think in this example there are a lot of really good people in physics with coding skills who would be trying to simulate these things. So I think Codex is just really good now.
Okay. Yeah, nice. I think we're at a special time now where, at least in some directions, AI has become superhuman, at least on certain tasks. That's what led to these recent papers, which resolve a problem that had puzzled expert physicists in the field for over a year. They were unable to resolve it, and AI was able to do so very quickly. I think that's a certain milestone we've passed, and I'm glad that you guys are bringing attention to this because, for the average person on the street who doesn't care about theoretical physics, this is not very noticeable. But I think it's a very profound change, and we've really passed some kind of threshold.
Let's specifically focus on the gluon paper in the physics part, and we can get to the AI part later.
Alex Lupsasca
Okay, so in physics there are 2 basic principles of nature that we think every law, or every theory, should respect. On the one hand, there's the principle of relativity, which, at a very high level, declares an absolute law that cannot be broken: you cannot transmit information faster than the speed of light. But then there's another principle, the uncertainty principle that underlies quantum mechanics, which says that everything is a little fuzzy. Your position and velocity have a little fuzziness to them.
You can see immediately, at this level of description, that there's a tension between these 2 principles, because one is an absolute law declaring that you cannot go faster than the speed of light, and the other is saying, “It's a little bit fuzzy.” This is just to give a sense of how, when you try to write down these principles mathematically, the equations don't really play nicely with each other. It's been a real struggle to come up with physical theories that can reconcile both principles simultaneously to describe the physical world around us.
I would say that the great achievement of 20th-century physics, which is really one of the greatest triumphs in human thought as far as I'm concerned, is the elaboration of this framework called quantum field theory. It's a general framework that can describe the physical forces of nature in a way that accommodates both of these principles.
In quantum field theory, which is our best theory to date, I will say it gets a little technical, but again, I'll try to keep it pretty high-level. What you're trying to compute or describe are the probabilities for certain events to occur. Because you're in this quantum-mechanical setting, you can't say with certainty what's going to happen when you have a certain experiment, but you want to predict probability distributions.
In quantum mechanics, probability distributions are obtained by squaring certain complex quantities. By complex, I don't mean complicated; I mean they're not real numbers. They're real plus imaginary numbers, which we call quantum amplitudes. So the goal of a theory is to predict quantum amplitudes, which are these complex objects whose squares give the quantum probabilities, and that's the most you can say about the outcome of an experiment.
These quantum amplitudes, in particular, include a variety called scattering amplitudes, which describe the following scenario. Suppose you have a bunch of particles that you throw at one another. This is what happens in particle colliders like the LHC at CERN in Geneva. You take a bunch of particles, smash them together, and stuff happens. They interact via the physical laws of nature. Various processes occur, and then other particles come out at the end of the interaction.
A scattering amplitude is the object that describes the probability for a particular type of interaction: we have some particles coming in with some energies and momenta, and some other particles coming out with other energies and momenta. These scattering amplitudes are functions of all the data describing the particles coming in and the particles going out.
In general, you can have arbitrarily many particles involved in an interaction. This is one of the hallmarks of quantum field theory: particles can be destroyed, so you don't necessarily have the same number of particles at the end as you had in the beginning. Particles can be created. Lots of things can happen.
In general, you want to describe all the possibilities, so you want to have an amplitude for an arbitrary number, n, of particles. That's called an n-point amplitude, because there are n particles coming in and out. It turns out in quantum field theory that if you have a particular force and you're able to compute the n-point amplitudes—these functions of the n parameters that square to the probabilities—then you know everything about the theory, more or less. There's always an asterisk, but it's basically the entire content of the theory.
Brandon
So, wait. If you have a theory that tells you any number of particles can come in and go out, then I can say, “I can declare anything about that system.”
Alex Lubyansky
Exactly. Then you know everything. Importantly, these amplitudes aren't just numbers; they're functions, because the probabilities that they compute depend on how much energy the particles have, what their momenta are, and also their polarizations.
A lot of particles, like the photon—the particle of light—have something called polarization. When you look at the surface of a lake and have polarized sunglasses, you can turn your head and see more or less sunlight reflected off the lake. That's because a photon, which you can think of as a little particle of light, carries, as it propagates, a little arrow perpendicular to the direction of propagation. That's called the polarization.
This polarization has a direction, and sunglasses can selectively let in light with one polarization and not the other. As light travels, its polarization can rotate. It can wind. It can do its own thing.
In general, if it winds in a right-handed way—if, as the particle travels, the polarization winds to the right—we call that positive helicity, or right-handed polarization. If it winds in the other direction, we call that left-handed helicity, or negative helicity. These amplitudes, which are the fundamental objects in quantum field theory and contain all the information there is to know about physical forces, depend not just on the energies and momenta, but also on the polarizations.
Now, I've told you about how there are 2 basic principles of nature: relativity and quantum mechanics. They come together in this framework of quantum field theory. I keep talking about forces, and there are 4 fundamental forces of nature.
There's electromagnetism, which is responsible for basically the properties of atomic elements in the periodic table, and therefore chemistry, biology, and everything that you see, touch, or feel. Pretty much all of it is due to electromagnetism: textures, colors, and so on. This force is mediated by the photon, which is the particle of light. That one is the most familiar to us.
Then there's gravity, which is another force that we feel very much because it keeps us to the ground. Then there are 2 nuclear forces: the weak and strong nuclear forces, which we don't really notice directly in our daily lives. The weak nuclear force is responsible for radioactive decay and other such processes, while the strong force, which is the strongest of them all, is what binds the nucleus together.
You learn in high school that like charges repel. If so, then why do protons stick together inside the nucleus of the atom? They should repel one another. Indeed, that's the case, but if you bring them really close, the strong force kicks in and overwhelms the relatively weaker electromagnetic force.
The strong force is mediated by the exchange of particles called gluons, because they're what glue together the nucleus of the atom. So, gluons are the particles of the strong force. Gravity is mediated by gravitons.
Brandon
I think the gluon paper was sort of the starting point for this, maybe—not. The gluon paper had a really specific result, right?
Yeah, absolutely. Maybe let me just flash the paper itself. We put this on the arXiv a little over a month ago. Here's the paper. Let me explain in a few sentences, now that I've given a lot of background, what the title means.
The title says, “Single-minus gluon tree amplitudes are nonzero.” This might sound forbidding, but I think we can unpack this for the audience. Gluons are the particles that carry the strong force, and gluon amplitudes are functions that describe the quantum probabilities for gluons to interact via the strong force.
The word “tree” here is a little bit of a technicality. It means we're only considering processes where no gluons are created or destroyed. If gluons are created or destroyed, then you get loops, which we can explain later, but this is just a technicality. We're considering special interactions where the same gluons that come in also come out.
Brandon
For anyone who's ever fit a polynomial, you can think of trees as being like a linear term, and loops can be higher-order terms.
Exactly. It's way more complicated than that, but conceptually, it's like the lowest order in a series.
Single-minus—now I have to explain that. Remember I told you earlier how particles have polarizations? When you try to study gluon amplitudes, this is a whole industry of physics. It's a very complicated field. People have written thousands of papers over the decades.
You always want to try to understand the simplest examples first. That's why you start with the tree amplitudes, or the leading effects, and then you worry about the loop corrections. You might think that the simplest example to start with is one in which all the particles have the same helicity. Say they're all right-handed, or, that is to say, they're all plus-helicity particles.
It's been known for a long time that, in that case, the amplitude is just 0, which means the interaction is forbidden and cannot happen. That's one way to think about it: it's just a symmetry that explicitly forbids this.
Brandon
So, you don't have to worry about calculating anything. You just know.
Yeah, just dimensional analysis. It's a very general argument.
Brandon
You don't need to do very much work.
It's true that it's the simplest example, but it's so simple that nothing happens. The answer is trivial. You might ask, what about the next level up? What if one of them has the opposite helicity, while all the others have plus helicity? That's what we would call a single-minus amplitude.
If you look at the lecture notes and textbooks that have been written on this, the same argument that rules out the all-plus amplitudes also appears to rule out the single-minus amplitudes. They're too simple. They can't really interact. There's nothing to see here; move on.
Then you might ask, okay, what about the next thing, where there are 2 particles with minus helicity and all the others have plus helicity? If there are n of them, there are n − 2 others that have positive helicity. These would be double-minus amplitudes.
People in the ’80s studied and computed these amplitudes. They're not 0. In particular, 2 physicists, Parke and Taylor, found this beautiful result. They did a lot of really hard work and computed these amplitudes—a very technical, difficult calculation. At the end, you get all these terms and you have to sum them all up. Almost all of them cancel, and at the end, you're left with this very simple formula that fits in half a line. It's now known as the Parke–Taylor formula for these amplitudes.
These amplitudes are now called MHV amplitudes, which stands for maximally helicity violating, because they have the largest—or so we thought possible—asymmetry between the plus- and minus-helicity particles. That's the most asymmetry.
Now, let's get to this paper, which came out last month. This is a paper written with Alfredo Guevara, who's a postdoc at the Institute for Advanced Study; David Skinner, a professor at Cambridge University; Andrew Strominger, a professor at Harvard who used to be my advisor; and also Kevin Weil, who studied as a particle physicist in a previous life.
How did this happen? Maybe we'll get into how I ended up at OpenAI a little bit later, but I ended up at OpenAI and started to improve the models' abilities to do physics. The models got really, really good at physics, and I thought, “Okay, it's so good now. We should try to solve some actual research problems at the frontier.”
I called up Andy, who used to be my advisor, and said, “Hey, Andy, do you want to come here to San Francisco, visit OpenAI, and we can try to solve one of your problems in physics?” I thought, “It's probably not going to work, but if it doesn't work, at least we'll figure out why it doesn't work. I can do this with a different physicist every month, and eventually something will work. In the meantime, we'll learn how to improve the models, so it's all fun and useful.”
Andy was the first one I invited to do this.
And he said, “Well, I have this perfect problem that I’ve been thinking about with Alfredo and David for the past year.” I’ll explain the problem, but the amazing thing is that we decided to start working on it using AI a little bit before Andy was scheduled to come, like the week before. In fact, using ChatGPT, we solved the problem before he even got off the plane.
swyx
Which was a huge surprise to him. [Laughter.]
To him. Yeah, I think to me, too, to be honest. We had not expected that. It’s a really cool story.
So, Andy, David, and Alfredo understood a year ago that this statement—that the single-minus amplitudes are zero—is not exactly correct, because the usual argument in the lecture notes and textbooks has a loophole. The loophole is that it assumes the particles are coming from generic directions. But in a certain regime where the particles are exactly aligned with one another—we say they’re collinear—the usual argument has a loophole, and it’s possible for the amplitudes not to be zero.
But then, if they’re not zero, what are they? Suddenly, these really simple amplitudes, previously thought to be zero, if they’re not zero, we should compute them, and they should be something really nice, simple, and special. Now, I’m sweeping a lot of details under the rug here. This has to work in some split-signature spacetime, and it connects to lots of other things they’ve been worrying about. We’re not going to worry about this.
I was actually hoping at the end that we could talk about what it means to be 2 dimensions in space and 2 dimensions in time, but I think part of this is doable.
swyx
Really mind-bending stuff.
The loophole is one about the alignment of the particles, but it’s also a loophole about the spacetime of the physics—the universe we’re living in.
So, they understood that they’re not zero, and they started to compute them. Alfredo is really, I think, the unsung hero of this story because he did a lot of really hard work to compute these things by hand. I’ll just show you an example.
In the paper, there’s a lot of formalism. Here is the beginning of the definition of the general answer. Then you have to define these vertex objects, V, and they’re complicated. They involve spinor-helicity functions of spinors, and then you have this recursive formula. It’s a whole mess.
Concretely, if you try to unpack this definition, remember these amplitudes are a function of the number of particles involved. There’s a 3-point amplitude where there are only 3 gluons in the interaction, and the answer is pretty simple. This is some function that we’ve defined here—not that complicated. Then this is the 4-point amplitude, where now there are 4 particles, and you can see that we go from 1 term to a sum of 2 terms here.
But then, once you get to 5 particles, you start to get a lot more terms. There are 8 of them being summed here. By the time you get to 6 particles, it explodes in your face. For those people not watching this on YouTube and listening, this equation takes up a quarter of the page. It’s 32 terms, each of which is a product of 4 terms, each of which is itself encapsulating a rather complicated formula.
swyx
Yeah, this is super nasty, and that’s as far as Alfredo got—or Andy, whatever. So Alfredo did it. Look, is this just an expansion of some sort? How hard is it to do this expansion?
Brandon
Very hard.
swyx
Okay.
David Skinner
Yeah. There’s a nice graphical way to understand this in terms of Feynman diagrams. I hadn’t planned to explain this, but it’s a visual subject.
The math is very complicated, and already back in the ’40s, Richard Feynman, who was one of the pioneers of quantum field theory, came up with this very visual way to organize our understanding of the subject. You can doodle these little cartoons that represent possible interactions.
The rules of quantum mechanics actually say that in these amplitudes, where you scatter a bunch of particles, you get to fix what comes in and what comes out because that’s the question you’re asking: What’s the probability for a certain interaction? But then everything that happens in between, you don’t get to choose, because the physical laws determine what happens.
In quantum mechanics, you’re supposed to consider all the possibilities—all the ways in which the incoming particles can interact and transform into the outgoing particles. You’re supposed to average or sum over all the possibilities to get the final amplitude for the process, as a sum over the amplitudes for each individual possibility for how you could get there.
swyx
So, just to be clear, there are incoming particles. They interact, and then there are all these different possibilities. They each have their own amplitudes, and then I select this one possibility and this one possibility, and I get one possible interaction. There’s an infinite number of those for each, and then I sum those infinite possibilities, I suppose, and I get the outcome?
Yeah. In principle, there are infinitely many pictures to sum over, but that’s why we organize them by how complex they are. It turns out that every time you get an interaction—every time there’s a vertex where 2 lines meet—that point interaction comes with a power of the coupling constant, which controls the strength of the interaction.
Every additional interaction makes the amplitude more suppressed, so it contributes less to the final answer. You want to first consider the diagrams with the fewest possible number of interactions because they will give you most of the final amplitude. Then, if you’re trying to get a more and more refined answer, you consider the more and more complicated cartoons with more and more interactions.
In fact, this is one of the ways in which the diagrams can get complicated: They can have loops. For instance, here you have a particle that decays into 2 particles, creating this loop because then they meet up again and disappear. In this interaction, you have intermediate particles being created and destroyed.
But whenever that happens, you get 2 extra vertices in your graph. These diagrams are suppressed because it’s less likely that you get these extra felicitous interactions, so you don’t need to worry about this as much. It’s like a small correction. Of course, in principle, you could keep going, but you’re never done except in very special circumstances.
swyx
Or higher-order powers in a polynomial or something. Or a Taylor series.
And so, to go back to the story back in the ’80s with the MHV amplitudes, which I think now is a bit of a misnomer, I would call them double-minus amplitudes because that’s where we’re going to get to in a second.
swyx
Right.
David Skinner
There was this heroic calculation where a lot of Feynman diagrams were summed. They were considering more and more interactions with more and more particles, and every time there were more and more terms, but they all canceled and in the end always gave a simple answer.
In fact, that’s what this PT term is. PT stands for Parke–Taylor. These formulas fit in a line, so it’s not that complicated, but it’s very surprising that such a messy calculation would clean up into such a simple result.
What Alfredo, Andy, and David did was understand that these single-minus amplitudes, in the special case where some of the particles are aligned, don’t have to be zero. You can do this very complicated Feynman diagram expansion to get the answer, which is not zero, but the problem is, if you do it this way, you can represent the answer in some horrendously messy, complicated way. If you unpack it, it’s extremely complicated.
When you consider the n-point amplitude—the probability of n particles interacting—the number of terms in your answer, which roughly corresponds to the number of diagrams you have to add up, grows factorially in n, the number of particles. Factorial growth is really bad. It’s superexponential; it goes faster than an exponential, so it blows up in your face. This is what you’re seeing here.
That’s because, roughly, you have to draw all the possible cartoons, and the possible combinations are a combinatorial problem. That’s where the factorial behavior comes from. But we know from the ’80s that in the actually more complicated double-minus case, Parke and Taylor found this miraculous simplification.
Andy, Alfredo, and David spent the last year chasing the analog of the Parke–Taylor formula—the very simple answer that was obtained in the ’80s for the double-minus amplitudes—but now for these single-minus amplitudes, which they understood are not zero. But then, what are they?
They were getting this really complicated answer. You never know in physics ahead of time if something will simplify. You have to believe in it to find the simplification. But because the double-minus one simplified, it felt like these should simplify, too.
We think they’re important for lots of things, and that these are somehow really important objects that are very fundamental. They should have a nice description. They spent a year looking for that.
swyx
There’s a funny—the next line, if you scroll down, is something like, “We need a simpler formula.”
Right. “A more concise formula is needed.” And this is where AI comes in. When I asked Andy, “Hey, do you have a problem in your pocket that we should use AI to target?” he said, “Well, I have just the perfect thing for you. We’ve been puzzling about this.”
swyx
It's really important. It's really interesting. It connects to all these things, and we don't know the answer.
When I was a grad student, if I had approached something like this, I probably would have plugged it into a computer algebra system, let it chug along, tried a few limiting cases, and seen if there were any magical simplifications. This type of thing is something where you oftentimes see, “We need a different approach.”
Exactly. Before Eddie even got here, we started to play with ChatGPT. Alfredo, Andy, and I were trying different things, with lots of different chats happening and going back and forth. David was involved as well.
The first thing that happened is that we fed the 5-point amplitude into ChatGPT and asked, “Can you simplify this?” It said, “There’s a special region, so there’s an extra assumption that you can make in which this answer simplifies to this one.”
So this assumption is equivalent to having 1 particle coming in and decaying into n − 1 of—
swyx
That’s one way to think about it, roughly. Okay, but we’re in 2 spacetime dimensions, so—
Yeah, it’s complicated. But basically, you can look at what we call phase space. It’s the entire space of possibilities for all the energies and momenta of incoming particles. There’s a special region in that phase space where 1 particle has a different sign of its frequency compared to the others.
In that region, there’s a big simplification that happens, which ChatGPT found. I should say this was the public model, but the Pro version that thinks really hard.
swyx
Was that a known fact that it was just able to relate to the problem, or was that something it put together?
As far as I know, it put that together. It said, “This 5-point function, which is a sum of 8 terms, each one of which is a product of 3 terms—they’re all pretty complicated.” It said, “Actually, this simplifies to this product of only 3 terms.”
We stared at this thing and thought, “Wow, that’s really nice. We didn’t know this.” In hindsight, once you know it, you can rederive it, but it takes a while to understand where this comes from. I think that was a leap of insight that the AI had.
At some point, it said, “I wrote Python code and ran through all 5,000 possibilities, and I deduced this.” It’s the equivalent of running a computer algebra system, but it just decided to do it on its own and came up with a huge simplification.
Brandon
Great. Yeah, awesome. Was this after making the assumption? Was this after the 1-particle decay assumption?
Alex Lubyansky
Yeah. It figured out there was some region in which things simplified. This was very experimental; we were talking about it a lot. It figured out there was a special region in which things simplified, and then ChatGPT came up with that simplification as well. All of them.
Then we were like, “Okay, let’s give it the 6-point function, which Alfredo heroically computed.” We didn’t have the 7-point function. I don’t think anybody could use the heat identity to expand it; it would be disgusting.
ChatGPT did its little thing, and then it was like, “Yep, simplifies to this.” We thought, “Whoa, okay, that is really nice.” Instead of 32 terms, it reduces to just 4 terms. It’s not a sum of 32 terms; it’s a product of only 4 terms.
Then we asked ChatGPT, “Can you guess the general formula for all n?” You could imagine using some programming language or symbolic manipulation software to do these reductions in specific examples. But to tackle the general case, I don’t know how to use a computer to do that. ChatGPT said, “Yeah, this is the answer in the general case.” Boom.
swyx
How long does that take?
Using Pro, it thinks for 20 minutes at a time. You go back and—
swyx
But it wasn’t like 6 days or something?
No, no, no. It was just over the course of several interactions.
The amazing thing is that the formula it proposed, instead of having this factorial growth—which is superexponential, where the number of terms blows up as you consider n, the number of particles, increasing—is actually linear. If you double the number of particles, you only double the number of terms. It’s the nicest possible behavior you could imagine.
This is the equivalent, I think, of the Parke–Taylor formula for the double-minus amplitudes that was known back in the ’80s, but now for the single-minus amplitudes. This was guessed by GPT-5.2 Pro, but it couldn’t quite derive it. So I said, “Hmm, it looks like this, but I don’t know how to prove that.”
Alex Lubyansky
Yeah, I think the model was not quite strong enough to prove it. Part of my work at OpenAI has been to develop stronger physics capabilities in the models. A lot of people have been adding lots of things; it’s not just my singular contribution. There’s a lot of great research happening, and it all comes together. It takes a village.
We had this internal model that could think for a very long time and was extra strong at physics. We gave it the whole problem from scratch without actually giving it this formula. We just formulated the problem in a very sharp way and asked the model to find the amplitude in the general case in this region, because we had identified that this was the special place to look.
It took 12 hours, which is a long time, but it came back with the same formula, which we had not given it. It rediscovered the correct formula. This time, it also found the proof that the formula is correct and derived it.
In fact, the remainder of the paper after we state the equation is devoted to the proof, which is basically what came out of the AI. We say, “The rest of this work is devoted to proving that the conjecture is correct.” There are 3 steps: first, you show this; second, you show blah; and third, you show blah. This is basically what the AI came up with.
Now I can finally summarize the paper. The title is *Single-minus gluon tree amplitudes are non-zero*. These are special interactions between gluons where only 1 of them has a different helicity from the others. They were previously thought never to occur, but these interactions can actually happen. The amplitudes are non-zero. That’s the main claim of the paper.
I think it’s quite surprising. I think it’s a really nice paper. The final result, I guess, has 2 parts. One is understanding that it’s not zero. That came from the humans a year ago. They were trying really hard to find a simple answer for what the amplitude is, and they were stumped for a year.
They were able to get this indirect representation, which is extremely complicated in terms of Feynman diagrams. But they were looking for the simple formula that is analogous to the Parke–Taylor formula from the ’80s for the more complicated amplitudes. That was done with the AI, and I think that’s a really interesting result.
Brandon
Yeah, it totally changes the way you should think about where we are in physics and how AI is going to change that. It’s a result that top researchers in this field were thinking about for a year, and then the AI solved it.
There are several things about the story that I think people didn’t understand on Twitter. Maybe scroll down to equations 35 to 38. I would say most people, even introductory grad students, would look at equations 35 to 38 and say that 39 is actually a very natural extension of this. I don’t think that’s that surprising. I think it’s interesting.
I didn’t know until just now that—[clears throat]—when you proved 39, that was a fresh session. That was without the limiting cases. You started from scratch.
Alex Lubyansky
Yes. I did it that way because it’s an extra way to be confident in the answer. If a different model independently comes up with it from scratch, then you’re not just spoon-feeding it the answer that you think is correct. That’s an extra confirmation.
We thought a lot about how to put this out into the world, and there’s no perfect way to do this. We could clearly have done a better job of communicating it. One thing that was important to us was not making this paper about AI, because I think this is a really interesting physics result. People will keep reading this paper, I hope, for a long time.
We didn’t put AI in the abstract because this is a physics result that stands on its own. There’s 1 paragraph about AI where we say, “The final formula was first conjectured by GPT-5.2 Pro and then proved by an internal OpenAI model.” That’s what happened. It’s true, but we didn’t really want to get into it because I don’t think that’s the point of the paper.
It’s really interesting how it happened, but the result stands on its own. If you read a paper today that was written 20 years ago and used a computer to do some critical step in the argument, and it had a whole discussion of how, “I loaded MS-DOS 3.1, it had 5 floppy disks, and I had to swap my floppy disk,” you wouldn’t care. That’s not why you’re reading the physics paper today.
We didn’t really want to go into that in the paper. We talked a little bit about it in the blog post that we released with OpenAI, which is this one. On Twitter, there were a lot of questions, and I wrote some tweets that I think clarified it. There was also a physicist who wrote a great blog post about actually understanding the story.
The Economist also put out a great article about it, and they really understood what happened. I thought it was great coverage. *Science* magazine also wrote about it. Harvard and the Institute for Advanced Study put out press releases. I think it got a lot of attention, but it’s kind of a subtle thing to explain.
It took us an hour to go through what happened and what was done, so it’s hard to explain. I think it would have been a distraction from the physics point of the paper to go into that.
swyx
Okay, let’s talk about the physics, then. Give us a sense, because my theoretical physics on the frontier comes from PBS Space Time, right? It’s a great channel.
Yeah, it’s a great channel, but it gives you a great high-level picture. It’s hard to know how this sits in the pantheon of papers that represent the cutting edge of theoretical physics.
swyx
Not exactly that. I want to understand: It seems like you’re comparing it with a previous result that is pretty significant, highly cited, and very important. How does this compare with that?
Okay, you’re putting me in a bit of a tough spot. I will say I think the result is surprising. That’s why the title is what it is: “Single-minus gluon tree amplitudes are non-zero.” If you’re somebody who works in this field, that should catch your attention.
Ultimately, it’s very hard to know in science, when you release something into the world, how it’s going to be received and how impactful it will be. I think the true value of a paper can only be assessed decades into the future, based on how much future work it leads to and what developments it opens up.
swyx
Maybe a better way of asking is: My understanding is that the previous paper opened up a whole line of thinking about—
Yeah, I think this is a great segue to the second paper that came out just 3 weeks later.
Brandon
Perfect. Then let’s talk about—
So, it got its own blog post. This is March 4, so I guess 2 weeks ago now. We were talking earlier about how there are 4 forces: the strong force, mediated by gluons, and gravity, which is mediated by gravitons. We can produce gluons at the LHC and measure their effects fairly directly. Gravitons, we think, are also around us, being produced all the time, even as I move my hands, but we’ve never done an experiment that directly measures gravitons.
They’re supposed to be the quantum of gravity, so they’re really interesting from a theoretical standpoint. Going back to RJ’s question earlier, what is a graviton? There are different answers we could give. Ultimately, the correct answer depends on what the theory of quantum gravity is, which we don’t know yet.
swyx
Yeah.
Alex Lubyansky
If you just naively try to take all of the tricks from field theory that we know from the Standard Model and apply them to gravity, things just break down. The theory is not self-consistent. There are various problems.
swyx
Yeah.
Guest
Just like in this room there’s light flowing around, there’s some indivisible bit of light that you eventually can’t break up into smaller bits. That’s the quantum of light. We call that the photon. The gravitational force is mediated by the exchange of gravitational force or gravitational waves.
If you try to take a gravitational wave and break it up into smaller and smaller pieces, at some point you get a quantum that you can’t break up anymore, and that would be the graviton. That’s how we understand them.
swyx
Okay, so the idea is that you can’t—you get to a certain point, and you can’t have less gravity than that. You either have some or none. Right? That’s one way to think about it?
Alex Lubyansky
Yeah. We wrote this paper, which is called “Single-minus graviton tree amplitudes are nonzero.” It’s almost the same title, except with “graviton” instead of “gluon.” That’s on purpose, because we wanted to extend the result.
It’s the same story in the sense that it was thought that all single-minus amplitudes are zero, but actually that’s not true for gravity either. Gravity is a lot more complicated, though, so if you want to compute the graviton amplitudes, it’s potentially a lot harder.
Brandon
Do gravitons have phase the same way that gluons do? Is that it?
They actually have spin 2 rather than spin 1. It’s getting into the weeds, so the numbers you have to use to describe them are a little bit different. They’re doubled in some sense.
Brandon
Okay, so their polarization is more complicated.
I see.
The special region in which the final answer simplifies has 2 labels because it’s a spin-2 particle, whereas in the gluon case there was only 1 label because it was a spin-1 particle.
Brandon
So this is like—
It’s not the same math. Gluons and gravitons do have some structural similarities compared to other types of particles.
Well, yes, in the sense that they’re particles of force.
Brandon
Yeah, yeah, but they’re sort of doubled.
Yeah, they’re sort of doubled. I guess the people watching this podcast probably like to geek out on this.
The modern definition of a particle in quantum field theory, which is our best-verified framework for nature, is that particles are irreducible representations of the Poincaré group.
swyx
We just lost 90% of our audience right now.
Yeah, okay. Maybe we cut this. There are mathematical representations, and they’ve all been classified. All the possibilities are known by Wigner, actually, a brilliant physicist.
It turns out that the representations of possible particles are completely labeled by the mass, spin, and charge of the particle. These are the 3 quantum numbers. Particles of long-range forces like gravity and electromagnetism have 0 mass. They have to have integer spin. Spin 1 is 3 of the 4 forces, and spin 2 is gravity. And then that’s it.
But let’s set that aside. The really cool thing about this paper is that, first of all, it came out 3 weeks after the first one, which is really fast. I think this is a great example of AI accelerating science.
In fact, we could have put this paper out 3 days after the first one, because that’s how fast we got the answer out of ChatGPT. But it took us 3 weeks because we wanted to check very carefully that it was correct. Most of the time was spent verifying the answer, not writing, which is insane, actually, if you take a step back.
If you told me a year ago, “You’re going to have this AI that just does really hard calculations for you, and then most of the human effort goes to verifying the answer,” I would have thought you were crazy. So, it’s very surreal.
We also had to write it up as a nice paper, which involved putting in the citations and references. That takes some time. I also had a baby in the meantime, so we lost some time there.
But we did this really fast. I think it’s an example of accelerating science. Another really cool thing is that, for this paper, we didn’t have to use an internal OpenAI model that had to think for hours. This was all done using the publicly available GPT Pro.
In fact, we shared 1 of the main prompts that we used. If you go to the blog post “Extending Single-minus Amplitudes to Gravitons” and scroll down to the text, there’s a link to 1 of the chats that we used. You can see we used GPT-5.2 Pro.
The amazing thing about this is that we gave it the gluon paper as a seed. We said, “Read and understand the paper. Make sure you understand the manipulations in the appendices, because that’s where most of the hard work goes.”
It comes back and says, “Yep, I understood the paper. Let me focus on the appendices. Here’s what happened.” Basically, the punchline is that GPT Pro, with the gluon paper as an anchor, was able to do the graviton calculation, which is really different mathematically, completely on its own—not from scratch, I guess, but from the gluon paper. It’s just a different thing, and it was strong enough to do it completely.
swyx
So it took the conceptual leap from the previous paper and just said, “Okay, what math do I need to make that same conceptual—”
Yeah, and it’s different math. That’s an important thing to emphasize. In particular, there’s a crucial application of something called the directed matrix-tree theorem.
Alfredo and David—we’ve been thinking about these things for a very long time—we were like, “Whoa, that’s really cool. That’s surprising. We hadn’t thought of that or seen that before.” That was known math, but maybe because it has such a broad understanding of math and physics, it’s able to say, “Oh, this is a good thing to apply in this case.”
Yeah, exactly. Here it understood the gluon paper, and then we said, “Okay, well, the task is to generalize this paper to the gravity case. Here are 2 key changes, but otherwise the manipulations should be similar.”
We tweaked some things at the get-go. Then we said, “Good luck. You’re a brilliant theoretical physicist.” We gave it 2 paragraphs. We gave it the gluon paper, a couple of paragraphs, and said, “Good luck.”
It thought for 20 minutes, and boom, it starts to think. It starts at the beginning and works through the implications. Really interesting stuff. Then it says, “Here’s what I would do next to turn this into the gravity paper. If you want, I can do blah.” And so we said, “Yeah, go ahead.”
Another thought for 31 minutes.
swyx
Thought for 31 minutes.
Yeah, this exchange is 110 pages. But I think it’s hilarious. I would describe this as vibe physics.
Because you can see its reasoning as it goes, it does a lot of hard work. It goes through lots of equations. It starts to do the—okay, now you have to use this different math; you have to use these tree calculations and loop-reduction formulas. There’s a lot happening: sums over trees and concrete checks.
One of the things I love is that it’s able to do the same things that a human would do: check some basic cases as a sanity check and to get intuition. It comes back every 3 minutes and says, “Here’s what remains to finish the full gravity paper.” Then there’s a list: “If you want, I can write the gravity analog.” “Yes, do that. This is the first step.”
It goes back and thinks for 34 minutes. Half-collinear support—it starts to do stuff. These formulas actually made it into the paper in some form. This is all correct. There’s a bunch of stuff.
At the end, it says, “If you want, the next most useful thing I can do is this.” And we’re like, “Yeah, verify this by performing the explicit check.” It goes on, and, just to cut to the end, finally we say, “Okay, write up the paper.” You can see the paper that it writes, and it’s very close to the final thing we actually put on the arXiv.
Brandon
Did it make suggestions that were not what you would have suggested as the next steps?
Alex Lubyansky
It’s very smart. It knows kind of where to go, and it’s useful to steer it. If you compare what it came up with with the actual paper that we put in, the intro—the abstract and introduction—was written by Andy, who’s an amazing writer. I think he gave this wider perspective on the problem, how it fits into physics, and how it connects to other things that the AI didn’t do. The intro it wrote was more generic.
But AI could write really well. We didn’t really try to make it. The other thing is that we added Section 2, which was not part of that initial exchange. It’s about how these graviton amplitudes transform under certain symmetries of physics.
That’s something that we’re really, really interested in because we eventually want to understand quantum gravity, as I mentioned earlier. Typically, the first step to uncovering a new theory is to understand what its symmetries are. That’s something that gives you some kind of ground to stand on.
In particular, Andy has been pushing this program of celestial holography, which is a whole thing we could get into, but it’s an exploration of the symmetries of quantum gravity. He really wanted to understand this. There’s a separate chat—we didn’t share that one—where we led the AI to explain how these answers fit into the symmetries that we know the theory should have. That’s something that went in there.
Actually, I think from Section 3 onward, it’s pretty much very close to what the AI wrote. I would say this is really remarkable. It’s a real, solid result in quantum gravity that was done pretty much completely by an AI, with humans steering it and asking the right questions.
All the math was derived by ChatGPT Pro, the public model you can access. Most of the time spent by us humans was checking everything and writing it up. That’s really wild. [laughter]
Brandon
As a physicist, you find yourself where a lot of coders have found themselves, where there’s a fundamental, maybe epistemological, question here. If, as a physicist, I could have done that—maybe I needed a little more background, but a lot of it was, “Yeah, go ahead.” Take this paper and give it some prompt.
You guys obviously prompted very well, but there wasn’t much more than that. Maybe an undergraduate in physics could have come up with a lot of it. How does the undergraduate in physics now learn when they don’t have to do the hard calculation themselves?
Alex Lubyansky
You’re opening up many different strands of conversation, which are all super interesting. Let’s try to unpack that a little bit. The most direct thing you asked is: How does the next generation learn?
Brandon
Yeah.
Alex Lubyansky
That’s a really good question. I think about this a lot. Now that a lot of senior physicists in the field are coming to grips with these new capabilities, one of the questions that comes up very quickly is, “How do we train the next generation?”
The way we were trained was by going through these difficult rites of passage, where you have to do these really arduous calculations. That’s how you build confidence in your own abilities and test your knowledge. It’s not just about what you’re capable of doing; it’s about knowing that you’re capable of doing it, proving it to yourself, and building that self-confidence. That is important. We don’t have a good answer. This is something that academia is going to have to grapple with.
One thing that is especially difficult is that, as a professor, I have graduate students. The gap between where classes take you—even graduate courses—and where research begins is actually huge, and it’s growing wider. Classes go very far, but only so far.
Usually, as a professor, when you take on new students, you keep in your pocket a few easy problems, in the sense that you know they’re going to work. There are some questions that you know, in principle, you could work out—not that they’re that difficult—but you give them to a student so that they go through the exercise of learning everything around the question and developing the technology.
You know enough about the problem that you’re sure there’s an answer, that the student can get there, and that you can advise the student in the process of discovering it. I think the issue is that many such problems now, I would say, these models can probably crush.
These are problems that we usually take—again, the time scale for a theoretical physics paper is 6 months to a year. That’s pretty typical. So if you tell a student, “Go away and think for 6 months about this one question. You have to work really hard, learn a lot of stuff around it, and do lots of calculations,” would even the most determined students go 6 months without asking ChatGPT for help? That’s a little bit weird.
It’s also an opportunity. I remember that time in my graduate school career. In my second year of grad school, I had taken all my graduate courses in my first year, and my second year was my first project. It was actually the hardest time for me in graduate school: traversing the desert from where classes take you to the research frontier.
It’s very hard, and there’s a lot of time spent banging your head against the wall. All the time, you’re confused and you don’t understand things, just because you need to absorb so much knowledge. AI can totally help you with that. It’s the best teacher and knows everything. It can unpack any complicated fact to any desired level of detail.
Actually, my experience as a trained professional physicist working on my own research using GPT now is that there are 2 key ways in which my research has completely changed. One is that I spend much less time being confused. I’ll do a calculation, get an answer, and think, “How does this fit in with this other fact that I know? How do I reconcile these things in my mind? I’m confused.”
Alessio Fanelli
Yeah, I do that all the time.
Nima Arkani-Hamed
In research, usually you take a step, and then you hit a roadblock or an obstacle. You’re confused, and you have to think for a few days. Maybe you go for a walk, work on another project, come back, and get a new idea. You spend a lot of time confused. That’s the nature of research.
With GPT, I’m like, “Hey, I just did this. I found this. How does this mesh with this other thing?” Then it’s like, “Oh, well, you forgot this thing,” or, “Oh, you didn’t quite think about it correctly,” or, “This is the standard fact.” The amount of time you spend confused dramatically shrinks, and you move so much faster. That’s one of the accelerating effects.
The other accelerating effect is that I only have so much free time and energy. Especially when you become a professor, you have to teach, you have students, and you have grants to administer. There are a lot of things you have to do. Your free time to think about research without distractions shrinks, and you only have so much energy to do hard calculations.
What you would usually do is, if you have a problem, you’re at point A and you want to get to point C, you think about the route. You think, “Oh, I have to go through point B first.” Actually, maybe there are multiple points, and you try to plot in your mind the course that you’re going to take before you start doing the hard work. You try to think really hard about where you’re going and chart a course.
With AI, you can launch 10 instances of a chat and have each one try a different route. You can send it as a scout that moves very fast into the unknown, pushing outwards, and very quickly get some feedback. You can see which approaches are not promising and which are much more promising.
If you follow them, there’s a huge difference between being the first to push into the unknown and following someone ahead of you. Even if ChatGPT doesn’t always get everything right, just having a scout that signposts some key steps along the way, which you can use to anchor your own movement, is extremely helpful.
Those are 2 concrete ways that AI has changed the way I work. If you’re entering research, having an assistant that can help you find your way to where you’re trying to go can be very good. It’s inevitably going to change how we work, how we operate, and how we train students.
Part of what’s exciting about my job is trying to figure out how all of this works. It’s not just a job for OpenAI; it’s actually a job for every researcher and professor more generally to think about this.
I think the future is very bright. We have some challenges to overcome, but on balance, this is such an amazing tool. I think it’s going to give human physicists AI superpowers because of what I just described. You can do so much more.
The kind of skill that is really useful to get great results out of AI is very similar to the kind of skill that you develop as an academic collaborating with other humans. This is like a collaborator. If you’re a professor who’s been advising students and postdocs, you know that a lot of what being a professor involves is knowing, for each student or postdoc you’re working with, exactly what question to give them.
It’s matching the problem to the person and knowing how to give them the question—with how much detail, and at what level of detail. Not too much and not too little. That’s actually what you have to think about when you interact with ChatGPT. It’s a transferable skill, and people who are good at this are about to get AI superpowers.
Alessio Fanelli
What you just described reminds me of several conversations we’ve had on the podcast so far, which keep coming back to this concept of taste. One thing that, especially in theoretical physics—high-energy physics—has maybe had a problem with, although I’m not sure if you want to describe it that way, is that it can be very trendy. Certain things become in fashion because maybe right now we’re in a world where we don’t have the data to define new directions, to really guide or constrain where we’re going.
I’m curious: how does something that is superhuman, in that it has basically all known physics, interact with a field where, at its core, what can often become popular—what people start working on—is based more on general aesthetics or what the community collectively thinks is cool at the time?
I can imagine it could vibe with so many different worlds. For example, just using Klein space, using this sort of 2+2-dimensional setup for this, was already an assumption that I think is actually kind of important in some ways and does provide feedback to our world.
But you could have asked ChatGPT to solve this problem in all sorts of ways, and maybe it could come up with all sorts of things that don’t really align with the useful taste of the community. How do you actually deal with a proliferation of really interesting results when it’s not clear where the field should go?
Nima Arkani-Hamed
You’re getting at the heart of what it means to make progress in theoretical physics and research. This is a hard question, and there is a simple answer: if there were, it would be research.
Let me say a couple of things. The first one is that when you go to graduate school in physics, it’s usually because you’re really interested in the big questions: Why are there 3 dimensions of space? What happened at the Big Bang? What’s inside a black hole?
These are the things I was thinking about because of science-fiction movies and books. What you realize is that, actually, even though these questions are really cool and exciting, they’re not really the most fruitful scientific questions. At any given time, there’s an edge of knowledge, and the role of scientists is to expand the edge of knowledge—to push into the unknown.
To do that, you want to find the questions that are right at the edge or just beyond the edge, but not so far that you can’t grapple with them. The question of why there are 3 dimensions of space is a really cool question, but I don’t know of anyone who has said anything really compelling about that. It’s just a question that’s beyond the edge.
As a professional physicist, I don’t spend my time thinking about this because I just don’t know of any pathway to solving the question. The process of training as a physicist involves coming to grips with what the edge of knowledge is, because that’s where the interesting, fruitful questions to make progress on as a scientist live.
Often, when you go through graduate studies, you worry, “Oh, my God, I have to learn about Feynman diagrams, all this math, and all these calculational methods.” It’s true; that’s a really hard thing to learn, and it takes a lot of work.
But in some sense, once you become a professional physicist, you should feel like you can learn any tool. You can pick up any tool that is needed for the task at hand. You should develop that confidence, and that’s what makes you a competent physicist.
A competent physicist is someone who can learn any new mathematical tool, piece of code, or whatever is needed to solve the problem at hand. In graduate school, it’s daunting—you have to learn a lot—but by the end, you should have a lot of skills in the toolkit and the confidence to pick up any new one as needed.
The difference between a good physicist and a great physicist is knowing what the right question is to ask. That’s actually the hardest part of being a scientist: knowing what the next fruitful question to tackle is. I think AI right now is a very good physicist—in fact, maybe superhuman—when it comes to certain computations.
It’s like an extremely technically skilled graduate student. You can give it a sharp, well-posed question, and it will do incredibly hard calculations correctly and come back to you with the answer. It’s super competent. But one of the things it doesn’t quite have yet is knowing what the right question to ask is.
Just like with humans, that’s actually the hardest skill to pick up. It’s the one that comes last.
Brandon
I know you’re not working directly on AI so much—I don’t know exactly how much you do—but do you get a sense that you can imagine a future where you just do better reinforcement learning, or change the architecture of the model completely so that it’s something other than a transformer, and the trajectory just keeps going like this?
It’s been a very rapid increase since o1 in terms of reasoning capabilities. Or do you get a sense that we’re getting near the edge of the frontier of knowledge now, so that the ability of the model to recombine knowledge in somewhat novel ways is kind of it?
I don’t want to play down any of these results, but it seems like a lot of what it did was recombination of known facts. Do you have any reason to believe that will continue, or are we going to say, “Okay, we know how to recombine stuff really well, and we can’t push beyond that,” without getting too philosophical?
Alex Lubyansky
I’m not sure that any of us are anything more than recombination of a stack machine working with GPT 4 on this problem. Me working with GPT-4 on this problem feels like working with a creative collaborator. It did things I didn’t know; I found them surprising. I’m not sure there’s a qualitative difference. I think it’s just a matter of degree.
As we continue scaling the capabilities, which is certainly happening, I don’t see why it’s going to stop. We definitely have a bunch of things in the pipeline that are going to keep coming this year. My horizon for seeing into the future is not that good beyond the year, but definitely we’re going to keep scaling up this year.
I don’t see any reason why it’s going to stop, and I think that’s going to make these models display feats of insight that look to us like real creativity. I would say this already happened in this project, at least. What is creative insight is a bit in the eye of the beholder, right?
I mean, AlphaGo, right? It came up with moves that were very surprising. I talked to Terry Tao a couple of weeks ago at UCLA.
We had an OpenAI event with IPAM, which is the Institute for Pure and Applied Mathematics. I talked to Terry Tao, and he said that, in his view, all of the proofs that he’s seen AI come up with in math—even the ones that at first seemed creative and surprising—were later tracked down and found to have really pulled facts out of some obscure reference.
I don’t want to put words in his mouth, but my understanding was that Terry Tao has not yet been impressed by a creative move in math. Terry Tao is a unique individual, though. I’ve been impressed. I consider myself—my bar is lower.
As we keep scaling this up, I can’t go into the details, but there’s a lot of effort at OpenAI. There are a lot of really smart, hardworking people who are pushing very hard to take this next step, and I think it’s going to come eventually. Just look at the trajectory that we’re on.
A year ago, I was a black hole physicist in academia, not really paying too much attention to AI. I thought, “Yeah, it’s cool for emails, but it’s not going to do what I do, which is special.” o3, which was really the first strong reasoning model, came out and was able to do a calculation for me that would have taken me days. It did it in 11 minutes.
I thought, “Wow.” That was shocking to me. We could go into the details if we have time. I could show the example because I saved it. It was really surprising to me.
Then I thought, “Okay, I’ve got to really start using this tool. There’s no other software that can do this kind of calculation, as far as I know. It’s really surprising and really cool.” Then GPT-5 came out 6 months later, and it was able to reproduce one of my hardest calculations, which I think the number of people in the world who could do that could be counted on your hands.
Brandon
When you say “reproduce,” do you mean this has been published or not published? Was it secret or internal?
Alex Lubyansky
Last summer, in June, I put out this paper, which I really like. It’s called “Why Is There No Love in Black Holes?”
Love is actually a technical term. It refers to Augustus Love, a British mathematician who studied the tides. When you have an object like the moon going around the Earth, it exerts tidal forces on the oceans. You can measure the tidal response of the Earth and its oceans to the moon via some coefficients that encode the strength of the tidal response, and these are called Love numbers in reference to Augustus Love.
Famously, black holes do not experience tides, so they have no love. There’s been a resurgence of interest in this fact in the last 5 years because people understood that it can be connected to a symmetry principle.
In physics, whenever something is zero—why should black holes never experience tides?—that’s surprising. Oftentimes, the answer is that there’s a symmetry principle at work that forbids the existence of tides and protects the structure of the black hole.
I found these new symmetries. These are differential operators that act on solutions to this equation, which describes perturbations of a black hole. These generators are symmetries because if you act on the solution to this equation, you get a new solution.
I thought this was very beautiful. I liked it very much, and it came out in June on the arXiv. In August, GPT-5 came out. The cutoff date for its training set precedes the release of this paper, so GPT did not see this paper during training.
When it came out, I thought, “Okay, I’m going to meet Mark Chen, who’s chief research officer at OpenAI.” He said, “Give GPT Pro a really hard problem. Let’s see how good it is.” I thought, “You want a hard problem? I’ll give you a hard problem.” I had just solved this problem and written a paper. I was very excited about it. I thought, “This is really deep and cool.”
I gave GPT the equation and said, “What are the symmetries?” I didn’t tell it that there were symmetries, because the default assumption should be that there aren’t any. It thought for 5 minutes and said, “Yeah, there are no symmetries,” which is what usually happens. That was wrong.
Mark Chen was visibly crestfallen. He said, “Oh. Well, okay, what if you give it an easier question?” Then I gave it the same question, but not for a black hole spacetime—for an empty, flat spacetime, which is a simpler problem. That’s actually how I approached this problem myself: you warm up on the easier question first.
I gave it the flat-space question, which is also in this paper. It’s this equation, which looks much simpler. This also has 3 symmetry generators, which are shown here. This is not new; these equations have been studied for 200 years. Everything in flat space has been known forever.
GPT-5 Pro thought for about 9 minutes, and it came up with the answer: a very beautiful, perfectly structured, perfectly correct answer. At the time, I also tried the other models from our competitors, and none of them could get this. GPT Pro was really ahead, and I think it continues to be the best model for this kind of mathematical physics work.
Mark Chen said, “Okay, this is great, but now that it’s done the warm-up problem, in the same chat instance, try the full problem again. Now that it’s been primed.” I thought, “Okay, why not?” I gave it the same question as before: “What are the symmetries of this equation?”—now the full black hole problem.
This time, it thought for 18 minutes, which I had never seen before, and it came up with the answer. Basically, in under 30 minutes, with 1 hint—which is the obvious warm-up problem to prime the model on first—it completely solved this problem. It was one of the nicest calculations that I’ve ever done, and that really blew my mind.
That was my Move 37, the Evan moment.
Brandon
Yeah, that’s how we call it in the AI world.
Alex Lubyansky
Once I saw that, I thought, “Okay, we’re on this crazy trajectory.” Eighteen months ago, it wasn’t useful. A year ago, it could do really hard calculations that would take me days. Eight months ago, it could reproduce some of my best work in under 30 minutes. In the last month, it solved these questions that we’ve discussed at length, which world experts had spent a year thinking about without being able to get to the answer.
I think it’s just going to keep getting better. Where are we going to be in 6 months or a year? I don’t see any reason why it would stop. I think we’re going to be having a very exciting year.
Brandon
Going back to these thoughts about scientific discovery and what these models can do versus just being superhuman at solving physics, people keep asking this question: hypothetically, could we train a version of ChatGPT where it’s never seen anything after 1904, and could it rediscover relativity?
I think there’s a very analogous question we could ask right here, which is a new conceptual result about single-minus gluon amplitudes that was sparked by human insight. There were some very specific assumptions that went into this, like understanding that working in Kerr spacetime is something that people have been thinking about and that has some useful, transferable insight. People have been thinking about maximally helicity-violating amplitudes for quite some time.
Have you ever tried using a model whose cutoff date was right before this paper and asked, given a Kerr metric, whether there’s anything interesting with regard to helicity violation? Or maybe turning it around and saying, “It’s long been thought—or long been known—that, with the exception of some set of measure zero due to Witten, there are no single-minus nonzero amplitudes.”
Have you tried either of these directions and asked it to discover a new insight, push the boundary as you were just talking about, and make a leap—in addition to not just solving a problem that you can give it, but actually getting that intuition?
Nima Arkani-Hamed
Yes.
Shawn Wang
You have tried this?
Alex Lubyansky
Not exactly the counterfactual version that you’re describing. I personally haven’t done that, but pushing the models at the frontier to try to make this type of leap is something that we’re very focused on.
I don’t want to talk about the internal research we’re doing, but I can say something publicly, I think. You can take this page from the paper and feed it to ChatGPT Pro—the best model we have right now—and ask it, “What should I do next? Give me the top 3 follow-up questions to ask based on this paper.”
I’ve done this experiment, and the top 3 questions it comes up with are my top 3 questions for what I should do next. I think the models are smart enough now and have enough background knowledge that, for this paper, GPT is about as good as me at finding the next thing to ask. That’s really interesting, and it opens up a lot of possibilities.
Brandon
Can you just do the agent loop where you say, “Okay, what’s the next question? Go ahead and solve that. What’s the next question?”
I guess this goes back to the question I was asking before: if you do that—and you probably have tried it, or something that OpenAI has tried—do you eventually get to some plateau where you’re not pushing the boundary of knowledge anymore? Or is the plateau just money, and if you had more money, you could go further?
Alex Lubyansky
Just to be very explicit, because I haven’t said this quite out loud: I think we now have models that can really turn out papers that are as good as human-written papers.
In fact, this is a bit of a problem because when a professional physicist uses this tool, steers the model, and checks the answer, they can get amazing results. But there are also people who feed it wrong questions that go off the deep end, and then submit that to arXiv. This is a problem that the academic community is trying to come to grips with now: AI slop in science. This is something we have to figure out.
With proper steering, you could probably turn out a paper a day now. If you give a question to ChatGPT, it’ll solve it if it’s not that hard of a question or if it’s a similar calculation to stuff that’s already been done. It can totally do it in 30 minutes, and then you could say, “Write it up as a paper,” and send it to arXiv.
I think we’re already in this moment. We’ve passed that threshold. This is the new reality, and more and more people are catching on to this all the time. Some of them are doing this, and this is why arXiv is now inundated with submissions.
So what’s the correct response to this? We put out these 2 papers in very fast succession. We could spend the rest of the year writing 30 more papers like this, but I don’t think that’s what we should be doing. Instead, now that we have this new tool that gives us AI superpowers, I think we should just raise the bar for what it means to write a good paper. We should aim higher, basically.
One thing that I’m excited about is that I think these single-minus amplitudes papers open the way to a whole direction of research, which I think is a line of attack on really interesting questions in quantum gravity. To go back to the start of the session, this is the missing piece of the puzzle of fundamental theoretical physics. I think we have a pretty clear line of attack through a series of questions, all of which I think will be amenable to solution with AI.
I’m excited to spend a good part of this year trying to follow this path and solve harder and harder problems. This paper gave an answer to a question that had stumped Andy, Alfredo, and David, who are experts in this, for a year. But we haven’t seen an AI solve a question that has stumped an entire community of physicists for decades. That hasn’t happened yet.
Given the trajectory that we’re on, at some point—hopefully not too far in the future—we should see that. I think that’s the exciting thing to try to move toward: pushing the envelope of what can be done.
Brandon
We wanted to ask a question: if you could remove 1 bottleneck for your domain—in this case, maybe it’s AI for physics, maybe it’s physics, or maybe it’s mostly AI—what would that be, and why?
Alex Lubyansky
Well, off the top of my head, I spend so much of my time writing papers. The way I think now is so far from papers that it just feels like they’re not the right way, somehow, to store and communicate knowledge.
I think an extreme version of this, which makes the problem more apparent, is math—especially certain parts of math where papers are very terse and take 4 pages. I had this experience when I was learning algebraic geometry in graduate school. I went to a mathematician and said, “What’s going on in this 4-page paper?” It was just very terse notation.
He said, “Forget what’s in the paper,” and took me to the blackboard and started to draw pictures. He said, “This is how you should think about it.” Then I was like, “Oh wow, this is amazing.” But none of that is in the paper.
Mathematicians, I think, have this cultural norm that they hide the messy work and write these beautiful, short, pristine papers. It depends on the subfield, but oftentimes that’s the case. The way they actually think about the subject as a living, breathing entity is very different from the way in which it’s recorded in papers.
Some of that is also true for physics. I love doing calculations, coming up with questions, and finding the answer. I would say the huge bottleneck is writing it up. Somehow, it feels like papers are not quite the way of the future—or at least the way we currently operate: I write it up, send it to a journal, and it takes 6 months. I don’t know. Why are we doing all of this? It feels like maybe there should be something better.
If you want to understand this paper, one thing you can do is upload it into ChatGPT and ask it to explain it to you. You can keep unfolding the complexity into more and more detailed explanations. If we move to a world where we use AI to do the calculation and get the result, then we have the step of condensing it into a paper. Then I send the paper to Brandon, and he puts it back into an AI. I mean, why are we doing this?
Brandon
Yes, right. That’s a little bit funny.
Alex Lubyansky
If you ask me whether I’d be confident that in 20 years we’ll have these sorts of static documents in which we publish our results as papers, I would think not. That doesn’t seem like the best thing we could be doing.
Maybe some kind of interactive paper that lives in an LLM. Maybe your whole paper is a ChatGPT page, and there’s a chatbot attached to the paper. You can say, “Explain the big picture,” or “Zoom into this fact.” I think we’re going to head in that direction. That would be a cool thing to see.
Writing a paper, though, is a useful exercise because it forces you to condense your thoughts and make them really clear. I’m not saying it’s a bad thing to do in general, but the way we do it is very slow.
Maybe another answer is that, in this project—the graviton paper—we got to a draft of a paper extremely fast, and then we spent most of our time checking the answer. I think that will effectively be a big bottleneck, maybe the next big bottleneck.
That’s one of the things the models are missing. If you ask me what we can really improve in the models for scientific research, I think we’ve touched on the 2 big things already, but just to spell them out: one is creativity, the spark of invention, and really taking the next step. I think that will come as we scale up the intelligence. We’ll see, but I don’t know that there’s something missing inherently. I think it’s just starting to make these leaps for me.
Maybe we should encourage the models to try to make bigger leaps, because large language models, after all, are trained to give you the middle-of-the-road answer. If you ask an AI like ChatGPT, “Write me an email about blah,” you want it to give you the expected answer, not sample from the tails—a wacky email. You want it to give you a reasonable thing.
For most tasks, you want that. But for scientific research, sometimes you want the idea that comes out of left field, the thinking outside the box, or really sampling far out of the distribution. That’s something we could do in principle, but that’s not how the models are set up. We’re not really favoring that, so we might have to make tweaks of this kind to enable the models to take bigger leaps.
The second thing is verification. We’re now in this new regime where the models are so capable that, for very hard computations at the frontier of knowledge, they can just do the whole thing. But is it correct? In this case, it was correct.
Sometimes I get emails from people saying, “I did this really long calculation, but there was a mistake somewhere.” Disappointing. The calculations are getting more and more complicated, longer and longer, but sometimes they mess up.
I think improving verification—or even just having the model indicate more directly how confident it is in its answer—is important. I think they’re smart enough to know whether they’re very confident in the answer versus when they’re just guessing in some step. Getting the AI to be more explicit about that is, I think, a way to improve it for research. That verification step is going to become maybe a bigger bottleneck this year.
Brandon
Yeah, Keren Hong from Axiom would agree with you emphatically. Formal verification is their thing, right?
Alex Lubyansky
Yeah. It’s interesting: a year ago, I would have said it was super important to have formal verification. Then the models got so smart that I thought, well, if Brandon and I talk about a mathematical proof and go over it, we’re not going to formalize it in set-theoretic notation, or reason the way Lean—which is this language for formal verification—does. We reason through the proof in natural language. We use words.
If a model is really smart enough, then it should be able to do the same thing. We’ve been seeing this huge increase in capability for mathematical reasoning and developing proofs using natural language. For a while, it looked like that wasn’t the thing to really focus on.
But now that we’re in this regime where you can just get ChatGPT to tackle thousands of questions at the same time, and it will return proofs for a significant fraction of them, the onus is back on humans to verify all the outputs. If that becomes a bottleneck, I think formalizing math and automating verification will become more valuable. That’s something we’re thinking a lot about as well.
Brandon
Thanks. What do you want the audience to take away from today? Is there 1 message that you want them to leave with?
Alex Lubyansky
Yeah, I think it’s important to get the word out that the models we’re developing at OpenAI are becoming really capable in scientific research.
I myself was a bit of an AI skeptic a year plus ago because I thought the models were very good at writing tasks but not mathematical tasks. That changed with o3, the first strong reasoning models. And then GPT-5 was able to do some of the hardest calculations that I can do and reproduce them correctly. Recently, in the past month, we’ve seen models solve open questions in theoretical physics. And now they’re solving problems in quantum gravity and quantum field theory.
So if you just extrapolate that into the future, imagine where we’re going to be in 6 months or a year. I think it’s kind of surreal to live through this time, but it’s really happening. It’s really amazing. And I think we’re going to see a lot of big changes happening in research.
So, yeah, pay attention to this space. Let’s stay tuned.
Brandon
That’s awesome. Thank you so much for taking the time. I learned a lot from our discussion, and I’m definitely going to keep up with what you’re up to.
Alex Lubyansky
Thank you. It’s been great to be here.
Brandon
Thank you. Thank you.