# Recursive Language Models — Alex Zhang, MIT PhD

Latent Space · 2026-10-02 · 101 min · https://www.youtube.com/watch?v=kog7mwsDqnk

## Transcript

Alex Zhang

I just love it. Claude Code, I love Codex. I love Pi. No, I mean, I love Oh My Pi. I love Prime.

swyx

Fair to say, I think they are all the same. Your models are actually capable of more if you make an effort, so it’s a skills issue.

### RLMs in the Wild

Alex Zhang

Yes. Yes. In the main, I think there is one thing I want to see. I appreciate the focus on intelligence, because it gives us the big picture of what we can do if we really set ourselves a goal.

But I would like—and maybe someone in academic circles should have done this—to really sit down and think: if I took a task, even a modern advanced model is not good enough to perform a certain task consistently and qualitatively over the course of a month, that’s 10,000 agents for 88 hours. That’s 130 billion public tokens, estimated at approximately $40 million at public prices. Maybe you would have to pay about $40 million to get a result.

It’s exciting. I will say this: it is amazing what we have. Is there a possibility at all to direct $40 million toward problem-solving?

swyx

### RLMs Explained

So we’re here in the studio with Alex Zhang, probably best known for RLM, but you have several other connections. Welcome to the show.

Alex Zhang

Thank you for inviting me. Probably GPU Mode as well.

swyx

Yes, GPU Mode too. Well, I brought you here, Alex Zhang. Not everyone gets this reception.

Alex Zhang

Yes, I’m very close to everyone in GPU Mode. We often work together in different roles, even outside the boundaries of GPU Mode itself.

swyx

Can we explain this for those who don’t know?

Alex Zhang

This is a Discord server. It used to be focused on CUDA Mode, and then it expanded a little. It was created by Mark Saroufim. To me, it’s like a talent pool for the PyTorch community.

swyx

And then you left PyTorch?

Alex Zhang

### GPU Mode, KernelBench, and AI-Written Kernels

Yes. Yes. Yes.

I think it appeared when I was in college, sometime in 2023. I think Mark started it with Andreas and Jeremy Howard. The primary idea was that it was a Discord server for training and writing GPU kernels, where people gave lectures. That was basically it.

I was interested because I wrote GPU kernels, and it happened purely accidentally. I was an intern at Snap and Recurse, and I was very bored at Recurse. They had a project they wanted me to write something for. It was an article entitled “Infinite Attention.” It seems this was an article from Google.

swyx

Yes, we discussed it in our Paper Club.

Alex Zhang

I was wondering whether it was possible to write specialized kernels for this at Snap. Nothing came of it, but I joined GPU Mode around that time. It was called CUDA Mode then. I think, for legal reasons, they changed the name, but I met Mark, Matt, and many other people who were very involved in the community.

Then Mark proposed the idea called Popcorn, which is now known as the leaderboard. Generally, the idea was that we all had an intuitive understanding that GPU programming is very similar to competitive programming, if you’ve done it. It’s not that I’m saying those are completely transferable skills. You do a little code golf.

There is a surprisingly limited set of optimizations that people use, and there aren’t really that many kernels people want to optimize. We thought: what if we had enough data, similar to Codeforces, where there are millions of tasks? If it were possible to do that with GPU code, you could scale and automate the development of GPU kernels, and that would be a huge step for researchers.

One of the biggest problems—for example, take Mamba—is that they publish the paper together with the kernels, because otherwise you won’t be able to use them effectively. Not everyone on the team has the expertise. So we’re very interested in this. KernelBench emerged from the same ideas: can we force LLMs to automate GPU kernel code?

I had a lot of fun with this work. It was between college and graduate school, and I really enjoyed doing things in GPU Mode. Now I sometimes help with lectures. I’m not as involved anymore, and I think we’re not spending as much time on competitions as before, but I still keep in touch with all the participants.

swyx

Do you have a friendly competition? To me, the previous community that was doing this was MLPerf. Is it the MLPerf generation, or what’s happening?

Alex Zhang

The beauty of GPU Mode is that it’s a community. Many lectures are available, beginners can ask questions, and the competitions aren’t necessarily about being the best; you can participate for the sake of learning.

I think many benchmarks, such as MLPerf, mostly take only serious participants from laboratories and companies, at least at a serious level.

swyx

At least that’s how I understood it. Outside of GPU Mode, it’s very exciting that many more sites and people have appeared who are organizing these competitions. For example, there’s a site called LeetGPU, or something like that. It’s like LeetCode for GPU tasks.

There are also others. We’ve seen them appearing, and people have been talking about them in GPU Mode. This is very exciting because earlier, GPU programming was something extremely specialized.

The reason I’m interested in this is that Tri Dao gave a talk at Princeton because he applied for a teaching position—he’s now teaching there. I listened to his FlashAttention talk in 2023 and thought, “Wow, this is the coolest thing in the world.”

Alex Zhang

Me too. I thought that was what everyone was working on. vLLM and other things came out, and we were like, “Oh, we have to write kernels.” But now everyone writes kernels. I think that, in a certain sense, this industry is almost oversaturated.

swyx

Are there any interesting tips for people who want to do this?

Alex Zhang

I think one of the loudest pieces of news was that GPT-5.6 wrote more efficient kernels, so Terra and Luna can become 80% cheaper. We also saw other competitions where people set records, claiming that they used a certain kind of autoresearch, even though these were people without experience writing kernels.

You’re right. Even on the GPU Mode leaderboards, if you look at many recent tasks, almost all the solutions are AI-generated. However, you’ll notice a guy named Gorst, who is a very consistent GPU Mode participant. We’ve known for a long time that he is an extremely capable GPU kernel developer.

We found out one thing: on this leaderboard, he also used AI to help with his solutions, but mostly he directed the work with clues in a certain direction. We found that his kernel was practically the only one in the top 10 that actually worked reliably in real, complex systems.

swyx

This raises the question, because GPU kernels have a problem with verification. We knew that was a problem with KernelBench. There are a lot of system hacks that get rewarded.

You also notice that the lines of code are much, much shorter. Is that noticeable?

Alex Zhang

Yes. No, it is definitely very important.

swyx

I think it’s very interesting that there is still a big advantage to being a professional.

Alex Zhang

Of course there is.

swyx

I think this concerns many AI systems too. Even with the latest mathematical proofs, that doesn’t necessarily mean mathematicians have become outdated. These companies are still hiring mathematicians for data labeling, or even just for directing models in solving problems. Deep knowledge in these areas still provides a big advantage.

Is it just knowledge, or is it also about better planning? Is there a new planning approach that works more effectively?

Alex Zhang

I think it’s a combination. Perhaps this is what you mean by intuition about how to solve the problem. For example, I always draw diagrams of my code.

swyx

Yes, that’s right.

Alex Zhang

If there’s a part of the diagram that I don’t understand, I work on it until I figure it out. Otherwise, it’s unacceptable.

### Human Expertise vs. Brute-Force AI Search

I think it’s a combination of these things. People who know how to look at these problems and how to solve them also know how to use AI for this, because you become a very strong verifier if you have the knowledge or know what to do.

swyx

I also think we found out from all these agent swarms and similar things that when you direct enough computational capacity at a problem, you can explore enough ways to solve it. But it often happens that you can spend 100 billion or 1 trillion tokens on something, and if you involve someone who understands the problem, they can suggest something to the model that will prevent that unnecessary expenditure of 1 trillion tokens.

I have you in mind here, although I’m not quite sure what the trends are. I think there are now so many problems that we want to solve, and we can’t always direct as many computational resources at them as possible. So the efficiency aspect of all these things remains extremely important.

Is there a theoretically correct answer that can be calculated purely from physics, and then approached as a physical limit?

### Open-Endedness and Sakana AI

Alex Zhang

Yes. For GPU cores, this is not always easy to calculate, depending on the complexity of the task. For example, for matrix multiplication, it’s very easy to calculate the speed-of-light assessment—that is, the theoretically possible kernel performance. This also depends on whether all your data is initially located on the central processor or in the DRAM on the video card. Data transfer and everything else change these indicators a little.

It’s not quite clear whether, in many cases, it’s possible to achieve this theoretical indicator at all. You need perfect overlap with data transmission, and maybe there’s some bottleneck that’s impossible to avoid. Often, the kernels we write don’t even come close to this indicator, so it’s not necessarily a significant concern.

swyx

Does it matter if it’s just speed? Or do you also worry about memory, which affects speed? Do you care about energy consumption? One of our best podcasts was with Jeff Dean a year ago, and he said, “Actually, I just track microjoules, or maybe nanojoules—picojoules.”

Yes, pico. This is optional, pico. Are you concerned about that?

Alex Zhang

I probably don’t.

swyx

But it all boils down to speed, right? Nobody counts joules.

Alex Zhang

There are caveats. I think there’s speed in the context of a separate kernel and speed in the context of a broader task—for example, a model from A to Z.

That’s precisely why it’s important to discuss what “speed of light” means. Let’s start with the assumption that everything is in HBM, right? But you can imagine a whole model from A to Z. Maybe between 2 layers, you want to sacrifice the speed of the 1st operation so that you can leave the data in the cache for the 2nd operation.

These are things that are impossible to fully understand by considering kernels in isolation. People call this the problem of operator fusion, or “megakernels.” This usually applies only when you’re limited by memory bandwidth. But as these models improve, the question becomes: do we just need to create these megakernels? Is that what we want?

What’s your opinion?

Vibhu

I think that’s very difficult because it requires data. I haven’t yet seen a real-world example of using this approach to solve a really difficult class of tasks without any examples.

Alex Zhang

Me too. I think there’s another reason why this might not be so interesting. At the level of a separate kernel, the kernels aren’t that complicated, and you’re reasonably sure that there aren’t that many structures within a single kernel.

With a megakernel, I would rather believe that a compiler would be better suited for this. For example, a compiler for high-level operations makes sense, because, in general, I think megakernels consist of individual kernels. There are certain areas where you might want to make unusual fusions, but in general, I think these are cases that a compiler will probably handle. There’s a company that does this, as far as I understand, and it has also given several talks about GPU Mode.

swyx

### Research Taste and Taking Big Bets

Yes. I actually want to group all the GPU Mode discussion here, because obviously there are other parts we need to move on to. I think it’s worth recommending. You put out a lot of really good lectures. They’re all on YouTube, so people can follow them, and you lead a significant part of them.

Alex Zhang

I used to lead them; sometimes I still do. I think it’s mostly Mark. Mark is the one who usually conducts them. Mate does sometimes, too.

swyx

I highly recommend them. It’s an extremely good resource. I think it’s crazy how many people there are sharing information.

Vibhu

Yes, because if you’re there, you’re exactly the right audience, you know?

swyx

Yes, exactly. It’s not going to reach the mainstream, but there are also many introductory materials that we’ve put out, which I think are useful for people. KernelBench was quite influential.

What other current work do you want to talk about? Do you remember which people are worth paying attention to, since you’re obviously involved in this field?

Alex Zhang

I can. I’ll give you the background on why I’m so involved in benchmarks, or why I was, before graduate school. It started because I was at Princeton, where I worked with the SWE-bench team. John and Carlos are there. They’re all wonderful, and I love them.

I don’t think people understand how many benchmarks come from the same group at Princeton.

swyx

Do you know Shunyu?

Alex Zhang

Yes.

swyx

He was already on our podcast. Now he runs a company, too. He’s a real star.

Alex Zhang

I met him when he was a research scientist and a friend of my boss, Michael Tang, who is currently working at Anthropic. They collaborated a lot. We were 2 students in Karthik’s lab.

swyx

Shunyu is awesome. I didn’t realize he was such a star until he left. But there are very few graduate students like you. Once a year, we see someone whose dissertation was right from start to finish. There aren’t that many of them. Shunyu was one of them, and Jack Morris was, too.

We talked before about writing about research taste. Some graduate students get incredibly lucky with their careers: everything they do turns out to be relevant and important. For others, nothing does. I think this applies even to people in industrial laboratories; graduate students are just much more visible. So are people genuinely lucky, or is it a mixture of luck, high intelligence, and other things?

Alex Zhang

I think research taste develops through opportunities, at least in my case. I was really lucky to work at Princeton and then to meet my scientific adviser, Omar Khattab, at MIT. He’s a fantastic adviser.

I’ll say this anyway: the most successful graduate students, or the most successful people in an academic environment, appear when they focus on problems that most of industry doesn’t pay attention to. I think the problem is that many graduate students work on things that look attractive to industrial laboratories. For example, they work on benchmarks.

Benchmarks in 2023 were a completely different story from what they are now. Many people are working on testing systems, meta-systems, and specialized tools for performing certain tasks. If you think about it, the reason why someone works on this is that there are clear goals around the models we have today. They say, “This is what I want to see.”

I’ll give you an example: the paper on RLM, or Recursive Language Models, because I think its idea is extremely simple. When it came out, there were a lot of people who, upon seeing something like that, asked, “What’s the whole purpose of this? It’s just some subagents or something like that, right?” I think that reaction is actually a good sign, because it shows that people still aren’t thinking about the real purpose of the idea.

I’ll give you another example: SWE-bench. OpenAI likes to tell this story. When SWE-bench appeared, nobody was interested. Everyone said it was an impossible task. Why should we consider it a benchmark? Only after Devin came out did everyone say, “Wow, this is what we want to succeed.”

You often see many ideas like that. My favorite example is Eric Zelikman’s work on STaR and Quiet-STaR. When you read that paper—at least when I first saw it—I thought, “Isn’t this an obvious idea?” Or maybe not. I thought, “It seems very simple.” Chain-of-thought is similar, and ReAct is the same: “Of course.” But when you really think about it, the question is: what’s the value of these papers?

I think the value is that they tell a particular story about what you think this area should look like. That’s very difficult to do in academic circles. If you look at all these papers—Quiet-STaR, ReAct, RLM, and SWE-bench—none of them released a GPT-6-level model. They aren’t cases where everyone says, “Oh my God, I’m going to use this now. It’s the best thing in the world.”

The academic environment simply doesn’t allow that, at least not yet. There are a whole bunch of reasons why that should change. But as far as I’m concerned, if you’re a graduate student, you’re in a unique position where you can work on mostly whatever you want.

If you don’t use that opportunity to work on things that people don’t care about, or that people think are trivial—“I was thinking about this, but I don’t know”—then the research will never be that interesting. You have to make big bets if you stay in academia. Otherwise, why not go to an industrial laboratory, where there are many more resources and talented people? Why limit yourself in academia, where you have fewer resources and not as many people around?

It all comes down to big bets. You just have to make big bets, and many of them will fail. But that’s your biggest advantage as a graduate student over anyone else in another laboratory: you don’t need to deal with bureaucracy and all the other constraints.

swyx

Right. Yeah. I’ve asked this question to many people, and usually they just brush it aside. So I appreciate such a thoughtful answer. The answer they give is no: this is your disadvantage, because everything else is, in fact, opposed to you.

### GEV and Rethinking the Language Model

Alex Zhang

Yes. Yes, that’s right. And so, honestly, I’ll give JEPA as an example, because it’s not an academic project. I want to remember this, because this also happened with RLM and many other models: when things become overvalued to a certain extent, something gets excessive hype, and then people say, “Why is this so overrated?” as if it were trivial or pointless.

I saw the same thing with JEPA. It seems that the release was like this. There’s this whole story about academics—or, if they’re not an academic group, people still have to engage in branding and somehow promote their research—so I understand that. But I think there was a lot of talk about what JEPA is, even though this is something we’ve known about for years, and I believe that misses the point of why such a system is interesting.

Why isn’t it just some stupid MLP, the classifier that we did in introductory machine-learning courses or something like that? What’s interesting about JEPA is that it violates the question of whether these are even language models. In the sense that, in the form in which they exist, can we consider another design space besides text-to-text?

They’re essentially saying, “I’m going to use this as the basis of language models. I know that it contains a lot of information about language, but I’m going to change the output model space to offer you a compromise. I’m going to perform inference very quickly if you have some specific a priori knowledge about the task.” Let’s say I need to do binary classification. Would I ask my language model to do this and pay 400 times more? No, that’s ridiculous, right?

For a long time, because laboratories are the only places that have control, you could never use anything other than GPT, GPT-6, or Fable, because they’re the best models. But because of this, people have gotten used to the idea that a language model is simply an autoregressive decoder. We accepted it.

I think when JEPA appeared, it was the same. One of the most frequent criticisms I received was, “This is not a language model,” or, “When I looked at this, I thought it was a new architecture, but actually, no.” My answer to this is: a language model is just modeling language. It doesn’t have to be a Transformer decoder, you know?

JEPA is very interesting in that it gives us a new parameter to configure: what is the latent space, and how does that affect inference latency? I think we can actually begin to use different parts of a language model. It’s a similar idea that received a lot of attention, with people saying, “This is a silly idea. Why would anyone need this?” But it’s a simple idea that opens up a whole series of new questions, which I think, especially if you’re a graduate student, you want to answer.

For example, there’s a lot we don’t know about JEPA. We don’t know how far we can go with recurrent Transformers. We also don’t know how far we can go in the other direction: what if we recurrently use only part of the model? What if we route data through a router to different parts of the model? Is it possible to reproduce what you do in a harness inside the model? What is possible if you combine different choices of bindings with model-selection architecture?

These are all questions that, in my opinion, are opened up by such models. That’s actually very fascinating. JEPA, in particular, when I saw it, I thought: this is really useful for RLM. I think it’s logical, because the biggest bottleneck in RLM, Swarms, or similar systems is latency.

When you’re constantly making numerous language-model calls, you’re distributing computation incorrectly. Maybe there’s something trivial that you want to entrust to a simple model, but you can’t, because your language model is just this bulky thing, you know? So I’m very excited. I think we’ll begin to see the emergence of new types of models beyond ordinary, standard autoregressive models. With these new trade-offs, you can do so many things.

I also want to note Thinking Machines and their model innovations.

swyx

Yes. Yes. Another great example.

Alex Zhang

Yes. So, essentially, they’re attempting to go beyond the sequence-to-sequence, decoder-only paradigm and do something else.

swyx

Yes. I mean, if you try to compete, you can’t compete with a frontier lab that creates an autoregressive decoder when it has that amount of computing resources. Even Thinking Machines probably can’t. Maybe they’re one of the few who could, but you really don’t care what you can do in that regard.

I don’t know too much about Thinking Machines, and I don’t want to say anything. But if their strategy was simply copying OpenAI or Anthropic, that would be a terrible strategy. They just don’t have the resources. It’s worth thinking about this from the perspective of someone who has advantages. If you’re going to use the same approach as everyone else, I’m sure you won’t win. If you still choose the same path, you’re essentially competing on the things you can control—that is, data and compute—and they obviously won’t be able to compete with the frontier labs.

Alex Zhang

Yes, it makes sense that if you’re a new laboratory—although I don’t know if I consider them a new laboratory—I think you have to do something different. Apparently, they’re somewhat unusual. They’re on this list. They released Tinker, so they have to be taken into account.

Anyone other than OpenAI or Anthropic—possibly Meta and Google DeepMind—you have to do something different. It’s a sad reality, but I think that’s good. I’m very glad that this is happening and that these companies will continue doing it, because it paves the way for new players if they find something really interesting.

I have doubts that this is seriously happening at frontier labs, because why would they do it? Why take the risk and spend a large amount of compute on new approaches when the old bet is already working? I think that generates a lot of new laboratories.

If you’re making a secondary bet, you don’t receive much compute and say, “Okay, I’ll go.” Your example of potential success is something like JEPA, which is 100 times cheaper. Maybe it’s a language model; maybe in this case it’s simply something else.

There are no firm assumptions about what JEPA really is. I have guesses about what it could be. I saw some people say, “Oh, that’s something like diffusion.” I think it’s parallel decoding.

swyx

Yes, parallel decoding.

Alex Zhang

Fair enough. I think that regardless of what it really is—and I saw an open replication of it—what fascinates me is what they’ve done. I’m not absolutely sure what their optimization objective was or how they trained it.

I think this is the same thing with RLM that we were thinking about: RLM is a very simple idea. If I release this paper, everyone will be able to use it. But the true value of RLM is whether you can train the system correctly and whether you can build an architecture around it to make it genuinely high quality.

That’s what I’m actively working on, but I think they found a way to train a system, and that’s a completely nontrivial task. I don’t even know how they did it. I saw comparisons online where some people claimed that they used Qwen or retrained it. But any Qwen model from open source would be worse, because whatever they did for training clearly works very well. It’s really exciting.

There’s one element of calibration, which is a fairly rare topic—people don’t even know about it or understand it. We discussed this with Clémentine Fourrier from Hugging Face, who used to be responsible for evaluation there. The point is that models are trained to give the most probable next token, but they’ll lie to you when you ask, “How sure are you?”

They’ll just give you the most probable answer instead of actually calibrating: “No, I’m 50% sure. I’m 20% sure.” Let’s try to calibrate that. I would say it’s fairly easy to generate synthetic data, because you can see the truth. You synthetically generate a heap of answers, make the model classify them, and then compare them with real data. That would be my reverse engineering.

I think calibration is underrated. People simply use models as very fast classifiers, but they don’t even use probabilistic or calibration scores. I think we still don’t understand this correctly. When you ask a model how sure it is, it just gives you, say, 43%. The big difference is that this is a calibrated classification, right?

swyx

Yes. Yes.

Alex Zhang

I’m looking forward to what people will do with this model. Will it solve every problem? Of course not. But I think it solves a class of problems with which we’ve traditionally had difficulties: low-latency tasks.

I really like the examples with games. I have a benchmark for language models that play video games. I’ve always been fascinated by the question of whether it’s possible to create an intelligent system that can play new games and things like that. So I think it’s really cool that they found a unique way to do this by combining language understanding with a fast—very, very fast—model.

swyx

Well, while we’re on this topic, is there anything you want to draw attention to?

### Video Game Agents and the Harness Problem

Alex Zhang

Oh, yes. I mean, you did it anyway: the benchmark for language models playing games.

swyx

I think these numbers are already very outdated. Many models are very, very different now. I’ve seen enthusiasts launching newer models in these games, and that’s really cool.

Alex Zhang

I think the general idea of this benchmark was to check whether vision-language models were suitable for integration into games, taking delays and other limitations into account. This happened right after Claude Plays Pokémon appeared—about 2 years ago, which apparently is now considered ancient history. But I think what’s really cool about this set of tasks is that it’s diverse in terms of game types. Most of these games are ones people know or have seen somewhere.

I saw the Japanese agent that plays Doom. To be honest, I think it probably played only very simple levels, but most models still don’t cope with these games. I don’t think any model can meaningfully complete them, although there are games they can handle. I think I saw that Astra was able to complete the Kirby game.

For this benchmark, we intentionally designed a very minimalist harness. I’ll come back to the topic of harnesses, because I think it’s worth discussing their value. Why does a model need a harness at all? In general, I very much hope to see new models beat all of these records soon.

It’s interesting that the old Claude Plays Pokémon read the game-state memory and saw which tiles you could walk on and which you couldn’t. We recorded a podcast with them a long time ago.

swyx

Got it. Got it. Yes, yes, it’s similar. Before, Claude didn’t have vision, so we had to submit the game state to it, and everything was so different.

### Why Claude Code, Codex, and Pi Are So Similar

Let’s go straight to the harness theme, since you already alluded to it. Harnesses for language models and compositional generalization—it’s hard to read. Explain it.

Alex Zhang

Okay. I’m a little dissatisfied with how people perceive harnesses. They say, “I love Claude Code, I love Codex, I love Pi—no, I love Prime Agent.” Honestly, I think they’re all the same. The majority of their design solutions—or, say, their design choices regarding these frameworks—are identical. Maybe Prime Agent is a little different because it is essentially an RLM, but overall, I think we can be much more creative with harnesses.

What I mean is that if you look at what a harness does for a model, then when you’re trying to solve a problem and want to use a language model for good performance on a difficult task, we find that next-token prediction is a very inconvenient form for many such tasks. Take SWE-bench, for example. When you’re navigating a codebase, can you understand how to do all of this with a single call to a language model? Something simple like, “Fix the code,” or, “Answer my query in this codebase”? No. Therefore, we rely on a harness that helps us do these things.

A harness is a very specific program for how you want the language model to adapt to a problem. I remember this because when we’re thinking about what makes a harness, we should seriously consider what choices the harness allows the language model to make, and whether we can really just have a model do it by itself.

If you think about when recurrent Transformers appeared, I think the most interesting thing is that you could model a recurrent Transformer in the same way and use it as a harness, right? You just loop the model. Now you can say, “Oh, I don’t need to decode,” so this is a little different, but in general, for some reason, we’re forever stuck with the same choice of model architecture, and there are many arguments for that.

But obviously, we’re now training models around the deployment of harnesses, and that’s why there’s this very inconvenient way of learning based on harnesses: we teach a language model to act within a harness. But now it’s really long—perhaps a multi-agent deployment—and there are very inconvenient, hacker-like ways to do it.

This blog is about one way to understand what’s happening here: consider a harness as an aid to a model in solving a specific task. Can different design options do something a little heavier than just giving you a list of tools to help you view your codebase?

The idea for this blog arose together with the RLM concept. We just didn’t formalize it then, that’s all. I think that’s really true. There are many other ideas around RLM that we’ll reveal later, but they were there from the very beginning.

I like that in RLM there were many iterations and versions of various abstractions that I wanted to implement, and eventually RLM turned out to be the most appropriate. There are many reasons why that’s true that aren’t public yet. You’ll see for yourself.

In this example, one of the things we see in RLM is that if you sufficiently offload the context and ask the model to write code based on that context, you get a very strange but useful property. When the model learns how to solve the problem, it turns out that the solution for many tasks is very similar, even for tasks where it isn’t obvious that the solutions should be similar.

In this example, we have a search task and an aggregation task, and they have very different requests. The task types are completely different. When you train an ordinary language model on these 2 tasks, the learning trajectories look very different. When you rely on something like Pi or Claude Code, these trajectories are too tightly tied to the problems, even though the solutions are the same.

One thing we discovered when teaching RLM is that the model sees these problems as the same. The reason it sees them as the same is that the subagent sees different problems, but the subagent solves an easier subtask. You assume that it’s smart enough to deal with that. At the level of the overall strategy, the problems eventually look the same.

When you’re training on the task on the left, the model can immediately solve the task on the right. If you look at our schedules, you’ll notice one thing: when you naively train your RLM on these tasks, it naturally learns to generalize to longer tasks, because the basic strategy is the same. You just change a certain length parameter.

This also applies to tasks that are different, not just to differences in length. These are completely different tasks—for example, mathematical tasks and writing tasks—but the solution, at the high level, is the same. When you train RLM on one of them, it propagates this behavior to the other. There’s no magic here.

swyx

### Harnesses as Compositional Generalizers

How would you describe in words what they’re learning? I think what you’re saying is that if you train on short tasks, they generalize to tasks 8 to 30 times longer. They learn to solve these types of problems—or, more importantly, what are they really learning?

Alex Zhang

Yes, they’re learning to solve these types of problems at a certain length, and it turns out that when you take this learned strategy, it transfers directly to greater lengths. It’s as if this is actually the same program. That’s very exciting.

What this means is that when you have a data corpus or environments on which you train a model, the hope is, first, that you can train on a smaller number of environments and generalize to a greater number than existing models can through naive training. But secondly, you want to use all the data you have. So when you train on these tasks, we hope it extends to a broader class of problems.

Why is this even more fascinating, at least in the context of RLM or any recursively composed system? This argument acts inductively. For example, I claim that competitive programming and GPU optimization require very similar skills. The model can potentially—you may have to give it a little push—understand, “Okay, the way I solve this GPU programming task is very similar to what I learned for competitive programming.”

Therefore, I’ll list a set of solutions, create subagents to search for promising options, and then write a loop to walk through and check these solutions, perhaps develop them, and check them with a verifier. Although these 2 tasks look the same, what the subagents do may be unique and specific to each one. But isn’t it possible that what the subagents solve actually has the same shape? It’s a recursive argument.

What I’m trying to say in this blog post is that it’s worth rethinking the role of the harness, because harnesses can significantly increase models’ ability to generalize and process data efficiently. This applies not only to RLM. I believe there’s a wide class of harnesses that haven’t yet been discovered and that can provide similar properties.

Developing this argument further, I think you can draw this conclusion: if I’m looking at RLM, which RLM components are actually necessary? Can I really just teach the model to do it directly? Can I teach the model to act as an RLM implicitly during the forward pass?

This is a strange thought, because you can say, “Ah, code is non-differentiable,” and so on, but there are many approximations of this behavior that we’ll start to reveal. I think we’ll go beyond simply saying, “I’ll create a new harness for encoding that uses a special form of compression.”

I think we can be much more creative about this. We can do a lot of things with these language models. I think we just don’t do it yet.

Speaker 2

I’m very inspired by this, because I think we can.

### Introduction

Alex Zhang

To get substantial benefits from very purposeful, high-quality design harnesses that are more scalable. I have in mind that RLM, for example, has a very primitive inductive bias. There is nothing extraordinary in its design, except that it differs very much from what we do now. But this could potentially scale much better with the data and environments that we have at our disposal.

swyx

The obvious counter-question people could ask is: modern harnesses are very universal when it comes to coding, and people see that they work for many areas. Claude Code is used for design, presentations, and everything else. Manus, Genspark, and Grok are very simple, neutral harnesses that work well with code and are also scalable. What’s an example of how we can improve them?

Alex Zhang

Yes. Let me mention another article that came out very recently. It’s an article about a taxonomy of runtime environments, and it seems to be from Arena. I really like this article because it makes assumptions that I had, and it confirms most of my assumptions exactly. The majority of environment selections are irrelevant because they’re all the same. But I will say that Grok, for example, is actually quite different, as far as I’m concerned, from how some other environments were developed. And I like it. I really like it.

At least here, it’s clear that I’m almost certain Anthropic and OpenAI train exclusively on their own environments. They probably don’t train on competitors’ environments. I would assume not, because I don’t know—why would they do that? But that’s what we see in the open models, right? For example, Qwen is very strong in open source. They need to train in environments. Older Gemma models were known for their weakness in this. The models are good, but they need to be trained in environments.

I still think that, to the extent that models become smarter or better, this difference becomes less important. In that sense, if you take Astra and put it in open source, it won’t go crazy, because it is, in my opinion, good enough. I say this because the only advantage between these different environments is mostly cost.

What I mean is that if you connect these models to RLM, they’re all not yet very good. They’re normal. I think this is mainly because of the types of environments in which we’re training. This is a class of environments. This cycle, I call it “trajectory as a prompt,” which simply means that you keep the entire execution trajectory as context for your base model. Even if you use subagents, that’s all it looks like.

I think that if we want to explore new environments, there should be teams engaged in conducting meaningful experiments on scaling new environments—for example, scaling after training on different environmental designs. I think we can get very significant knowledge or benefits from things like that, regardless of whether it’s RLM or something else.

And that’s exciting, because I think, for example, if you train on Fable, Fable was the best model for RLM because it had dynamic worker processes, and it was quite obvious that this was, to some extent, an acquired ability. Even if the model was a little imperfect, in the RLM environment it still worked much better than other models. Astra now also copes well with these tasks.

swyx

But is that where we are? We see that Fable is best for RLM. Is there any other model?

Alex Zhang

Oh no, that’s probably all internal results. Yes, I have them now. They’re not publicly accessible. But in general, I think you can very easily understand that we still haven’t optimized workflows for RLM. If we obtain models that do this right, they’ll be significantly more effective. You can just consider why this is, right?

swyx

I think this is the moment when it’s worth giving a 10-second explanation of what RLM is, because there are many listeners here whom we assume have that knowledge. But you’ve already set a certain context, so can we get a clear definition, taking everything we’ve said into account? RLM?

Alex Zhang

Yes. Good. I want to return to the blog about compositional generalization. I’d like to have it in the original article.

RLM is, in essence, a design for an environment where the only tool is code. It’s a software-subagent call in which the model has the opportunity to call itself as a tool, as well as other tools, but they’re all functions in the code. The context with which it works is always stored in the memory of this software environment. This can be a file system.

For example, I’ll give you Prime Agent. In the Prime Agent trajectory, even when you compress the context and do all these things, it’s stored on disk, so the model can always return to its initial context, even if it’s compressed. All its tools are launched, say, in the Python REPL or Bash REPL.

This is a very primitive abstraction, and I would say that most environments differ in that context unloading doesn’t happen. I mean, if you take a look at Prime Agent—by the way, Prime Agent does unload context, but not fully. It still keeps the standard Claude Code and Codex cycle, where the trajectory is the thing you compress, but it has an additional feature: the context is unloaded. The uniqueness of Prime Agent is that its only tool is IPython. So it’s kind of a very general abstraction around RLM.

swyx

Specifically, what is the main problem that RLM is trying to solve?

Alex Zhang

To some extent, I think when they first appeared, this was about context, but I’ll skip that part. The initial problem was that the frameworks, or harnesses, were very difficult to work with over a long context. Usually, they were used only for specific things, for example, code. As you know, they could work with your codebase because they were trained for it.

But now it’s more about what’s in this blog: compositionality. I think we want systems of language models that have much more control over the actions they perform at each step. By that I mean that the current harnesses are very tool-limited because they need to call a tool at every turn. You need to call tool A, then tool B, then tool B, and there’s no central context that you can contact. RLM is specifically designed around composition and the availability of a central context from which you can always draw information. This context is developed around existing language models.

Another design example that’s very similar is agent swarms—for example, the incident with Hugging Face. Such agent swarms have a message board through which they communicate, and this board is, in a sense, the common context in which they act. RLM, in essence, says that the best way to communicate here is code: you write code to realize this. That’s because these models are so skilled at writing code that we want to use that ability. And then, for this compositional thing, you create your own harness specific to a specific task.

swyx

That’s right. That’s right. Let’s get to this drawing.

Alex Zhang

Okay. So we’re talking about the idea of locally in-distribution tasks for the framework. This is an idea that complements in-distribution tasks when we’re thinking about language models. An in-distribution task is a simple task where a request is something the model has already seen before, or has seen some version of. Most systems work like that, right? They constantly add the trajectory as a query.

Ultimately, if you’re not Anthropic or OpenAI and you don’t train on these custom trajectories, most of these things, for the most part, go beyond distribution. But locally in-distribution—this is, in essence, a compositional argument: if an RLM breaks its calculation into a kind of meta-system or program that includes subagents for local problem-solving, then every separate language-model call during the task remains in distribution, even if the task as a whole is outside it.

This is a very desirable property, in my opinion, for obvious reasons. For example, if each task for every separate language-model call is in distribution, you’re most likely to get the right response.

The logical boundary of RLM is RLMs where you don’t just write a system, but also train your own model and collect data. It’s like a completely automated AI researcher inside your system. We’ll see where RLM training moves. I’m not working on this at MIT, at least not at scale, because I can’t afford to. But there are companies that are working on this.

I think Prime Intellect and Scale AI clearly are working on this, and this is very cool. I really want to see—maybe we’ll observe better results from scaling after training when you train the system around a smart harness. Maybe we’ll even see the emergence of smarter systems that work better on the basis of such principles.

swyx

What is a “smarter system”? It doesn’t mean anything. You just said that they’re all the same.

Alex Zhang

No, I mean that Claude Code, Codex, Pi, and so on are all the same in the sense that when you disassemble their logic, they’re actually identical. Two calls in a cycle, and that’s it. But with RLM and other system abstractions, this looks quite different. And right here, I think you really see the difference.

swyx

It’s just like with the choice of language-model architecture: many architectural solutions eventually become similar when you scale them. These differences turn out to be quite insignificant. For the laboratory, it’s not a trifle, but maybe one model performs a little better than the other—just a little.

Alex Zhang

But in general, if the scaling limits from pretraining only apply because the choice of architecture is somewhat stable, then if you completely changed the architecture, those scaling limits probably would not work, or the power law would look completely different. It’s exactly the same with harnesses. It seems to me that all the harnesses we have now mostly look approximately the same, but there are certain exceptions.

Actually, one of the things that makes me most interested in this, especially when we talk about graduate students who are willing to take big risks, is that people are also investigating another side: scaling laws from pretraining don’t work if you change the data. Right now, it’s simple, unstructured text—basically, data from the internet. What if you had a better data distribution for training? Then, yes, your scaling law would change, too.

swyx

There’s architecture, there’s data, and there’s everything else you can think of. I just want to return to this: it all makes sense. I’m wondering how you’re moving up and down the stack, from very conceptual things to where we are now. Of course, you can scale up and down, probably.

### Prime Agent and Persistent Subagents

I’m wondering how you started working with Prime Intellect. Does Prime Intellect do more work on this? Is this their response to Hermes agents? You mentioned that Grok is a bit different. I just wanted to list all these guys and know your opinion about everyone.

Alex Zhang

I joined Prime Intellect after they published a blog post—not related to me at all—about how they believed RLMs were kind of the future. I had a friend who worked there, a friend from GPU Mode, and we got in touch. I think I agreed with many of the researchers there, and I was very affected by what they believed about design principles.

I think they understood the purpose of the articles about RLM, which is not necessarily that we solve problems with long context, but that we need more specialized framework designs. There’s always the result of an article that you choose to highlight, in contrast to the real essence.

swyx

I would really like to talk about incentives in the academic environment, about why it’s imperfect, and about all the problems. We’ll come back to that.

Alex Zhang

Aside from that, I really like the guys at Prime Intellect. After we decided to work together, we decided to study RLMs and also create an RLM framework to see where this would lead. That’s how Prime Agent appeared.

I think Prime Agent performed quite well. The only thing I was worried about was that none of the models, at least by the time we created it, was particularly skilled at performing RLM tasks. This was before the advent of Prefabs and Astra[?].

swyx

Can you take a step back and explain what Prime Agent is and how it differs from traditional Claude Code, or what people expect from a framework?

Alex Zhang

Prime Agent is essentially a framework on top of Pi, specifically pi-mono, which is a minimalist framework. I use Pi as a standard for everything because I think every other framework is actually just Pi.

The difference is that we clearly restrict IPython as the only available tool. Any other tool can be loaded as a Python module or as a Bash script that the model can run. Therefore, it uses the main abstraction of RLM on top of Pi. It also has something called a continual harness. This was developed by Seth, another graduate student. He worked a lot with Joel, who is the same person whose Gemini plays Pokémon.

The continual harness is what he used to make frontier language models play games. It’s also very simple, and I quite like it. It’s essentially a design principle: which parts of the framework do you allow the model to change?

There are certain elements that you allow it to modify—for example, its own skills, the subagents available to it, the system prompt for the models, and the continual harness itself, which is available virtually as a tool inside the IPython kernels. That’s why Prime Agent is developed around this. Everything else in Prime Agent is standard.

A lot of the new advanced models work very well inside it. Even many open-source models work great, at least some of the newer ones.

There’s one more thing about Prime Agent worth highlighting: our very specific communication system between agents, or, let’s say, a framework for that communication. RLMs tend to create many subagents, and we want the subagents to be able to communicate with the root agent or with one another. So there are certain design decisions about which agents each agent is allowed to communicate with and how it does so.

It all happens in code. The agent writes code for this type of communication, which is, in my opinion, very, very clean. There are also persistent subagents. They may exist longer than the standard context window of the original agent, and you can go to such a subagent and give it more prompts. So you have more visibility and flexibility regarding what’s going on.

This is my main problem with Codex right now: its subagents are very fleeting, and it actively refuses to let you use them for long-term processes.

swyx

That makes sense.

Alex Zhang

The whole secret is to take everything out to the filesystem. That’s the main secret. You also force everything to work through code. You trust a model that can write code, and it will write all the necessary tools for itself.

swyx

Is Prime Intellect studying post-training custom models for this? Is this a one-time collaboration between you?

Alex Zhang

As of now, yes, they train models internally. I think they spoke openly about this back in March. Obviously, that’s their business. They show that they can train a model to use their own stack for training.

But no, I’m not involved in model training. The main reason is that, within the PhD, I have other things to do that I want to work on. I think there are many other big bets worth taking beyond RLMs themselves.

I don’t know as much as I could tell you right now, but in general, I think one of the advantages of studying in graduate school is that it gives you the opportunity to make a lot of big bets. Most of them probably won’t lead anywhere, but this is a very interesting time for research. In my opinion, most of the progress in this industry has been somewhat boring.

I’m not saying the results are boring, but the process of implementing these things is usually quite boring. So the question is: what do we want to do more of? We can talk about that later.

swyx

I want to summarize your research, and then we can move on to something else, because after the release of RLM there was a lot of excitement around it. Are there any third-party projects you want people to remember? Projects where you’d say, “You should look at this”?

Alex Zhang

Harvey, the legal AI company, published a blog post. They aren’t related to us, but they ran RLM on their legal work, which often involves a lot of document review and searching for diverse, specific information that may not be easy to extract with simple search systems. They showed very good results.

This is fascinating. I was shocked that they worked on it. They don’t like me, they said. So when this came out, I thought, “Oh, this is great.” This is what I think is super cool about what they’re doing.

There also seems to be a collaboration with Base10.

swyx

This is B. Headlong[?]—what is it?

Alex Zhang

It’s a Ludwig Institute instrument. It’s their system that works constantly. It’s very cool. They renamed it; they used to call it something different. It was something like AutoTerminus[?].

swyx

Yes, I know—they’ve gone through this. It’s Andy’s big project, Kwinski[?].

Alex Zhang

This is super, super cool. I love Andy. I don’t want to belittle what they’re doing because they use the RLM abstraction, but they’re doing something much cooler than just RLM. They have a system that is constantly thinking. So even when you don’t make a request, it has a way to think about problems that are in its context.

swyx

It’s like the system is always on—something that always works, but not very expensively.

Alex Zhang

They control the token costs to make sure it doesn’t burn through all your funds. That’s very cool.

I’m trying to remember. If you go to the GitHub RLM page, there are a bunch of really cool things that people have done. Axe is another very cool project, and it seems to be from a single developer. It’s like the DSPy toolkit and RLM. DSPy also has RLM.

The last thing I want to mention is ARC-AGI-3. I think there were a lot of tools for its competition on Kaggle—official tools, not publicly selected tools—on which people conducted an evaluation. They all claim to use, or refer to, for example, Tufa[?], some kind of form or inspiration from the RLM abstraction in their tools. That’s very, very cool.

swyx

I think that’s it. This is an environment where you would see a lot of benefits from composition and code usage, combining neurosymbolic systems with artificial intelligence and so on. We love a good link to neurosymbolism.

You also mentioned ARC-AGI-3. OpenAI is coming out and saying that they reached 99.9% on this. They also say they solved the Navier–Stokes equation. They just threw it at the model. There’s some debate about whether it’s a simple model.

Mhm. Do they use RLMs? Do you know?

Alex Zhang

I mean, I would assume that they absolutely do not, but don’t quote me. I’ll be careful here, because people argue that there is RLM and whatnot. To a certain extent, it is clear that they used some kind of swarm of agents working from a shared context, such as a shared filesystem.

That is very much in the spirit of RLM, but I think a lot of more cunning things were done that may not be related to RLM itself. I partially agree that the toolkit wasn’t necessary for what they did. I would put it this way: in my opinion, a model like GPT-6 Astra is technically smart enough, given the appropriate information, to find a proof for these very complex tasks.

But how do you get that information? That’s the big question. In their case, it probably came down to a very long search among many subsystems or agents in a swarm, possibly with the participation of researchers. I’m actually not sure about that part, but they may have added their intuitions about what was worth investigating and things like that.

In the end, it worked out. Some agent was able to use the right information to complete the proof. So, in that sense, was the scaffolding important? No. I think this indicates that the specific details of the scaffolding don’t actually have much value.

What exactly does this indicate? It’s the idea in my article about a “tax” on the scaffolding: apart from the user experience of the system, all that really matters is how you combine these agents in a meaningful way to get to the final answer. Perhaps that is the point of swarms and things like that.

Therefore, at least from my point of view, if we start thinking about custom scenarios, what do we want from scaffolding? We want to take the best of it: Claude Code, Codex-type models, and the stream that people love to see. But under the hood, it can work as a very strange and complex swarm of agents that ultimately yields a response.

The user doesn’t want to see all of that, of course. That is incomprehensible information. There is still one thing in the spirit of RLM: recursive language model. It sounds like a language model, but it is not a language-model architecture.

The reason is that, as I think we’ll start to see in the future—maybe one day, and I wrote about this on the blog—what we think of as a language model, the object we appeal to, could actually be a swarm, a framework, or some strange scaffolding design that scales well, but that the user never sees. Ultimately, everything the user sees is just some interface for this system.

swyx

Hmm, and this perhaps takes us back to the basic limitations of the Transformer. Obviously, if you just take a basic Transformer and tell it to solve the Navier–Stokes equations or something like that, it won’t do it. We all know that this isn’t going to happen.

Alex Zhang

But I think the more interesting part—and perhaps the more relevant question about whether it has meaningful complexity—is whether simply directing models or agents around a context of information might be enough to solve very complex problems.

swyx

I can believe that. Although there are a lot of things I want to say about the amount OpenAI spent, it is semi-public: 10,000 agents over 88 hours, using 130 billion tokens, estimated at approximately $40 million at public prices. Surprisingly, that’s actually less than I thought.

Alex Zhang

Yes, it’s not really that much. That was 130 billion tokens for the output, for the final result, but as you said, a lot was passed through the context. It’s more than twice as much if you count the total number of messages between agents.

swyx

Yes. I think one thing I should mention is Cursor, which refers to swarm systems. This is a bit older information—it’s from February, which is already ancient—but if you scroll to the very end, the final architecture of the multi-agent system they arrived at was actually an organizational diagram of a conventional software-development team.

One thing I’m talking about—I think that’s what’s there—is a “gather all” function, which you perform with subagents. This is actually very similar to programming for GPUs.

Alex Zhang

Yes, and therefore it is a bottleneck. There is one main agent coordinating all the subordinates. Then everything needs to be gathered again, and then it has to be re-coordinated. This is slow, and it’s bad. A real swarm should consist of independent agents.

swyx

Well, I agree with that, but I still think there is a question worth asking. Let me give an analogy: when do you use compaction, and when do you use an RLM? There are many situations where an RLM can solve a more difficult task than compaction, but you would prefer compaction in most cases because it is cheaper and faster.

Alex Zhang

I think something like that is happening in the context of swarm agents. I’m pretty sure that about 95% of the swarm is absolutely useless, or that what it investigates is just burning tokens. Although, in this system, maybe that’s not exactly the case.

Everyone has a job. There’s a board in Jira. I think that, at a certain level, precisely formulated problems work like this, right? If you have a search task and create a pile of subagents to implement it, there will be a lot of useless information.

There is one answer that you receive and branch out from. But that is deliberate and understood, right?

swyx

Yes, but there is a question about the appropriate use of these systems: which problems should be solved with which design? In theory, OpenAI could package this into an API called “Swarm,” give it to you, and say, “Direct this at any problem, and we’ll give you a solution.”

Perhaps it would be called Pro. Maybe you would have to pay $40 million for the result, but that’s fascinating. I will say that it is very exciting to be able to throw $40 million at solving problems.

Still, we need to do a lot of research in this field to understand what is necessary and what we want to do. Which design do we need? We probably don’t want everything to be a swarm, but where do we draw the boundary? Can the agent design it itself and make that decision?

Vibhu

And also, just a quick question: do you have an opinion on this? Have you investigated open-endedness as a general category of problems? I mean, without hints—just “go forward.”

Alex Zhang

A little. I was at Sakana during the summer, right before graduate school, and they are very much working on this. I think many people, even in recursive superintelligence, are thinking about it.

swyx

Oh, yes. Richard’s company—and then, wait, maybe that’s one and the same company. I’m not sure I remember. And Tim Rocktäschel too?

Alex Zhang

Okay. He’s the main co-founder. He’s head of open-endedness at Google.

I think the problems of open-endedness are similar to unresolved mathematical problems. Maybe that is a strange way to phrase it, but I think many of the methods people use to approach them are similar. Evolutionary search is very similar to launching swarms of agents in the hope that they will come up with something interesting. That is what AlphaEvolve and some other work from a year or two ago did.

Vibhu

But maybe not—not at all. I’m not sure I understood you. In open-ended search tasks, do we formulate them all as unsolved, very difficult problems, or as problems where the goal is something else? If not, I still don’t understand where their value lies.

Alex Zhang

Maybe you have a different opinion. I don’t have much of a view on this, but at least when I was at Sakana, the impression was that, in the end, we still wanted to approach things like OpenAI approached Navier–Stokes. We still wanted to push in a certain direction to reach the point where we got something interesting.

My view is that there is a possible division between fundamental and applied science. Fundamental science is research for the sake of research: you just want to understand things better. I have no idea whether any application will come from it at all. With applied science, you have a target; you are trying to minimize losses somehow.

I really think this applies to the large bets that people make. What if there wasn’t a prompt? You just choose. You’re in this swarm of things and say, “Hey, what’s up, guys?” Then you say, “I like what you’re working on,” and decide, “I see, this is an interesting problem.”

The biggest problem we had with open-endedness was how to make a choice at the end. How do you isolate the interesting things? When there is no goal, maybe the agent invents one, but in many cases there is something called Fugu.

It seems like something similar to routing models, inspired at least in part by the idea, “Let’s choose a task where we can find the best solution for something.” In this case, it is like choosing the best model for a task.

swyx

You are the first person who has connected routing models with open-endedness.

Alex Zhang

No, no—yes. But I remember this because I think the problem with open-endedness is usually that, among a huge body of waste, you have to filter and find the real pearls. The solution for an open-ended goal is simply to allow models to work forever and generate data at extremely high productivity, asking questions like, “Maybe that’s cool.”

For me, it is very interesting as a counterweight to almost everything in machine learning, where you have a goal, and here you have a lack of purpose.

Yes. Or maybe an undefined goal. You think, “What about this goal?” and you're like, “Well, okay, possible.” Then you explore more and find that this is an interesting goal. I think finding your own target functions—as you said, Jeff found an objective function that was interesting and that no one had investigated—is a message to graduate students: stay in university if you want to do open-ended work. If you want to maximize profit and avoid permanent underclass status, then you go to a laboratory.

swyx

Yes. This is funny. I don't think I do that often; I hear these discussions. I'm from the East Coast, and the atmosphere is completely different there. But when I come here, this is always the main topic of discussion. You won't be able to pay rent without doing this. You're being pushed out because of prices, guys.

Alex Zhang

Okay, so, yeah, that's it. I don't know—do you want to? Is this enough? It's important to talk about the “poorly managed geniuses” you mentioned. I remember that.

swyx

No, no, it's just one from your blog, but I'll touch on this topic.

Alex Zhang

Yes, yes. They did it in their blog post; I also talked to them about it. One interesting result overall is that they actually gave it freedom to automate scientific research with minimal human intervention, or without it at all. It just does it.

swyx

Yes, so this is AI Scientist, which is a little more open. There are different levels of openness, and I agree with that.

Alex Zhang

Yes, this is different from autoresearch with a specific purpose. Here, it's just “do something.” But that's cool. They're working on it for people who need it.

swyx

Since we're already talking about what they discussed, what is your opinion of Sakana AI? What do they do, apart from being from Japan?

Alex Zhang

I really like the people who work there. I think they have a large and very smart team. That's quite logical. They split off from the previous team at Google DeepMind, which was probably also doing similar research in an evolutionary, open-ended style.

One thing I liked from my experience there is that they had a failure. I don't remember when, but I think it was related to the scientists.

Vibhu

No, it was the GPU kernel.

Alex Zhang

Yes, the GPU kernel. People can't forgive them for that. Regarding their AI Scientist, there is some criticism that I can't really comment on because I don't work on it; I don't have an opinion on that.

In general, what do I care about them? I like that they are more like a research laboratory. They definitely don't work on the same plane as OpenAI or Anthropic. I think it was quite obvious, at least when I was there, that they didn't have a large-model competitor that everyone would use.

But to me, it seems they function to some extent as a PhD laboratory, and that's cool. David Ha, to me, is really very reasonable. I think he understands well that the Japanese market differs somewhat in terms of AI, and their target audience is a little different from what we're used to here.

But yes, I like that they take their research in a slightly strange direction. It's strange to me when people watch them, you know, and I like it. I think there should be more eccentricity.

swyx

Yes, that's right. That's right. When you say “another market,” is it more of a corporate segment?

Alex Zhang

The way everything works there is slightly different—how deals are made, and so on. I think their page is primarily in Japanese, but they have a model specialized for the Japanese language. Didn't you know that?

swyx

I didn't know that you knew.

Vibhu

I also have personal friends on this team, and I know it.

Alex Zhang

So, from the point of view of how you communicate culturally, the answers are tuned specifically to it. It's not something more advanced than benchmarks; it's a culturally adapted model for them. They have chat and everything else.

swyx

Supporting your opinion, there is also the question of what education should look like. Someone wants to do this work, and it's very similar to a PhD-format laboratory. Do your thing—why not? We have money; do your research.

Vibhu

Oh, they say Kimi is fine, too. That's not bad. I'm glad for them.

Alex Zhang

Yes, speaking of Kimi, another one was a graduate student who separated off and said, “I have this Kimi Delta Attention that I want to work on.” Somehow he managed to make a one-shot. I still don't understand it.

swyx

Yes, yes. Well, he is incredibly cool, at least as far as I understand. I think in general many Chinese laboratories have done really cool work.

### Kimi vs. OpenAI Agent Swarms

Alex Zhang

Yes. What are your thoughts on swarms—Kimi agents?

swyx

One thing I would say is that OpenAI wouldn't have done anything with a swarm of agents if it were clearly the right way. We need to think about this. Nothing in particular works without a very, very smart design framework, and I don't think anyone has been given one like that, easy and free.

For example, designing a swarm of agents is not something that can be taken for granted. It's not as if GPT-6 or Astra is simply super-smart and then it just has swarms of agents. They are clearly trained. I have in mind that Hugging Face was training its system so that it behaved like a swarm.

Well, I think they clearly did something very good. If it's possible to spend $40 million to resolve an unsolved problem, then that's what you do. And with this thing about swarms of agents from Kimi, at least from what I read, it looks interesting, but I don't know whether this can solve something new.

Alex Zhang

They just said that it creates spreadsheets. They simply said, “Here is a swarm that does something, and that's cool.”

swyx

But I will say this specifically about dynamic workflows. Actually, I think the release of dynamic workflows was kind of a failure.

Speaker 1

I don't know what you mean by that. Do you think people are using it wrong?

Speaker 2

As far as I understand, people are often using it wrong, or it's just too expensive.

Speaker 1

This is too expensive. This is Claude Code. It's basically, “Take my code under control.” It doesn't work. I tried it, and it doesn't work.

Again, I'm not OpenAI's biggest fan, but I think what they did is very impressive. They somehow managed to force this swarm to move toward goals, and yes, that's very difficult.

Speaker 2

Of course. So the efficiency of a multi-agent swarm is the objective function here, yes? How much useless work is there? A bunch of junk.

Speaker 1

I think so too. I think we take it for granted that the swarm will arrive at a certain answer. That's not something that can be taken for granted.

Speaker 2

Yes. We recorded a podcast with Noam Brown, and his whole concept came down to the fact that we worked on many competitive agents. Now we're working on collaborative agents, and this is now called a swarm.

Speaker 1

Yes. I would like to quickly ask: what are you thinking about Gemini?

Speaker 2

I think they were the gold standard. There was a time when they forced agents to reason for a long time.

Speaker 1

Is that too much?

Speaker 2

No, no. I think Gemini is a little overestimated. Again, I must point out that I didn't work at any of these companies, so accept this skepticism, okay? I mean, you have your own opinion about agent systems, and you know that.

But let me tell you first: this work was truly impressive. I think they showed that it took time when the models were not yet so accurate. They managed to approach wisely what this toolkit does.

I remember at least—for example, AlphaGeometry. It seems that was a year before that, but it was very cool. They took it, squeezed the maximum out of it, and developed something. To be honest, I am very impressed by Olympiad mathematics.

It seems to me that Google DeepMind is a bit sad. Everyone I've talked to at Google DeepMind has approximately the same opinion: there is too much bureaucracy. They have the talent and resources to make almost anything, but I don't know when they'll understand that.

I have nothing against Antigravity, for example, but I don't know anyone who has used it. I tried it once and don't see any reason to move on to it, and I think some of the reasons for this are internal fighting.

Speaker 1

Yes. Well, you know, many people criticized Meta for a long time, until they started producing products, and I think Google is now experiencing exactly this phase. You need to stay afloat.

Speaker 2

I want to go back to your thoughts in general. We can talk about speculative tool calling, about “geniuses in an uncontrollable environment,” or just throw it all away and talk about anything else.

Speaker 1

### Capability Overhang and Speculative Tool Calling

Okay, let's talk a little about “geniuses in an uncontrollable environment.” Just one comment on speculative tool calling: this is a very simple idea. It's almost obvious that this had to be done, and there's not much to discuss. I think it's just worth using for any software-agent calls, like RLM or CodeAct. It's a no-brainer—the obvious thing.

For those who haven't read it, what's the point in one sentence? It's simple: while the model writes code, or even after it has finished, a lot of tools work sequentially, or you have to wait for them. So you need to launch them in advance. If you can compile this code, you can probably understand—even despite the variables and everything else—how to statically analyze its assumptions.

Speaker 2

Yes, someone pointed out to me that scientists, especially experts in programming languages, have very cool ways to do this. Maybe someday I'll work on it. Mostly, it's necessary to change the language so that JavaScript or Python can't let you do this.

Speaker 1

Yes, yes, yes. For example, Haskell. That’s normal; it’s not Lisp. Or Lisp, or OCaml—any functional language where you can do this to convey effects, and therefore use effect types if you want to work with TypeScript.

Okay, we can move on to the diagram. What is, in essence, your core thesis? It looks like this to me, paraphrasing—but maybe I’m missing something: work on better tools, and your models are actually capable of more if you make an effort. So this is a problem of skills.

Speaker 2

Actually, there’s one thing I want to say. I appreciate the great focus on “uneven,” or jagged, intelligence, because it draws a wide picture of what we can do if we really focus. But I would have wanted—and maybe someone in academia will do this—someone to just sit down and think: If I took Astra, even the current leading model, and it wasn’t good enough at executing some work over the course of a month, consistently and qualitatively, I think that’s a silly problem.

I sincerely believe we can solve this. You don’t have to be an advanced laboratory or do all these quirky things for your IPO. I think these models are so smart that, even if it’s a silly approach, we can make this work. It’s a skills question: you can force the model to work equally well, like, let’s say, an 18-year-old schoolboy performing a task. I believe it’s absurd that we can’t make this happen.

Part of the reason is that the language-model format isn’t very well adapted to this. But I think you can build tools around it and do it. Ignoring RLM and all sorts of questions about which abstractions to use, I just think someone can create a tool that will do this. You ask it what to do—perform a long but simple task—and it does it reliably.

Speaker 1

So is this different from, say, your favorite company, Harvey, which uses an LLM for legal work? What’s the difference?

Speaker 2

I guess I think it’s something like this, except that the narrow part isn’t certain legal knowledge or something like that. Honestly, I don’t know. Let’s say this: people can take models, build pipelines or whatever, and force an agent to perform whatever tasks they want, continuously, right?

Speaker 1

Yes, yes.

Speaker 2

For example, if I wanted a general system that I could communicate with as I would with an intern, and simply ask it to investigate some small thing, that would be useful. An example might be simple automated research. It doesn’t have to search for extremely novel solutions; it could at least optimize all the easy parts of a problem.

Maybe this isn’t very clear, but in workflows there are many moments when it would be easy to quickly write code for a specific tool to automate something. One example is searching scientific articles. Usually, people create a specialized agent to help them with this, or they quickly write some code and run it as a Slack bot or something.

But it seems to me that there should be some standard format for a tool that you just connect. It doesn’t need to be developed specifically for searching articles; you just tell it, “Find this for me,” and connect it to the right environment. In my opinion, there are many simple things that can be automated.

Speaker 1

Is this a hypothesis or an argument about an abundance of opportunities—that even if we took a break from improving current models, they would still have a big influence?

Speaker 2

In a sense, yes. I think what I’m presenting has this form in its simplest version, but it also indicates that we have uneven intelligence in many areas. For example, models are disproportionately good at programming and mathematics. That means we can transfer these abilities to many other areas.

If we take someone who has won a gold medal at the International Mathematical Olympiad and apply their abilities to many different industries and problem-solving tasks, they can figure out almost anything. I don’t really know whether this applies to models. For example, in GPU code optimization, there’s a very interesting question: If you removed all the model’s data about GPU programming, but it remained as capable as a state-of-the-art model, would it still be able to optimize GPU kernels?

Could it study what it needs in context and then use some kind of pipeline, or come up with a solution, to perform optimization tasks? I think there’s a discrepancy. If you take a person who is as smart as a state-of-the-art model, there’s a gap between what a human can do and what the model can do, perhaps because of the tools available to it. I think we can actually do much more to approach human abilities.

Speaker 1

To me, it sounds very similar to the problem of continuous learning. Is that what you mean?

Speaker 2

This is the best example.

Speaker 1

Yes. Why didn’t you just say that?

Speaker 2

I guess I was wondering if I could boil it down to a few words. I try to be careful.

Speaker 1

I actually know where you studied as an undergraduate. You were a mathematician?

Speaker 2

Yes, exactly. I studied mathematics—at least, I wanted to. I was interested in abstraction and category theory, where you think in terms of categories and then have to go down to the specifics, while actually caring more about the category itself.

Speaker 1

That’s a mistake in communication, because everyone is waiting to hear something specific, while you’re trying to convey something general.

Speaker 2

Yes, which is quite difficult. I don’t know. Perhaps you can use some abbreviated notation, like, “I’m on the second level,” then, “I’m moving on to the third level,” and, “I’m going back to the second one.” We should have some epistemological shorthand for this, because it’s difficult when you’re trying to compress a lot of meaning into a sequence of words.

Should we switch to neurocode? Do you know if there’s some better way to encode data than English, Python, or JavaScript?

Speaker 1

### Neuralese, Future Research, and AI for Science

Not that I know. It’s a bit similar—just kidding—but people have thought about what the native language is and what language they want to give a model. Some say binary code. This is a topic for Mark and Jason, I suppose. PTX, maybe—a mixture of English and Python.

I only say this because the capabilities of a model are, to some extent, a reflection of what we train it on. So we still want—

Vibhu

I don’t really believe in the argument about binary code. I think I understand it, but it’s like—you want to model the world in a certain way.

I always remember one thing in conversations like this: the Sapir–Whorf hypothesis. If you choose English, you’re limited by how long the language has existed—let’s say 500 years, which isn’t very long. In fact, the language you speak limits your thinking. If you learn another language, for example, Chinese—

swyx

I don’t really know Chinese, and I’m talking in Chinese.

Vibhu

Oh, I knew about that, but my Chinese isn’t very good. Or, let’s say, Japanese. I’ve forgotten which language it is, but in Korean, when you’re talking to someone, you have to consider their social status, and that’s a different dimension from sex.

swyx

Exactly. It’s simple, but it affects everything you do.

Vibhu

When I studied linguistics, I remember there was some language in Africa with a “vegetable” grammatical gender.

swyx

Yes, exactly. Or the absence of words for “snow,” or something like that.

Vibhu

The language you choose affects your thinking. If you choose to express your train of thought in English, you shift the focus toward what English emphasizes. I don’t know what the a priori assumptions of the English language are.

swyx

That’s interesting. I hadn’t thought about it that way. To some extent, you’re right: in many models, the chain of thought also varies depending on the language. An obvious example is Chinese models that are speaking in English but can still think in Chinese.

At the same time, most models are very capable of multilingualism. We see that it’s possible to add languages without making a great effort to study the whole language; they can all think interchangeably.

Vibhu

Yes, we’re all autoregressive.

swyx

Take German, for example. The verb has to be placed at the end, which is very annoying. It’s a known fact, right? You don’t know what someone is saying until, at the end, this pile of nouns appears and then the verb.

The most classic example, which many people have heard, is the film Arrival, where the heptapods perceive the world as something flat. Therefore, they think and utter complete sentences all at once. This is very similar to the difference between autoregression and diffusion.

We’re speaking with the help of autoregression. What if we could speak by diffusion, where everything gradually becomes clear over time?

Vibhu

I understand. But the whole idea arises immediately.

swyx

I understand. So this is a radically different language, but it’s still a language.

Vibhu

It really is interesting. In this language, they can talk to machines, which we probably will never be able to do, but cars don’t care.

swyx

Maybe this is a big digression, but is there anything that, in its own right, is essentially a chain of reasoning that is autoregressive?

Vibhu

Sometimes, yes. The first thing happens, and then something else happens. Anything in code, for example, must be causal and sequential—at least usually, it must be.

Alex Zhang

Well, no. Then you don’t study enough theory of functional logic programming, where everything is purely functional and fully relational, and you abstract away from a solver that transforms these eternal truths into code. So, yes, I feel that this may be a little outside my competence.

But I love languages, right? Programming languages or human languages, and I think a lot about how they affect thinking and the boundaries of our opportunities. I don’t think so. It’s worth delving into this more, but I don’t know. Do you have any other thoughts?

swyx

My final question was about all these research directions you take. Do you want to work—you know, you had a GPU Mode phase, and then there was the RLM stage. You probably have something different planned, so you don’t focus only on this.

By the way, I noticed that it was interesting how you started from the GPU side and then moved on to zero-gradient, as it’s called. Doesn’t it seem less serious than working with GPUs?

### OpenAI Swarms and the Future of Language Models

Alex Zhang

Yes, probably in the sense that you mentioned. I love thinking about things from a mathematical point of view, and sometimes it’s very uncomfortable to work on scaffolds and agents because it seems that all the scaffolds are the same, with few new ideas in scaffolding. This is also simply empirically difficult: it’s hard to check a lot of conclusions, at least with the computational resources we have.

But the reason I switched to many of these problems is that, in my opinion, this is still where most of the innovations are. GPU work is a tool for learning about other ideas. You want to, for example, write kernels—or even automate their writing—for the sake of a broader goal: “I want to explore ideas where I’m not restricted by system calls.”

In this sense, I think that a lot of what is written concerns scaffolding, but I’m also interested in things at the model level. I’ll stop there.

swyx

Okay, that’s a good hint. If people want to apply to you, what exactly are you looking for? Who do you want to collaborate with? Any calls to action?

Alex Zhang

Probably not anything like that. There’s no one I feel the need to work with, except perhaps a company that can provide computational capacity, or just people to discuss this with. But I never mind working on ideas with other people. Students often apply to me, or even other podcasters.

swyx

Podcasters.

Alex Zhang

Podcasters. And usually I get emails like, “I really like RLM. I want to work together.”

I feel that this is an unfortunate approach, right? The worst thing is when someone asks, “Is it possible to use your knowledge?” I’m like, “Why exactly? Read my article, dude.” Typically, they’ll say, “I read your article,” in quotes, “about Recursive Language Models, or about a project such as the RLM Toolkit for self-improvement,” or something like that.

I really like people who have their own opinions, even if we don’t agree. I think that if you have solid beliefs and are able to think about why you think those thoughts are correct—not because it’s usually difficult to actually say—but if you have stable views on some problems, I’m always happy to chat and maybe even work together.

I have no restrictions about whom or what I would like to work with.

swyx

In the era of agents, I think more work can be done around bandwidth. In general, I’m hard to impress, but I think it takes a little effort to know what you want.

When you see something new appear—well done, a nice, simple idea—it immediately attracts your attention, right? It’s actually not that difficult to attract the attention of all the people from advanced labs, because they’re looking for you. You just need to put yourself out there, right?

Alex Zhang

That’s right. Yes. But it’s true. I will say that human attention is very scarce now, and I really struggle with the number of projects I lead. I don’t know how to cope with this.

swyx

I don’t think agents are generally helpful. I guess I just create a prompt, create a thing, and then never watch it.

Alex Zhang

Yes, which is very common. And that sucks.

swyx

Probably one of the small differences is that, in my mind, you may be engaged in research. I’m not sure.

Alex Zhang

But for me, at least, I might have around 10 or 15 different ideas that I want to implement, but most of them are unsuccessful. This could also be true of someone who comes to me. Maybe an idea is actually unsuccessful, but if it gets me interested, we can spend a little time studying it. If we feel that there is something there, it’s worth dedicating the following few weeks to taking it seriously and doing something with it.

That’s my style. That’s why I love being in graduate school, by the way, because there are moments when I’m just thinking about problems—for example, while jogging or playing tennis, or doing something else. It’s like I’m not working, but these are the most interesting moments.

Then, when I’m really into something, I just drop everything and take it on. I spend all my time reflecting on and working on this problem. When you reach a stage where you can just conduct experiments, everything becomes quite easy and goes like clockwork.

swyx

So, yes—sorry, that’s it. The penultimate question is about what many people now view as the next milestone in science: physical science, biology, or even mathematics. How do you differentiate, let’s say, projects that are applicable in industry from projects that are purely scientific?

Alex Zhang

I was actually working on artificial intelligence in biology even before starting graduate school. Since then, this area has changed a lot, I have to say. Previously, it was purely theoretical, like, “Of course, you have something in mind. Here, I have one path, and I chose it.”

But now many people are moving into this area, so we created a research department to cover these topics, because many engineers say, “Actually, there are solvable problems.” You know, I’ll start with this: my understanding of many of these topics is probably quite limited, but if I find out about something—or someone contacts me, or I see a problem myself—and I think, “Hey, some design principles that we use, or are thinking about now, are actually well-suited to this task,” I admire these things too, but I think it’s more difficult.

I—I don’t know. I think science, especially empirical science, whether it is applicable, has very long feedback cycles. Yes, it is. It turns into...
