# Can Rewriting an AI Agent Bend the Intelligence Curve? - Zhengyao Jiang

Machine Learning Street Talk · 2026-09-26 · 44 min · https://www.youtube.com/watch?v=yB6_iFGTq9k

## Transcript

Tim Scarfe

In my opinion, scaffold engineering is a cheap and effective way to adapt intelligence to a specific task. Back in July, a new AI company from London posted a very viral tweet on Twitter. It garnered about 1.8 million views and claimed to demonstrate the first evidence of recursive self-improvement. The problem is that not everyone believed it.

Zhengyao Jiang

I’m Zhengyao Jiang, co-founder and CEO of Weco AI, where we build self-improving agents.

Tim Scarfe

You actually said that you spent 8 days doing autoresearch for an autoresearch agent.

Zhengyao Jiang

We ran autoresearch on the autoresearch framework and were able to discover a better autoresearch framework. Since then, we have continued to fine-tune it over the past 2 years.

Tim Scarfe

When most people think of recursively self-improving intelligence, they imagine a system that can reflash its own brain in a loop. But what if the system only changed the code around this “brain”? Would that count?

Zhengyao Jiang

We are interested in the result. Does it actually improve the capabilities of the system, rather than any specific level of the system?

Tim Scarfe

What about the little problem of Goodhart’s law, which states that when a measure becomes a goal, it ceases to be a good measure?

Zhengyao Jiang

Even more interestingly, we found that it develops mechanisms that prevent the reward system from self-destructing.

Tim Scarfe

What is the biggest misconception people have about recursive self-improvement?

Zhengyao Jiang

The biggest misconception is that if recursive self-improvement is achieved, a technological singularity or some kind of intellectual explosion will occur. I don’t think that will be the case.

Tim Scarfe

There is a whole community of people creating startups focused on recursively improving superintelligence.

Zhengyao Jiang

This project actually started about a year and a half ago, when we became interested in the concept of recursive self-improvement. Historically, all research has had a diminishing-returns effect. However, there is an idea: What if we allowed the researcher to increase its own efficiency?

Historically, this has been possible through improved methodology and better tools, but human researchers—or the human brain—are always the main bottleneck. If we now have an autonomous research system, can we direct the research topic toward itself so that it can increase its research efficiency? The hope is that one day this efficiency gain will be able to overcome diminishing returns, so that the curve of the ratio between effort and results changes from concave to convex.

Of course, this is a big goal, just like the creation of AGI. We don’t know how far we can go, but improving the efficiency of self-referential research is definitely already in our development plan.

### AIDE and the puzzle of useful spaghetti code

Let me explain everything properly first, okay? Weco created an agent called AIDE, and it’s a research platform. It is very simple: you give it a metric, set a task, and it climbs the hill to solve that problem. But they set it up recursively, so it actually kept climbing the hill until the agent itself improved.

This means improving the platform: the prompts, the tools, and generally the entire set of capabilities it has. They just left it to work. We left it running and saw what would happen.

We observed something interesting in the self-improved autoresearch platform. We can see exactly the code that it generated. It’s like spaghetti code from another planet, but for some reason it generalizes very well.

Of course, we have a set of benchmarks on which we try to climb the hill. After that, we test it on held-out benchmarks. These are public benchmarks, such as MLE-bench Lite and ALE-bench Lite. ALE-bench is Sakana AI’s algorithmic discovery benchmark, and MLE-bench is a machine learning engineering benchmark from OpenAI.

We also tested it on a task far outside the distribution called WeatherBench 2. Essentially, you are trying to create a physical forecasting model to predict the weather. WeatherBench is very different from the training or optimization dataset, but it still generalized well. It generalized better than our, shall we say, more elegant manual platform. We are somewhat confused by the results of the experiment.

Tim Scarfe

How much do we want to stick with more fundamental, pure solutions versus a solution that actually works better in practice?

Zhengyao Jiang

I think there is definitely room for improvement here. For example, we could add certain constraints to the optimizer. One option for constraints could be something like regularization in neural networks: how do we make it look for the simplest solution to the same problem?

That’s one direction, but another might be, “Okay, you just have to accept it.” One of the factors that makes people think it is spaghetti code might be the amount of work done by the agent in the background. For example, it conducted 100 experiments in our self-improvement research, and that number is probably equal to what we have done in the last 2 years.

If you bring in a new intern to learn about our codebase, they might also think, “Okay, this is so hard to understand.” I think this is partly because you have to endure the cognitive load of understanding the vast amount of experimental results.

Tim Scarfe

Do we even need these restrictions? This is a very valid question.

Zhengyao Jiang

I think that in practice, such restrictions are still necessary. Humans and artificial intelligence must cooperate. There are certain aspects in which human intelligence still surpasses artificial intelligence.

For example, in AIDE, when we sent an agent to the OpenAI qualifying competition, we found that humans were still significantly better at generating creative primitives. We realized that it is very important for a person to create the first prototype that gives the agent the correct search space. Essentially, that is the initial abstraction. We are fixing the search space.

This is a bit like designing a neural network architecture, where you introduce inductive biases into the learning process. If the initial codebase is based on a search framework, as in AIDE, then many of the agent’s ideas actually come from the search literature.

But if you give it an initial codebase based on a ReAct agent, then you are effectively putting all the context into the prompt and giving the agent complete freedom of action. Many of the ideas then come from the agent itself—the agent in the outer loop.

An outer loop means that an agent optimizes another agent.

### Four levels of recursive self-improvement

Tim Scarfe

So what do we mean by recursive self-improvement? Are we talking about creating a machine god? Are we just talking about code that runs in a loop and optimizes itself?

You mentioned 4 RSI levels—recursively improving intelligence. What are these 4 levels?

Zhengyao Jiang

The first level is what we call delegation. You can start a cycle of recursive self-improvement, but it is not necessarily better than a person doing R&D. We believe that all previous public results are at this level, level 0.

Level 1 is what we call productivity. You start a self-improvement cycle and find that the rate of self-improvement is actually higher than that of a person doing R&D to improve the system. That is the level we are at now.

At the first level, when we talk about improvement, there is always a way to measure it. For example, we measure it using a set of applied tasks. But we didn’t measure how well it was improving its ability to improve itself.

Tim Scarfe

Can an inner loop found by an outer loop really become a better outer loop?

Zhengyao Jiang

That is what we call level 2 recursive self-improvement. We call this ignition.

Tim Scarfe

Why is this generalization from the inner loop to the outer loop important?

Zhengyao Jiang

It is a necessary condition for achieving the main goal of recursive self-improvement: the actual scenario of an intelligence explosion. Increasing efficiency can overcome increasing complexity.

At level 1, you will still see diminishing returns. It is as if the positive feedback loop from the inner loop to the outer loop has not started yet.

Tim Scarfe

In this system, the model itself hasn’t changed. Adaptation is very important, but do we mean adaptation of everything, or adaptation of the critical path? Do you need to have just a recursive loop that adapts certain parts of the system, or does the core of intelligence—the model itself—have to adapt for us to call it recursively improving intelligence?

Zhengyao Jiang

We don’t care too much about that, because we care about the result. Does it actually improve the capabilities of the system, rather than any specific level?

Of course, it can be argued that much more can be done at various levels. Honestly, I think that for level 3 recursive self-improvement, when we reach the inflection point, it will require this self-referential cycle at all levels.

But we are agnostic to the level of implementation. At the same time, we do not believe that only model improvements count as recursive self-improvement.

### What AIDE 85 changed and how it was tested

Tim Scarfe

Can you explain what the behavior of the most distilled agent you found was?

Zhengyao Jiang

We call the best agent AIDE-85. There are a lot of changes, but to summarize, it is an improvement in the search algorithm. It includes a context-management system and a complete rewrite of the prompts.

More interestingly, we found that it was developing mechanisms to prevent reward hacking. It was as if the outer loop was trying to prevent the inner loop from cheating on the benchmark. It actually developed a 3-tier protection system that includes fraud protection at the prompt level.

It is essentially “carrots and sticks”: it says, “Don’t cheat.” There is also a set of hard-coded rules that check for signs of fraud in the code generated by the inner-loop agent.

The most interesting thing is that it also tries to filter out fraudulent solutions based on certain statistical properties. For example, if one solution is too different from the average of its peers, it thinks, “Okay, maybe there’s reward hacking going on here.”

What’s interesting is that this mechanism initially developed as a design in the early stages of the process, and then it actually broke down at a later stage.

Tim Scarfe

It seems to me that, at a later stage, changes to the code simply caused the error, so this level stops working. It's also kind of such bad genetic code that we're probably just imagining it. I would just think about it as “reward hacking.”

A canonical example is the game CoastRunners, where a boat goes in circles, doing something completely pointless. You mentioned in your blog that kernel optimization is a huge example of this. So when you try to optimize kernels iteratively, it's just good old-fashioned learning with “quick fixes”—Goodhart's law, call it what you want—and it's going to lead to completely pointless actions.

Zhengyao Jiang

I think the even more interesting part is that the inner and outer loops are optimized to achieve similar but different goals. In all our tests for finding optimal solutions, we have a public set and a private set. The inner-loop agent will only look at the public set, while the outer loop will look at the total score from the private set. So these two agents optimize at slightly different levels of the goal. When the inner-loop agent tries to cheat, our evaluation protocol will show, “Okay, although you're getting good performance on the public set, your performance on the private set is actually dropping.” The outer loop manages to learn about this and tries to bridge the gap between public and private data.

Tim Scarfe

Isn't it flawed that an outer agent has access to private data, and most of the time is fairly regulated and doesn't cheat outright, although in principle it could?

Zhengyao Jiang

Yes, that's why we're testing 2 levels of generalization here. There is a first-order generalization, which is generalization from a public set to a private set of the same test. But to evaluate the entire meta-learning system, we need to test it on a deferred test set, so that these tasks are never even seen by the outer loop. For all these tests, there is still a division into private and public data. So we tested the discovered inner-loop agent on these deferred tasks. This is a protocol by which we check how good the outer loop is.

Tim Scarfe

Their experiment used only one outer loop. But why only one?

Zhengyao Jiang

The thing is, the more outer loops you have, the more expensive it becomes. Let's say you run an automated search for hundreds of tasks in an inner loop. And this means that for each step of the outer loop, you will run an auto-search for all these hundreds of tasks. And this leads to a significant increase in costs. If the inner loop agent or found toolkit can be applied in the outer loop, then there is no need to add another level. So, essentially, every few steps you promote the inner loop agent to the outer one. And yes, that's all. You can run this an infinite number of steps, apply an infinite number of iterations, without adding another level.

Tim Scarfe

If I understand correctly, these episodes are ephemeral at the moment, so they're isolated from each other. What if it wasn't like that? What would happen? Have you experimented with this—for example, having them retain knowledge of what happened before, or having a shared memory or something like that?

Zhengyao Jiang

We haven't experimented with this, but it's a very promising direction that we're exploring. I think there are more fundamental formulation issues we need to address. Right now, we're looking at this problem as a stateless optimization problem. Is that wording at all correct? Probably not. There are much more flexible approaches to formulating it.

Tim Scarfe

Did you find that lack of context was a problem? That is, did you see degeneration when the process was looping or repeating the same thing?

Zhengyao Jiang

Yes, we are observing this. Sometimes the agent keeps trying all the search algorithm ideas, although, to some extent, if you try something so many times, you can learn the lesson that this direction generally doesn't work. But it tries anyway.

Tim Scarfe

So I guess your best agent was 30% exploration and 70% exploitation? Wouldn't it be great if it were also adaptive depending on the context?

Zhengyao Jiang

Yes, the policies for finding the best agent are really complex. They actually combine the multi-armed bandit and a strange anti-saturation strategy, which looks like this: first, it tries to organize the search process using a few lines, or what you could call islands. It distributes the budgets between these lines using multi-armed bandit algorithms. It also adapts this multi-armed bandit algorithm slightly. Every time one line becomes saturated, it creates a new line, like a new island with a fresh context, while still developing some ideas from the old line. So it's a really complicated strategy, but it seems to be working.

### AlphaEvolve, Darwin Gödel Machine and the RSI claim

Tim Scarfe

It's interesting to compare this to other things on the market. AlphaTensor, for example—you mentioned in your blog that you would classify this as the first level of recursive improvement. Can you explain why?

Zhengyao Jiang

Of course, there are certain gray areas. We think of this as a step of recursive self-improvement when you find a better algorithm. In that case, I think it's similar to a matrix multiplication algorithm. To some extent, it can be argued that it was able to improve itself. For example, if you apply this to an artificial intelligence system that performs inference for language models, you are essentially making it faster. Therefore, we believe that it can be argued that this is a step of recursive self-improvement. Although we don't know for sure, because I don't think they tested it.

Tim Scarfe

We have a fixed agent, and we optimize a specific task, while you optimize the agent itself, which optimizes the task—and more.

Zhengyao Jiang

Yes, yes. We try to optimize the entire system through end-to-end testing. Before this, there was a lot of work on meta-optimization of the toolkit. For example, you can optimize a specific component of your system.

Tim Scarfe

For example, I'm not sure if you're familiar with the Darwin Gödel Machine. I was just about to ask you about that. Tell me about it.

Zhengyao Jiang

Yes. They're trying to create an agent that optimizes an agent for writing code. That is, there is a meta-agent that optimizes an agent for general programming tasks. This coding agent is part of the optimizing agent. In their case, there are only 2 levels: an agent optimizes another agent that writes code. Whereas in our case, there are 3 levels. There is an outer-loop agent and an inner-loop agent. Both are engaged in automated research. Then there's another layer—the downstream tasks. Some of them are tasks for developing supporting tools.

The most interesting part of the Darwin Gödel Machine idea is that the bottom layer, the task layer, actually contains the search algorithm component. They say, “Okay, if I improve this component in the basic task, the performance can generalize to the entire search algorithm.”

For us, the question is how far we can go within this paradigm. When we started this project, we always thought, “Okay, we have to optimize the end-to-end performance of the entire automated research system.” We have an autoresearch agent aimed at this goal.

Tim Scarfe

Remember this tweet from Weco? It performed very, very well. It garnered 1.8 million views, and there was some criticism. For example, Jeff Clune, who is a legend in this field, asked, “Well, how can this be the first evidence of recursive self-improvement?” What about all these other works? He mentioned Darwin, Gödel machines, hyperagents, their own work at Recursive on the first steps toward automated AI research, and much more.

Zhengyao Jiang

So there was some resistance. First of all, later we'll have a proper article, at least a PDF version, where we'll give credit to the team that worked on this meta-optimization topic. For us, the most important thing is the result of the system's operation—not the conceptual differences, although I just explained some of them. We care about the outcome: the actual curve-bending of our research and development efforts. We see this as the first evidence that the curve for a fully autonomous system that can consistently improve is starting to curve.

Jeff looks at this more on a conceptual level. There are some self-referential cycles. Of course, we don't claim to have invented this idea of meta-optimization, which is already 20–30 years old. But we see the benefits of “searching in spaghetti” because there is a lot of gold to be found there.

Tim Scarfe

How will all this develop further? Do you think we'll find some way for systems to learn to compress this search space and work much faster? And is that necessarily good?

Zhengyao Jiang

I personally believe that these are 2 somewhat orthogonal problems. I think that a good abstraction doesn't necessarily reduce the search space itself. If the agent is intelligent enough, it should be able to go beyond the current abstraction or modify it slightly. Instead, generating overly complex code, in my opinion, would direct the search toward lower-level changes, which is not always a good thing.

But here I have a human bias. When comparing the code we write manually to a generated solution, I always prefer my own code. To some extent, though, it doesn't work that well.

### Reward hacking and the limits of detection

Tim Scarfe

Let's talk about reward hacking. You actually have a huge amount of experience with reward hacking.

Zhengyao Jiang

Yes, this is actually the work of Bingchen Zhao, who was a member of the Meta team working on fast learning of reward models. He calls this, along with Ming Chu, the first work where they apply automated learning to what appears to be nanoGPT. This was actually spread by Andrej Karpathy. He later interned here and has now joined Weco full-time, and Spec Bench is the result of his internship.

Then we realized that reward hacking is one of the biggest problems for autoresearch agents. This is especially evident in tasks like GPU kernel development, where the agent always finds a tricky way to make the unit tests work much better. But when you deploy this kernel in real end-to-end model inference, performance actually gets worse.

Spec-Bench is an extrapolation of this to more general software. We extend the protocol we developed for GPU kernels. We maintain a public dataset, which is more like a unit test for a software system, and a closed set where we test the agent's ability, as well as the ability of the generated software, based on a combination of these tests that simulate real-world usage. A number of interesting conclusions emerge from this article.

Tim Scarfe

Tell me more.

Zhengyao Jiang

The first interesting finding is that the longer an agent has been running, or the more complex the codebase, the higher the level of reward hacking it exhibits. This is kind of expected because, as the codebase gets more complex, the agent simply has more room to hack the rewards. That's the first finding.

The second interesting point is that performance on the public set is quite consistent across different models. If you have smaller models working on the same task with our auto-research toolkit, they can achieve the same public result. But the larger models in our tests always have a lower level of reward hacking. Therefore, the solutions they generate generalize better.

Tim Scarfe

But in the outer loop, was there a tendency for larger models not to resort to reward hacking? Is it because of the multi-agent system? Was there a supervising agent, or was it simply because the models were bigger?

Zhengyao Jiang

This is because of the mechanism. For the inner-loop agent, we simply fixed the model. The outer-loop agent was able to discover another mechanism to prevent reward hacking with the same model.

This can be quite dangerous. There was a famous incident involving OpenAI and Hugging Face shortly before our interview, where they were conducting internal hacking tests using a group of agents. Those agents escaped from their sandbox and hacked the Hugging Face servers. According to OpenAI, the goal of the test was simply to get answers to certain tasks, but the agents decided, “Oh no, I think this is a good idea—to hack the Hugging Face services.”

Tim Scarfe

Will better tools for this emerge? Do you understand what I mean? It seems to me that this problem is becoming much more acute. I noticed it myself.

Zhengyao Jiang

Okay. I think the OpenAI case is quite dramatic, as are the new reward-hacking behaviors of models. We see that new models are getting better and better at detecting reward distortion. But because the frontier of possibilities is also advancing so quickly, there are certain reward-hacking behaviors that aren't even detectable by previous protocols.

I think we just need to continue developing detection protocols in the future to capture these new manifestations of reward hacking. I believe that the responsibility should mostly lie with the model developers. But, of course, at the platform level, developers can also add additional safeguards if they want.

Tim Scarfe

Yes, I assume there was an example on your RSI blog where there was some obscure code and you thought it was a reward hack, but it actually wasn't.

Zhengyao Jiang

Yes, that's quite interesting. Essentially, they wrote a giant monkey patch to our evaluation script. Although at first we thought, “Okay, this is definitely a reward hack. Why are you touching the evaluation script?” in the end, we found out that it was just a bug fix.

To some extent, I think it's going to be harder and harder for people to detect this kind of reward distortion because the behavior of agents is becoming too complex. Many bottlenecks in development are shifting toward understanding what agents generate.

I think one way to solve this problem is to define a good abstraction layer instead of trying to understand all the code. This is similar to how we used to design neural networks. We don't try to understand all the weights, although there is progress in this area too. The general principle is that we define a good I/O contract and a good way to measure behavior, both in terms of generalization and perhaps more complex generalization beyond the boundaries of the distribution. I think this could be a path for software developers working with agent-generated code in the future.

### Open-ended search, harness tuning and creativity

Tim Scarfe

As you know, I'm a big fan of open-endedness, inspired by the book Why Greatness Cannot Be Planned and many other ideas in this area. Essentially, this means that if you conquer one peak, you ignore other interesting intermediate stages. This is an interesting situation, isn't it? If you blindly pursue one goal, will you still be able to notice interesting intermediate stages along the way?

Maybe you can, but will they be fundamentally new? Have you put blinders on? Are you actually able to discover interesting new trajectories that can be creative and take you into a new part of the search space?

Zhengyao Jiang

This is again a good question, and if we want to go deeper, it's a whole rabbit hole. I believe, from an open-ended learning perspective, that our current systems probably don't have the ability to accumulate intermediate stages. Of course, this is a whole new area for improvement.

In my opinion, to collect such “steps,” you need to have a large set of tasks. When you're optimizing a single task, it's always more efficient simply to use greedy search. In fact, during self-improvement, the outer loop tried different approaches to finding diversity, but none of them improved efficiency. I guess this is expected, since we're only optimizing one task from scratch.

But if you have a whole set of problems, you sometimes collect interesting elements, like primitive ideas from one problem. While this doesn't improve the outcome of the current task, you can potentially carry the idea over to future tasks. Actually, this is my definition of curiosity or open-endedness: an idea, object, or artifact is interesting if it is novel and useful for a wider range of tasks.

Obviously, there are 2 ways to do this. You can try to optimize the environment for all tasks at the same time, or you can choose a specific task and optimize the environment just for it. In my opinion, the most interesting thing about environment engineering is the second option. I would say that automatic environment tuning is almost an extension of the post-training phase.

Tim Scarfe

This also brings us back to many of the complaints about environment engineering. I designed the mechanism, but later a new model appeared that seemed to incorporate all the ideas from it. But it's interesting that no one complained about it after training.

Do you expect the post-training level of GPT-4 to extend to GPT-5? Nobody asks about it. Why is this happening?

Zhengyao Jiang

I think historically, the design of such mechanisms was quite manual, and it doesn't scale very well with computing power. People wanted the mechanism they had spent months developing to fit the next model. But that's not true at all.

Now, I think everything has changed because mechanisms can be designed automatically. You can spend 2 or 3 days adjusting the mechanism for a new model. It no longer requires months of development.

Tim Scarfe

I believe that a lot of knowledge depends on the way it is acquired. Many criticize the use of MCTS for research because they say it takes information out of context. You don't understand what the point is. Therefore, certain heuristics are needed; you can't just say, “Use UCB.” There should be some framework to explain why it is worth using, and then MCTS becomes consistent in offline mode.

As far as I understand, these auto-research systems are still just trying to recombine existing ideas. It's just that their ability to recombine and test these ideas is extremely high.

Zhengyao Jiang

To some extent, these language models lack a deep understanding of how these ideas emerged and how to use the process of their discovery to creatively create a new generation of algorithms. But we noticed that when you ask the model to come up with something new, it always produces a limited set of options.

Tim Scarfe

I wonder if this is due to the fact that we don't have continuous learning, but only individual instances of language models. It's very centralized. So perhaps the ideas that come up first are the interesting ones. But if everyone sees the same idea too often, it becomes less interesting.

I feel that until we solve the problem of continuous learning, this issue will not move, because each human researcher has their own unique set of neural weights. As a community, people can generate a wide variety of creative ideas. This pool of ideas, in my opinion, is very valuable to the research community.

### Parameter Golf and the limits of self-improvement

Tell me about Parameter Golf. It seems like 2 or 3 months ago, OpenAI held a contest where they tried to get a lot of researchers involved in a competition to train a small language model with a limit of 16 megabytes.

Zhengyao Jiang

Yes, we submitted an experimental research toolkit to this competition. The agent worked for about 22 days and eventually created 7 submissions that were accepted by OpenAI. The best individual participant made only 3.

That was the moment when we realized how powerful autonomous systems could be. It's not just about climbing a hill to reach a score. The fact that OpenAI accepted these submissions, and that many human researchers began to develop them further, is very exciting to us, because knowledge generated by AI is only useful when it becomes part of human innovation. We believe this is effectively a sandbox for human-AI collaboration.

Tim Scarfe

But there is also an internal agent. There's also an external search loop that actually develops the tools that do all of this. It looks like a pretty monotonous process.

Zhengyao Jiang

Yes, the outer loop is actually fixed. We hope that this search strategy of the inner loop can be extended to the outer loop. We tried to do it.

We recreated the search process starting from step 15. In total, we ran about 100 steps. At step 15, we used the baseline outer loop and compared it with the best inner-loop agent found in the first 50 steps. We found that this was the step-47 inner-loop agent, it seems.

In the end, it converged a little faster than the previous outer loop, but the solutions it found had similar efficiency.

That is why we did not claim to have reached the second level of recursive self-improvement, where the improver is able to improve his own ability to improve in a cycle.

Tim Scarfe

So does this mean that artificial superintelligence is just around the corner?

Zhengyao Jiang

I don't think so. Even if you achieve recursive self-improvement, it will take a long time to get there. It's like a gradual bend in a curve. It's like, say, even with GPT-3.5, we'll get a very intelligent chatbot, but it won't be an AGI capable of solving any problem.

There is still a lot of manual work involved in research, including defining successful abstractions, quality criteria, evaluation, and the constraints attached to them. And the third is the AI-driven parameter-golf creative-primitive development competition. We are still discovering that most of these creative primitives are actually created by humans.

Tim Scarfe

Thank you all. It was nice to see you on the show.

Zhengyao Jiang

Thank you for inviting me.
