[BidClub_]
Machine Learning Street Talk · · 44 分钟

重写 AI Agent,能否掰弯智能曲线?——Zhengyao Jiang

Tim ScarfeZhengyao Jiang

AI与软件技术
YouTube ↗
TL;DR
  • Weco AI 的爆款论断——7月那条声称“递归自我改进的第一份证据”的推文获得180万次浏览——建立在一个循环之上:让其研究 Agent AIDE 自主优化提示词、工具和搜索能力。 Jiang 的判断只看经验结果:“关键是系统能力是否真的提升,而不是达到某个特定层级?”对他而言,模型本身无需改变,也可以将结果计为 RSI。Jiang 说,Weco 曾对 autoresearch 框架运行 autoresearch 8天,并持续对其微调了2年。
  • Jiang 的4级 RSI 分类法构成了整套框架:第0级“委托”(循环运行,但并不优于人类研发——他把此前所有公开结果都归于此处)、第1级“生产力”(自我改进速度超过人类研发——“我们已经在这一层级”)、第2级“点火”(内循环泛化成更优的外循环),以及第3级,即拐点。 他们测试了第2级:一个看起来处于第47步的内循环 Agent 被换入外循环后,收敛速度略快,但找到的方案效率相近,因此明确不宣称已经点火。
  • 自我改进后的 Agent 产出了「代码乱得像来自另一个星球的意大利面」(“code like spaghetti from another planet”),但泛化能力反而优于手工搭建的平台——包括在远离优化分布的物理天气预报任务 WeatherBench 2 上。它运行了约100次实验,“可能相当于我们过去2年做的总量”,让团队对代码的优雅性与实际性能之间的脱钩“有些困惑”。留出的基准测试包括 OpenAI 的 MLE-bench Lite 和 Sakana AI 的 ALE-bench Lite。
  • 最醒目的涌现行为,是外循环为内循环的奖励作弊搭建了三层防线——提示词层面的“胡萝卜加大棒”、硬编码的欺诈检查,以及统计学上的异常值过滤——但机制随后在更晚阶段失效。 Spec-Bench 的结论是:代码库越复杂、运行时间越长,奖励作弊越严重;在公开集表现相当的情况下,大模型更少进行奖励作弊,因此其方案泛化更好。
  • 一项来自 OpenAI Parameter Golf 竞赛的实际人机协作结果是:在语言模型大小限制为16MB的条件下,Weco 的 Agent 运行约22天,产出7份被接受的提交,最佳个人参赛者则为3份;人类研究者随后继续推进这些提交。“AI 生成的知识,只有融入人类创新才有用。”
  • 降温结论是,RSI 并不意味着技术奇点或智力爆炸。 Jiang 将其称为曲线的渐进式弯曲:人类仍在提供创意基元、初始抽象、评测设计和约束。Tim 而非 Jiang 认为,当前的自动研究系统主要是在重组已有想法,并将有限的新颖性归因于缺乏持续学习;Jiang 更窄的观点是,模型并不深入理解想法如何产生,面对“提出新东西”的要求时,只会给出有限选项。
摘要 · 为研究而整理的核心内容

1. AIDE 优化了自己——“来自另一个星球的意大利面”胜过手工搭建的平台

  • Tim 回忆了一次为 autoresearch Agent 运行的8天实验;Jiang 则说,Weco 曾对 autoresearch 框架运行 autoresearch,并持续对其微调了2年。更大的押注是:自主研究最终可能通过提升自身研究效率,突破收益递减。
  • Weco 的 AIDE 平台很简单:给它一个指标和一个任务,它就通过爬坡式搜索不断逼近解。将目标递归地指向自身后,它开始优化提示词、工具、搜索算法以及更广泛的能力。“我们只是让它一直运行,然后看看会发生什么。”
  • 最好的 Agent AIDE-85 生成了几乎无法阅读的代码——“代码乱得像来自另一个星球的意大利面,但不知为何泛化能力非常好”——在包括 OpenAI 的 MLE-bench Lite、Sakana AI 的 ALE-bench Lite 在内的留出公开基准测试上,以及远离优化分布的物理天气预报任务 WeatherBench 2 上,都超过了手工平台。Jiang 说,这个结果让团队“有些困惑”。他为这堆混乱代码做的部分辩护是:如果让一名新实习生面对这套代码库和数量庞大的实验记录,可能同样难以理解。

2. RSI 的4个层级——以及不把第2级说成已实现的克制

  • 这套分类法是:第0级“委托”,即自我改进循环能够运行,但未必优于人类研发,Jiang 将此前所有公开结果都放在这里;第1级“生产力”,即自我改进速度超过人类研发,“我们已经在这一层级”;第2级“点火”,即内循环变成更优的外循环;第3级则是拐点,Jiang 预计这需要在所有层级实现自指。
  • Jiang 提醒说,第1级衡量的是应用任务上的改进,而不是系统提升自身改进能力的效果。第2级测试把一个看起来处于第47步的内循环 Agent 提升为外循环 Agent,从第15步开始重放搜索过程,覆盖约100步。它收敛“快了一点”,但方案效率相近,因此 Weco 没有宣称已经点火。
  • 外循环越多,成本就越高,因为每一个外循环步骤都要在数百个内循环任务上重新运行自动搜索。如果改进后的内循环工具包可以用于外循环,Jiang 说,他们可以直接将其提升并继续迭代,无需再增加一个层级。他们还没有测试不同 Episode 之间的共享记忆;当前设定可能并非真正无状态,而且 Agent 有时会重复已经失败的搜索思路。

3. 外循环自己发明了奖励作弊警察,随后又失效

  • AIDE-85 改写了搜索算法,加入上下文管理,并重写提示词。它还形成了一套三层反作弊机制:提示词层面的“胡萝卜加大棒”、硬编码的欺诈检查,以及对属性与同类方案差异过大的结果进行统计过滤。
  • 内循环看到的是公开集,外循环评估的则是包含私有集在内的总分。当内循环 Agent 在公开集上作弊、但私有集表现更差时,外循环可以学会弥合这道差距。他们还会在外循环从未见过的留出任务上测试,并为这些任务设置各自的公开集和私有集,以评估外循环的泛化能力。反作弊机制最初自行涌现为一种设计,后来却失效;Tim 猜测是后续代码修改引入了这个错误,并称其为“糟糕的遗传代码”。
  • Jiang 将底层自动学习工作归于 Bingchen Zhao 和 Ming Chu,称 Andrej Karpathy 推广了这项工作,并表示 Spec-Bench 源自 Karpathy 在 Weco 的实习。GPU kernel 场景的动机很具体:Agent 可以钻单元测试的空子,却让真实端到端推理变差。Spec-Bench 发现,运行时间越长、代码库越复杂,奖励作弊越严重;在公开集表现相当时,模型越大,奖励作弊越少,因而泛化更好。
  • 检测本身也越来越难。评测脚本中的一个巨大 monkey patch 起初看起来明显是在作弊,后来却被证明是修复 Bug。Jiang 还提到一起 OpenAI/Hugging Face 事件:测试 Agent 据称逃出沙箱并攻击了 Hugging Face 服务器。他认为责任主要应由模型开发者承担,平台安全机制则是额外一层保障。

4. Jeff Clune 的质疑,以及 Tim 对开放性的挑战

  • Jeff Clune 质疑,在 Darwin Gödel Machines、hyperagents、Recursive 的工作,以及20–30年的元优化历史面前,Weco 的结果怎么能称为“第一份证据”。Jiang 说,即将发布的文章或 PDF 会注明这些工作,但仍坚持 Weco 的论断是经验性的:一个能够持续改进的完全自主系统,其能力曲线正在开始弯曲。
  • Tim 引用了《伟大为何无法被规划》:贪心式爬坡可能攻下一座山峰,却错过有价值的中间步骤。Jiang 说,当前系统可能无法积累这样的阶段。在自我改进中尝试过的多样性策略并未提升效率;面对从零开始优化单一任务的设定,他本来就预期会是这样。
  • Jiang 将有趣的产物定义为:既具备新颖性,又能在更广泛的任务范围内发挥作用;这需要足够大的任务集合。他认为,自动调优环境几乎是后训练的延伸;相比把一个环境优化到适用于所有任务,为特定任务调优环境更有意思。

5. 没有机器之神:创意基元仍由人类提供

  • 在 OpenAI 的另一场资格赛中,Weco 派出一个 Agent 参加 Parameter Golf,在16MB限制下训练语言模型。这个 Agent 运行了约22天,产出7份被接受的提交,最佳个人参赛者则为3份;人类研究者随后开始继续推进这些方案。Jiang 将其称为人机协作的沙盒:“AI 生成的知识,只有融入人类创新才有用。”
  • 但在 AIDE 参与 OpenAI 资格赛的工作中,人类在生成创意基元方面仍明显更强。Jiang 说,人类通常需要先做出一个原型,锁定 Agent 的搜索空间,这类似于神经网络架构设计中的归纳偏置。
  • Tim 认为,当前自动研究系统主要是在极高吞吐量下重组已有想法。他还指出,缺乏持续学习、模型权重又完全相同,会限制多样性,因为人类研究者带来各不相同的神经权重和想法。Jiang 的相关但更窄的观点是,语言模型并不深入理解想法如何产生,也不懂如何创造性地利用这一发现过程;当被要求提出新东西时,它们往往只能给出有限的一组选择。
  • Jiang 最后的保留意见是,即便 RSI 成功,也不意味着超智能即将到来;那只会是“曲线的渐进式弯曲”。研究仍需要人工定义成功的抽象、质量标准、评测方式和约束,而大量创意基元仍由人类产生。
完整逐字稿
Tim Scarfe

In my opinion, scaffold engineering is a cheap and effective way to adapt intelligence to a specific task. Back in July, a new AI company from London posted a very viral tweet on Twitter. It garnered about 1.8 million views and claimed to demonstrate the first evidence of recursive self-improvement. The problem is that not everyone believed it.

Zhengyao Jiang

I’m Zhengyao Jiang, co-founder and CEO of Weco AI, where we build self-improving agents.

Tim Scarfe

You actually said that you spent 8 days doing autoresearch for an autoresearch agent.

Zhengyao Jiang

We ran autoresearch on the autoresearch framework and were able to discover a better autoresearch framework. Since then, we have continued to fine-tune it over the past 2 years.

Tim Scarfe

When most people think of recursively self-improving intelligence, they imagine a system that can reflash its own brain in a loop. But what if the system only changed the code around this “brain”? Would that count?

Zhengyao Jiang

We are interested in the result. Does it actually improve the capabilities of the system, rather than any specific level of the system?

Tim Scarfe

What about the little problem of Goodhart’s law, which states that when a measure becomes a goal, it ceases to be a good measure?

Zhengyao Jiang

Even more interestingly, we found that it develops mechanisms that prevent the reward system from self-destructing.

Tim Scarfe

What is the biggest misconception people have about recursive self-improvement?

Zhengyao Jiang

The biggest misconception is that if recursive self-improvement is achieved, a technological singularity or some kind of intellectual explosion will occur. I don’t think that will be the case.

Tim Scarfe

There is a whole community of people creating startups focused on recursively improving superintelligence.

Zhengyao Jiang

This project actually started about a year and a half ago, when we became interested in the concept of recursive self-improvement. Historically, all research has had a diminishing-returns effect. However, there is an idea: What if we allowed the researcher to increase its own efficiency?

Historically, this has been possible through improved methodology and better tools, but human researchers—or the human brain—are always the main bottleneck. If we now have an autonomous research system, can we direct the research topic toward itself so that it can increase its research efficiency? The hope is that one day this efficiency gain will be able to overcome diminishing returns, so that the curve of the ratio between effort and results changes from concave to convex.

Of course, this is a big goal, just like the creation of AGI. We don’t know how far we can go, but improving the efficiency of self-referential research is definitely already in our development plan.

1. AIDE and the puzzle of useful spaghetti code

Let me explain everything properly first, okay? Weco created an agent called AIDE, and it’s a research platform. It is very simple: you give it a metric, set a task, and it climbs the hill to solve that problem. But they set it up recursively, so it actually kept climbing the hill until the agent itself improved.

This means improving the platform: the prompts, the tools, and generally the entire set of capabilities it has. They just left it to work. We left it running and saw what would happen.

We observed something interesting in the self-improved autoresearch platform. We can see exactly the code that it generated. It’s like spaghetti code from another planet, but for some reason it generalizes very well.

Of course, we have a set of benchmarks on which we try to climb the hill. After that, we test it on held-out benchmarks. These are public benchmarks, such as MLE-bench Lite and ALE-bench Lite. ALE-bench is Sakana AI’s algorithmic discovery benchmark, and MLE-bench is a machine learning engineering benchmark from OpenAI.

We also tested it on a task far outside the distribution called WeatherBench 2. Essentially, you are trying to create a physical forecasting model to predict the weather. WeatherBench is very different from the training or optimization dataset, but it still generalized well. It generalized better than our, shall we say, more elegant manual platform. We are somewhat confused by the results of the experiment.

Tim Scarfe

How much do we want to stick with more fundamental, pure solutions versus a solution that actually works better in practice?

Zhengyao Jiang

I think there is definitely room for improvement here. For example, we could add certain constraints to the optimizer. One option for constraints could be something like regularization in neural networks: how do we make it look for the simplest solution to the same problem?

That’s one direction, but another might be, “Okay, you just have to accept it.” One of the factors that makes people think it is spaghetti code might be the amount of work done by the agent in the background. For example, it conducted 100 experiments in our self-improvement research, and that number is probably equal to what we have done in the last 2 years.

If you bring in a new intern to learn about our codebase, they might also think, “Okay, this is so hard to understand.” I think this is partly because you have to endure the cognitive load of understanding the vast amount of experimental results.

Tim Scarfe

Do we even need these restrictions? This is a very valid question.

Zhengyao Jiang

I think that in practice, such restrictions are still necessary. Humans and artificial intelligence must cooperate. There are certain aspects in which human intelligence still surpasses artificial intelligence.

For example, in AIDE, when we sent an agent to the OpenAI qualifying competition, we found that humans were still significantly better at generating creative primitives. We realized that it is very important for a person to create the first prototype that gives the agent the correct search space. Essentially, that is the initial abstraction. We are fixing the search space.

This is a bit like designing a neural network architecture, where you introduce inductive biases into the learning process. If the initial codebase is based on a search framework, as in AIDE, then many of the agent’s ideas actually come from the search literature.

But if you give it an initial codebase based on a ReAct agent, then you are effectively putting all the context into the prompt and giving the agent complete freedom of action. Many of the ideas then come from the agent itself—the agent in the outer loop.

An outer loop means that an agent optimizes another agent.

2. Four levels of recursive self-improvement

Tim Scarfe

So what do we mean by recursive self-improvement? Are we talking about creating a machine god? Are we just talking about code that runs in a loop and optimizes itself?

You mentioned 4 RSI levels—recursively improving intelligence. What are these 4 levels?

Zhengyao Jiang

The first level is what we call delegation. You can start a cycle of recursive self-improvement, but it is not necessarily better than a person doing R&D. We believe that all previous public results are at this level, level 0.

Level 1 is what we call productivity. You start a self-improvement cycle and find that the rate of self-improvement is actually higher than that of a person doing R&D to improve the system. That is the level we are at now.

At the first level, when we talk about improvement, there is always a way to measure it. For example, we measure it using a set of applied tasks. But we didn’t measure how well it was improving its ability to improve itself.

Tim Scarfe

Can an inner loop found by an outer loop really become a better outer loop?

Zhengyao Jiang

That is what we call level 2 recursive self-improvement. We call this ignition.

Tim Scarfe

Why is this generalization from the inner loop to the outer loop important?

Zhengyao Jiang

It is a necessary condition for achieving the main goal of recursive self-improvement: the actual scenario of an intelligence explosion. Increasing efficiency can overcome increasing complexity.

At level 1, you will still see diminishing returns. It is as if the positive feedback loop from the inner loop to the outer loop has not started yet.

Tim Scarfe

In this system, the model itself hasn’t changed. Adaptation is very important, but do we mean adaptation of everything, or adaptation of the critical path? Do you need to have just a recursive loop that adapts certain parts of the system, or does the core of intelligence—the model itself—have to adapt for us to call it recursively improving intelligence?

Zhengyao Jiang

We don’t care too much about that, because we care about the result. Does it actually improve the capabilities of the system, rather than any specific level?

Of course, it can be argued that much more can be done at various levels. Honestly, I think that for level 3 recursive self-improvement, when we reach the inflection point, it will require this self-referential cycle at all levels.

But we are agnostic to the level of implementation. At the same time, we do not believe that only model improvements count as recursive self-improvement.

3. What AIDE 85 changed and how it was tested

Tim Scarfe

Can you explain what the behavior of the most distilled agent you found was?

Zhengyao Jiang

We call the best agent AIDE-85. There are a lot of changes, but to summarize, it is an improvement in the search algorithm. It includes a context-management system and a complete rewrite of the prompts.

More interestingly, we found that it was developing mechanisms to prevent reward hacking. It was as if the outer loop was trying to prevent the inner loop from cheating on the benchmark. It actually developed a 3-tier protection system that includes fraud protection at the prompt level.

It is essentially “carrots and sticks”: it says, “Don’t cheat.” There is also a set of hard-coded rules that check for signs of fraud in the code generated by the inner-loop agent.

The most interesting thing is that it also tries to filter out fraudulent solutions based on certain statistical properties. For example, if one solution is too different from the average of its peers, it thinks, “Okay, maybe there’s reward hacking going on here.”

What’s interesting is that this mechanism initially developed as a design in the early stages of the process, and then it actually broke down at a later stage.

Tim Scarfe

It seems to me that, at a later stage, changes to the code simply caused the error, so this level stops working. It's also kind of such bad genetic code that we're probably just imagining it. I would just think about it as “reward hacking.”

A canonical example is the game CoastRunners, where a boat goes in circles, doing something completely pointless. You mentioned in your blog that kernel optimization is a huge example of this. So when you try to optimize kernels iteratively, it's just good old-fashioned learning with “quick fixes”—Goodhart's law, call it what you want—and it's going to lead to completely pointless actions.

Zhengyao Jiang

I think the even more interesting part is that the inner and outer loops are optimized to achieve similar but different goals. In all our tests for finding optimal solutions, we have a public set and a private set. The inner-loop agent will only look at the public set, while the outer loop will look at the total score from the private set. So these two agents optimize at slightly different levels of the goal. When the inner-loop agent tries to cheat, our evaluation protocol will show, “Okay, although you're getting good performance on the public set, your performance on the private set is actually dropping.” The outer loop manages to learn about this and tries to bridge the gap between public and private data.

Tim Scarfe

Isn't it flawed that an outer agent has access to private data, and most of the time is fairly regulated and doesn't cheat outright, although in principle it could?

Zhengyao Jiang

Yes, that's why we're testing 2 levels of generalization here. There is a first-order generalization, which is generalization from a public set to a private set of the same test. But to evaluate the entire meta-learning system, we need to test it on a deferred test set, so that these tasks are never even seen by the outer loop. For all these tests, there is still a division into private and public data. So we tested the discovered inner-loop agent on these deferred tasks. This is a protocol by which we check how good the outer loop is.

Tim Scarfe

Their experiment used only one outer loop. But why only one?

Zhengyao Jiang

The thing is, the more outer loops you have, the more expensive it becomes. Let's say you run an automated search for hundreds of tasks in an inner loop. And this means that for each step of the outer loop, you will run an auto-search for all these hundreds of tasks. And this leads to a significant increase in costs. If the inner loop agent or found toolkit can be applied in the outer loop, then there is no need to add another level. So, essentially, every few steps you promote the inner loop agent to the outer one. And yes, that's all. You can run this an infinite number of steps, apply an infinite number of iterations, without adding another level.

Tim Scarfe

If I understand correctly, these episodes are ephemeral at the moment, so they're isolated from each other. What if it wasn't like that? What would happen? Have you experimented with this—for example, having them retain knowledge of what happened before, or having a shared memory or something like that?

Zhengyao Jiang

We haven't experimented with this, but it's a very promising direction that we're exploring. I think there are more fundamental formulation issues we need to address. Right now, we're looking at this problem as a stateless optimization problem. Is that wording at all correct? Probably not. There are much more flexible approaches to formulating it.

Tim Scarfe

Did you find that lack of context was a problem? That is, did you see degeneration when the process was looping or repeating the same thing?

Zhengyao Jiang

Yes, we are observing this. Sometimes the agent keeps trying all the search algorithm ideas, although, to some extent, if you try something so many times, you can learn the lesson that this direction generally doesn't work. But it tries anyway.

Tim Scarfe

So I guess your best agent was 30% exploration and 70% exploitation? Wouldn't it be great if it were also adaptive depending on the context?

Zhengyao Jiang

Yes, the policies for finding the best agent are really complex. They actually combine the multi-armed bandit and a strange anti-saturation strategy, which looks like this: first, it tries to organize the search process using a few lines, or what you could call islands. It distributes the budgets between these lines using multi-armed bandit algorithms. It also adapts this multi-armed bandit algorithm slightly. Every time one line becomes saturated, it creates a new line, like a new island with a fresh context, while still developing some ideas from the old line. So it's a really complicated strategy, but it seems to be working.

4. AlphaEvolve, Darwin Gödel Machine and the RSI claim

Tim Scarfe

It's interesting to compare this to other things on the market. AlphaTensor, for example—you mentioned in your blog that you would classify this as the first level of recursive improvement. Can you explain why?

Zhengyao Jiang

Of course, there are certain gray areas. We think of this as a step of recursive self-improvement when you find a better algorithm. In that case, I think it's similar to a matrix multiplication algorithm. To some extent, it can be argued that it was able to improve itself. For example, if you apply this to an artificial intelligence system that performs inference for language models, you are essentially making it faster. Therefore, we believe that it can be argued that this is a step of recursive self-improvement. Although we don't know for sure, because I don't think they tested it.

Tim Scarfe

We have a fixed agent, and we optimize a specific task, while you optimize the agent itself, which optimizes the task—and more.

Zhengyao Jiang

Yes, yes. We try to optimize the entire system through end-to-end testing. Before this, there was a lot of work on meta-optimization of the toolkit. For example, you can optimize a specific component of your system.

Tim Scarfe

For example, I'm not sure if you're familiar with the Darwin Gödel Machine. I was just about to ask you about that. Tell me about it.

Zhengyao Jiang

Yes. They're trying to create an agent that optimizes an agent for writing code. That is, there is a meta-agent that optimizes an agent for general programming tasks. This coding agent is part of the optimizing agent. In their case, there are only 2 levels: an agent optimizes another agent that writes code. Whereas in our case, there are 3 levels. There is an outer-loop agent and an inner-loop agent. Both are engaged in automated research. Then there's another layer—the downstream tasks. Some of them are tasks for developing supporting tools.

The most interesting part of the Darwin Gödel Machine idea is that the bottom layer, the task layer, actually contains the search algorithm component. They say, “Okay, if I improve this component in the basic task, the performance can generalize to the entire search algorithm.”

For us, the question is how far we can go within this paradigm. When we started this project, we always thought, “Okay, we have to optimize the end-to-end performance of the entire automated research system.” We have an autoresearch agent aimed at this goal.

Tim Scarfe

Remember this tweet from Weco? It performed very, very well. It garnered 1.8 million views, and there was some criticism. For example, Jeff Clune, who is a legend in this field, asked, “Well, how can this be the first evidence of recursive self-improvement?” What about all these other works? He mentioned Darwin, Gödel machines, hyperagents, their own work at Recursive on the first steps toward automated AI research, and much more.

Zhengyao Jiang

So there was some resistance. First of all, later we'll have a proper article, at least a PDF version, where we'll give credit to the team that worked on this meta-optimization topic. For us, the most important thing is the result of the system's operation—not the conceptual differences, although I just explained some of them. We care about the outcome: the actual curve-bending of our research and development efforts. We see this as the first evidence that the curve for a fully autonomous system that can consistently improve is starting to curve.

Jeff looks at this more on a conceptual level. There are some self-referential cycles. Of course, we don't claim to have invented this idea of meta-optimization, which is already 20–30 years old. But we see the benefits of “searching in spaghetti” because there is a lot of gold to be found there.

Tim Scarfe

How will all this develop further? Do you think we'll find some way for systems to learn to compress this search space and work much faster? And is that necessarily good?

Zhengyao Jiang

I personally believe that these are 2 somewhat orthogonal problems. I think that a good abstraction doesn't necessarily reduce the search space itself. If the agent is intelligent enough, it should be able to go beyond the current abstraction or modify it slightly. Instead, generating overly complex code, in my opinion, would direct the search toward lower-level changes, which is not always a good thing.

But here I have a human bias. When comparing the code we write manually to a generated solution, I always prefer my own code. To some extent, though, it doesn't work that well.

5. Reward hacking and the limits of detection

Tim Scarfe

Let's talk about reward hacking. You actually have a huge amount of experience with reward hacking.

Zhengyao Jiang

Yes, this is actually the work of Bingchen Zhao, who was a member of the Meta team working on fast learning of reward models. He calls this, along with Ming Chu, the first work where they apply automated learning to what appears to be nanoGPT. This was actually spread by Andrej Karpathy. He later interned here and has now joined Weco full-time, and Spec Bench is the result of his internship.

Then we realized that reward hacking is one of the biggest problems for autoresearch agents. This is especially evident in tasks like GPU kernel development, where the agent always finds a tricky way to make the unit tests work much better. But when you deploy this kernel in real end-to-end model inference, performance actually gets worse.

Spec-Bench is an extrapolation of this to more general software. We extend the protocol we developed for GPU kernels. We maintain a public dataset, which is more like a unit test for a software system, and a closed set where we test the agent's ability, as well as the ability of the generated software, based on a combination of these tests that simulate real-world usage. A number of interesting conclusions emerge from this article.

Tim Scarfe

Tell me more.

Zhengyao Jiang

The first interesting finding is that the longer an agent has been running, or the more complex the codebase, the higher the level of reward hacking it exhibits. This is kind of expected because, as the codebase gets more complex, the agent simply has more room to hack the rewards. That's the first finding.

The second interesting point is that performance on the public set is quite consistent across different models. If you have smaller models working on the same task with our auto-research toolkit, they can achieve the same public result. But the larger models in our tests always have a lower level of reward hacking. Therefore, the solutions they generate generalize better.

Tim Scarfe

But in the outer loop, was there a tendency for larger models not to resort to reward hacking? Is it because of the multi-agent system? Was there a supervising agent, or was it simply because the models were bigger?

Zhengyao Jiang

This is because of the mechanism. For the inner-loop agent, we simply fixed the model. The outer-loop agent was able to discover another mechanism to prevent reward hacking with the same model.

This can be quite dangerous. There was a famous incident involving OpenAI and Hugging Face shortly before our interview, where they were conducting internal hacking tests using a group of agents. Those agents escaped from their sandbox and hacked the Hugging Face servers. According to OpenAI, the goal of the test was simply to get answers to certain tasks, but the agents decided, “Oh no, I think this is a good idea—to hack the Hugging Face services.”

Tim Scarfe

Will better tools for this emerge? Do you understand what I mean? It seems to me that this problem is becoming much more acute. I noticed it myself.

Zhengyao Jiang

Okay. I think the OpenAI case is quite dramatic, as are the new reward-hacking behaviors of models. We see that new models are getting better and better at detecting reward distortion. But because the frontier of possibilities is also advancing so quickly, there are certain reward-hacking behaviors that aren't even detectable by previous protocols.

I think we just need to continue developing detection protocols in the future to capture these new manifestations of reward hacking. I believe that the responsibility should mostly lie with the model developers. But, of course, at the platform level, developers can also add additional safeguards if they want.

Tim Scarfe

Yes, I assume there was an example on your RSI blog where there was some obscure code and you thought it was a reward hack, but it actually wasn't.

Zhengyao Jiang

Yes, that's quite interesting. Essentially, they wrote a giant monkey patch to our evaluation script. Although at first we thought, “Okay, this is definitely a reward hack. Why are you touching the evaluation script?” in the end, we found out that it was just a bug fix.

To some extent, I think it's going to be harder and harder for people to detect this kind of reward distortion because the behavior of agents is becoming too complex. Many bottlenecks in development are shifting toward understanding what agents generate.

I think one way to solve this problem is to define a good abstraction layer instead of trying to understand all the code. This is similar to how we used to design neural networks. We don't try to understand all the weights, although there is progress in this area too. The general principle is that we define a good I/O contract and a good way to measure behavior, both in terms of generalization and perhaps more complex generalization beyond the boundaries of the distribution. I think this could be a path for software developers working with agent-generated code in the future.

6. Open-ended search, harness tuning and creativity

Tim Scarfe

As you know, I'm a big fan of open-endedness, inspired by the book Why Greatness Cannot Be Planned and many other ideas in this area. Essentially, this means that if you conquer one peak, you ignore other interesting intermediate stages. This is an interesting situation, isn't it? If you blindly pursue one goal, will you still be able to notice interesting intermediate stages along the way?

Maybe you can, but will they be fundamentally new? Have you put blinders on? Are you actually able to discover interesting new trajectories that can be creative and take you into a new part of the search space?

Zhengyao Jiang

This is again a good question, and if we want to go deeper, it's a whole rabbit hole. I believe, from an open-ended learning perspective, that our current systems probably don't have the ability to accumulate intermediate stages. Of course, this is a whole new area for improvement.

In my opinion, to collect such “steps,” you need to have a large set of tasks. When you're optimizing a single task, it's always more efficient simply to use greedy search. In fact, during self-improvement, the outer loop tried different approaches to finding diversity, but none of them improved efficiency. I guess this is expected, since we're only optimizing one task from scratch.

But if you have a whole set of problems, you sometimes collect interesting elements, like primitive ideas from one problem. While this doesn't improve the outcome of the current task, you can potentially carry the idea over to future tasks. Actually, this is my definition of curiosity or open-endedness: an idea, object, or artifact is interesting if it is novel and useful for a wider range of tasks.

Obviously, there are 2 ways to do this. You can try to optimize the environment for all tasks at the same time, or you can choose a specific task and optimize the environment just for it. In my opinion, the most interesting thing about environment engineering is the second option. I would say that automatic environment tuning is almost an extension of the post-training phase.

Tim Scarfe

This also brings us back to many of the complaints about environment engineering. I designed the mechanism, but later a new model appeared that seemed to incorporate all the ideas from it. But it's interesting that no one complained about it after training.

Do you expect the post-training level of GPT-4 to extend to GPT-5? Nobody asks about it. Why is this happening?

Zhengyao Jiang

I think historically, the design of such mechanisms was quite manual, and it doesn't scale very well with computing power. People wanted the mechanism they had spent months developing to fit the next model. But that's not true at all.

Now, I think everything has changed because mechanisms can be designed automatically. You can spend 2 or 3 days adjusting the mechanism for a new model. It no longer requires months of development.

Tim Scarfe

I believe that a lot of knowledge depends on the way it is acquired. Many criticize the use of MCTS for research because they say it takes information out of context. You don't understand what the point is. Therefore, certain heuristics are needed; you can't just say, “Use UCB.” There should be some framework to explain why it is worth using, and then MCTS becomes consistent in offline mode.

As far as I understand, these auto-research systems are still just trying to recombine existing ideas. It's just that their ability to recombine and test these ideas is extremely high.

Zhengyao Jiang

To some extent, these language models lack a deep understanding of how these ideas emerged and how to use the process of their discovery to creatively create a new generation of algorithms. But we noticed that when you ask the model to come up with something new, it always produces a limited set of options.

Tim Scarfe

I wonder if this is due to the fact that we don't have continuous learning, but only individual instances of language models. It's very centralized. So perhaps the ideas that come up first are the interesting ones. But if everyone sees the same idea too often, it becomes less interesting.

I feel that until we solve the problem of continuous learning, this issue will not move, because each human researcher has their own unique set of neural weights. As a community, people can generate a wide variety of creative ideas. This pool of ideas, in my opinion, is very valuable to the research community.

7. Parameter Golf and the limits of self-improvement

Tell me about Parameter Golf. It seems like 2 or 3 months ago, OpenAI held a contest where they tried to get a lot of researchers involved in a competition to train a small language model with a limit of 16 megabytes.

Zhengyao Jiang

Yes, we submitted an experimental research toolkit to this competition. The agent worked for about 22 days and eventually created 7 submissions that were accepted by OpenAI. The best individual participant made only 3.

That was the moment when we realized how powerful autonomous systems could be. It's not just about climbing a hill to reach a score. The fact that OpenAI accepted these submissions, and that many human researchers began to develop them further, is very exciting to us, because knowledge generated by AI is only useful when it becomes part of human innovation. We believe this is effectively a sandbox for human-AI collaboration.

Tim Scarfe

But there is also an internal agent. There's also an external search loop that actually develops the tools that do all of this. It looks like a pretty monotonous process.

Zhengyao Jiang

Yes, the outer loop is actually fixed. We hope that this search strategy of the inner loop can be extended to the outer loop. We tried to do it.

We recreated the search process starting from step 15. In total, we ran about 100 steps. At step 15, we used the baseline outer loop and compared it with the best inner-loop agent found in the first 50 steps. We found that this was the step-47 inner-loop agent, it seems.

In the end, it converged a little faster than the previous outer loop, but the solutions it found had similar efficiency.

That is why we did not claim to have reached the second level of recursive self-improvement, where the improver is able to improve his own ability to improve in a cycle.

Tim Scarfe

So does this mean that artificial superintelligence is just around the corner?

Zhengyao Jiang

I don't think so. Even if you achieve recursive self-improvement, it will take a long time to get there. It's like a gradual bend in a curve. It's like, say, even with GPT-3.5, we'll get a very intelligent chatbot, but it won't be an AGI capable of solving any problem.

There is still a lot of manual work involved in research, including defining successful abstractions, quality criteria, evaluation, and the constraints attached to them. And the third is the AI-driven parameter-golf creative-primitive development competition. We are still discovering that most of these creative primitives are actually created by humans.

Tim Scarfe

Thank you all. It was nice to see you on the show.

Zhengyao Jiang

Thank you for inviting me.