[RL/推理现状] IMO/IOI 金牌、OpenAI o3/GPT-5 与 Cursor Composer——Ashvin Nair,Cursor
- Ashvin Nair 的核心判断是,今天的 RL 是一种“峰值很高但很窄”的工具:它可以“彻底杀穿训练分布”,却很难在分布外实现强泛化。 因此,近期价值池属于那些能把工作者的真实上下文纳入分布的产品——代码库、终端、文档、Slack 和积累下来的实验——并围绕这套工作流共同设计模型;瓶颈可能并不在模型本身的能力上。
- 奥赛金牌已经无法直接映射到经济自动化:Ashvin 曾以为 IMO 金牌意味着“大家都可以直接去度假了”,但“生活还是老样子”。 他与主持人将这种错位归因于基准选择和社区层面的过拟合:模型可以从普通开发者能力跃升到顶尖竞赛表现,却仍然缺少处理日常工作所需的上下文。
- 尽管叠衣服演示令人印象深刻,机器人仍比软件 agent 更早期,也更难变现。 主持人将机器人类比为“GPT-1 到 GPT-2 时代”;Ashvin 强调其中已经出现泛化的迹象,但认为机器人更像是对一个团队的投资,而非已经验证的技术。他预计,在 AI 机器人甚至可能还没达到100亿美元市场规模之前,LLM agent 就会成为万亿美元市场,因为实体系统仍需证明其用途、可靠性、维护经济性和泛化能力。
- 当预训练模型跨过足够的能力门槛后,推理 RL 以一条平滑的内部扩展曲线出现,而不是靠某个奇迹般的单次发布。 2023年的一个小型原型产生了异常准确的推理轨迹,并在小模型上拿到出人意料的数学成绩——否则就需要更多预训练才能实现的表现;到2024年初,Ashvin 认为这套配方拿下 IMO 和 IOI 已经可以预期,而公众看到的“跃升”掩盖了层层叠加的实验和逐月积累的提升。
- 在 OpenAI,内部领先公开发布的时间已经从约6个月缩短到可能只有1到2个月,而各家实验室也正在收敛到“差不多”的 RL 配方。 Ashvin 没有从 DeepSeek 身上看到重大的内部教训——OpenAI 当时已经有更好的模型,更聪明的模型依然有价值——但他注意到,连 Anthropic 的 Opus 4.5 在 ARC-AGI-2 上的曲线都与 OpenAI 的表现相似。
- Cursor 的切入点是紧密的产品—模型协同设计,代表性做法是大约每2小时更新一次 Online Tab 策略,以及由20—25人组成的 Composer 机器学习团队。 Composer 已经“足够聪明”而可用,同时又足够快,不会迫使用户频繁切换上下文;更大的目标,是自动化完整的软件工程闭环——写代码、检查 Datadog、形成假设、重新运行并持续学习。
- 最可能带来断层式变化的,不是又一次静态基准胜利,而是具备类似人类数据效率的持续学习。 模型即使在同一上下文中也会重复犯错,而人只要看到别人碰一次热炉子就能学会。Ashvin 认为这个方向极其有意思;主持人猜测它可能在一年内改变范式,但 Ashvin 明确表示自己“完全不知道”这种变化会是什么。
- 无论 AGI 是两年后还是10年后到来,治理问题都仍未解决。 在 OpenAI 的“短暂危机”期间,Ashvin 签了员工联名信,但如果能展开真正的治理讨论,他愿意把股权放到一边;他想知道,由公众公司广泛持股是否可能比7名非营利机构董事更民主,同时也承认资本主义已经无法妥善处理社交媒体和不健康食品问题。
1. 软件 agent 应远早于机器人实现变现
Ashvin 从 Berkeley 的机器人研究走到 OpenAI、再到 Cursor,听起来跨度很大,实际并没有那么断裂:这两个领域都要求研究者检查大量杂乱数据,并在系统拒绝工作时坚持下去。机器人研究会塑造出“特别能扛事的人”,因为与仿真不同,现实世界没有退路,只能正面面对。
Ashvin 没亲眼看过周日展示的机器人,但觉得报道中的演示“挺酷”;主持人则亲眼看过 Physical Intelligence 的机器人在一间普通客厅里叠衣服。Ashvin 认为,真正的拐点可能与 GPT-2 所代表的那类泛化能力有关,但也强调细节很重要:机器人目前仍更像是押注一个团队,而不是已经验证的技术。
他的市场判断非常鲜明:“LLM agent 将成为万亿美元市场,而机器人可能连100亿美元市场都还没到”(“LLM agents are going to be like a trillion-dollar market before robotics is maybe even like a $10 billion market”)。Agent 已经在创造价值;机器人则仍需证明能完成有用任务、硬件足够可靠、维修可行,并具备成立的单位经济性。
2. 奥赛金牌暴露出基准领先几乎不保证什么
Ashvin 曾把 IMO 金牌视为终点:“我原本会直接认为,大家都可以去度假了——AI 已经解决了”(“AI is solved”)。结果真的到来后,普通生活却几乎没有变化。国际象棋和围棋也带来过同样的意外,但每次出现这种错位,冲击感依然很强。
主持人的解释是,人们会不断移动 AGI 的终点线;Ashvin 部分认同这种做法,因为整个社区都会集体优化它选择的任何基准。主持人将其与语言模型的表现并置:模型在日常工作中看起来大致处于初级到高级开发者水平,却可以赢下顶级编程竞赛;Ashvin 认为这种错位很可疑。
他在2017—22年间亲历的 RL 领域就是一个警示。学术研究看起来通过离策略学习、价值函数和新的数学工具不断推进,但 RL 寒冬到来后,一些建立在这些前提上的创业公司最终放弃了。回头看,研究者创造出大量可调旋钮,并在社区层面默认把它们调向共同的基准。
学术界又放大了这个问题:相比“真正有效的简单想法”,它更奖励数学上有趣的新颖性。真正奏效的方法往往更简单、旋钮更少,也更依赖算力,但能包装成论文里的“秘方”更少——这套激励机制并不利于发现哪些方法能够泛化。
3. 当预训练跨过正确门槛,推理 RL 才真正奏效
Ashvin 将这条路线归功于 OpenAI 长期积累的传统,包括 Ilya Sutskever 和 Jakub Pachocki 在内的一批人“骨子里就有 AGI”。Dota 已经包含了这个模板:单纯复制互联网内容最终会触及平台期,而 RL 可以生成更强的智能。RLHF 只是一个有限的分支,因为人类反馈无法承接相当规模的算力。
转折点出现在2023年前后。即使在小模型上运行 RL,也能产生异常准确的推理轨迹和出人意料的强劲数学成绩——否则就需要更多预训练才能达到的表现。这个原型就像 GPT-3 之前的 GPT-2:在原始性能让机会变得显而易见之前,必须先有人从第一性原理出发相信它。
在 OpenAI 内部,进展随后呈现连续状态:实验带来增量提升,不确定的想法层层叠加,资源也随之转向这条新路线。到2024年初,Ashvin 认为这套配方能“真正横扫” IMO 或 IOI 已经可以预期,尽管外部观察者将每次发布体验成突然出现的突破。
主持人提到,推理方向周围大约有300人,而最初 o1 视频里只有十几张熟悉面孔。Ashvin 不认同这种算法:早期贡献者可能有50—200人;随着 o3 变成产品,安全、评测和其他职能都扩大了参与范围,直到他“已经数不清人数”。
4. 工作自动化要求把完整工作流纳入分布
Scaling 还没有走到尽头,但 Ashvin 认为它的形态已经改变。当前 RL 确实能实现某种程度、且颇有意思的泛化,但仍然“峰值很高却很窄”:不需要太大代价就能在训练分布内占据压倒性优势,却无法足够广泛地迁移到工作自动化。
因此,关键要求是把具备经济价值的任务纳入训练分布。GDPval 大致符合这个方向:覆盖128项任务,来自合计占 GDP 超过5%的白领工作,并尽量贴近原始文档,而不是使用经过清洗的模型输入。Ashvin 还没有仔细研究其轨迹,无法判断会计工作究竟包含什么、需要哪些上下文;这恰恰说明,产品必须呈现真实工作流,而不只是提供一个评测。
Ashvin 在 OpenAI 的研究工作说明了缺失的上下文是什么。他写的代码相对较少,却花了一年时间跑 sweep、研究超参数之间的相互作用,并从大量图表中积累知识,而这些知识大多“只存在我的脑子里”。编码模型可以写出脚本,却无法在没有这套累积上下文的情况下复现这份工作。
主持人把 GPT-5 和 GPT-5 Codex 两条产品线解读为“一套模型适配所有场景”正在消亡的证据。Ashvin 的反驳是:“OpenAI 有调整组织架构的倾向。”专业化可能反映的是谁掌握数据,而不是模型能力的边界;如果拥有所有相关数据,联合训练仍可能带来有价值的跨领域泛化。
5. 前沿竞争同时压缩预测和发布窗口
在 Curve 大会上,预测者预计到2027年前后,Epoch AI 的 FrontierMath 和 Humanity’s Last Exam 得分只有10%—20%;而 Ashvin 当时已经见过超过这些预期的内部模型。同一批人中,有些还在考虑2035年建造戴森球:可能“短期过度悲观、长期过度乐观”。
不过,他仍然尊重这个群体,因为他们会提前登记预测,而不是事后宣称“我一直都看到了”。他们较早对能力的预测,在方向上强于当时认为 AI 是骗局的主流观点;Ashvin 则把人类水平的智能放在“2030年前后”。
主持人说,OpenAI 的内部模型一度领先公开发布约6个月;Ashvin 估计当前领先时间大约只有1到2个月。主持人以 Nano Banana Pro 为例,说明领先优势可以多快转化为差异;Ashvin 则强调,如今的发布窗口已经“极其短暂”。
DeepSeek 给 Ashvin 带来的意外更多在市场影响,而非技术信息:他认为 DeepSeek 说明 NVIDIA 芯片比此前想象的更有用,但 NVIDIA 股价却下跌了。OpenAI 当时已经有更好的模型,各家实验室很快也收敛到相似的 RL 形态;Anthropic 的 Opus 4.5 甚至在 ARC-AGI-2 上呈现出与 OpenAI 相似的曲线。
6. OpenAI 危机留下的核心治理问题仍无答案
Ashvin 在感恩节期间和2位 OpenAI 朋友一起得知 Sam Altman 被解雇,最初以为这是个玩笑。经历了“疯狂的、充满上下起伏的周末”后,他和约95%的人一样签署了联名信,但他的理由远不只是简单的忠诚。
无论 AGI 是两年后、10年后还是更远的未来到来,他都认为治理很重要。危机期间,他准备“把股权抛到一边”,认真讨论治理结构;他也在想,一个类似 Microsoft 的董事会,其利益相关者可能通过养老金覆盖公众,是否会比由7个人运营一家非营利机构更民主。
当主持人后来问,OpenAI 的非营利机构是否拥有一个更好的“秘密影子董事会”来决定 AGI 时,Ashvin 无法回答。他坦率的结论是,社会“根本没有解决治理问题”。如果资本主义激励机制在食品和社交媒体领域都已经无法带来健康结果,我们就没有多少理由相信它能妥善治理 AGI。
7. Cursor 押注产品与模型的距离胜过实验室规模
主持人的质疑很直接:OpenAI 拥有近乎无限的资源、充足的数据和 Codex,为什么还要离开?Ashvin 的回答落在组织形态上。Cursor 提供了一个小而专注的环境,让产品和机器学习团队坐在一起,能够有意识地把产品的测试分布拉进 RL 训练。
Online Tab 是最典型的例子:Cursor 大约每2小时就能更新一次策略。Ashvin 不同意这只是因为自动补全使用了更小的模型、所以更容易迭代;真正起作用的是一个足够紧凑的组织,能快速连接用户行为、产品决策和训练过程。
Cursor 的机器学习团队只有20—25人,Ashvin 对 Composer 的质量感到惊喜。它的吸引力不只是智能,还在于延迟:它已经“足够聪明”,让用户愿意使用;同时又足够快,让用户能够留在闭环中。更慢但更聪明的模型会迫使用户频繁切换上下文,“多少有点像把人搞成 ADHD”。
目标并不止于回答提示词。Cursor 希望模型把软件工程当成一套流程来执行:写代码、检查 Datadog、诊断行为、形成假设、重新运行系统并持续迭代。内部工具允许研究人员通过 SSH 进入用户环境,贴近真实数据;在 Ashvin 看来,这正是让模型与产品保持在一起的持久优势。
8. 持续学习可能是下一个真正的范式转变
主持人质疑朴素的在线学习:如果把每个用户动作都吸收进去,模型可能会被带向平庸行为。Ashvin 用“碰热炉子”重新定义了问题。人类不只是过滤掉一个坏例子;人类大概拥有某种价值函数,能在观察一次后就避免重复犯错。
模型在这种数据效率上仍落后“几个数量级”。它们可能反复引入同一个代码 bug,甚至在同一上下文中也是如此;而人类通常应该只犯一次错,并把教训保留到其他上下文中。Ashvin 提到了“带无限记忆的上下文学习之类的东西”,也就是让上下文中的一次经历进入模型权重,使模型不再重复同一个错误。
他认为眼下并不存在明显的容量问题:部署过程中可能向一个原本用万亿级 token 训练的模型再增加几千、甚至几百万 token,“只是水桶里的一滴水”。更深层的不确定性在于,权重的主要作用究竟更像是存储事实的硬盘,还是能够通过可复用电路计算事实的 CPU。
主持人猜测,持续学习可能“在下一年左右改变范式”;Ashvin 认同其中确实有些极其有意思的东西,但表示自己完全不知道这种变化会是什么。Cursor 正在重点招聘代码数据和奖励方向的人才,并偏好2天的工作试做,而不是 trivia 式问答;不过,“为什么 off-policy RL 不稳定?”仍然是他最能看出候选人水平的面试问题。
Okay, we're here at NeurIPS. We're recording a special Latent Space episode from NeurIPS, and we're here with Ashvin from Cursor. Welcome.
Hi. Yeah, thanks for having me.
So I guess Cursor is a new identity. I didn't even know if I should say that, because you only joined Cursor 3 months ago. Before that, you were at OpenAI, where you worked on o3; before that, you did a Berkeley PhD in RL, but focused on robotics.
Robotics. Yeah.
Is it weird switching from robotics to language models?
Okay, this is kind of interesting because a lot of people have been doing this. I mean, OpenAI—yeah, robotics. I was actually at OpenAI in 2017, also working on robotics. I was interning right before my PhD, where I worked on robotics there.
2017? He was famously OpenAI's first intern.
Oh, really? Okay, then he might have been before me. But yeah, there were 15 interns. It was a very different company. It was just robotics, Dota, and 15 interns that summer, all having pretty exciting individual projects. That set of interns—if you look at where they are now, it's kind of cool.
Yeah. Was there anyone from that class that you would shout out?
There were just a lot of cool papers that came out. Lerrel Pinto is now at NYU.
Yeah.
The person who leads reasoning at xAI—I forgot his name.
Well, he left. Eric?
But yeah, I forget his name. He worked on K-FAC and stuff, I think.
The vision dude, Greg.
Not Greg, but yeah. It was an exciting time to be there. I think robotics is a pretty good fit for LLMs because the switch ends up being pretty similar. You want to look at a lot of data, and it's hard to get stuff working in the robotics world. I think it builds very gritty people who look at data a lot, that kind of thing.
So, for whatever reason, I think that transfers. It's happening a lot, and I think it makes a lot of sense.
One of my NeurIPS highlights so far was having dinner yesterday with Lex Fridman. It was a small group dinner, and Lex used to be in robotics. He gave his assessment of robotics people: robotics people are the best to talk to at NeurIPS because they're the most well-rounded.
Because they don't have a choice. They work with the real world and real-world-looking data. The most unhinged, the most detached from reality, are the simulation people.
I see. Yeah, yeah, yeah, yeah.
I think I agree. I actually did a little bit of both during my PhD. I worked on prototyping ideas in simulation and then getting them working on real-world robotics. Robotics is probably where you feel AGI the least, because it's just so far away from working.
Over the last year, there have been demos that have been super interesting from Physical Intelligence and Sunday and stuff. I'm starting to think, “Okay, this kind of stuff—”
Have you seen the Sunday robots themselves?
I haven't seen them. Apparently, they've been doing demos, and I'm pretty keen to see them.
Yeah, I've seen the Physical Intelligence ones live, and it's pretty impressive. Just in someone's living room, folding laundry and stuff—you can just toss it in there. Everybody must be mesmerized.
Yeah.
Okay, one last thing on robotics, and then you can pivot to OpenAI. OpenAI is restarting a robotics team. Is that serious?
I actually know very little about it, because I was in a pretty different part of the org.
Yeah, I mean, I think it's serious. I think there's a ton of excitement around robotics right now. I'm actually curious what drives it, because I don't think I fully understand. There have been crazy raises recently for robotics companies, right?
I guess my own view on it is that when I left robotics in 2022, I thought I would actually come back to robotics. But my view on it now is that LLM agents are going to be a trillion-dollar market before robotics is maybe even a $10 billion market.
This is because LLM agents already create value out in the world. With robotics, it's hard to make the case that AI robotics does anything that useful yet. Once it does something useful, you have to make the unit economics work out, and I think that's also quite hard. Reliability, fixing these robots, and those kinds of things—it all makes it difficult.
I would say the market is kind of efficient in that software LLM companies are raising tens of billions, and robotics companies are raising hundreds of millions.
I think very recently it's been single-digit billions.
Oh, really?
Yeah. I think that's the surprising thing to me: it feels ahead of where it's actually at.
I would say robotics is in kind of the GPT-1 to GPT-2 era right now. I haven't worked on robotics. What task would qualify as, “Oh, that's the inflection point”?
It's a little bit like—you know it when you see it. I thought the Sunday demos were kind of cool. Maybe it's starting to get there, where the details matter a lot. It can't just be in a new scenario, one that you haven't seen before, and maybe on—
Generalization?
Yeah, exactly. I think that was kind of what GPT-2 was too, right? You start to see hints of cool generalization. It doesn't have to work out of the box, but at this point, especially, it still feels like in robotics you're not exactly investing in a technology; probably you're just investing in a team.
Yeah, yeah. I'm not in the space whatsoever, but that's kind of my impression.
It's actually nice when you're not in there, because you're as informed as basically everyone else. So we just kind of speculate.
Exactly. There's a robotics team at OpenAI.
So, coming back to language models, did you join OpenAI, or were you—
I joined right before ChatGPT, in September 2022.
Yeah.
I was pretty burnt out from my PhD, and I thought, “Okay, I'm going to go to this chill research lab.” Then ChatGPT happened, everything blew up, and a lot of stuff got refocused.
What did they tell you they were looking for you to do? Obviously, ChatGPT surprised OpenAI, but—
I joined on the CodeGen team.
The Codex team?
Exactly. It was the team that shipped Codex, but by the time I joined, we were more so working on the model, doing tool use and those kinds of things. It was very related to ChatGPT; we were kind of a sister team to the team that made ChatGPT.
We were working on making the models smarter through programming competitions, how to do SFT for that, and that kind of stuff.
And IMO gold has felt reachable in that time?
Oh yeah, crazy. If you told me that we could have gotten IMO gold, I would have assumed that we could all just go on vacation. It's all over—AI is solved. There's no point in working anymore.
We got it.
Yeah.
It feels like nothing's changed that much. Life is still the same.
Yeah. I think that's super interesting. I don't have a great way to explain it, but that's actually what I spend a lot of time thinking about: why is that the case? You see this again and again in AI, with solving chess, and then it doesn't really matter, solving Go, and—
Yeah, so you keep seeing it, but it surprises you every single time. I think, first, we keep moving the goalpost.
Yeah, we're very good at that.
And second, I think our definitions of what constitutes AGI are bad. We don't actually mean what we say when we say, “When we have achieved this, then we have AGI.” Clearly, when we've achieved IMO gold with a language model, we have AGI. That's wrong.
Yeah, and I think shifting the goalpost to some extent is correct. We keep Goodharting whatever goalpost we have—
And I think it's kind of hard to—
To say “Goodharting” is too negative. It's like, “I will cheat to do what you asked me to.”
Mhm.
But I don't think it was cheating.
It was just scaling test-time compute at a meta level. I think the community is not cheating, but it makes a lot of implicit decisions to go after the eval benchmarks that matter the most. And so, yeah, for sure.
Yeah, exactly.
But yeah, hopefully I’m not that Goodharted.
Well, it kind of clearly is to some extent, right? Most programmers in the world cannot do AI at any decent level, but we’re still struggling to automate most programming jobs. There’s a lot of stuff left to do, so language models are here at the junior-to-senior developer level, and then suddenly for AI, you’re like—
Exactly, and there’s something suspicious about that.
Yeah. Okay.
I kind of saw this at a meta level with RL research also. I did my PhD with Sergey Levine at Berkeley from 2017 to 2022, and that era of RL research was super interesting because it was super hyped, starting from DQN in 2015. A lot of the methods that people were really excited about were off-policy learning, value functions, and these kinds of things.
Somehow, that stuff hasn’t really panned out, I would say. It’s not exactly clear why, but in the academic literature, we thought we were making a ton of progress. In retrospect, I have to say that we probably overfit to the benchmarks pretty heavily.
The way I see this in retrospect is that we gave ourselves a lot of new knobs to tune and then implicitly tuned those to fit the benchmarks. Everyone knew at some level that we were doing that, but I think it’s hard to appreciate that it’s not just happening for a single paper at a meta level—it’s happening for the whole community, too.
Yeah. And I think the result is that a lot of the RL research that came out of that era isn’t used that much. I think it’s for a similar reason: basically, we were doing benchmark maximization.
I will flat-out say there was an RL winter. Entire startups were founded based on the premise at the time and basically gave up.
Mhm. Some of them died, some of them pivoted, whatever.
Yeah, yeah, yeah. I think because I was in academia, there was still quite a lot of excitement over it. But it still felt quite academic, and I was a little bit frustrated in that era because I felt like one of the pitfalls of academia is that it doesn’t really reward simple ideas that work. Instead, it tends to reward math-y ideas.
Those math-y ideas also give you implicit knobs to tune that allow you to overfit, while the things that actually work tend to be simple ones that have fewer knobs and just generalize to many things.
There’s just less secret sauce to it apart from throwing a lot of compute.
Exactly. Exactly. But those are things that tend to—
It’s not intellectually interesting.
Yeah, exactly. From an academic point of view, it’s like, “Why am I sitting in school?”
Yeah, yeah. I think for a lot of people who do PhDs, they’re wired in a way where they want to think about interesting new stuff.
And yeah, the scaling era kind of speaks to that.
Scaling era. Is the scaling era over since we’re up?
Well, I think I’ve just been pulled into that from an interview. I don’t think it’s over, but there’s definitely something interesting happening, right? The thing I was saying about AIME and IMO—I think we’ll still continue more or less on the same track.
Clearly, these labs are releasing their new pretrained models, and they’re still doing much better than before. So I think scaling is still happening, but I think it’s happening in a different way.
It’s worth seriously interrogating why we’re not just automating all jobs right now.
I think my view is that RL, the way it’s applied to LLMs right now, is kind of a weird, funny tool where it doesn’t really generalize beyond the training distribution that much. It generalizes to some extent, and it generalizes in interesting ways, but it’s very peaky, right?
It can kind of kill the training distribution completely. It can be the best in the world at it with not that much effort, really, but it doesn’t really generalize. So I think what we have to do is bring the world of economically useful tasks into the distribution for RL.
If we commit to using RL as a tool—and it might be the case that maybe there’s some cool continual-learning thing or something that shifts the paradigm next year—
But—
It really feels like if RL is a tool, then a big thing that needs to happen is that the intelligence of the models isn’t the bottleneck. It’s more that you need to have products that bring the entire context of what someone wants to do into the product so that the LLM can see it, and then you need to do RL on top of that.
Yeah.
Have you seen GDPval?
Yeah, I’ve seen it. Yeah, yeah.
Is that basically what you’re envisioning?
Yeah, I haven’t looked at GDPval closely. Actually, I haven’t seen exactly—roughly, to recap, it’s 128 tasks across white-collar jobs that account for more than 5% of GDP, right? They basically created all the context for the eval and evaluated every model.
Famously, OpenAI’s evals—whoever runs that one—always finds that Anthropic’s are the best for—
Yeah, yeah. It’s, uh—yeah, props to them for publishing it.
Doing that. It’s actual science.
I think it’s good.
But in a sense, generalizing beyond coding competitions to economically useful tasks—
That is it. I think that is what’s more important for GPT-6.
Yeah. What I’d like to do is—I just haven’t read the GDPval traces closely. It’s not clear to me what the job of an accountant entails and what kind of context needs to be in the product so that you can do it.
PDFs. I see. I see.
They try to go as close to source documents as possible.
I see. I see. Yeah, yeah.
So yeah, I think roughly operating in this kind of thing is what I envision, because it can’t be an artificial thing like, “Oh, let me clean up this data for you to make it easy for the LLM to process.” No.
PDFs in an agent loop. Yeah, I think that’s roughly the right shape of the thing.
I guess how I imagine this being operationalized is that you’d want to co-design the product and the model, so that the product—whatever it is—can provide the right context.
Coding is maybe the easiest first step, because most of the context that you care about is just your codebase, being able to run stuff in the terminal, and that kind of thing. And still, we’re not that close to automating it necessarily.
But for all the other jobs, the context is insane, right? It’s all the conversations you’ve had with your coworkers, your Slack messages. For my work at OpenAI, I was working on hyperparameter-scaling research, and I actually wrote not that much code.
Grid search or neural architecture search?
No, more like understanding the scaling laws of deep learning in 2020, where it was, “Oh, you have to initialize the layers in a particular way to get good scaling”—kind of the analog for that for RL.
Okay.
The thing is, I didn’t write a ton of code. Writing code wasn’t the bottleneck. It was more that, over the course of a year, I would run sweeps, look at the interactions between different hyperparameters, and build up that knowledge through a year of looking at different graphs.
To do my job, the model would also need all those things in context to successfully automate my job. You would want a product that allows you to bring all that context in.
Did you have to build it for yourself, or is there an existing one?
No. Those graphs were just sitting in my head.
Yeah.
Right. So I think it would be pretty hard to automate that job. What you need to do is build a product that brings that context in, and then you want to do RL on top of that to teach the models to use that context.
Yeah. Another conversation that has really come to a head this year is the death of “one model fits all.”
I feel like the point of the G in AGI is one model fits all.
Mhm.
I think OpenAI has clearly abandoned that this year.
Oh, what do you mean?
Fidji wrote a blog post titled “We Are No Longer Doing One Model Fits All.”
Okay, interesting. Okay.
And I think Mark Chen, or one of the other senior people who aren’t Sam, also said this in a podcast. Basically, the idea was that you started with Codex, someone else was doing InstructGPT, and then we launched GPT-4—
I guess o1.
And o1 was kind of supposed to be a reasoning, one-model-fits-all system—
And we merged the GPT-4o and o1/o3 lines into GPT-5—
And now we’re splitting it out into GPT-5 and GPT-5-Codex again. It’s just a weird—
Well, OpenAI is very guilty of that. I don’t think you should interpret those as scientific facts about the universe.
It’s just more like OpenAI has a tendency to shift the org chart, basically.
Yeah. Right. The world has a tendency.
Yeah, exactly. So I think a lot of it is related to that. But yeah, I see what you mean by the current reasoning paradigm fitting itself to this kind of peaky-in-certain-areas thing, right? I don’t think it’s so much a matter of model capacity, though.
It’s just more of an organizational thing: if you care a lot about coding, you probably don’t have the data to do all the other stuff. I don’t think it’s so much a matter of—if you had all the data, you would probably benefit from just training on all of it, and you’d get some generalization between these—
But it’s hard to find one organization that cares about all of these at once.
Yeah. Yeah.
Yeah. So before I double-click on the o-series in OpenAI, I do like to ask OpenAI people who were there: do you have a favorite blip story?
Yeah, the blip was crazy for me. I was at Thanksgiving—
Everyone remembers where they were and what they were working on.
Yeah, exactly. I was at Thanksgiving with 2 OpenAI friends, actually, and one of them, on Friday afternoon, was like, “Oh, Sam Altman just got fired.”
We were just working together, so I was like, “What? Oh, haha, good joke.” And then, yeah, it was crazy. It was just a crazy weekend of ups and downs. We thought—
You signed a letter?
Yeah, I did. About 95% of people signed it.
Yeah, yeah, yeah. I thought you would move on to Microsoft, or—
Well, I think maybe I had a slightly more complicated reaction. I actually do think that governance feels really important to me, because it does feel like, no matter if we hit AGI in 2 years or 10 or whatever, it’s not clear that we have a good structure for the governance of it.
Okay.
And so it is a question that we probably should spend more time on. During that period, I was pretty willing to be like, “You know what? Let’s forget about the equity and stuff. I think it’s good and healthy to have a conversation about how exactly the governance should work.”
Okay. You care about this.
Uh-huh.
Right. So now the OpenAI nonprofit has this secret shadow board of members that determines when we’ve reached AGI.
Yeah, yeah.
Is that better?
Yeah, I don’t have an answer. It’s just—it’s not fair—
Above my pay grade, but—
Yeah, and even back then I was kind of like, well, I don’t care. I—
I do care quite a lot. When the blip happened, one of my reactions was, well, you know, this nonprofit board stuff—if it takes somewhat surprising, maybe erratic actions, maybe you’d rather just have a thing like the Microsoft board, which is probably all the pensions of the world. Why? Serious people, but also the stakeholders are kind of the whole world, because everyone is, through their pensions or something, invested in it.
Maybe that is a bit more of a democratic way to run things than having 7 people run it. But yeah, I don’t really know. It feels like we haven’t solved governance at all, though, right? Forget AI—even stuff like unhealthy food or social media. It kind of feels like whatever the capitalistic incentive is doesn’t actually capture good outcomes for society, maybe.
So about the transition into reasoning: you shocked me by mentioning that the reasoning team is 300 people.
It’s kind of like, now that o3 was shipped as a product, I think it just gets larger and larger, how many people work on it. I’ve lost track of the numbers, but a lot of people contribute to the different aspects of safety and evals.
Yeah. The original o1—I saw the video, and it was like a dozen people.
Yeah, well, even then, if you look at all the contributors, it was probably more like 50 to 200 people.
Okay. So let’s tell that story from your point of view: figuring out what RL means there. And I guess, was this a branch of any prior work that you want to credit?
Yeah. Setting the scene, I guess, in 2023 people were talking about whether scaling laws were dead, that kind of stuff.
Every year, every year.
Yeah, yeah. But especially that year, it felt pretty serious.
In general, OpenAI is really good about having conviction in something and just, from first principles, going after it. I think the people who are probably most responsible for that are Ilya Sutskever and Jakub Pachocki. I think Dota 2 was more or less the same template in some ways, right? That was in 2017, and a lot of the people there have this AGI-in-their-bones kind of point of view.
They were basically convinced that RL would be the way to get there. For a long time, people had been convinced that something like that should work, and it’s just that it started to work once the pretraining got good enough.
Okay.
Yeah. I think human feedback is a bit of a side branch, because you can’t really pour that much compute into it, right? You take the model and elicit it to be a little bit better in terms of personality.
But the people there were really convinced that, at some point, it’s not about copying the internet. You can do RL, and that’s the path to getting much better intelligence. So there’s kind of a long line of returning to RL in different ways, and then around 2023 is when it started really clicking.
It was interesting because the initial models didn’t perform way better than the existing models, because they were smaller scale. But people were very good at saying, “This is kind of interesting. The reasoning trace that you see here is not something that you’ve really seen be so accurate in other models.”
It’s similar to how a lot of people didn’t really think of GPT or GPT-2 as something that was super compelling. I personally didn’t like GPT-2 that much. I was like, “Okay, whatever,” and then GPT-3 happened, and I was like, “Oh, wow. I feel a lot of FOMO sitting in my PhD.”
It’s kind of that. I think it takes a bit of first-principles conviction to decide that there’s something here and that we should really scale it up. OpenAI is really good about, once you decide that something is good, scaling it up all the way.
Was there an internal prototype before o1 that was like, “Okay, this is the thing. We’ll fund it to scale it up”? There usually is.
Yeah, exactly. It was running RL on even a pretty small model, producing very interesting reasoning traces and getting surprisingly good scores on math in a way that we couldn’t have done without a bunch more pretraining.
Once that looked good, more and more resources went to scaling up that new line, as well as adding things like tool use.
I think a lot of people make a lot of headlines about the large models, but I think the minis are very underappreciated—how well this solution works. Any comments or discoveries there?
Yeah, nothing much to say there. I was also not super involved in the mini stuff. Maybe one thing, not exactly related to that, is that externally people are like, “Research seems to come in these big leaps.”
Okay.
But internally at OpenAI, it feels very smooth. You have a bunch of experiments. Some of them have inconclusive results, but maybe you stack them.
Yeah, exactly. You stack them and just keep scaling. You keep having different runs that get a little better each time.
Okay.
So I think that’s maybe one other aspect that’s a little underappreciated.
I don’t know. In the media, there are just these wild swings between, “Oh, we’re—”
Yeah, exactly. Internally at the big labs, it’s just kind of like, “We’re chugging along.” Maybe this month is a little better than last month or something, but it’s not as crazy up and down.
I think the question is that there used to be more of this, and now I know there’s less. The stuff we’ve released—we’re internally about 6 months ahead. Part of the reason why people at OpenAI weren’t that excited about ChatGPT’s launch was because they already had GPT-4. They were like, “Oh, we just put this out. We’re already way ahead.”
I think now people are just releasing things as they have them, like—
I think, yeah, especially because there’s some competitive pressure, right?
I think people are probably pretty worried that if you let a lead linger for too long, that’ll grab a lot of market share. I don’t know—Nano Banana Pro right now is probably—it’s like a month…
So I would say the lead time from internal to external is about 1–2 months.
Yeah, which is exactly—
Tiny. Pretty short.
Yeah. Tiny. Anything else on the reasoning side? I guess you can talk about the work on coding. Was there anything that surprised you, or is there an external misconception about the o1/o3 side before we go to Cursor?
Not really. I think it felt by early 2024 like, “Oh, wow, this recipe really works, and we can see how far we can take it.” It was very steady progress, and by that point it was probably pretty predictable that we could really smash things like the IMO or IOI.
One funny thing that happened while this was going on is that I went to a conference called The Curve.
Yeah.
It was about AI progress—
And Joseph Gordon-Levitt—
Yeah. I went last year. This was before the o1 stuff was released.
Yeah.
I went to this thing where people were making bets on where we would be on Epoch AI’s FrontierMath exam and Humanity’s Last Exam and things like that. Their estimates were, “Oh, we’ll be at 10–20% in 2027.” At the time, there were already models internally that were better than their estimates, so they were off by about 2 years or something.
The interesting thing is that these are also people who were predicting that there would be Dyson spheres by 2035 or something.
Okay, so you see what I’m saying? Their current estimate is way under. Are you too pessimistic in the short term and too optimistic in the long term?
I don’t know. There might be Dyson spheres by 2035. I don’t regard it as impossible. But I think that is one interesting aspect: people still seem pretty miscalibrated in different ways. I do really appreciate how that community makes predictions, because most of the rest of the world just cynically says, “I saw this the whole time.” I appreciate that they actually make predictions.
Is this EA-adjacent?
Yeah, exactly. It’s that group.
Yeah. I like that they register their opinions ahead of time.
I think broadly, the people who’ve made capabilities predictions in that group have been broadly correct if you look from 2015 to 2020 or something. A lot of people thought AI was a sham, or that it wasn’t really going to be useful for a long time, and actually it is. It’s somewhere in the 2030-ish range that it will probably reach human-level intelligence.
Yeah, it’s weird. I feel like a skeptic when I keep saying that everyone always predicts AGI happens in their lifetime. That’s very convenient for whoever is making the prediction, and we have a consistent view of history where you can see people in the 1800s and 1900s making predictions. It somehow always lands in their lifetime, whatever the thing is.
But this time it might happen—almost surely, right? I don’t know.
I’m pretty sure here.
Yeah. It’s an interesting observation: how different are we from our predecessors in terms of developing a technology? Did the DeepSeek moment earlier this year—which was also crazy—change anything internally?
Not really. I think what surprised me more was that it created such a moment. It was kind of confusing, right? DeepSeek shows that NVIDIA chips are actually more useful than previously thought, and NVIDIA’s stock goes down a bunch. It was kind of—
I think the steelman of that side is that you don’t need the top-of-the-line NVIDIA chips. You can use the previous generation, or the restricted ones they sell to China, to do an equivalent amount of work for a reasoning model.
I see. Yeah. But I think the feeling at OpenAI was that we had a better model already at the time, right? Smarter models were clearly quite valuable, so you wanted to be at the frontier.
Okay. I wasn’t quite framing this as a race-dynamics thing between labs. It was more, “Were they right? Were their approaches right?”
They had R1, which was a really cool branch. So it’s more commentary on what we learned about RL this year.
Yeah.
It does seem like a lot of the labs have converged on some similar way of doing RL, and they’re all kind of back at the same level of being at the frontier again. Even the Anthropic models, like Opus 4.5, have this kind of ARC-AGI-2 plot that looks exactly like the OpenAI ones, right?
I think everyone seems to be converging on a pretty similar form of RL. It’s interesting. People basically figured out, in one way or another, how to achieve more or less the same thing.
Yeah. Let’s talk about the move to Cursor.
Yeah.
Why is Cursor accumulating and drawing in so many cool RL people?
From Cursor’s perspective, it’s nice not to be so dependent on external labs for everything. I think there are also unique opportunities to co-design the product with the model in ways that we couldn’t unless we actually built the model ourselves and had access to making it good.
Yeah.
That’s broadly why Cursor is so excited.
Okay, I’ll push back a little bit. OpenAI has infinite resources.
Infinite data, and it has Codex. You could have just stayed.
Yeah. Yeah. Well, actually, right around when I was leaving is when people started really using Codex a lot. That happened right after I left, so that was kind of funny. Mostly, people were using Cursor internally—
Maybe a bit of Windsurf because it was left over from the previous thing.
Sure. Yeah, exactly. So it wasn’t that obvious. But more to the point, RL is kind of a tool that doesn’t generalize that well. What you want to do is bring the entire test distribution inside your training distribution. I saw the opportunity to do that directly at Cursor.
I think the Cursor folks are also really excited about that vision. It’s a small place where the product people sit right next to the ML people, and I think there’s a lot of potential there.
You can see that recently Jacob Jackson had a blog post about online Tab, where we’re doing policy updates—
Which is every 2 hours.
Exactly. A policy update every 2 hours or something. That’s the type of thing that’s a little hard to do. It’s very hard to imagine that at OpenAI, for example, because the product is this complicated thing, and the product people and RL people are on different sides of the organization.
I think if you put your mind to it, you would. Tab is an autocomplete. It’s a smaller model, and it’s not as complex, I guess, as—
Yeah, but I don’t think that’s really why Cursor was able to do it. It’s more about the organization itself being smaller and a bit more focused.
Yeah. Since you’re indulging this, I think the question about continual learning—which has always been a big theme and is an even bigger theme this year—is: don’t you need to curate your data? You can’t just chuck whatever your users are doing straight in, because that tends to get you toward the middle of the distribution. You actually want to spike it.
I guess it depends how you’re thinking about continual learning. Humans are quite good at dealing with bad data too, right? You can see someone doing something dumb and decide that you’re not going to do it.
Filter it out.
Yeah, but it’s not even actually filtered out. You presumably have some kind of value function that says, if you see someone touch a hot stove, you’re not going to do it. It’s not just filtering it out; you’re actually not going to do it.
You could rediscover hot stoves from first principles.
Yeah, but you don’t need to. I think there’s something pretty deep there. It seems like we’re a few orders of magnitude of data efficiency away from that kind of thing. You do something once, or you make a mistake—if you introduce a bug in your code, you’re not going to do it again. But the models will happily keep doing it, even within the same context, and definitely across contexts.
Yeah. I think there’s something interesting and deep there. I suspect that it’ll be paradigm-shifting in the next year or something, but I have no idea what it might be.
Yeah.
So primarily, you’ve worked on Composer, Tab, and maybe Search?
So I’ve actually just worked on Composer, and that’s kind of the main focus of the company, basically—or at least the ML group—is shipping a better Composer.
Can you describe, I guess, the impressive brag a bit about the ML group?
Yeah. I think the ML group is great. It’s just 20 or 25 people, and I was honestly very pleasantly surprised at how good Composer is, given the size of the group. It’s not like a big research lab yet, and I think it’s a really good model.
You can see that in the reception, and I think it’s the start of hints of co-design with the product in some ways. One of the reasons people really like it is that it’s smart enough that you actually want to use it. It’s also fast, so you stay in the loop with the model while you use it, because all the other smart models are slow enough that you want to context-switch away and come back. That sucks. As a programmer, it just sucks to context-switch; it kind of gives you ADHD. It’s really terrible.
I agree. It’s one step in the direction of being able to be more in sync. I think the whole company is full of people who want to code. Even the co-founders actually code, and the co-founders are often some of the best high-taste testers, which also gives you a lot of reassurance that you’re going to ship good stuff.
Is there any example of a task that Composer doesn’t solve yet but that you’re really motivated to solve?
Ironically, I feel like I’m actually a low-taste tester in some ways, because I just write slow machine-learning code and think about algorithms all day.
Yeah.
I think more broadly, I’m super excited about co-designing the product so that you can actually—not just right now, where we’re getting better and better at answering user prompts. I think that’s why Composer 1 is quite good.
What we’re really aiming for is to automate software engineering as a process, where you write code, go look at Datadog, look at what’s happening, then come back and maybe have some hypothesis about what’s better and rerun things. That’s the type of thing that we actually want to make the model do.
I do think that Cursor is uniquely positioned to do that. If a lot of what a software engineer does ends up in the product, I think we can use that to get better and better at not just writing code, but the whole job.
I think that’s very inspiring. Just to double-click on any sort of RL insights: Sasha and Lee have talked a lot about the internal tooling that you’ve had for all the cluster visualizations. Is that helpful? Is that what every lab has?
Yeah, I think the tooling at Cursor is actually really good. People are just down to vibe-code stuff and test their own stuff, so we have a lot of good tooling where you can have an SSH session into our own user environment or something and see whether code runs the way that users got it to run. I think that’s quite nice.
One of the big lessons in ML in general is that you want to be really close to your data and understand your data well. I think we’re again uniquely positioned to do that well, especially because all the internal tooling is just internal—you’re not buying anything.
Yeah, it’s just internal. Part of it is that we’re also working on a product where you can understand it really well because it’s a code product. If I were to look at a biology question at OpenAI, I’d have no idea what it was about.
Yeah. Interesting. I think that’s a good overview of everything. Other than OpenAI and Cursor, is there interesting RL work that other people are doing that you’re still mulling over, that’s influential to your thinking? Good papers, anything like that?
Unfortunately, I’ve gotten in the habit, especially at OpenAI, of not reading that much external work and just reading people’s internal Slack posts as the main way to learn new stuff. No super-inspiring recent things have popped out to me.
I do think that this vibe of continual learning feels like there’s something super interesting there. It feels like maybe even people in academia could make a big crack at it.
And continual learning specifically means kind of what Tab is doing?
Maybe what Tab is doing, but also just in-context learning with infinite memory or something. Once you experience something in context, it should just be in your weights, and you shouldn’t have to make that same mistake again. That kind of thing.
Why do you think there’s—okay, but it should be in your weights, and there’s a finite capacity for the weights to remember things. You will forget things if you do that too much, right?
Not necessarily. You start out by memorizing or learning from trillions of tokens. Now you’re going to experience thousands or maybe millions of tokens.
You only need one epoch.
Yeah, exactly. So crazy. It feels like if you could learn enough about those million tokens that you’re actually deploying on, I don’t think you should need to worry about overloading the capacity of your model. I don’t think there’s a risk of that, right?
Because you can train on a trillion tokens and it’s fine, right?
Right. So proportionately, it’s a drop in the water bucket unless you run it for years and, at some point, it starts—
Maybe.
Yeah. So basically, I find it very curious.
I’ve only had one podcast on the information theory of language models. What is the theoretical capacity? How much are we using? You should probably track that.
Yeah, that’s a good idea. Treat the weights as a hard drive. If you want to store things in the weights, treat them as a hard drive. What’s the capacity of the hard drive? How much can be stored in there? We know the capacity—it’s the number of bits represented by the parameters.
Yeah. You physically cannot store more than that.
Yeah. I’ve heard that someone recently at Cursor, Jacob, brought up this view. I don’t know if it’s a more public view, but there’s kind of a hard-drive view of neural networks and a CPU view of neural networks. Is what’s happening in the weights just memorizing stuff, or are you having a few circuits that do a lot of work? This kind of thing.
I would love to explore it. There are so many of these more science-y questions that I’d love to explore sometime, but they conflict with empirical work. Unfortunately, at any given moment, it doesn’t seem like the most fruitful thing for improving something, especially in the short run or even in the next couple of years, is understanding some of these questions.
I guess this is technically supposed to be the world of academia, but it’s also hard to explore those ideas there without enough compute. I would love to return at some point to exploring these fundamental science ideas.
Okay, this is something I was kind of springing on you, so you can take some time. What’s a good RL interview question that, if somebody can answer it, means they should join Cursor immediately?
Ooh, it’s a hard question.
I assume you do interviews.
Yeah. At Cursor, we do work trials. They’re 2-day work trials, and I actually think that’s more representative, because you plug in and see how they behave.
I think that’s more valuable. This is honestly less about how well you understand RL and a bit more about whether you were around in the 2017-to-2022 era, but why is off-policy RL unstable? That’s a good question to dive into.
I don’t actually know, so I’ll have to dig into it.
Cool. Thank you. That was a great conversation.
We are definitely hiring at Cursor. If you're interested in working especially on data and rewards for code, I think that's a huge need. Please get in touch. That's it.
Yeah, thank you. Sweet.