[BidClub_]
Matthew Berman · · 44 分钟

如何在所有人之前看懂下一波 AI|Tibo 访谈

Matthew BermanTibo

YouTube
TL;DR
  • Codex 用户数已达到 2000 万,Berman 形容其增长曲线在 Anthropic 曾经“吸走房间里所有氧气”的阶段之后突然变得近乎垂直。 Tibo 不愿直接参与这场竞争,只说“我通常不会太关注竞争对手”,并将增长归因于把 Codex 并入 ChatGPT,让 coding agent 直接触达 ChatGPT 庞大的存量用户。
  • Tibo 那条“Codex 2-3 个月后会显得原始”的推文,指向的下一个约束其实是笔记本电脑本身(“Codex will seem primitive in two to three months”)。 笔记本电脑原本围绕人的限制设计——打字速度、注意力和同时打开的窗口数量;“模型没有同样的限制”,未来或许可以“同时完美处理 100 个应用”,这意味着模型最终需要访问超出单台笔记本资源范围的能力。
  • Berman 估算,ultra-fast mode 的 token 速度约为正常速度的 10–14×,这可能把单人开发者从同时运行 10–15 个并行 agent,带回到保持工作流连贯的状态。 Berman 认为这可能意味着一次只用 3 或 4 个 agent,Tibo 则强调要围绕人的注意力来设计。生成占主导的工作体感可能快约 10×,但工具调用密集的轨迹只有“3×或4×提速”,因为瓶颈转移到了网络和 agent 开销。Tibo 预计,未来“可能 1-2 年内”,这种速度“即使不是默认配置,也会非常接近默认配置”,同时更高成本的档位始终会高一档。
  • Berman 所说 Luna 降价 80%,既反映了提前规划的算力容量,也反映了效率提升;Tibo 没有量化两者占比。 前沿模型帮助 OpenAI 重做 serving stack,使其在“同一算力范围”内实现“非常显著的吞吐提升”——这已经是递归自我改进的早期形态。非 ultra-fast 速度比 3 个月前快约 60%,而 Soul 的 token 效率“显著高于 Terra”。Tibo 的承诺是:“不会只是把这部分收益装进自己的口袋。”
  • 针对 Berman 所说 Sam Altman 暂停“RL 的绝对前沿”训练一事,Tibo 解释称,团队需要先“加固系统的所有部分”,再“完全掌控地”重启训练。 安全团队通过他所说的辩论与探索过程,形成了“一套相当清晰的原则”,与此同时,alignment 投入出现“巨大增长”。
  • Usage reset 按钮如今已经是一个实体物件,但本质上仍是一种刻意保持非正式的善意做法。 Berman 将提供 reset 的能力与容量规划联系起来;Tibo 说:“这不是和市场或财务部门合作完成的……我想什么时候按就什么时候按。” 他强调的区别是:“你可以口头上说你在乎,也可以像我们一样,真的在乎。”
  • ChatGPT 与 Codex 合并的终点,是让所有人使用同一个自适应界面:“你和你妈妈会用同一个东西”(“you and your mom will use the same thing”)。 它会成为每个人的“个人 AGI”。软件工程师、设计师等职业标签,“只是人类为处理抽象概念而发明的概念”;界面应该适配用户,而不是逼用户先给自己分类。
  • DeepMind 的往事构成了这期节目的文化论点:LaMDA chat 早在 ChatGPT 大约 1 年前就存在,但“DeepMind 并不是为推出产品而设立的”。 OpenAI 的自下而上文化几乎没有“叫停项目的阻力”,也愿意自我颠覆,“即使这意味着要从主业重新调配资源”。Berman 将其视为 Google 无法做出同样动作的对比;Tibo 承认 Google 当时有计划,但表示那不是适合他的地方。
摘要 · 为研究而整理的核心内容

1. DeepMind 早 1 年做出了 LaMDA chat,却没能把它推向市场

  • Tibo 回忆 ChatGPT 之前的时代:DeepMind 的语言模型团队内部已经有了“ LaMDA chat”,并计划在 ChatGPT 上线前大约 1 年公开发布;按他自己推文中的说法,Google 当时过于紧张,而 DeepMind 被挡在了发布可能颠覆 Google 的产品之外。“DeepMind 并不是为推出产品而设立的”;相比之下,OpenAI 的“研究与产品协作得非常紧密”,并且“明显偏向发布”。当被问到那是否是特殊时刻时,他说:“第一次意识到自己可以生成连贯文本……一开始更多是好玩,而不是有用。”
  • 他从这组对比中提炼出的创业者建议是:坚定判断、根据真实用户快速迭代,以及“愿意自我颠覆……即使这意味着要从主业重新调配资源”——“这非常难,但极其重要”。他也承认应该公平看待 Google:“他们有一个计划……一切都是大计划的一部分”,但那不是适合他的地方。
  • 自下而上快速发布也有制衡:“不能让产品变成功能大杂烩”。简单、对质量有自豪感仍然重要,ChatGPT iOS app 被视为“市面上最好的应用之一”,这就是标准。

2. “Codex 2-3 个月后会显得原始”

  • 对 harness 的批评是:高级用户已经把产品的笨拙当成常态——skill files“长期很难维护”,记忆并不完美,sub-agent 网络也会在不同环节戳破那层“幻觉”。目标应该是“真正深度理解你的东西”,一个主动出手的“完美小伙伴”,在不打破这种幻觉的前提下帮助用户处理日常工作。
  • 关于笔记本电脑是约束,Tibo 的逻辑是:笔记本电脑围绕人的吞吐量设计,包括打字速度、思考速度,以及用户需要同时打开多少个应用。“模型没有同样的限制”;它最终或许可以“同时完美处理 100 个应用”,所以未来模型“需要访问超出你的笔记本资源范围的东西”。
  • Tibo 将行业问题拆成两类:一类是个人 AGI,核心是“深深植根于对你这个人的理解”;另一类是彻底自动化——根据生产日志自动修复回归问题,或者在网络安全场景中,把扫描器发现漏洞到漏洞修复之间的窗口压到“接近 0”,人类只需批准高风险操作。
  • 语音已经改变了 Tibo 自己的行为:新的 voice model 支持调用工具,所以“早上我只要拿着手机坐在那里,口述几件要 ChatGPT 做的事,它就会自己去完成。它可以访问我的所有工具。”

3. Ultra-fast 让多 agent 工作流重新回到心流

  • Berman 的现状是:当前速度迫使他同时运行 10–15 个并行 agent,每轮等待 30–45 分钟,由此产生“相当显著的认知开销”。他认为 ultra-fast 可能让自己改为一次使用 3 或 4 个 agent。Tibo 的回答聚焦于注意力:ultra-fast 加上语音后,“它的运行速度可以和你一样快,甚至比你更快”,因此“我以前同时处理 10 个 agent 的方式,我不太想回去了”。设计原则是对人的注意力“足够友好”,让技术适应用户。
  • Tibo 对提速效果给出了明确限定:当工具调用不多时,效果“惊人地好”——原型开发网站或电子游戏,体感可能快约 10×;但工具调用密集的轨迹只有“3×或4×提速”,因为开销转移到了网络或 agent stack 的其他环节。
  • 内部资源分配也很说明问题:ultra-fast 会优先给故障期间的 incident commander,因为“每一秒都很重要”,也会给“认为自己正在做非常关键事情”的团队;但“绝大多数额度留给客户”。OpenAI 员工受到限制,否则他们可能“把我们的生产 GPU 全部吃光”。
  • 定价背后的效率来自推理硬件、token 效率和 stack 工程的复合改进,所以“可能 1-2 年内,这种速度即使不是默认配置,也会非常接近默认配置”;但“总会有高一档”,对应不同且更昂贵的硬件取舍。

4. 合并的终局:一个界面,“你的个人 AGI”

  • 为什么不顾用户反弹,仍要把 Codex 与 ChatGPT 合并?因为“我们的未来模型希望我们合并……底层是同一套技术,同一套 harness”,而且会高度多模态、以语音为先。软件工程师、设计师等标签,“只是人类为处理抽象概念而发明的概念”;最终状态已经说得很明确:“你和你妈妈会用同一个东西。它会是你的个人 AGI。”它连接的是不同工具,并根据每个人的生活进行适配。
  • Tibo 反复提到的“幻觉”,指的是一种无处不在、扎根于人的计算体验:在办公室白板上写下内容,系统也应该能够理解。新版 ChatGPT voice 上线后,语音交互“增长非常快”;他的概括是:“每次你转向更自然的方式,人类都会选择阻力最小的路径。”
  • 谈到 Anthropic 时,Berman 提到在局面发生变化前,Anthropic 曾“吸走房间里所有氧气”。Tibo 不愿把注意力放在竞争上,而是关注“我们能独特做好的事情”。他强调的增长驱动,是让产品经理、设计师、销售、市场和传播团队都能使用这项技术,再通过已有大量用户的 ChatGPT “非常快速地”完成分发。

5. Reset 按钮:由算力托底的非正式善意

  • 这个做法起源于迭代把系统弄坏时补偿用户:“我们刚好把它弄坏了 30 分钟,这是一些额外用量。”如今它已经被制度化,但刻意保持非正式:“背后其实没有太多审查……也不是和市场或财务部门合作完成的。我想什么时候按就什么时候按。”现在确实已经有一个实体按钮。
  • Berman 用 Amazon 的退货政策类比 reset 如何建立信任;Tibo 的说法是:“你可以口头上说你在乎,也可以像我们一样,真的在乎。” Reset 也可以用来标记产品时刻:“去探索这个新东西……你还没用过 ultra,给你一些额外用量。”
  • Berman 认为 reset 需要容量规划;Tibo 则说“我们很早就提前规划了算力”。2 年前,OpenAI 曾因过度投资算力受到质疑;Berman 称这是“那种疯狂但极其正确的下注”,Tibo 表示认同。

6. Soul 优化 Luna、RL 暂停与效率飞轮

  • Berman 提到 Luna 降价 80%以及 Terra 的一次降价,随后追问其中有多少来自算法改进、多少来自战略性容量规划。Tibo 表示算力很早就已规划,而前沿模型帮助 OpenAI“服务、重构或重新设计”整个 stack,带来显著的效率和性能提升。结果是在“同一算力范围”内实现“非常显著的吞吐提升”,几乎谈不上牺牲。除 ultra-fast 外,速度已经“比 3 个月前快 60%”;公司给出的承诺是分享收益,而不是“只是把这部分有意思的增益装进自己的口袋”。
  • Tibo 认为,递归自我改进不应局限于模型开发模型:模型也可以构建“处在使用这些模型关键路径上的基础设施”。Berman 提到推理 stack、优化、CUDA kernels、产品和 cloud agents;Tibo 认为这“从某种意义上说”确实是递归自我改进,但“更多是基础设施”,并且可以让这项工作反过来作用于自身。“如果我们不这么做,我觉得那会相当愚蠢。”
  • 对于 Berman 所说 Sam Altman 暂停“RL 的绝对前沿”训练一事——问题中还顺带提到一个没有解释的“Hugging Face incident”——Tibo 表示,这次暂停“有必要让团队和个人真正理解并加固系统的所有部分”,之后再“完全掌控地”重启训练。他没有给出单一的重启指标,而是说安全团队通过“一场辩论和探索过程”形成了“一套相当清晰的原则”,与此同时,alignment 投入出现“巨大增长”。
  • Tibo 最后向对 AI 感到焦虑的人推介:Luna 在 6 个月前还会处于前沿水平,如今已经“便宜得离谱”。Berman 提到与 Replit 有关的免费模式;Tibo 也表示,这类能力正通过免费模式提供。因此,“现在处于前沿的任何东西,6 个月后运行成本都会低得多、低得多”。对他个人而言,他使用健康和金融工具后,在去看医生时“掌握的信息更多了”。
Tibo

I can press the button whenever I want, whenever it feels right. I don't tend to look at the competition that much. I really look at what we can do uniquely well, what our values are, and how we can maximally accelerate toward that.

Maybe in a year or 2, these speeds will become, if not the default, very close to the default. And you look at the cost of Luna, right? It's phenomenal. Technology has a way of becoming very, very efficient over time. We're very focused on broad access, and we're optimizing for the utility that you get out of it directly.

Matthew Berman

I heard there's an actual physical button now.

Tibo

Yes, there is. I'll show it to you. It's very, very cool.

Matthew Berman

Tibo, thank you so much for joining me.

Tibo

Of course. Yeah, glad to be here.

1. Lessons from Google & DeepMind

Matthew Berman

I'm really excited to talk to you. I want to actually start with your time at Google. You were on the DeepMind team, and before ChatGPT, Google had something called LaMDA chat. You had tweeted, “Google was too nervous to release it. DeepMind was blocked from shipping products that could disrupt Google.” I think about that a lot. What were you thinking at that time while you were working on these products, well before ChatGPT really changed the world?

Tibo

Yeah, it was a very exciting time. DeepMind was a very creative place. I was mostly focused on infrastructure, and my specialty was products to accelerate research. There was obviously a group working on language models and scaling them, and they had gotten pretty good results. It was natural to think, “Hey, can you turn that into something that you can chat with and use for various things?”

The idea of something like LaMDA chat naturally emerged. It was internal at first, and then there was an ambition to make it into a publicly available tool.

Matthew Berman

What year was this?

Tibo

This was roughly a year before ChatGPT.

Matthew Berman

Okay.

Tibo

Yeah. We were also building all sorts of other things that I'm not going to talk about, but it was a very creative place.

DeepMind was not set up to ship products. OpenAI is a very, very different place in that sense. Research and products collaborate extremely closely together. We ideate together, we co-design a lot of things, and we have a big bias toward shipping. We also have a big bias toward making things available to people, which I really love. That's what drove me here: the mission, the people, and the talent density. There are so many great things about OpenAI, really.

Matthew Berman

Did you know at the time, when you were involved in LaMDA chat, that it was something special or would become something special?

Tibo

It felt very special. It was the first time you realized that you could get coherent text and something helpful. Initially, it was more funny than helpful, and then gradually it became more and more helpful.

Matthew Berman

You say you think about that often, and I understand that. In a lot of ways, Google got in its own way. What are some of the lessons that you learned there that you took to OpenAI?

Tibo

Yeah, this is why I think about it often. I think about it in terms of the culture that I have on the team, the culture of OpenAI itself, the good parts to preserve, and what not to do.

OpenAI has a very bottom-up culture. It's a very empowering culture. People can come up with all sorts of ideas, get together, and then very quickly ship something. There's very little stop energy in general for new product ideas, which is exhilarating and fun. It's all about impacting the world in positive ways, so preserving that is very important to me.

The other thing that is also important is not making it a mess. You don't want to have a hodgepodge of features with no overall direction or coherence. That's counterbalanced with a sense of simplicity and being proud of the quality of the product.

I think the ChatGPT iOS app is one of the best apps out there. We want to keep that. We're investing a lot in things like delight, performance, efficiency, and simplicity. There are these overall principles while still empowering everyone to try new things and ship very quickly.

2. Building OpenAI’s Culture

Matthew Berman

If you were to give advice to a founder about how to develop that kind of culture, what are some of the more tangible elements or practices that occur inside OpenAI that could give advice to a founder?

Tibo

I think having conviction, finding a way to get users, and iterating very quickly based on feedback are important. You also have to be willing to disrupt yourself. That's not as relevant for a founder, but it is relevant for companies like OpenAI. We come up with new research and new ideas all the time, and being able to identify when it's the right moment to invest in them, even though it might mean reallocating resources from the main gig, is super important. It's very hard, but it's super important to be able to do that.

Matthew Berman

Yeah. I mean, that's the exact thing that you were describing at Google. They kind of weren't able to do that. That's great. I mean, does that—

Tibo

They had a plan, to be fair. It was all part of a big plan. But to me, it wasn't the right place.

Matthew Berman

At OpenAI, or at any company as it matures, does that become more difficult to maintain—that culture of shipping and willingness to disrupt yourself? Especially when you have a cash cow that's printing money, and you have this other new thing over here that might be something cool and innovative?

Tibo

We are very, very forward-looking. The future of AI, what it will all look like, and how humanity benefits doesn't really wait or care about whatever you have established here over the next month or 3 months. I think it's very important to lean in and just be open-eyed about where it's all going.

Matthew Berman

And then figure out how to position yourself so that you do catch that wave.

Tibo

Even for OpenAI, we train models and then discover their capabilities. Benchmarks don't tell you everything. We have to play quite a bit with the models themselves to realize, “Maybe we haven't thought about benefiting from it in this specific way,” or, “It can do this.” Then you're just like, “Oh, that is a shift in how we think about the product.”

For example, we launched the new voice, and it's super delightful to talk to. It's very natural now, and it's capable of tool use as well.

Matthew Berman

Yeah.

Tibo

That changes things. Now I spend a lot more time just talking to it. Another thing I do all the time is dictation because the quality of the dictation is so good. It's much more efficient than typing the prompt.

In the morning, I just sit there with my phone and dictate a couple of things to do for ChatGPT, and then it just goes and does it. It has access to all my tools.

And that was not possible before we had really good voice models, and so that completely changes suddenly how you think about the product.

3. The Future of AI Agents

Matthew Berman

Yeah. Let's continue talking about new models and new harnesses. A few weeks ago, I'm going to start with another one of your tweets because these are bangers: “Codex will seem primitive in 2 to 3 months. We're about to go through another major evolution. The next generation of models need more than your laptop.”

What areas—let's start with the harness first—are still ripe for innovation as a model gets better?

Tibo

So many. If you're a sophisticated user of Codex and any other coding agent, you've gotten used to a little bit of the clunkiness. You have to manage skill files, and this is a way to teach it things, but a lot of people have realized that it's kind of hard to maintain over time.

Memory is a thing, but it doesn't always remember everything. If you have subagents, you have to care about the subagents, and it constructs a little network. The illusion gets broken at various points when you interact with it.

What you really want is something that deeply understands you, understands your goals, understands your day-to-day, and understands what your team is up to as well. Ideally, it reacts and is also proactive, and just helps you in your day-to-day. It doesn't break that illusion of a perfect little partner that you have. That's what we're working toward.

Another thing that you realize when you have very, very powerful models is that your laptop becomes a constraint in and of itself. The laptop was designed for humans. It's designed roughly to be able to absorb the amount of work that you can produce, how fast you can type, how fast you can think, and how many applications you need to have open. All of these are human constraints.

The model doesn't have the same constraints. The model can, for example, handle 100 applications opened at the same time perfectly fine, maybe in the future. In terms of access to resources, it's very clear that models of the future will need access to more than the resources of your laptop.

Matthew Berman

I mean, I'm guessing you're talking about cloud agents. All of a sudden, when you have things like ultrafast—which we're going to talk about in a little bit—and token speeds that are 10 to 14 times faster than what fast is, the bandwidth changes—or, sorry, the bandwidth constraint changes. The CPU now becomes the bandwidth; literally, tool calls, network tools, and any kind of overhead in the stack become the limiting factor.

Tibo

But then you can compensate by doing multiple things concurrently as well. You can think about exploring on one end, writing tests, compiling, and testing a new hypothesis all at once. Then you're shifting the bottleneck around because you're able to do more concurrently, and the model can think very efficiently and very quickly through it.

4. Pausing Frontier AI Training

Matthew Berman

With current token speeds, I find myself kicking off 10 or 15 agents in parallel, and that becomes a pretty significant cognitive overhead for me because of the context switching. You're kicking them off, and you can expect 30 to 45 minutes before a task comes back. Now, with ultrafast speed, that workflow changes significantly, and I don't think I would be able to have 10 or 15 agents—and that might be a good thing. Maybe it's 3 or 4 at a time. How do you see the workflow of a solo developer changing over time?

5. How AI Changes Developer Workflows

Tibo

Yeah. I think managing your attention and being much more friendly to your attention is something that we care a lot about. After all, we're trying to build for humans. We're trying to build the technology that's the most empowering for humans, and that requires building around your ability to multitask. How do you want to manage your attention, and do you want something brought up now, or is it better to bring it up in 30 minutes?

6. OpenAI vs. Anthropic

When you have ultrafast speeds combined, maybe, with voice, suddenly you're like, "Okay, this thing can operate at the same speed, if not faster, than you." You stay in the flow. You get to ideate, see prototypes, and build little reports in real time, and that just feels really good. Suddenly, you're like, "Oh, yeah. What I was doing before, multitasking 10 agents—it's like, I don't really want to go back to that."

Matthew Berman

Yeah.

Tibo

And so we're trying to bring that sort of experience that is really natural but also feels built for you, where you don't have to adapt. The technology adapts to you.

Matthew Berman

So there have been a number of agentic coding techniques discussed over the last few months. Loops were popular and still are popular. Now I'm hearing about graphs. Are these all techniques to allow the solo developer to manage, or be friendly to, their attention, as you said? I like that term.

Tibo

Yeah. So I think about 2 different categories of problems. The first one is building the very best personal AGI, or the personal agent that will be in the flow with you, proactively raise important new ideas when it can find them, and be very, very efficient at doing exactly what you want. It doesn't matter whether it's a technical problem or it's more like research or advice—it can do it all, and it's super, super tailored to you. This is a very important thing, and it's deeply rooted in the understanding of you as a human, as an individual who is unique.

The other category of problem is full-on automation, where you're building intelligent systems that can take care of a very complex process. Maybe something that does require intelligence and seems very complex—for example, going and looking at production logs and automatically doing performance optimizations, or looking at regressions and automatically patching them. In cybersecurity, we're seeing this as well. You have a scanner that comes up with a vulnerability. Can you automatically patch it and reduce the window where you have that open vulnerability to almost zero?

That would be without a human in the loop, or with very minimal involvement, where you only need to approve a high-risk action. It's mostly an automated system, but it's also not as important for you to be in direct control of it.

Matthew Berman

Okay. I want to slightly change topics. ChatGPT and Codex have been on this merge path over the last few months. So I guess, first, I wanted to ask you: How has that been going? How does it feel internally? What's the feedback you've been getting from your customers?

Tibo

It's really been a boon. The feedback we had initially was, "Why do you merge them? Do you really have to do it?" And it's like, well, our future models want us to be merged. So we're just going to do it because it is the simple and proper thing to do.

We're building this very personal, super-capable agent that can help you in all sorts of ways. This is the same technology under the hood. It's the same harness. It's the same way that we think about it. It's highly multimodal, voice-first, super-efficient, and it doesn't matter if you're trying to code or not. This agent is capable of it all, and it's the most efficient at it.

Then the interface that you want should tailor itself to your needs. You shouldn't decide, "I'm a coder, so I want a coder interface," or, "I'm not technical, so I want a nontechnical interface." There's a spectrum of people. We come up with labels like software engineer or designer; these are just human concepts that we have invented to deal with abstractions because reality is too complex for us to handle. But individuals are individual, and they're somewhere on the spectrum. We're trying to build the perfect interface that adapts for everyone. It doesn't matter if you're technical or not; it adapts based on your specific individuality. So that's why we went and did this.

Matthew Berman

But does that mean, inevitably, it's going to end up with a singular interface? No dropdown selecting between products. It's kind of wild to think that my mom might use the same exact interface as me, and then obviously it'll customize to my needs. Maybe I'll need more information if I'm doing more sophisticated work. What is the end state for you?

7. The Future of Human-AI Interaction

Tibo

That's right. It's the same thing. You and your mom will use the same thing. It will be your personal AGI. You will have very different kinds of tasks and utility that you get from it. You will connect it to different tools in your life. You will bring different ideas and different needs. Then it will continue to tailor itself to maximally benefit you, and it will do so with your friends and with everyone else.

Matthew Berman

So I want to go back to something you said. You used the word "illusion" a couple of times in that kind of end state. What is that perfect illusion for the typical user? If you can envision us a few years from now, what does the interaction between AI and a human look like?

Tibo

Yeah. To me, it's something that is very, very tailored to humans. This is why large language models are also a success: it's natural language. Natural language is a human concept, right? We're used to speaking to each other. If you write me a letter tomorrow, I'll be able to read it.

We know each other quite a bit now, so I'll be able to decipher a little bit of the emotion or maybe a little bit of the nuance behind the letter if you wrote me one. All of that is deeply human. The technology that we're building is rooted in humanity and in the way that humans communicate and get things done.

There shouldn't really be a situation where you're like, "Oh, you misunderstood me because you didn't quite decipher the nuance in my tone, or you didn't quite understand the text or how I meant it." That's something that we're trying to avoid. We're trying very much not to have you adapt, but to have the technology be perfectly created as a natural extension of how humans already act in the world.

Matthew Berman

When I think about communication between humans, so much of it is nonverbal—the way I move my hands, my facial movements, and so on. How much of that do you see being sensed or read by artificial intelligence in the future, maybe through vision? Is that even important?

What you're describing now is text-only. For those of us who grew up online, we're very used to communicating over text and adding subtleties to that text to convey what we really mean and what our tone is. Is it still important to have AI be able to read our facial expressions, hand gestures, and so on?

Tibo

I think so. When I think about the future of what we're building, it's very ambient.

It’s very natural.

Matthew Berman

Mm-hmm.

Tibo

If tomorrow I go to my office and write something on the whiteboard and have an idea, it should be capable of being there as well and understanding. Maybe I tell it, “Hey, what about this thing?” and then we just have a natural conversation over voice.

Since we shipped the new ChatGPT voice, the number of users who interact with ChatGPT just through voice has been growing very fast. I think the lesson is that every time you lean into something more natural, humans just choose the path of least resistance. As you said, typing on a little box is natural maybe for some of us, but not for everyone. When you get something that’s a little bit easier, a little bit better, you tend to go and use it.

8. ChatGPT & Codex Merging

Matthew Berman

Yeah. Okay. First of all, congratulations. I saw that you posted this morning: Codex reached 20 million users. I’ve seen the graph, and for a while it was like this, and then all of a sudden it went vertical.

So congratulations. I want to talk a little bit about that competition with Anthropic because, of course, a lot of people think OpenAI and Anthropic are the 2 major competitors in the industry right now. There was a period of time in which Anthropic was sucking all the oxygen out of the room, right? They were really dominating, and then all of a sudden something changed. First of all, what’s your read on the market today?

Tibo

Right now, we’re focused on building the most capable models, building models that are highly efficient, and taking a lot of pride in building products for everyone. I think this is something that OpenAI does really well: caring about the world and caring about how we’re taking this very powerful technology and putting it in the hands of as many people as possible.

This is what we did as well with merging Codex and ChatGPT. It was this desire of, “We have this technology. We can make it safer and easier to use for everyone.” Whether you’re a product manager, a designer, or in sales, marketing, or comms, you should be able to use all of it.

Then we distributed it very quickly through ChatGPT, where we already have a ton of users. That’s been really driving this growth curve as well that you mentioned. I don’t tend to look at the competition that much. I really look at what we can do uniquely well, what our values are, and how we can maximally accelerate toward that.

Matthew Berman

Okay. I want to dig a tiny bit more into that because I know you’re not thinking about Anthropic all that much, but a lot of other people do. They’re thinking, “Which product do I believe in? Which product do I want to give my $2200 to?”

When you look at the market position, the branding, and the tone from OpenAI, and just the way that it interacts with developers and the broader audience, how do you see that comparing to the way Anthropic does?

Tibo

Yeah. Maybe, again, what I care a lot about is community building for the world—bringing everyone along. I think you can feel that in the way that we’re super transparent about things and take a lot of ideas from the community. It’s also so much fun, to be honest, because we get so much energy from it as well.

The technology that we’re building isn’t just for ourselves. We’re not just building it to accelerate OpenAI. The mission is super important, and this is where we also get our energy from. It feels very grounded, it feels fun, and then good things happen as a result of that.

9. Why OpenAI Keeps Resetting Limits

Matthew Berman

Well, let’s talk about some of those good things. I want to talk about the resets for a second, too. I know everybody is following your every tweet because of this. Specifically, looking at that growth curve of Codex, maybe this is a silly question: how much of those resets is a boon toward marketing and growth, or was it just goodwill for the developer community?

Tibo

I think maybe it’s counterintuitive, but OpenAI is a place where you can just do things. It felt right initially to compensate when we were iterating and breaking things, or maybe we had misconfigured something and it wasn’t quite as good as we wanted.

So it was like, “Hey, thank you for trying this product. We know we’re trying very hard to build it; it’s early days. Here’s some extra usage because we happened to break it for 30 minutes, and we understand this is really important and you rely on it. Thank you for being a user.”

This is how it started, and this is how I still treat it. If we break it, or if the experience is suboptimal and we don’t fully understand why, we will compensate for that. We will reset the usage limits.

It turned into quite the thing, obviously. There’s a whole reset button now, but there isn’t really a whole lot of scrutiny behind it. It’s not done in partnership with marketing or finance. I can press the button whenever I want, whenever it feels right. We have these principles: we’re trying to build something amazing, and when it is not, we will make up for it.

Matthew Berman

Yeah. I still think there’s a piece of it that has built so much goodwill in the community and maybe contributed to the growth, at least in a small part.

Tibo

I think caring for your users does a lot, right? You can pay lip service and say that you care, or you can be like, “We actually care, and if we break it, we’re really sorry about it. Here’s how we make up for it.”

Matthew Berman

It kind of reminds me of Amazon’s return policy. If you’re not happy in any sense, go ahead and send it back. You’re kind of building that same culture, or that same perception, of OpenAI. It’s like, “Hey, if we make a mistake, go ahead, use those tokens again,” or, “Here’s a fresh batch of tokens for you.” I really appreciate it.

Tibo

There are also good moments when we just want to celebrate and mark the moment. There isn’t really something that we can give that is more meaningful at times like this. We always ship new features, and we will ship them as broadly as we can.

But something to share with the entire community is, “Hey, go explore this new thing. You just haven’t used ultra yet? Here’s some extra usage. Try it.”

Matthew Berman

And I heard there’s an actual physical button now.

Tibo

Yes, there is.

Matthew Berman

Yeah. Okay. You’ll have to show me that afterward.

Tibo

I will show it to you. It’s very cool.

10. AI Efficiency & Compute

Matthew Berman

But with all of these resets, you can really only do that if you’ve done significant compute capacity planning. You have to have enough compute to give all of these resets.

I want to start to talk a little bit about self-improvement because, speaking of capacity, a few weeks ago there was an article you put out, and it stated that Soul had optimized Luna efficiency. You dropped the price of Luna by 80%. There was also a price drop for Terra.

How much of an efficiency gain were you able to eke out of Luna versus how much of it was, “We did really great compute capacity planning, our margins are great, and we still want people to use it”? How much of it were algorithmic gains versus strategic planning?

Tibo

We planned compute way ahead. You know, I think if you look back 2 years, OpenAI was questioned for why there was so much investment in compute.

Matthew Berman

One of those crazy good bets.

Tibo

Yes. And now we’re very happy to have it. A very large fraction of the compute is used for research, where we invest in our future and ever-better models, and then also in the efficiency of the models that we have.

The amazing thing that’s happening is that when we push the frontier of capability for the most advanced models we have, we can use these models to figure out very quickly how to serve, restructure, or re-engineer our stack in order to gain very significant efficiency or performance gains.

We haven’t just improved cost efficiency; we’ve also improved speed efficiency. Outside of ultrafast, things have gotten significantly faster over time. They have—

Matthew Berman

Like, if you plot it, the amount of speed that you get now is roughly—

Tibo

60% faster than what it used to be 3 months ago. We’re going after every part of the stack and really making sure that we design and engineer it optimally for the kind of workloads that we have.

The most powerful models that we have are the ones that really make it possible for us to do this with a very small team. Whenever we come up with very significant efficiency gains and cost efficiency, our commitment is to keep things at the frontier of performance and cost, and not just pocket that interesting gain, but make it something that we share with our customers and users. That’s what we did with Luna.

Matthew Berman

What do the discussions look like internally when you’re trying to decide compute allocation toward researching new models, efficiency gains on existing models, and inference? What does that tension look like internally?

Tibo

We usually look at things from first principles, and we have an allocation for research, an allocation for product, and then within product we make different kinds of trade-offs. But this one was almost not even a trade-off, because the efficiency gains were there. We were pretty much able to use the same compute envelope in order to serve this very significant increase in throughput.

11. Recursive Self-Improvement

Matthew Berman

When I saw the blog post a few months ago, prior to the price-drop blog post, you guys were talking about one model training the next model or helping optimize the next model. Then you see these efficiency gains that were achieved by Soul looking at how Luna was running. It seems to me like recursive self-improvement in the very early innings. What are your thoughts there? Is that what is happening?

Tibo

Yeah, I think recursive self-improvement is obviously a huge topic right now, and it’s most often, I think, applied to research—models developing other models. But what we’re seeing a ton of success with is using those models to develop the infrastructure that’s on the critical path of using those models, which is also a form of recursive self-improvement. It’s all one big system.

Matthew Berman

There’s the inference stack, the optimization, and the CUDA kernels that we use. There’s also developing new products and new ways to interact with those models that are more efficient. You talked about cloud agents. If we really crack cloud agents, suddenly you become much more productive as well. Is that a form of recursive self-improvement, because then you have a better ability to get utility from them?

Tibo

I think it is in some sense, but it’s much more infrastructure, and then being able to take that and point it back at itself.

Matthew Berman

Yeah.

Tibo

And so, of course, we’re doing that. If we were not doing that, I think that would be pretty silly.

Matthew Berman

As we’re on the topic of recursive self-improvement, Sam Altman talked about pausing the absolute frontier of RL right now, I believe. Can you talk a little bit about that? I know we talked about the Hugging Face incident briefly, but what went into that decision? What did those discussions look like?

Tibo

Yeah, this is something very much within research, where OpenAI has always been able to invest its resources where it matters most. As we increase the capabilities of our models, it is very obvious that the alignment and safety aspects are ever more important. Having a tremendous amount of investment there is a very natural thing for OpenAI, and something that OpenAI is very committed to.

We’re seeing a huge surge in investment in this, and the pause was sort of necessary to allow the teams and individuals to really understand and harden all parts of the system, to ensure that we could restart training with full command. This is something that I believe OpenAI will always continue to do when necessary. I’ve not seen us internally not able to make such decisions very efficiently.

Matthew Berman

Was there some set goal in place where it was very clear you needed to reach a certain point before unpausing, or was it, “We’ll know it when we see it”?

Tibo

Yeah, this is something that sits within the safety team, and it’s very much a debate and a discovery process as you go. But they did reach a fairly clear set of principles that, when reached, would put us in a good position.

12. What Ultra Fast Unlocks

Matthew Berman

I want to go back to ultra-fast mode, because I think people don’t appreciate what that kind of speed unlocks. Let’s start with what use cases you’re using internally that were not possible prior to having that kind of tokens-per-second rate.

Tibo

We see it used a lot when the stakes are high. For example, when we have an outage, the incident commander and the response team get access to ultra-fast, because every second matters. High-stakes scenarios are where we’re really using ultra-fast.

It’s also kind of a fun thing where teams that are either working on something very critical or believe they are working on something very critical will always request ultra-fast as well.

Matthew Berman

Does Pets fall under that?

Tibo

Pets?

Matthew Berman

Yeah, Pets.

Tibo

Pets is not quite hypercritical. But I love my pet. It’s always on my screen. When you walk around, you see people’s pets on their screens, and when they dial into a video call, it’s just always there. I think it’s very delightful, and it brings me joy every time I see it.

Pets is not quite critical right now. We do maintain it, and we take good care of our pets. But say someone is working on a new idea and they’re like, “I really think this could be something special, and we have to try it, but we have to make a decision on Monday on whether we include this in DevDay or not.” It’s like, okay, of course, use ultra-fast.

People have different kinds of preferences on whether they like to be monothreaded, as we talked about, or multitask a lot. For folks who like to multitask a lot, you don’t benefit as much from ultra-fast.

Matthew Berman

But some people

Tibo

don’t like to change context all the time.

Matthew Berman

Where do you fall on that spectrum?

Tibo

I have ADHD, so I context-switch all the time.

Matthew Berman

It’s funny, because I also have ADHD, and I actually don’t want to context-switch all the time. That’s really hard for me. I want to focus on 2 or 3 things, and that’s why I was so excited about ultra-fast. It’s fascinating that you’re the opposite there.

Tibo

I thrive in context-switching and making lots of little decisions. Sometimes I do want to just stay focused on one thing, and then ultra-fast is delightful because it keeps you right there in the flow.

The thing with ultra-fast is that it works amazingly well when there aren’t that many tool calls involved, or when there’s a lot of context generation. For example, if you’re trying to prototype a website or a video game and you just need it to write a lot of code, then it will do it so quickly—10 times more quickly. But if there are a lot of tool calls, the overhead is somewhere else in the network or somewhere else in the agent trajectory, and then you’ll only feel a 3× or 4× speedup.

Matthew Berman

Yeah. You won’t get that full 14×. I know OpenAI employees get unlimited tokens, and I can imagine that if I had unlimited tokens, I would always set it to max thinking, 5.6 solar, or whatever the latest model is. I would think similarly about ultra-fast: when cost isn’t on my mind, I’m like, “Okay, max it out.” Is that how it is internally?

Tibo

We don’t give ultra-fast to everyone. We reserve a lot of our capacity for external users and customers. OpenAI employees have the ability and capacity to gobble up all of it—all of our production GPUs and all of ultra-fast. We would use all of it, but we restrict it in a way where we look at how much is reasonable for us.

That way, we use it to understand the product, keep improving it, and benefit from recursive self-improvement, while the vast majority is reserved for customers.

Matthew Berman

Okay. Yeah, that’s good. Thanks. What latency-sensitive use cases outside of OpenAI are you most excited about that get unlocked by that kind of speed?

Tibo

It’s interesting. One thing that I’m very excited about in general is non-text interactions. Can you operate on a shared canvas? Can you create things? Can you generate ideas and different images, then select one—choose your adventure—and have a very quick mockup of a prototype that you can then steer in real time, either through voice or through text, and just see it right there?

It’s this very creative process, which I think these speeds allow. As an engineer, sometimes you sit back and think, “I need to design this whole system. I need to think about the trade-offs and the requirements.” But maybe you can just create it in 1 minute and see how it actually does, and then be more in the flow and understand things better. I think these speeds allow that.

13. Will Ultra Fast Become the Default?

Matthew Berman

Yeah, and I’m assuming the ultra-fast price is going to be significantly higher than normal speeds. Do you think ultra-fast speeds are going to become the standard, or are they always going to have a premium price point?

Tibo

That’s interesting. I think, in the same way technology usually goes, it will become more broadly accessible over time. The speeds at which agents get things done will continue to improve. We’re seeing massive improvements month after month.

This is not just inference speed. It’s also how token-efficient the models are. Soul is significantly more token-efficient than Terra. The next model will be significantly more token-efficient than Soul, as you might expect. We’re always pushing on that, so things just get faster over time. Inference hardware also continues to improve and get faster.

I do think that, in maybe 1 or 2 years, these speeds will become, if not the default, very close to the default. But I also think you will always have 1 tier up, where you can use more hardware and make different trade-offs that are more costly but give you something extra.

14. Reassuring People About AI

Matthew Berman

So, Tibo, the last question I usually like to end on is for a broader audience. There are a lot of people out there who are quite nervous about AI, whether it’s job automation, environmental impact, or just this thing that’s happening that feels quite foreign. What words of encouragement would you give to the broader audience?

Tibo

We really build for the world with ChatGPT, and we’re very much investing in how efficient it is. This is directly aligned with broad access and broad utility. The cheaper it is to serve, the more you can do with it, the more you get out of it in your daily life.

It has gotten very, very efficient. If you look at Luna, for example, it’s a much smaller model. It is incredibly efficient, but if you rewind 6 months, it would have sat at the frontier.

Matthew Berman

Yeah.

Tibo

And you look at the cost of Luna, right? It’s—

Matthew Berman

It’s crazy cheap.

Tibo

It’s phenomenal. Right.

Matthew Berman

You just did that thing with Replit. You’re giving it away for free now.

Tibo

They’re giving it away in this free mode, right? It’s just like, wow. This access to incredible intelligence will become ubiquitous. It’s only possible when you push efficiency month after month, year after year.

Whatever is at the frontier now will become way, way cheaper to run in 6 months. This is how I would answer this question: technology has a way of becoming very, very efficient over time. We’re very focused on broad access, and we’re optimizing for the utility that you get out of it directly.

15. Why Everyone Should Try AI

Matthew Berman

How about for people who are apprehensive about even trying AI for the first time? What are you telling them, and how can you paint them a picture—a vision of the future in which AI is helping the world?

Tibo

Yes. I think you don’t have to look very far. ChatGPT helps people in very personal and deep ways. A lot of our users use it for help with writing, but also for personal advice or medical advice. We launched Health and Finance, and I use them regularly. I feel like I get a lot of support that would otherwise be hard for me to get, and it allows me, for example, to be more informed when I go see my doctor.

You don’t need to go very far to see the utility it can provide. Talking to others and getting inspired by how others use it and benefit from it is a great way to start considering how you could benefit from it.

Matthew Berman

Well, Tibo, thank you so much. Thank you. I appreciate your time.