[BidClub_]
BG2 · · 69 分钟

OpenAI企业业务内部:前沿部署工程、GPT-5及更多 | BG2嘉宾访谈

Apoorv AgrawalSherwin WuOlivier Godement

YouTube
TL;DR
  • OpenAI的企业业务早于ChatGPT——“OpenAI最初真正的产品并不是ChatGPT,而是B2B产品,也就是API”——如今公司还在运行一种Palantir式的前沿部署工程师模式,将工程师嵌入T-Mobile(由OpenAI模型实时处理语音客服)、Amgen(GPT-5在药物开发文书领域的头部客户)和Los Alamos(将o3物理部署到与Lawrence Livermore、Sandia共享的气隙Venado超级计算机上)等客户。
  • 面对MIT“95%的AI部署都无法奏效”的标题,Olivier从约200次企业部署中总结出的规律是:自上而下的支持加上自下而上的“虎队”,并且先做评测;“客户拿不出好的评测时,目标就一直在移动。” 企业知识大多存在于“人的脑子里”,而不是SOP中,从46%爬升到99%是“艺术,有时甚至更像艺术而非科学”。
  • 本期最锋利的判断是:尽管安全门槛更高,2025年的“物理自主性领先于数字自主性”。 因为自动驾驶有10-15年时间,以及道路、红绿灯等脚手架;而“AI agent就像被直接丢在荒郊野外”。Agent只诞生于o1-preview的推理范式,而“我认为这条斜率陡得惊人”;按收入计算,它们可能已经越过Waymo。
  • GPT-5的差异化在于“工艺——模型的风格、语气和行为”,而不是已经饱和的基准测试。 一位主持人称,某项评测显示幻觉“基本降到了零”;尚未解决的核心取舍是推理token与延迟——GPT-5 Pro能“一次解决”过去无法解决的问题,但需要等待10分钟。反直觉的教训是:指令遵循变得过于字面,旧式“请保持简洁”的提示词反而让输出失灵。
  • 强化微调(RFT)——Olivier说这是OpenAI创造的术语——比SFT“强一个数量级”,也把推介重点从定制模型转向利用专有数据打造“世界上最好的模型”(金融服务领域的Rogo、在TaxBench上达到SOTA的Accordance)。 Olivier的判断是:在前沿能力上,“RFT基本会成为行业常态”。
  • 多空交易上,Sherwin做空“围绕AI产品的整个工具链”——评测产品、框架、向量数据库,以及如今的RL环境创业公司——因为技术栈迭代太快,抽象层很难跨越一代模型存活。 Olivier做空以记忆为核心的教育(“知识token”),做多医疗——“可能是最受益于AI的行业”。
  • 值得记录的软信号包括:即便工程师人数的含义仍不明确,未来仍将需要多得多的软件工程;“世界上存在巨大的软件短缺”(OpenAI的产品经理如今直接交付写好的原型,而不是PRD);Codex CLI与GPT-5的使用正在加速;而“那次插曲”(董事会政变)让OpenAI变得反脆弱——“更抗打,也能更快恢复”。
摘要 · 为研究而整理的核心内容

1. API先于一切——企业业务是OpenAI“分发AGI”的方式

  • Sherwin提醒说,他3年前加入时,API“实际上是我们唯一的产品”——ChatGPT后来才出现。如今ChatGPT大约是“全球第5大网站”,但平台——API、政府/公共部门产品,以及正在成形的直销企业业务——才被视为使命更完整的表达:它能触达企业内部的应用场景和终端用户,而消费级应用做不到这一点。
  • Olivier的说法是,B2B之所以是核心,是因为“有一大类应用场景只能通过B2B实现”:每年多生产10倍的药物、更好的教育和公共服务——“真正推动现实世界运转的,就是这些企业”。他明确将其称为Palantir命题。

2. T-Mobile:用OpenAI模型处理语音客服,加上Palantir式FDE

  • 在嵌入客户约1年后,T-Mobile应用内的客服通话“实际上由后台的OpenAI模型处理”——包括文字和语音;在实时API支持下,效果听起来“超自然,无论延迟还是质量都像真人”。Olivier说,实时API大约在前一周正式GA。
  • 模型之上的部分才是真正的工作:前沿部署工程师(“我们借用了Palantir的说法”)负责把模型与CRM及企业工具编排起来,而这些工具“甚至没有API”——先搭建网关,再定义黄金评测集。音频评测“尤其难打分”:一次5分钟的转接电话, “你到底怎么知道正确的事情发生了?”
  • Sherwin指出了飞轮效应:T-Mobile的经验直接反馈进新的实时GA模型快照——深度嵌入企业客户,同时也是模型研发。

3. Amgen与Los Alamos:药物文书和气隙环境中的o3

  • Amgen是“GPT-5的头部客户”,将医疗需求分为两类:基于海量数据的纯研发,以及常被低估的行政和文档撰写工作。一旦研发“基本锁定了药物配方”,监管申报与审查就会变成“大量工作”。Apoorv补充说,如果药物能更快上市,影响范围可能达到数亿人的生命。
  • Los Alamos获得了o3在气隙Venado超级计算机上的定制化本地部署——“我们确实得把模型权重物理带进他们的超级计算机”,而该设施禁止携带手机。OpenAI能看到的信息有限,但反馈显示,o3正在加速实验、处理数据笔记本;推理模型还带来了一项新能力:充当实验设计上的“思想伙伴”,这是老模型“很难真正做到的”。该系统与Lawrence Livermore、Sandia共享。

4. 为什么95%会失败:先组虎队,再做评测,然后爬坡

  • 针对那份“震动市场几天”的MIT报告,Olivier根据“几百次”部署总结出的模式是:自上而下的支持,加上自下而上的虎队;团队同时具备技术能力和机构知识,因为即便在客服领域,“绝大多数知识都存在于人的脑子里”,而不是写进SOP。
  • 第二步是先做评测——“客户拿不出好的评测时,目标就一直在移动”。Olivier进一步强调,评测必须从实际操作者自下而上地产生;自上而下下达评测要求并不起作用。
  • 第三步是爬坡:“你已经有了评测,目标是达到99%。起点可能只有46%。”而抵达目标是“艺术,有时甚至更像艺术而非科学”;当模型遇到明确局限时,偶尔还需要OpenAI亲自对模型做微调。

5. 自主性的悖论:Waymo能跑,你的订票Agent却不行

  • Apoorv提出的问题是:物理安全的标准高于人类能力,毕竟牵涉生命,但“2025年的物理自主性领先于数字自主性”。听起来更容易的问题,为什么反而更难?
  • Sherwin给出两部分答案。第一是时间线:自动驾驶经历了10-15年发展,其中包括一段幻灭低谷;而Agent真正起源于不到1年前的o1-preview推理范式,“但我认为这条斜率陡得惊人”。第二是脚手架:汽车拥有道路、红绿灯和交通法规,而“AI agent就像被直接丢在荒郊野外”。
  • 可交易的推论是:失败的企业部署“很可能没有脚手架”。大量FDE工作其实是在搭建连接器、整理数据,让模型拥有可以标准化交互的对象。而交叉点可能已经出现:“如果现在AI agent产品的收入超过Waymo,我不会感到意外。”

6. GPT-5:工艺胜过基准测试,以及思考时间的取舍

  • Olivier的框架是:类似“SWE-bench”的基准测试分数很高,但“同样重要、同样有影响力的是模型的工艺——风格、语气和行为”。这是首个经过数月客户反馈闭环打造的大型版本;客户很多时候要的不是更高的智力,而是“在不知道答案时更可能说不知道的模型”。
  • Sherwin说,仍在持续迭代的最难取舍是推理token与延迟。GPT-5 Pro能“一次解决”其他模型无法处理的问题(Sam转发的Andrej推文),但“代价是你要等10分钟”;对商业用户而言,“完全不用等待但答案稍差”可能更能接受。思考过程应当变得更动态,下一个快照还会针对代码质量和惯用写法。
  • 在稳健性方面,一位主持人称某项评测显示幻觉“基本降到了零”——“并不完美,仍有大量工作要做”,但推理能力让模型更可能拒答,而不是编造。

7. 指令遵循的猴爪效应

  • 开发者要求更好的指令遵循,得到的是一个几乎“逐字照办”的模型。测试中“太简洁”的抱怨,后来被追溯到旧提示词:其中连续10行要求“保持简洁”,几乎是在乞求模型这么做。GPT-5严格执行后,只返回一句极短的回答。删掉这些历史包袱后,模型行为就恢复正常。“提示词工程仍然非常、非常重要。”
  • Apoorv还提到一个对投资人有意义的旁证:被投网络安全公司Expo在GPT-5上的表现提升如此明显,以至于“他们很快就需要一套新的评测”,而回应是:“一切都在于评测。”

8. 语音到语音胜过拼接,但逻辑层尚未统一

  • 实时API之所以优于“拼接模式”(语音转文字→思考→文字转语音),是因为额外环节会增加延迟,而且会“丢掉情绪”——电话中的停顿和语气都很重要。Apoorv反驳说,函数调用逻辑是用文字编写的,因此语音意味着“略有不同的架构”;跨模态统一编排仍在推进,许多客户仍采用拼接方案。
  • 语音模型擅长“轻松的日常对话——和教练、治疗师聊天”,但必须被教会具有经济价值的行为:知道什么是SSN;如果某一位数字听不清,“就必须重新确认,而不是猜一个答案”。
  • Sherwin感叹道:“你实际上是在拿一个人说话时的音频比特……然后生成另一段音频比特。在我看来,这项技术居然能工作本身就很疯狂”,更不用说还要处理口音、语气和实时客服通话。

9. RFT:在自己的数据上启动RL

  • 强化微调是Olivier所说由OpenAI创造的术语(“在我们宣布它之前,它根本不是一个真实存在的东西”),于去年年底推出,可能是在“圣诞节12天”活动期间,随后正式GA。它把RL引入微调:“复杂得多,也更难调,但比SFT强一个数量级”。它需要可评分的任务和客观评分器,而不是提示词—补全配对。
  • 这里的重新定义是:不是定制模型,而是利用专有数据,“为你关心的业务场景打造一个世界上最好的模型”。案例包括金融服务公司Rogo(团队来自DeepMind)处理金融文件任务;以及Accordance,Olivier称其在面向CPA工作的TaxBench上取得了他认为属于SOTA的结果。
  • Apoorv的判断是,基础模型已经足够好,没必要再通过微调来引导行为;但如果目标是推动前沿能力,“我的直觉是,RFT基本会成为行业常态”,并配合定制环境。所需数据必须来自底层:由领域专家提供复杂的任务知识。

10. 多空、编程未来,以及那次插曲

  • Sherwin看多电竞:亚洲已经具备体育场级别的规模,年轻人的注意力正在迁移过去,而经历COVID行情后,它仍“被低估”。他给出的“辛辣”做空是“围绕AI产品的整个工具链”——2年前是评测产品、框架和向量数据库,如今则是RL环境创业公司——因为“今天炙手可热的框架,下一代模型可能根本不会用”。Olivier的组合是:做空任何强调记忆的教育(LLM已经拥有的“知识token”),做多医疗——“可能是最受AI影响、受益最大的行业”,因为这里同时具备结构化数据、行政工作密集的文化,以及适合研发的企业。
  • 对于10年后软件工程师是否会从4000万-5000万人增长这个问题,两人的答案都是:软件工程一定会更多,只是岗位名称未必清晰。一个Reddit故事讲的是,一名男子用ChatGPT为无法用语言表达的兄弟打造定制工具,让他能够浏览互联网。Apoorv说,“世界上存在巨大的软件短缺”;Olivier说,OpenAI的产品经理如今几小时内就能交付写好的原型,而不是PRD。个人工具选择上,两人都投了Granola;至于Codex CLI与GPT-5,Sherwin说:“我感觉自己已经和模型心意相通。”
  • 荆棘与玫瑰: “那次插曲”(董事会政变)当天极其惨烈,却让OpenAI变得反脆弱——“更抗打,也能更快恢复”。此外,前一年11月或12月曾发生过一次总计3-4小时的中断,证明API“几乎已经像公用事业”。玫瑰则包括:以大规模token吞吐量发布GPT-5且没有中断,以及2023年11月的开发者日——现场演示成功后,Sherwin搭乘Waymo回家。
  • 两人都说自己已经被AGI彻底感染(“我觉得自己已经AGI-pilled了。”——“你绝对已经AGI-pilled了。”)。Olivier的关键时刻包括:2023年意识到自己“以后再也不需要手动写代码了”,以及模型能听懂他的法国口音;Sherwin则是在2022年9月、ChatGPT发布前于内部见到GPT-4,以及体验深度研究:“我会把某件事丢给模型,心想这东西不可能做得到。结果它却直接把球打出了公园。”

Apoorv Agarwal

In San Francisco, you could take a car from one part of San Francisco to the other fully autonomously. As opposed to the digital world, I can't book a ticket online right now. Physical autonomy is ahead of digital autonomy in 2025.

Sherwin Wu

I think AI agents are really in day one here. ChatGPT only came out in 2022. The slope, I think, is incredibly steep.

Apoorv Agarwal

I actually do think self-driving cars have a good amount of scaffolding in the world. You have roads. Roads exist. They're pretty standardized. You have stoplights. AI agents are just dropped in the middle of nowhere.

We'll start with the long-short game. I'm short on the entire category of tooling and evals products.

Olivier Godement

Healthcare is probably the industry that will benefit the most from AI. I think I'm AGI-pilled.

Apoorv Agarwal

You're definitely AGI-pilled.

Hey, folks. I'm Apoorv Agarwal, and today, at the OpenAI office, we had a wide-ranging conversation about OpenAI's work in enterprise. I have with me the head of engineering and the head of product of the OpenAI Platform, Sherwin Wu and Olivier Godement.

OpenAI is well known as the creator of ChatGPT, a product that billions across the world have come to love and enjoy. But today, we dive into the other side of the business, which is OpenAI's work in enterprise. We go deep into their work with specific customers and how OpenAI is transforming large and important industries like healthcare, telecommunications, and national security research.

We also talk about Sherwin and Olivier's outlook on what's next in AI, what's next in technology, and their picks on both the long and short side. This was a lot of fun to do. I hope you really enjoy it.

1. OpenAI’s Enterprise Mission: Beyond ChatGPT

Well, 2 world-class builders—2 people who make building look easy. Sherwin, my Palantir 2013 classmate and tennis buddy, with 2 stops at Quora and Opendoor through the IPO before joining OpenAI, before ChatGPT, you've now been here for 3 years and lead engineering for all of OpenAI Platform.

Olivier, former entrepreneur, winner of the Golden Llama at Stripe, where you were for just under a decade, and now lead all of the product at OpenAI Platform.

Olivier Godement

That's right.

Apoorv Agarwal

Thanks for doing it.

Olivier Godement

Thank you. Thanks for having us.

Apoorv Agarwal

As a shareholder and as a thought partner, kicking ideas back and forth, I always learn a lot from you guys. And so it's a treat. It's a real treat to do this for everybody.

I'll open with this: People know OpenAI as the firm that built ChatGPT, the product that they have in their pocket, that comes with them every day to work and into their personal lives. But the focus for today is OpenAI for enterprise. You guys lead OpenAI Platform. Tell us about it. What's underneath the OpenAI Platform for B2B and for enterprise?

Sherwin Wu

Yeah. So this is actually a really interesting question, too, because when I joined OpenAI around 3 years ago to work on the API, it was actually the only product that we had. I think a lot of people forget this: The original product for OpenAI was not ChatGPT. It was a B2B product—the API—and we were catering toward developers. I've actually seen the launch of ChatGPT and everything downstream from that.

But at its core, I think the reason why we have a platform and why we started with an API comes back to the OpenAI mission. Our mission, obviously, is to build AGI, which is pretty hard in and of itself, but also to distribute the benefits of it to everyone in the world—to all of humanity. It's pretty clear right now to see ChatGPT doing that because my mom, maybe even your parents, are using ChatGPT.

But we actually view our platform, and especially our API and how we work with our enterprise customers, as our way of getting the benefits of AGI, of AI, to as many people as possible—to everyone in every corner of the world.

ChatGPT is really, really, really big now. It's, I think, the 5th-largest website in the world. But by working through developers using our API, we're actually able to reach even more people in every corner of the world and every different use case that you might have. Especially with some of our enterprise customers, we're able to reach use cases within businesses and the end users of those businesses as well. And so we view the platform as our way of fully expressing our mission of getting the benefits of AGI to everyone.

Concretely, what the platform actually includes today: The biggest product we have is obviously our developer platform, which is our API. Many developers—the majority of the startup ecosystem—build on top of this, as well as a lot of digital natives and Fortune 500 enterprises at this point.

We also have a product that we sell to governments in the public sector. That's all part of this as well. An emerging product line for us in the platform is our enterprise product. We might sell directly to enterprises beyond just a core API offering.

Apoorv Agarwal

Fascinating. And maybe to double down, I think B2B is actually quite core to the OpenAI mission. What we mean by distributing AGI benefits is this: I want to live in a world where there are 10x more medicines going out every year. I want to live in a world where education, public service, and civil service are increasingly optimized for everyone.

There is a large category of use cases that only go through B2B, frankly, unless you enable the enterprises. And we talk about Palantir—I think that's probably the same thesis at Palantir. Those are the businesses that are actually making stuff happen in the real world. So if you do enable them, if you do accelerate them, that's how, essentially, you distribute the benefits of AGI.

Yeah. Well, maybe we can double-click into that, Olivier. The reach for ChatGPT is obviously wide—billions of users. But for enterprise, maybe tell us about it. Maybe we go deep into a customer example or 2. What is an organization that we have helped transform, maybe, and at what layers?

2. Case Study: T-Mobile - Voice & Support

Olivier Godement

If I were to step back, we started our B2B efforts with the API a few years ago. Initially, the customers were startups, developers, indie hackers, and extremely technically sophisticated people who were building cool new stuff and taking massive market risk. We still have a bunch of customers in that category, and we love them and keep building with them.

On top of that, over the past couple of years, we've been working more and more with traditional enterprises and also digital natives. Basically, everyone woke up with ChatGPT, and those models are working. There's a ton of value, and they could see many use cases in the enterprise.

A couple of examples that I like the most. One, which is both very fresh and quite cool, is that we've been working a lot with T-Mobile, a leading U.S. telco operator. T-Mobile has a massive customer-support load: People asking, "Hey, I was charged that amount of money. What's going on?" Or, "My cell phone isn't working anymore."

A massive share of that load is voice calls. People want to talk to someone. And so for them, to be able to essentially automate more and more and help people self-serve—to debug their subscription—was pretty big. We've been working with T-Mobile pretty much for the past year to basically automate not only the text support but also voice support.

Today, there are features in the T-Mobile app that, if you call, are actually handled by OpenAI models behind the scenes. It does sound supernatural—human-sounding, latency-wise and quality-wise. So that one was really fun.

Apoorv Agarwal

Just on that, can I ask you a follow-up question? We've got text models. We've got voice models. Maybe even video models someday that are deployed at T-Mobile. But what above the models, or adjacent to the models, might we have helped T-Mobile with, for example?

Olivier Godement

Yeah, there's a ton we're doing.

The first one is—you have to put yourself in the shoes of an enterprise buyer. Their goal is to automate, reduce, and optimize customer support. You're going from a model—tokens in, tokens out—to that case, and it's hard.

First, there's a lot of design—system design. We do have forward-deployed engineers now who are helping us quite a bit.

Bill Gurley

Forward-deployed engineers.

Olivier Godement

We borrow the term from Palantir.

Bill Gurley

Yeah, it's a great term. Were you an FD at Palantir?

Olivier Godement

I was not an FD. I was on, I think they called it the dev side, right? It's like software engineering. I was also only an intern at Palantir. But, yeah, it's a great term. I think it accurately describes what we're asking folks to do, which is embed very deeply with customers and, honestly, build things specific to their systems. They're deployed onto these customers.

But, yeah, we are obviously growing and hiring that team quite a bit because they've been very effective at T-Mobile.

Bill Gurley

Four years of my life. Yeah, yeah, yeah. Forward deployed. But go ahead.

Olivier Godement

So, forward-deployed engineering—the sort of systems and integrations they're doing—is, first, about orchestrating those models. Those models are not just models; they know nothing about the CRM or what's going on. You have to plug the model into many, many different tools.

Many of those tools in the enterprise don't even have APIs or clean interfaces, right? It's the first time they're being exposed to a third-party system. So, there's a lot of standing up API gateways and tools, and connecting them.

Then you have to essentially define what good looks like. Again, to put in your exercise for everyone, defining a golden set of evals is easier than it sounds.

Bill Gurley

Harder than it sounds.

Olivier Godement

Yeah. And so, we've been spending a bunch of time with them.

Bill Gurley

Evals are important. Evals are super important.

Olivier Godement

Especially audio evals. Evals are extra hard to grade and get right. The bulk of the use case here is actually audio. We have, I don't know, five-minute call transcripts—how do you actually know that the right thing happened? It's a pretty tough problem.

Bill Gurley

Yeah, it's pretty tough.

Olivier Godement

And then actually nailing down the quality of the customer experience until it feels natural. Latency and interruptions are really important parts of that.

We shipped a Realtime API in GA. I think it was last week.

Bill Gurley

A couple of weeks ago, yeah.

Olivier Godement

Yeah, it was just last week, I think. It's a beautiful work of engineering. There was a really cracked team behind the scenes, which basically allows us to get the most natural-sounding voice experience without having these weird interruptions or lag where you can feel that the thing is off.

So, cobbling all that together, you get a really good experience.

Bill Gurley

Yeah, that's a lot more than just models.

Sherwin Wu

One really great thing that I think we've gotten from the T-Mobile experience is working with them to improve our models themselves. For example, in the last Realtime GA last week, we obviously released a new snapshot—the GA snapshot. A lot of the improvements that we got into the model came out of the learnings that we had from T-Mobile.

It brings in a lot of other changes from other customers, but because we were so deeply embedded into T-Mobile and were able to understand what good looks like for them, we were able to bring that to some of our models.

Bill Gurley

That makes sense. So, we are working with a large customer with tens of millions of users, if not hundreds of millions, and the before and after is on the support side—both tech support internally and then their customer support.

Yeah. Is there another one that you guys can share? I like Amgen a lot. Amgen, the healthcare business.

3. Case Study: Amgen - Accelerating Drug Development

Olivier Godement

Amgen, yeah. We are working quite a bit with healthcare companies. Amgen is one of the leading healthcare companies. They specialize in drugs for cancer or inflammatory diseases, and they're based out of L.A.

We've been working with Amgen to speed up drug development and the commercialization process. The north star is pretty bold. It's really interesting when you similarly embed pretty deeply with Amgen to understand what their needs are.

When I look at healthcare companies, I feel like there are 2 big buckets of needs. One is pure R&D. You're seeing a massive amount of data, and you have super-smart scientists who are trying to come up with and test out things. That's one bucket.

A second bucket is much more common across other industries. It's pure admin, document-authoring, and document-scribing work. By the time your R&D team has essentially locked the recipe of a medication, getting that medication to market is a ton of work. You have to submit to various regulatory bodies and get a ton of reviews.

When we looked at those problems and what we knew models were capable of, we saw a ton of benefits and a ton of opportunities to automate and augment the work of those teams. Amgen has been a top customer of GPT-5, for instance.

Bill Gurley

Wow. I mean, this could be hundreds of millions of lives if a new drug is developed faster.

Olivier Godement

Yeah, exactly. Huge impact. So that's, I think, one good example of the kind of impact you need to enable enterprises to create. And so, I think we're going to do more and more of those.

Frankly, on a personal level, it's a delight. If I can play a tiny role in essentially doubling the kind of medication that people get in the real world, that feels like a pretty good achievement.

Bill Gurley

Huge. Huge, huge. I know you had one as well.

4. Case Study: Los Alamos National Lab

Sherwin Wu

One of my favorite deployments that we've done more recently is with the Los Alamos National Laboratory. This is the government national research lab that the U.S. government runs in Los Alamos, New Mexico. It's also where the Manhattan Project happened back in the ’40s and ’50s, when it was a secret project.

After that, they ended up formalizing it as a city and a program, and now it's a pretty sizable national laboratory. This one is very interesting because, first, the depth of impact here is unimaginable to me. It's on the scale of Amgen and some of these other larger companies.

Obviously, they're doing a lot of actual new research there, so a lot of new science. They're doing a lot of work with our Defense Department and defense use cases as well. It's very intense stuff.

But the other thing that's very interesting about this one is that it's also a story of a very bespoke and new type of deployment that we've done. Because they're a government lab, they're so restrictive and high-security and high-clearance with a lot of their work, we couldn't just do a normal deployment with them. You can't have people doing national-security research just hitting our APIs.

And so, we actually did a custom, on-premises deployment with them onto one of their supercomputers, called Venado. This involved a bunch of very bespoke work with some FDEs, as well as with a lot of our developer team, to actually bring one of our reasoning models, o3, into their laboratory, into an air-gapped supercomputer—Venado—and deploy it and get it installed and working on their hardware, on their networking stack, and actually run it in this particular environment.

And so it was actually very interesting because we literally had to bring the weights of the model physically into their supercomputer, in an environment, by the way, that’s very locked down for a good reason: they’re not allowed to have cell phones or any electronics with them. So I think that was a very unique challenge.

The other interesting thing about this deployment is how it’s being used. Because it’s so locked down and on-premises, we actually don’t have much visibility into exactly what they’re doing with it, but they do give us feedback. They actually do have some telemetry, but it’s within their own systems.

We do know that it’s being used for a bunch of different things. It’s being used to help speed up their experiments. They have a lot of data-analysis use cases and a lot of notebooks that they’re running with reams of data that they’re trying to process.

They’re actually using it as a thought partner, which is something that’s pretty interesting to me. o3 is a pretty smart model, and a lot of these people are tackling really tough, novel research problems. A lot of the time, they’re using o3 and going back and forth with it on their experiment design and what they should actually be using it for, which is something that we couldn’t really say about our older models.

So it’s being used for a lot of different use cases for the National Lab. The other cool thing is that it’s actually being shared between Los Alamos and some of the other labs—Lawrence Livermore and Sandia as well—because it’s a supercomputer setup where they can all connect to it remotely.

5. Why 95% of AI Deployments Fail?

Apoorv Agarwal

Fascinating. We’ve just gone through 3 pretty large-scale enterprise deployments, which might touch tens, if not hundreds, of millions of people. But on the other side of this is the MIT report that came out a couple of weeks ago: 95% of AI deployments don’t work. There were a bunch of scary headlines that even shook the markets for a couple of days.

Put this in perspective. For every deployment that works, there’s presumably a bunch that don’t work. So maybe we can talk about that. What does it take to build a successful enterprise deployment, a successful customer deployment, and the counterfactual, based on all your experience serving all these large enterprises?

Olivier Godement

I think at that point, I may have worked with a couple hundred. So, okay, I’m going to pattern-match. What I’ve seen as a clear leading indicator of success is, number 1, the interesting combination of top-down buy-in and enabling a very clear group—a tiger team, essentially—at the enterprise, which is sometimes a mix of OpenAI and enterprise employees.

Typically, take T-Mobile. The top leadership was extremely bought in: “It’s a priority.” But then they let the team organize and say, “Okay, if you want to start small, start small,” and then you can scale it up, essentially. That would be part number 1: top-down buy-in and a bottom-up tiger team.

The tiger team is made up of people with a mix of technical skills and people who just have organizational knowledge—institutional knowledge. It’s really funny: in the enterprise, customer support is a good example. What we found is that the vast majority of the knowledge is in people’s heads.

You would think that, in customer support, everything is perfectly documented. The reality is that the standard operating procedures, the SOPs, are largely in people’s heads. Unless you have that mix of technical people and subject-matter experts, it’s really hard to get something off the ground.

That would be number 1. Number 2 would be evals first. Whatever we define as good evals, that gives people a clear common goal to hit. Whenever the customer fails to come up with good evals, it’s a moving target, essentially, whether you’ve made it or not.

Evals are much harder to get done than they look, and they also oftentimes need to come from the bottom up, because all of these things are in people’s heads—in the actual operators’ heads. It’s actually very hard to have a top-down mandate of, “This is how the evals should look.” A lot of it needs bottom-up adoption.

Apoorv Agarwal

Right. Yeah. Yeah.

Olivier Godement

And so we’ve been developing quite a bit of tooling for evals. We have an evals product, and we’re working on more to essentially solve that problem, or make it as easy as we can.

The last thing is that you want to hill-climb, essentially. You have your evals; the goal is to get to 99%. You start at 46%. How do you get there? Frankly, I think oftentimes it’s a mix of—I will say—almost wisdom from people who’ve done it before. A lot of that is art, sometimes more than science: knowing the quirks of the model and its behavior.

Sometimes we even need to fine-tune the models ourselves when there are some clear limitations, and be patient, get your score up there, and then ship.

6. Physical vs Digital Autonomy: Scaffolding & Infrastructure

Apoorv Agarwal

Can we go under the hood a little bit? One of the things that we think about a lot is autonomy more broadly. What is the makeup of autonomy? On one side, in San Francisco, you could take a car from one part of SF to the other fully autonomously. No humans involved. You press a button. We love Waymo. They’ve done billions of miles. I think it was, what, 3.5 billion miles on Tesla FSD. I think they’ve almost done tens of millions of rides.

That’s a lot of autonomy in the physical world, as opposed to the digital world. I can’t book a ticket online right now. There are all sorts of problems that happen if I have my operator try to book a ticket.

It’s very counterintuitive because the bar for physical safety is so much higher. The bar for physical safety is higher than the human’s capability because lives are at stake. The bar for digital safety is not that high, because all you’re going to lose is money. Nobody’s life is at stake.

Yet physical autonomy is ahead of digital autonomy in 2025, which seems counterintuitive. Why is that the case at a technical level? Why is it that what should sound easier is actually a lot harder?

Sherwin Wu

Yeah, so I think there are 2 things at play here. I really like the analogy with self-driving cars because they’ve actually been one of the best applications of AI that I’ve used recently. One of them is honestly just the timelines.

We’ve been working on self-driving cars for so long. I remember back in 2014, it was kind of the advent of this, and everyone was like, “Oh, it’s happening in 5 years.” It turns out it took 10 or 15 years or so for this to happen.

There was probably a dark age back in 2015 or 2018 or something, when it felt like it wasn’t going to happen.

Brad Gerstner

A trough of disillusionment.

Apoorv Agarwal

Yes, yes, yeah. And then now we’re finally seeing it get deployed, which is really exciting. But it has been, I don’t know, 10 years, maybe even 20 years from the very beginning of the research.

Whereas I think AI agents are really in day 1 here. ChatGPT only came out in 2022, so around 3 years—less than 3 years ago. I actually think what we think about with AI agents and all that really started with the reasoning paradigm, when we released the o1-preview model back in late last year, I think.

And so I actually think this whole reasoning paradigm with AI agents and the robustness they bring has only really unfolded for about a year—less than a year, really.

And so, I know you had a chart in your blog post that I really like, where the slope is very meaningfully different now. Self-driving started very, very early. The slope seems to be a little bit slower, but now it’s reaching the promised land. But, man, we started super recently with AI agents, and the slope, I think, is incredibly steep, and we’ll probably see a crossover at some point.

But we really have only had about a year to explore these things. Do you think we haven’t crossed over already when you look at the coding work in particular?

Olivier Godement

Yeah, it’s a good point. Your chart actually shows AI agents below self-driving, but what is the Y-axis? If you look at some measures, I would not be surprised if AI-agent products are making more revenue than Waymo at this point. Waymo was making a lot, but just look at all the startups coming up. Look at ChatGPT, how many subscriptions are happening there, and all of that.

Maybe we have actually crossed over, and a couple of years from now, it’s going to look very, very different.

Apoorv Agarwal

Yeah, the Y-axis is tangible, felt autonomy. Don’t pick the objective—how do I feel about it?

Olivier Godement

Exactly. It vibes more than revenue. But revenue is a good one. We should probably redo that with revenue.

Sherwin Wu

There’s a second thing I wanted to mention as well, which is the scaffolding and the environment in which these things operate. I remember in the early days of self-driving, a lot of the researchers around self-driving were saying that the roads themselves would have to change to accommodate self-driving. There might be sensors everywhere so that the self-driving cars could interact with them, which, in retrospect, I think was overkill.

But I do think self-driving cars have a good amount of scaffolding in the world for them to operate in—not completely unlimited. You have roads. Roads exist, and they’re pretty standardized. You have stoplights. People generally operate in pretty normal ways, and there are all these traffic laws that you can learn.

Whereas AI agents are just dropped in the middle of nowhere, and they kind of have to feel around for things. Going off of what Olivier just said, my hunch is that some of the enterprise deployments that don’t actually work out likely don’t have the scaffolding or infrastructure for these agents to interact with, either.

A lot of the really successful deployments that we’ve made—a lot of what our FDEs end up doing with some of these customers—is creating almost like a platform or some type of scaffolding, building connectors, and organizing the data so that the models have something they can interact with in a more standardized way. My sense is that self-driving cars have actually had this to some degree, with roads, over the course of their deployment.

But I think it’s still very early in the AI-agent space. I would not be surprised if a lot of enterprises and companies just don’t really have the scaffolding ready. So if you drop an AI agent in there, it doesn’t really know what to do, and its impact will be limited.

Once this scaffolding gets built out across some of these companies, I think deployment will also speed up. But, again, to our point earlier, I think there’s no slowdown. Things are still moving very fast.

Bill Gurley

That’s great. Well, you know, I’ve thought about autonomy as a three-part structure. You’ve got perception, you’ve got the reasoning—the brain—and then you’ve got the scaffolding, the last mile of making things work.

7. GPT-5: Release, Benchmarks vs Behavior

Maybe we can dive into the second part, which is the reasoning—the juice that you guys are building with GPT-5 most recently. Huge endeavor. Congrats. It’s the first time you guys have launched a full system, not a model or a set of models, but a full system.

Talk about that—the full arc of that development. What was your focus? Honestly, the benchmarks all seem so saturated. Clearly, it was more than just benchmarks that you were focused on. What is a North Star? Tell us about GPT-5, soup to nuts.

Olivier Godement

It’s been the labor of love of many people for a long time. To your point, I think GPT-5 is amazingly intelligent. You look at the benchmarks, like [SWE-bench?] and the like, and it’s going pretty high.

But to me, equally important and impactful was the craft—the style, the tone, and the behavior of the model. So: capabilities, intelligence, and behavior of the model.

On the behavior of the model, I think it’s the first large-model release for which we’ve worked so closely with a bunch of customers, month after month, to better understand the concrete blockers. Often, it’s not about having a model that is way more intelligent; it’s about having a model that better follows instructions and is more likely to say no when it doesn’t know something.

That super-close customer feedback loop on GPT-5 was pretty impressive to see. I think all the love that GPT-5 has been getting in the past couple of weeks means people are starting to feel that from the builders, essentially. Once you see it, it’s really hard to come back to a model that is extremely intelligent but is essentially an exquisite academic.

Apoorv Agarwal

Are there trade-offs that you made as you were going through it? What were the hardest trade-offs you made as you were building GPT-5?

Sherwin Wu

I actually think a very clear trade-off, which I honestly think we’re still iterating on, is the trade-off between the reasoning tokens—how long it thinks—and performance.

This is something we’ve been working on with our customers since the launch of the reasoning models. These models are so smart, especially if you give them all this thinking time. The feedback I’ve been seeing around GPT-5 Pro has been pretty crazy, too. There are these unsolved problems that none of the other models could handle; you throw them to GPT-5 Pro, and it just one-shots them. It’s pretty crazy.

But the trade-off here is that you’re waiting for 10 minutes. That’s quite a long time. These things just get so smart with more inference time, but for the product builder on the API side, for some of these business use cases, it’s pretty tough to manage that trade-off.

For us, it’s been difficult to figure out where we want to fall on that spectrum. We’ve had to make some trade-offs on how much the model thinks versus how intelligent it should get. As a product builder, there’s a real latency trade-off that you have to deal with. Your user might not be happy waiting 10 minutes for the best answer in the world. They might be more okay with a substandard answer and no wait at all.

Brad Gerstner

Yeah, I mean, even between GPT-5 and GPT-5 Thinking, I have to toggle it now because sometimes I’m so impatient I just want it ASAP. I think there’s an ability to skip, right?

Sherwin Wu

Yeah, that’s right.

Brad Gerstner

And GPT-5 is like, “I’m impatient. I just want a simpler answer.”

Sherwin Wu

That’s right, that’s right.

Brad Gerstner

Well, 4 weeks in, GPT-5—how’s the feedback?

Sherwin Wu

I think feedback has been very positive, especially on the platform side, which has been really great to see. A lot of the things that Olivier mentioned have been coming up in feedback from customers.

The model is extremely good at coding and extremely good at reasoning through different tasks. Especially for coding use cases, when it thinks for a while, it’ll usually solve problems that no other models can solve. I think that’s been a big positive point of feedback.

Brad Gerstner

Yeah, yeah, yeah.

8. GPT-5 Feedback: Instruction Following, Hallucinations, Code Quality

Sherwin Wu

I think there’s an eval that showed that hallucinations basically went to zero for a lot of this. It’s not perfect—there’s still a lot of work to be done—but because of the reasoning in there, it just makes the model more likely to say no and less likely to hallucinate answers. That’s been something that people have really liked as well.

Another bit of feedback has been around instruction-following. It’s really good at instruction-following.

This almost bleeds into the constructive feedback that we're working on. It's so good at instruction following that people need to tweak their prompts; it's almost too literal. That's an interesting trade-off because when you ask developers what they want, they want the model to follow instructions, of course. But once you have a model that is extremely literal, that forces you to express extremely clearly what you want; otherwise, the model may go sideways.

That feedback was interesting. It's almost like the monkey's paw: developers and platform customers ask for better instruction following, and we're like, “Yes, we'll give you really good instruction following.” But it follows instructions almost to a T, so it's obviously something that the team is working through.

A good example of this, by the way, is that some customers would have these prompts. I remember when we were testing GPT-5, one piece of negative feedback we got was that the model was too concise. We were like, “What's going on? Why is the model so concise?” Then we realized it was because they were using their old prompts from other models. With the other models, they have to really beg the model to be concise.

There are 10 lines of “Be concise. Really be concise. Also, keep your answer short.” It turns out that when you give that to GPT-5, it's like, “Oh my gosh, this person really wants it to be concise.” The response would be one sentence, which is too terse. Just by removing the extra prompts around being concise, the model behaved in a much better way and much closer to what they actually end up wanting.

Brad Gerstner

Yeah, it turns out writing the right prompt is still important.

Sherwin Wu

Yes. Prompt engineering is still very, very important. On constructive feedback for GPT-5, there's actually been a good amount as well, which we're all working through. One of the things that I'm really excited for the next snapshot to fix is code quality and small code paradigms or idioms that it might use. I think there was feedback around the types of code and the patterns that it was using, which I think we're working through as well.

The other bit of feedback, which I think we've already made good progress on internally, is around the trade-off between reasoning tokens, thinking, latency, and intelligence. Especially for simpler problems, you don't usually need a lot of thinking. The thinking should ideally be a little bit more dynamic. Of course, we're always trying to squeeze as much reasoning and performance into as few reasoning tokens as possible. So I'd imagine that kind of going down as well.

Brad Gerstner

Yeah. Well, huge congrats. I know it's a work in motion for a bunch of our companies. They've had incredible outcomes with GPT-5. One of them is Expel, a cybersecurity business; it's a huge—

Sherwin Wu

Yeah, I saw the chart from that. It was pretty crazy.

Brad Gerstner

Huge, huge upgrade from whatever they were using prior to that. I think they're going to need a new eval soon.

Sherwin Wu

That's right. They're going to need a new eval.

9. Multimodality: Text, Voice, and Video

Brad Gerstner

It's all about evals. On the multimodality side, obviously you guys announced the real-time API last week. I saw T-Mobile was one of the featured customers on there. Talk about that. Obviously, the text models are leading the pack, but then we've got audio and video. Talk about the progress on the multimodal models. When should we expect to have the next big unlock, and what would that look like?

Sherwin Wu

It's a good question. The teams have been making amazing progress on multimodality: voice, image, and video, frankly. The last-generation models have been unlocking quite a few cool use cases. One piece of feedback that we've received is that, because text was leading the pack so much in intelligence, people felt like, in voice, the model was somewhat less intelligent. Until you actually see it, it does feel weird to have a better answer on text versus voice. That's pretty much the focus that we have at the moment.

I think we filled part of that gap, but not the full gap, for sure. Catching up with text would be one. A second one, which is absolutely fascinating, is that the model is excellent at easy, casual conversation—talking to your coach or your therapist. We basically had to teach the model to speak better in actual, economically valuable work setups.

To give an example, the model has to be able to understand what an SSN is and how to spell an SSN. If one digit is fuzzy, it actually has to repeat it rather than guess. There are lots of inflections like that that someone has in their voice that we are currently teaching the model. That's ongoing work with our customers. Until we actually confront the model with actual customer-support calls at actual scale, it's really hard to get a feel for those gaps. That's a top priority as well.

10. Audio: Realtime API vs Stitched Audio

Apoorv Agarwal

This is completely off script, but an interesting question that comes up in voice models, particularly the real-time API, is that previously people were taking a speech input, converting that to text, then having some layer of intelligence. Then you would have a text-to-speech model that would play it back. It would be a stitch of these 3 parts. The real-time API—you guys have integrated all of that. How does it happen? A lot of the logic is written in text. A lot of the Boolean logic or any function calling is written in text. How does it work with the real-time API?

Olivier Godement

That's an excellent question. The reason why we built the real-time API is that we saw a couple of issues with the stitched model. We call it a stitched-together model: speech-to-text, thinking, and text-to-speech. We saw essentially a couple of issues. One: slowness, with more hops, essentially. Two: loss of signal. With a stitched model, the speech-to-text model is less intelligent.

Brad Gerstner

Yeah, you'd lose the emotion.

Sherwin Wu

Exactly—pauses. When you're doing actual voice, like phone calls, those signals are so important. One of the challenges we have is what you mentioned: it means a slightly different architecture for text versus voice. That's something we're actively working on. But I think it was the right call to start with: let's make the voice experience sound natural to a point where you feel comfortable putting it in production, and then work backward to unify the orchestration logic across modalities.

A lot of customers still stitch these together. It's what worked in the last generation. But what we're interested in seeing is more and more customers moving toward the real-time approach because of how natural it sounds and how much lower the latency is, especially as we uplevel the intelligence of the model.

Bill Gurley

But also, even taking a step back, I will say it's pretty mind-blowing to me that it works. I think it's mind-blowing that these LLMs work at all: you just train them on a bunch of text, and they're autoregressively coming up with the next token, and it sounds super intelligent. That's mind-blowing in and of itself.

But I think it's actually even more mind-blowing that this speech-to-speech setup actually works correctly because you're literally taking the audio bits from someone speaking, streaming them, putting them into the model, and then it's generating audio bits back. To me, it's actually crazy that this works at all, let alone the fact that it can understand accents, tone, pauses, and things like that, and then also be intelligent enough to handle a support call or something like that. If you've gone from text-in, text-out to voice-in, voice-out, that's pretty crazy.

11. Model Customization & Reinforcement Fine-Tuning (RFT)

Apoorv Agarwal

We have a bunch of companies in our portfolio that are using these models: Parloa on the customer-support side, LiveKit on the infrastructure side. There are a bunch of use cases we're starting to see that a speech-to-speech model could address. A lot of the harder ones are still running on what you're calling the “stitched model.” But I hope the day is not far when it's all on the real-time API.

Sherwin Wu

It's going to happen at some point.

Brad Gerstner

Right, right, right. And actually, maybe that’s a good segue into talking about model customization because I suspect that you have such a wide variety of enterprise customers. I think you mentioned hundreds of customers, or maybe more. Each of them has a different use case, a different problem set, a different goal and envelope of parameters that they’re working in—maybe latency, maybe power, maybe others. How do you handle that? Talk about what OpenAI offers enterprises that need a customized version of a great model to make it great for them.

Olivier Godement

Yeah, model customization has actually been something that we’ve invested very deeply in on the API platform since the very beginning. Even before ChatGPT, we had a supervised fine-tuning API available, and people were using it to great effect. The most exciting thing around model customization, I think, obviously resonates quite well with customers because they want to bring in their own custom data and create their own custom version of o3, o4-mini, or even GPT-5, suited to their own needs. It’s very attractive, but the most recent development, which I think is very exciting, has been the introduction of reinforcement fine-tuning. It was something we announced late last year, I think during the 12 Days of OpenAI. We’ve since GA’d it, and we’re continuing to iterate on it.

Brad Gerstner

What is it? Break it down for us.

Olivier Godement

It’s called reinforcement fine-tuning. It’s actually funny—I think we made up the term reinforcement fine-tuning. It wasn’t a real thing until we announced it.

Brad Gerstner

It’s stuck now. I see it all the time. I remember we were discussing it, and I was like, “I don’t know about RFT.”

Olivier Godement

You’re not kidding. You’re not kidding. So, reinforcement fine-tuning introduces reinforcement learning into the fine-tuning process. The original fine-tuning API does something called supervised fine-tuning—we call it SFT. It does not use reinforcement learning; it uses supervised learning. Usually, that means you need a bunch of data, a bunch of prompt-completion pairs. You need to really supervise and tell the model exactly how it should be acting, and then, when you train it on our fine-tuning API, it moves the model closer in that direction.

Reinforcement fine-tuning introduces RL, or reinforcement learning, into the loop. It’s way more complex, way more finicky, but an order of magnitude more powerful. That’s what has really resonated with a lot of our customers. With RFT, the discussion is less about creating a custom model that’s specific to your own use case. You can use your own data and turn the crank on RL to create a best-in-class model for your particular use case. That’s the main difference here.

With RFT, the data set looks a little bit different. Instead of prompt-completion pairs, you really need a set of tasks that are very gradable. You need a grader that’s very objective that you can use here as well. That’s something we’ve invested a lot in over the last year, and we’ve seen a number of customers get really good results on it. We’ve talked about a couple of them across different verticals.

Rogo is a startup in the financial services space. They have a very sophisticated AI team—I think they hired some folks from DeepMind to run their AI program. They’ve been using RFT to get best-in-class results on parsing through financial documents, answering questions about them, and doing tasks around that as well. There’s another startup called Accordance that’s doing this in the tax space. I think they’ve been targeting an eval called TaxBench, which looks at CPA-style tasks as well. Because they’re able to turn it into a very gradable setup, they’re able to turn the RFT crank and get, I think, SOTA results on TaxBench just using our RFT product.

It has shifted the discussion away from just customizing something for your own use case to really leveraging your own data to create a best-in-class, maybe best-in-the-world, model for something that you care about for your business.

Apoorv Agarwal

Yeah, I feel like the base models are getting so good at instruction-following that, for behavior steering, you don’t need to fine-tune at that point. You can describe what you want, and the model is pretty good at it. But pushing the frontier on actual capabilities, my hunch is that RFT will pretty much become the norm. If you’re actually pushing intelligence in your field to a pretty high point, at some point you need to do RL, essentially, with custom environments.

Fascinating. And even going back to the point earlier around top-down versus bottom-up for some of these enterprises, a lot of the data that you end up needing for RFT requires very intricate knowledge about the exact task that you’re doing and understanding how to grade it. A lot of that actually comes from the bottom up. I know a lot of these startups will work with experts in their field to try to get the right tasks and the right feedback to craft some of these data sets.

12. Rapid Fire: Long/Short Picks

Without further ado, we’re going to jump into my favorite section, which is a rapid-fire question. We had a lot of great friends of ours send in some questions for you guys. We’ll start with Altimeter’s favorite game, which is a long-short game. Pick a business, an idea, or a startup that you’re long, and the same short that you would bet against because there’s more hype than reality. Whoever’s ready to go first: long, short.

Sherwin Wu

My long is actually not in the AI space, so this is going to be slightly different.

Brad Gerstner

Wow. Here we go.

Olivier Godement

My short is, though, in the AI space. I’m extremely long esports. By esports, I mean the entire professional gaming industry that’s emerging around video games. It’s very near and dear to my heart. I play a lot of video games, and so I watch a lot of this. Obviously, I’m pretty in the weeds on it. But I actually think there’s incredible untapped potential in esports and incredible growth to be had in this area.

Concretely, what I mean is a really big one: League of Legends. All of the games that Riot Games puts out have their own professional leagues. They have professional tournaments, believe it or not. They rent out stadiums now. But I just think that if you look at what the youth and younger kids are looking at, and where their time is going, it’s predominantly going toward these things. They spend a lot of time on video games.

Brad Gerstner

They watch more esports than soccer or basketball?

Olivier Godement

Yeah, yeah, yeah. A growing number of them do, too. I’ve actually been to some of these events, and it’s very interesting.

Brad Gerstner

He’s very committed to his long.

Olivier Godement

Yeah, yeah, yeah. I’m extremely long on this.

Brad Gerstner

And so they’re booking out stadiums for people to go watch esports.

Olivier Godement

Yeah, yeah, yeah. I literally went to Oracle Arena, the old Warriors’ stadium, to watch one of these, I think, before COVID.

Brad Gerstner

Before COVID? Wow, that’s 5 years ago.

Olivier Godement

6 years ago. So I’ve been following this for a while, and I actually think it had a really big moment during COVID. Everyone was playing video games. I think it’s kind of come back down, so I think it’s undervalued. No one’s really appreciating it now, but it has all the elements to really, really take off.

The youth are doing it. The other thing I’d say is it’s huge in Asia—absolutely massive in Asia. It’s absolutely big in Korea and China as well. We rented out Oracle Arena, I think, or the event I went to was in Oracle Arena. My sense is that in Asia they rent out entire stadiums, like soccer stadiums, and the players are treated like celebrities. Korean culture is really making its way into the U.S. as well, and I think that’s another tailwind for this whole thing. Anyway, esports is something you should keep an eye on because there’s a lot of room for growth.

Brad Gerstner

Very unexpected. Good to hear. Short?

Olivier Godement

My short is a little spicy: I’m short on the entire category of tooling around AI products. This encapsulates a lot of different things. It’s kind of cheating because some of these, I think, are starting to play out already.

But I think 2 years ago, it was maybe eval products, frameworks, or vector stores. I'm pretty short on those. I think nowadays there's a lot of additional excitement around other tooling for AI models. RL environments, I think, are really big right now as well. Unfortunately, I'm very short on those. I don't really see a lot of potential there. I see a lot of potential in reinforcement learning and applying it, but I think the startup space around RL environments is really tough.

The main thing is, 1, it's just a very competitive space. There's a lot of people operating in it. And then, 2, if the last 2 years have shown us anything, the space is evolving so quickly, and it's so difficult to try and adapt and understand what the exact stack is that will really carry through to the next generation of models. I think that just makes it very difficult when you're in the tooling space because today's really hot framework or really hot tool might just not get used in the next generation of models.

I've been noticing the same pattern, which is that the teams that build breakout startups in AI are extremely pragmatic. They are not super intellectual about the perfect world, et cetera. And it's funny because I feel like our generation basically started in tech at a very stable moment, where technology had been building up for years and years with SaaS, cloud, et cetera.

So we were, in a way, raised in that very stable moment where it makes sense at that point to design very good abstractions and tooling because you have a sense of where it's going. But it's so different today. You don't have a way to know what's going to happen in the next 1 or 2 years, so it's almost impossible to define the perfect tooling platform.

Brad Gerstner

Right. Right. Right. Well, that's—there's a lot of that going around right now. Yes. Spicy. A lot of homework there. Olivier, over to you, sir.

Olivier Godement

Long-short. I've been thinking a lot about education for the past month in the context of kids. I'm pretty short on any education that basically emphasizes human memorization at this point. And I say that having mostly been through that education myself, but I learned so much about history facts and legal things. Some of it does shape your way of thinking; a lot of it, frankly, is just knowledge tokens, essentially. And those knowledge tokens, it turns out, other AI models are pretty good at. So I'm quite short on that.

Brad Gerstner

You will need memory when strategy is bionic. You can just think about it straight into your head.

Olivier Godement

Exactly. Exactly. What am I long at? Frankly, I think healthcare is probably the industry that will benefit the most from AI in the next 1 or 2 years. I would say more. I think all the ingredients are here for a perfect storm: a huge amount of structured data—it's basically the heart of pharma companies—and the models are excellent at digesting and processing that kind of data.

There's also a huge amount of admin-heavy, document-heavy culture, but at the same time, companies that are very technical and very R&D-friendly, companies for which technology is, in a way, at the heart of what they do. And so, yeah, I'm pretty bullish on that.

Brad Gerstner

This is life sciences? So you mean life sciences research organizations that are producing drugs. Gotcha.

Olivier Godement

Exactly. It's almost like, over the last 20 or 30 years, these pharma or biotech companies have basically—if you look at the work that they're doing, only a small amount of it is actual research. So much of it ends up being admin and documents and things like that. And that area is just so ripe for something to happen with AI. I think that's what we're seeing with Amgen and some of these other customers.

Brad Gerstner

Exactly.

Olivier Godement

And it's also not what they want to do. I think it's good that we have some regulations there, obviously, but it just means that they have reams and reams of things to go through. And so when you have a technology that's able to really help bring down the cost of something like that, I think it'll just tear right through it.

And I think once governments and institutions realize that, if you step back, it is probably one of the biggest bottlenecks to human progress, right? You step back over the past decade—how many breakthrough drugs have there been? Not that many. How different would life be if you doubled that rate? Essentially. So once you realize what is at stake, then my hunch is that we're going to see quite a bit of momentum in that space.

Brad Gerstner

Wow. All right. Lots of homework there as well. Yeah. Next one: favorite underrated AI tool other than ChatGPT, maybe?

Olivier Godement

I love Granola.

Bill Gurley

Oh man, you stole mine. You stole my answer.

Olivier Godement

I use Granola so much.

Brad Gerstner

Two votes for Granola. Hey, what about ChatGPT Record?

Olivier Godement

I like ChatGPT as well, but there are some features of Granola that I think are really done well. The whole integration with your Google Calendar is excellent. And the quality of the transcription and the summary is pretty good.

Brad Gerstner

Do you just have it on? Because I know your calendar is back-to-back. Do you just have Granola on?

Olivier Godement

The funny thing is that I don't use Granola internally. I use Granola for my personal life mostly.

Brad Gerstner

I see. Yeah, I see. On dates. I'm joking.

Bill Gurley

I was going to say, yeah, Granola is actually going to be mine. So, 2 votes for Granola.

Olivier Godement

I was going to say the easy answer for me is Codex. As a software engineer, it's just gotten so good recently. Codex CLI, especially with GPT-5. Especially for me, I tend to be less time-sensitive about the iteration loop with coding, so leaning into GPT-5 on Codex, I think, has been really interesting.

Brad Gerstner

What about Codex has changed? Because Codex has also been through a journey. Codex has been around for a bit. I remember it launched more than a year ago. What's changed about Codex?

Olivier Godement

Yeah, I was actually going to say Codex CLI has been around for a bit. I feel like it's been less than a year for Codex.

Brad Gerstner

I feel like it's been less than a year for Codex. The time dilation is so crazy. It feels like it's been around for a year with GPT-5. That demo feels like ages ago, and it didn't even come out yet.

Olivier Godement

Probably because it hasn't happened yet.

Brad Gerstner

I think it was a naming thing, okay, but anyway—

Olivier Godement

Oh, there was a Codex model. That's what I'm thinking about.

Brad Gerstner

There was a Codex model. Also, I think the GitHub thing was called Codex.

Olivier Godement

That's right. Yes, yes. I'm talking about our coding product within ChatGPT, which is the Codex Cloud offering and then also Codex CLI. So, actually, maybe if I were to narrow my answer a little bit more, it's Codex CLI, which I've really, really liked.

I like the local environment setup. The thing that's actually made it really useful in the last, I'd say, month or so is, 1, I think the team has done a really good job of getting rid of all the paper cuts—the small product-polish and paper-cut things. It kind of feels like a joy to use now. It feels more reactive.

And then the second thing, honestly, is GPT-5. I just think GPT-5 really allows the product to shine. At the end of the day, this is a product that really is dependent on the underlying model. When you have to iterate and go back and forth with the model 4 or 5 times to get it right, to get it to do the change that you want, versus having it think a little bit longer and one-shot and do exactly what you want, you get this weird, bionic feeling where you're like, “I feel so mind-melded with the model right now, and it perfectly understands what I'm doing.”

Getting that kind of dopamine hit and feedback loop constantly with Codex has made it an indispensable thing that I really, really like.

Sherwin Wu

The other thing I’d say Codex is really good for, for me, is personal projects. I also use it to help me understand codebases. As an engineering manager, I’m not as in the weeds on the actual code, and so you’re able to use Codex to really understand what’s happening with the codebase, have it ask questions and answer them, and really catch up to speed on things as well. Even the non-coding use cases are really useful with Codex CLI.

Brad Gerstner

Fascinating. Sam had this tweet about Codex usage ripping, I think, yesterday. So I wonder what’s going on there, but you’re not alone.

Sherwin Wu

Yeah, I think I’m not alone. Just judging from the Twitter feedback, I think people are really realizing how great of a combination Codex CLI and GPT-5 are.

Bill Gurley

Yeah, I know that team is undergoing a lot of scaling challenges, but the system hasn’t gone down for me, so props to them. But we are in a GPU crunch, so we’ll see how long that goes.

Brad Gerstner

Awesome. All right, the next one: Will there be more software engineers in 10 years or less? There are about 40–50 million full-time, professional software engineers.

Sherwin Wu

That’s what you mean—full-time, actual jobs?

Brad Gerstner

Yeah, because it’s a hard one. I think without a doubt there’s going to be a lot more software engineering going on.

Sherwin Wu

Yes, of course. There’s actually a really great post that was shared, I think, in our internal Slack. It was a Reddit post recently, and I actually think that highlights this. It was a really touching story.

It was a Reddit post about someone who has a brother who’s nonverbal. They have to take care of him. They tried all these types of things to help the brother interact with the world and use computers, but vision tracking didn’t work because I think his vision wasn’t good. All the tools didn’t work, and then this brother ended up using ChatGPT.

I don’t think he used Codex, but he used ChatGPT and basically taught himself how to create a set of tools that were tailor-made for his nonverbal brother—a custom software application just for them. Because of that, he now has this custom setup that was written by his brother and allows him to browse the internet. I think the video was of him watching The Simpsons or something like that, which was really touching.

I think that’s actually what we’ll see a lot more of. This guy’s not a professional software engineer. His title isn’t software engineer, but he did a lot of software engineering—probably pretty good, and definitely good enough for his brother to use. The amount of code, the amount of building that’ll happen, is just going to go through an incredible transformation.

I’m not sure what that means for software engineers like myself. Maybe there’s—of course, more Sherwin.

Brad Gerstner

Of course, more Sherwin.

Sherwin Wu

More of me specifically.

Brad Gerstner

We need more of you.

Sherwin Wu

That’s right. But definitely, there’ll be a lot more software engineering at a lot of companies.

Apoorv Agarwal

I buy that completely. I completely buy the thesis that there is a massive software shortage in the world. We’ve been accepting it for the past 20 years. But the goal of software was never to be that super-rigid, super-hard-to-build artifact. It was to be customized and malleable.

So I expect that we’ll see more of a reconfiguration of people’s jobs and skill sets, where way more people code. I expect that product managers are going to code more and more, for instance. You made your PMs code recently, if I remember right.

Olivier Godement

Oh, yeah, we did that. It was really fun. We started essentially not doing PRDs—product requirements documents. Classic PM thing: You write 5 pages, “My product does that,” et cetera. And PMs have basically been coding prototypes. One is pretty fast with GPT-5 and Codex—just a couple of hours, I think.

Bill Gurley

Fricking fast.

Olivier Godement

And second, it sort of conveys so much more information than a document. You get a feel, essentially, for the feature: Is it right or not?

Brad Gerstner

Yeah, instead of writing English, you can actually now write the actual thing you want. That’s amazing. Advice for high school students who are just starting out their careers?

Olivier Godement

My advice is—I don’t know. Maybe it’s evergreen: Prioritize critical thinking above anything else. If you go into a field that requires extremely high critical-thinking skills—I don’t know, math, physics, or maybe philosophy is in that bucket—you will be fine regardless.

If you go into a field that turns down that thing, and again, it gets back to memorization and pattern matching, I think you will probably be less future-proof.

Brad Gerstner

What’s a good way to sharpen critical thinking?

Olivier Godement

Use ChatGPT and have it test you. That’s a tricky test. Having a world-class tutor who essentially knows how to put the bar about 20% above what you can do all the time is actually probably a really good way to do it.

Brad Gerstner

Nice. Anything from you, sir?

Sherwin Wu

Mine is—I think we’re actually in such an interesting, unique time period. So maybe this is more general advice, not just for high school students, but for the younger generation, even college students.

I think the advice would be: Don’t underestimate how much of an advantage you have relative to the rest of the world right now because of how AI-native you might be, or how versed in these tools you are. My hunch is that high schoolers and college students, when they come into the workplace, are going to have a huge leg up in how to use AI tools and how to actually transform the workplace.

My push for some of the younger high school students is, first, just really immerse yourself in this thing. Second, really take advantage of the fact that you’re in a unique time where no one else in the workforce really understands these tools as deeply, probably, as you do.

A good example of this is that we had our first intern class recently at OpenAI—a lot of software interns. Some of them were just the most incredible Cursor power users I’ve ever seen. They were so productive. I was shocked, by the way. I was like, “Yeah, I know we can get good interns, but I don’t know if they’d be this good.”

I think part of it is they’ve grown up using these tools, for better or worse, in college. But I think the meta-level point is they’re so AI-native. Even Olivier and I are kind of AI-native—we work at OpenAI—but we haven’t been steeped in this and grown up in this.

The advice here would just be: Leverage that. Don’t be afraid to go in and spread this knowledge and take advantage of it in the workplace, because it is a pretty big advantage for them.

I can’t remember who said this to us at Palantir, but every intern class was just getting faster and smarter, like laptops getting smarter every generation. You sure didn’t peak in 2013, when I was an intern.

Bill Gurley

That’s right. There’s a weird spike. That’s summer 2013.

Brad Gerstner

Well, lots happened here. A lot’s happened since you guys joined OpenAI, right? It’s been 3 years and almost 3 years. In your OpenAI journey, what has been the rose moment—your favorite moment; the bud moment, where you’re most excited about something but there’s still opportunity ahead; and the thorn, the toughest moment of your 3-year journey?

Olivier Godement

The thorn is easy for me: What we call the blip, which was the board coup. That was a really tough moment. It’s funny because, after the fact, it actually reunited the company quite a bit. OpenAI had a pretty strong culture before, but there was a feeling of camaraderie that was even stronger. But it was tough on the day.

Bill Gurley

It’s very rare to see that kind of antifragility. Most organizations, after something like that, break apart, but I feel like OpenAI got stronger. OpenAI came back.

Olivier Godement

It’s a good point. I feel it made OpenAI stronger for real, essentially, when they look at it after the fact. When they look at other news, like departures or whatever—bad news, essentially—I feel the company has built a thicker skin and an ability to recover way quicker.

Bill Gurley

I think that’s definitely right. Part of it, too, I think, is also just the culture. I also think this is why it was such a low point for a lot of people.

Olivier Godement

So many people at OpenAI care so deeply about what we're doing, which is why they work so hard. You just care a lot about the work. It almost feels like your life's work. It's a very audacious mission and thing that you're doing, which is why I think the blip was so tough on a lot of people, but also what I think helped bring people back together and how we were able to hold together and get that thick skin as well.

I have a separate worst moment, which was the big outage that we had in December of last year. You remember.

Brad Gerstner

I do.

Olivier Godement

It was a multi-hour outage. It really highlighted to us how essential—almost like a utility—the API was. I think we had a 3- or 4-hour outage sometime in November or December last year. It was really brutal, a pure shitshow. No one could hit ChatGPT. No one could hit the APIs. It was really rough.

That was just really tough from a customer-trust perspective. I remember we talked to a lot of our customers to postmortem with them what happened and our plan moving forward. Thankfully, we haven't had anything close to that since then. I've actually been really happy with all the investments we've made in reliability over the last 6 months. But in that moment, I think it was really tough.

13. Highlights and Lowlights @ OpenAI

On the happy side, on the roses, I think I have 2 of them. The first one would be that GPT-5 was really good. The sprint up to GPT-5 really showed the best of OpenAI: cutting-edge science and research, extreme customer focus, and extreme infrastructure and inference talent. The fact that we were able to ship such a big model and scale it to many, many, many tokens per minute almost immediately, I think, speaks to it.

Brad Gerstner

With no outages.

Olivier Godement

Yeah, really good reliability. I can remember when we shipped GPT-4 Turbo, like a year ago, a year and a half ago, we were terrified by the insane-scale traffic. I feel we've really gotten much better at shipping those massive updates.

The second happy moment for me would be the first Dev Day. It felt like a coming of age for OpenAI. We were embracing that we have a huge community of developers. We were going to ship models and products. I remember seeing all my favorite people, OpenAI or not, nerding out on what they were building and what was coming up next. It felt really like a special moment in time.

That was actually going to be mine as well, so I'll just piggyback off of that: the very first Dev Day, November 2023. I remember it. Obviously, a lot of good things have happened since then, but for me, it was a very memorable moment.

One, it was actually quite a rush up to Dev Day. We shipped a lot, so our team was just really, really sprinting. It was this high-stress environment going up. To add to that, of course, because we're OpenAI, we did a live demo during Sam's keynote of all the stuff that we shipped. I remember being in the back of the audience, sitting with the team and waiting for the demo to happen. Once it finished, we all just let out a huge sigh of relief. We were like, “Oh my God, thank you.”

For me, the most memorable thing was right after Dev Day. All the demos worked well, all the talks worked well, and we had the after-party. Then I was just in a Waymo, driving home at night with the music playing. It was such a great end to Dev Day. That was what I remember. That was my rose for the last few years.

Brad Gerstner

Love it. That's awesome. I assume you guys are, but please tell me if you're AGI-pilled, yes or no. And if so, what was the moment that got you there? What was your aha moment? When did you feel the AGI?

Olivier Godement

I think I'm AGI-pilled. You're definitely AGI-pilled? I am? I've had a couple of them.

The first one was the realization in 2023 that I would never need to code manually ever again. I'm not the best coder; I chose my job for a reason. But realizing that what I thought was a given—that we humans would have to write basically machine language forever—is actually not a given, and that the paradigm shift is huge.

The second feel-the-AGI moment for me was maybe the progress on voice and multimodality. Text, at some point, you get used to it: the machine can write pretty good text. Voice makes it real. But once you start actually talking to something that understands your tone, understands my accent in French, it felt like a moment—machines are going beyond cold, mechanical, deterministic logic to something much more emotional and tangible.

Yeah, that's a great one. Mine are—I do think I am AGI-pilled. I probably gradually became AGI-pilled over the last couple of years. I think there are 2, and for me, I actually get more shocked from the text models. I know the multimodal ones are really great as well. For me, I think they line up with 2 general breakthroughs.

The first one was right when I joined the company in September 2022. It was pre-ChatGPT, 2 months before, about the time GPT-4 already existed internally. I think we were trying to figure out how to deploy it. I think Nick Turley talked about this a lot early on with ChatGPT. But it was the first time I talked to GPT-4, and it was like going from nothing to GPT-4. It was the most mind-blowing experience for me.

For the rest of the world, maybe going from nothing to GPT-3.5 in ChatGPT was the big one, and then going from GPT-3.5 to GPT-4. But for me, and I think for a lot of other people who joined around that time, going from—not nothing, but what was publicly available at the time—to GPT-4 was just incredible. I remember asking it and throwing so many things at it. I was like, “There's no way this thing is going to be able to give an intelligible answer.” And it just knocked it out of the park. It was absolutely incredible. GPT-4 was insane.

I remember GPT-4 came out when I was interviewing with OpenAI, and I was still like, “Should I join?” And then I was like, “Okay. I mean, there is no way I can work on anything else at that point.”

That's true. Yeah, GPT-4 was just crazy. And then the other breakthrough was the reasoning paradigm. I actually think the purest representation of that for me was Deep Research.

Asking it to really look up things that I didn't think it would be able to know, and seeing it think through all of it, be really persistent with the search, get really detailed with the write-up, and all of that—that was pretty crazy. I don't remember the exact query that I threw at it, but I just remember that the feel-the-AGI moments for me are when I'll throw something at the model that I think, “There's no way this thing will be able to get,” and then it just knocks it out of the park. That is kind of the feel-the-AGI moment. I definitely had that with Deep Research, with some of the things I was asking.

Brad Gerstner

Well, this has been great. Thank you so much, folks. You guys are building the future. You guys are inspiring us every day, and I appreciate the conversation.

Olivier Godement

Yeah, thank you so much.

Thank you. Thanks for having us.