Apoorv Agarwal
In San Francisco, you could take a car from one part of San Francisco to the other fully autonomously. As opposed to the digital world, I can't book a ticket online right now. Physical autonomy is ahead of digital autonomy in 2025.
Sherwin Wu
I think AI agents are really in day one here. ChatGPT only came out in 2022. The slope, I think, is incredibly steep.
Apoorv Agarwal
I actually do think self-driving cars have a good amount of scaffolding in the world. You have roads. Roads exist. They're pretty standardized. You have stoplights. AI agents are just dropped in the middle of nowhere.
We'll start with the long-short game. I'm short on the entire category of tooling and evals products.
Olivier Godement
Healthcare is probably the industry that will benefit the most from AI. I think I'm AGI-pilled.
Apoorv Agarwal
You're definitely AGI-pilled.
Hey, folks. I'm Apoorv Agarwal, and today, at the OpenAI office, we had a wide-ranging conversation about OpenAI's work in enterprise. I have with me the head of engineering and the head of product of the OpenAI Platform, Sherwin Wu and Olivier Godement.
OpenAI is well known as the creator of ChatGPT, a product that billions across the world have come to love and enjoy. But today, we dive into the other side of the business, which is OpenAI's work in enterprise. We go deep into their work with specific customers and how OpenAI is transforming large and important industries like healthcare, telecommunications, and national security research.
We also talk about Sherwin and Olivier's outlook on what's next in AI, what's next in technology, and their picks on both the long and short side. This was a lot of fun to do. I hope you really enjoy it.
1. OpenAI’s Enterprise Mission: Beyond ChatGPT
Well, 2 world-class builders—2 people who make building look easy. Sherwin, my Palantir 2013 classmate and tennis buddy, with 2 stops at Quora and Opendoor through the IPO before joining OpenAI, before ChatGPT, you've now been here for 3 years and lead engineering for all of OpenAI Platform.
Olivier, former entrepreneur, winner of the Golden Llama at Stripe, where you were for just under a decade, and now lead all of the product at OpenAI Platform.
Olivier Godement
That's right.
Apoorv Agarwal
Thanks for doing it.
Olivier Godement
Thank you. Thanks for having us.
Apoorv Agarwal
As a shareholder and as a thought partner, kicking ideas back and forth, I always learn a lot from you guys. And so it's a treat. It's a real treat to do this for everybody.
I'll open with this: People know OpenAI as the firm that built ChatGPT, the product that they have in their pocket, that comes with them every day to work and into their personal lives. But the focus for today is OpenAI for enterprise. You guys lead OpenAI Platform. Tell us about it. What's underneath the OpenAI Platform for B2B and for enterprise?
Sherwin Wu
Yeah. So this is actually a really interesting question, too, because when I joined OpenAI around 3 years ago to work on the API, it was actually the only product that we had. I think a lot of people forget this: The original product for OpenAI was not ChatGPT. It was a B2B product—the API—and we were catering toward developers. I've actually seen the launch of ChatGPT and everything downstream from that.
But at its core, I think the reason why we have a platform and why we started with an API comes back to the OpenAI mission. Our mission, obviously, is to build AGI, which is pretty hard in and of itself, but also to distribute the benefits of it to everyone in the world—to all of humanity. It's pretty clear right now to see ChatGPT doing that because my mom, maybe even your parents, are using ChatGPT.
But we actually view our platform, and especially our API and how we work with our enterprise customers, as our way of getting the benefits of AGI, of AI, to as many people as possible—to everyone in every corner of the world.
ChatGPT is really, really, really big now. It's, I think, the 5th-largest website in the world. But by working through developers using our API, we're actually able to reach even more people in every corner of the world and every different use case that you might have. Especially with some of our enterprise customers, we're able to reach use cases within businesses and the end users of those businesses as well. And so we view the platform as our way of fully expressing our mission of getting the benefits of AGI to everyone.
Concretely, what the platform actually includes today: The biggest product we have is obviously our developer platform, which is our API. Many developers—the majority of the startup ecosystem—build on top of this, as well as a lot of digital natives and Fortune 500 enterprises at this point.
We also have a product that we sell to governments in the public sector. That's all part of this as well. An emerging product line for us in the platform is our enterprise product. We might sell directly to enterprises beyond just a core API offering.
Apoorv Agarwal
Fascinating. And maybe to double down, I think B2B is actually quite core to the OpenAI mission. What we mean by distributing AGI benefits is this: I want to live in a world where there are 10x more medicines going out every year. I want to live in a world where education, public service, and civil service are increasingly optimized for everyone.
There is a large category of use cases that only go through B2B, frankly, unless you enable the enterprises. And we talk about Palantir—I think that's probably the same thesis at Palantir. Those are the businesses that are actually making stuff happen in the real world. So if you do enable them, if you do accelerate them, that's how, essentially, you distribute the benefits of AGI.
Yeah. Well, maybe we can double-click into that, Olivier. The reach for ChatGPT is obviously wide—billions of users. But for enterprise, maybe tell us about it. Maybe we go deep into a customer example or 2. What is an organization that we have helped transform, maybe, and at what layers?
2. Case Study: T-Mobile - Voice & Support
Olivier Godement
If I were to step back, we started our B2B efforts with the API a few years ago. Initially, the customers were startups, developers, indie hackers, and extremely technically sophisticated people who were building cool new stuff and taking massive market risk. We still have a bunch of customers in that category, and we love them and keep building with them.
On top of that, over the past couple of years, we've been working more and more with traditional enterprises and also digital natives. Basically, everyone woke up with ChatGPT, and those models are working. There's a ton of value, and they could see many use cases in the enterprise.
A couple of examples that I like the most. One, which is both very fresh and quite cool, is that we've been working a lot with T-Mobile, a leading U.S. telco operator. T-Mobile has a massive customer-support load: People asking, "Hey, I was charged that amount of money. What's going on?" Or, "My cell phone isn't working anymore."
A massive share of that load is voice calls. People want to talk to someone. And so for them, to be able to essentially automate more and more and help people self-serve—to debug their subscription—was pretty big. We've been working with T-Mobile pretty much for the past year to basically automate not only the text support but also voice support.
Today, there are features in the T-Mobile app that, if you call, are actually handled by OpenAI models behind the scenes. It does sound supernatural—human-sounding, latency-wise and quality-wise. So that one was really fun.
Apoorv Agarwal
Just on that, can I ask you a follow-up question? We've got text models. We've got voice models. Maybe even video models someday that are deployed at T-Mobile. But what above the models, or adjacent to the models, might we have helped T-Mobile with, for example?
Olivier Godement
Yeah, there's a ton we're doing.
The first one is—you have to put yourself in the shoes of an enterprise buyer. Their goal is to automate, reduce, and optimize customer support. You're going from a model—tokens in, tokens out—to that case, and it's hard.
First, there's a lot of design—system design. We do have forward-deployed engineers now who are helping us quite a bit.
Bill Gurley
Forward-deployed engineers.
Olivier Godement
We borrow the term from Palantir.
Bill Gurley
Yeah, it's a great term. Were you an FD at Palantir?
Olivier Godement
I was not an FD. I was on, I think they called it the dev side, right? It's like software engineering. I was also only an intern at Palantir. But, yeah, it's a great term. I think it accurately describes what we're asking folks to do, which is embed very deeply with customers and, honestly, build things specific to their systems. They're deployed onto these customers.
But, yeah, we are obviously growing and hiring that team quite a bit because they've been very effective at T-Mobile.
Bill Gurley
Four years of my life. Yeah, yeah, yeah. Forward deployed. But go ahead.
Olivier Godement
So, forward-deployed engineering—the sort of systems and integrations they're doing—is, first, about orchestrating those models. Those models are not just models; they know nothing about the CRM or what's going on. You have to plug the model into many, many different tools.
Many of those tools in the enterprise don't even have APIs or clean interfaces, right? It's the first time they're being exposed to a third-party system. So, there's a lot of standing up API gateways and tools, and connecting them.
Then you have to essentially define what good looks like. Again, to put in your exercise for everyone, defining a golden set of evals is easier than it sounds.
Bill Gurley
Harder than it sounds.
Olivier Godement
Yeah. And so, we've been spending a bunch of time with them.
Bill Gurley
Evals are important. Evals are super important.
Olivier Godement
Especially audio evals. Evals are extra hard to grade and get right. The bulk of the use case here is actually audio. We have, I don't know, five-minute call transcripts—how do you actually know that the right thing happened? It's a pretty tough problem.
Bill Gurley
Yeah, it's pretty tough.
Olivier Godement
And then actually nailing down the quality of the customer experience until it feels natural. Latency and interruptions are really important parts of that.
We shipped a Realtime API in GA. I think it was last week.
Bill Gurley
A couple of weeks ago, yeah.
Olivier Godement
Yeah, it was just last week, I think. It's a beautiful work of engineering. There was a really cracked team behind the scenes, which basically allows us to get the most natural-sounding voice experience without having these weird interruptions or lag where you can feel that the thing is off.
So, cobbling all that together, you get a really good experience.
Bill Gurley
Yeah, that's a lot more than just models.
Sherwin Wu
One really great thing that I think we've gotten from the T-Mobile experience is working with them to improve our models themselves. For example, in the last Realtime GA last week, we obviously released a new snapshot—the GA snapshot. A lot of the improvements that we got into the model came out of the learnings that we had from T-Mobile.
It brings in a lot of other changes from other customers, but because we were so deeply embedded into T-Mobile and were able to understand what good looks like for them, we were able to bring that to some of our models.
Bill Gurley
That makes sense. So, we are working with a large customer with tens of millions of users, if not hundreds of millions, and the before and after is on the support side—both tech support internally and then their customer support.
Yeah. Is there another one that you guys can share? I like Amgen a lot. Amgen, the healthcare business.
3. Case Study: Amgen - Accelerating Drug Development
Olivier Godement
Amgen, yeah. We are working quite a bit with healthcare companies. Amgen is one of the leading healthcare companies. They specialize in drugs for cancer or inflammatory diseases, and they're based out of L.A.
We've been working with Amgen to speed up drug development and the commercialization process. The north star is pretty bold. It's really interesting when you similarly embed pretty deeply with Amgen to understand what their needs are.
When I look at healthcare companies, I feel like there are 2 big buckets of needs. One is pure R&D. You're seeing a massive amount of data, and you have super-smart scientists who are trying to come up with and test out things. That's one bucket.
A second bucket is much more common across other industries. It's pure admin, document-authoring, and document-scribing work. By the time your R&D team has essentially locked the recipe of a medication, getting that medication to market is a ton of work. You have to submit to various regulatory bodies and get a ton of reviews.
When we looked at those problems and what we knew models were capable of, we saw a ton of benefits and a ton of opportunities to automate and augment the work of those teams. Amgen has been a top customer of GPT-5, for instance.
Bill Gurley
Wow. I mean, this could be hundreds of millions of lives if a new drug is developed faster.
Olivier Godement
Yeah, exactly. Huge impact. So that's, I think, one good example of the kind of impact you need to enable enterprises to create. And so, I think we're going to do more and more of those.
Frankly, on a personal level, it's a delight. If I can play a tiny role in essentially doubling the kind of medication that people get in the real world, that feels like a pretty good achievement.
Bill Gurley
Huge. Huge, huge. I know you had one as well.
4. Case Study: Los Alamos National Lab
Sherwin Wu
One of my favorite deployments that we've done more recently is with the Los Alamos National Laboratory. This is the government national research lab that the U.S. government runs in Los Alamos, New Mexico. It's also where the Manhattan Project happened back in the ’40s and ’50s, when it was a secret project.
After that, they ended up formalizing it as a city and a program, and now it's a pretty sizable national laboratory. This one is very interesting because, first, the depth of impact here is unimaginable to me. It's on the scale of Amgen and some of these other larger companies.
Obviously, they're doing a lot of actual new research there, so a lot of new science. They're doing a lot of work with our Defense Department and defense use cases as well. It's very intense stuff.
But the other thing that's very interesting about this one is that it's also a story of a very bespoke and new type of deployment that we've done. Because they're a government lab, they're so restrictive and high-security and high-clearance with a lot of their work, we couldn't just do a normal deployment with them. You can't have people doing national-security research just hitting our APIs.
And so, we actually did a custom, on-premises deployment with them onto one of their supercomputers, called Venado. This involved a bunch of very bespoke work with some FDEs, as well as with a lot of our developer team, to actually bring one of our reasoning models, o3, into their laboratory, into an air-gapped supercomputer—Venado—and deploy it and get it installed and working on their hardware, on their networking stack, and actually run it in this particular environment.
And so it was actually very interesting because we literally had to bring the weights of the model physically into their supercomputer, in an environment, by the way, that’s very locked down for a good reason: they’re not allowed to have cell phones or any electronics with them. So I think that was a very unique challenge.
The other interesting thing about this deployment is how it’s being used. Because it’s so locked down and on-premises, we actually don’t have much visibility into exactly what they’re doing with it, but they do give us feedback. They actually do have some telemetry, but it’s within their own systems.
We do know that it’s being used for a bunch of different things. It’s being used to help speed up their experiments. They have a lot of data-analysis use cases and a lot of notebooks that they’re running with reams of data that they’re trying to process.
They’re actually using it as a thought partner, which is something that’s pretty interesting to me. o3 is a pretty smart model, and a lot of these people are tackling really tough, novel research problems. A lot of the time, they’re using o3 and going back and forth with it on their experiment design and what they should actually be using it for, which is something that we couldn’t really say about our older models.
So it’s being used for a lot of different use cases for the National Lab. The other cool thing is that it’s actually being shared between Los Alamos and some of the other labs—Lawrence Livermore and Sandia as well—because it’s a supercomputer setup where they can all connect to it remotely.
5. Why 95% of AI Deployments Fail?
Apoorv Agarwal
Fascinating. We’ve just gone through 3 pretty large-scale enterprise deployments, which might touch tens, if not hundreds, of millions of people. But on the other side of this is the MIT report that came out a couple of weeks ago: 95% of AI deployments don’t work. There were a bunch of scary headlines that even shook the markets for a couple of days.
Put this in perspective. For every deployment that works, there’s presumably a bunch that don’t work. So maybe we can talk about that. What does it take to build a successful enterprise deployment, a successful customer deployment, and the counterfactual, based on all your experience serving all these large enterprises?
Olivier Godement
I think at that point, I may have worked with a couple hundred. So, okay, I’m going to pattern-match. What I’ve seen as a clear leading indicator of success is, number 1, the interesting combination of top-down buy-in and enabling a very clear group—a tiger team, essentially—at the enterprise, which is sometimes a mix of OpenAI and enterprise employees.
Typically, take T-Mobile. The top leadership was extremely bought in: “It’s a priority.” But then they let the team organize and say, “Okay, if you want to start small, start small,” and then you can scale it up, essentially. That would be part number 1: top-down buy-in and a bottom-up tiger team.
The tiger team is made up of people with a mix of technical skills and people who just have organizational knowledge—institutional knowledge. It’s really funny: in the enterprise, customer support is a good example. What we found is that the vast majority of the knowledge is in people’s heads.
You would think that, in customer support, everything is perfectly documented. The reality is that the standard operating procedures, the SOPs, are largely in people’s heads. Unless you have that mix of technical people and subject-matter experts, it’s really hard to get something off the ground.
That would be number 1. Number 2 would be evals first. Whatever we define as good evals, that gives people a clear common goal to hit. Whenever the customer fails to come up with good evals, it’s a moving target, essentially, whether you’ve made it or not.
Evals are much harder to get done than they look, and they also oftentimes need to come from the bottom up, because all of these things are in people’s heads—in the actual operators’ heads. It’s actually very hard to have a top-down mandate of, “This is how the evals should look.” A lot of it needs bottom-up adoption.
Apoorv Agarwal
Right. Yeah. Yeah.
Olivier Godement
And so we’ve been developing quite a bit of tooling for evals. We have an evals product, and we’re working on more to essentially solve that problem, or make it as easy as we can.
The last thing is that you want to hill-climb, essentially. You have your evals; the goal is to get to 99%. You start at 46%. How do you get there? Frankly, I think oftentimes it’s a mix of—I will say—almost wisdom from people who’ve done it before. A lot of that is art, sometimes more than science: knowing the quirks of the model and its behavior.
Sometimes we even need to fine-tune the models ourselves when there are some clear limitations, and be patient, get your score up there, and then ship.
6. Physical vs Digital Autonomy: Scaffolding & Infrastructure
Apoorv Agarwal
Can we go under the hood a little bit? One of the things that we think about a lot is autonomy more broadly. What is the makeup of autonomy? On one side, in San Francisco, you could take a car from one part of SF to the other fully autonomously. No humans involved. You press a button. We love Waymo. They’ve done billions of miles. I think it was, what, 3.5 billion miles on Tesla FSD. I think they’ve almost done tens of millions of rides.
That’s a lot of autonomy in the physical world, as opposed to the digital world. I can’t book a ticket online right now. There are all sorts of problems that happen if I have my operator try to book a ticket.
It’s very counterintuitive because the bar for physical safety is so much higher. The bar for physical safety is higher than the human’s capability because lives are at stake. The bar for digital safety is not that high, because all you’re going to lose is money. Nobody’s life is at stake.
Yet physical autonomy is ahead of digital autonomy in 2025, which seems counterintuitive. Why is that the case at a technical level? Why is it that what should sound easier is actually a lot harder?
Sherwin Wu
Yeah, so I think there are 2 things at play here. I really like the analogy with self-driving cars because they’ve actually been one of the best applications of AI that I’ve used recently. One of them is honestly just the timelines.
We’ve been working on self-driving cars for so long. I remember back in 2014, it was kind of the advent of this, and everyone was like, “Oh, it’s happening in 5 years.” It turns out it took 10 or 15 years or so for this to happen.
There was probably a dark age back in 2015 or 2018 or something, when it felt like it wasn’t going to happen.
Brad Gerstner
A trough of disillusionment.
Apoorv Agarwal
Yes, yes, yeah. And then now we’re finally seeing it get deployed, which is really exciting. But it has been, I don’t know, 10 years, maybe even 20 years from the very beginning of the research.
Whereas I think AI agents are really in day 1 here. ChatGPT only came out in 2022, so around 3 years—less than 3 years ago. I actually think what we think about with AI agents and all that really started with the reasoning paradigm, when we released the o1-preview model back in late last year, I think.
And so I actually think this whole reasoning paradigm with AI agents and the robustness they bring has only really unfolded for about a year—less than a year, really.
And so, I know you had a chart in your blog post that I really like, where the slope is very meaningfully different now. Self-driving started very, very early. The slope seems to be a little bit slower, but now it’s reaching the promised land. But, man, we started super recently with AI agents, and the slope, I think, is incredibly steep, and we’ll probably see a crossover at some point.
But we really have only had about a year to explore these things. Do you think we haven’t crossed over already when you look at the coding work in particular?
Olivier Godement
Yeah, it’s a good point. Your chart actually shows AI agents below self-driving, but what is the Y-axis? If you look at some measures, I would not be surprised if AI-agent products are making more revenue than Waymo at this point. Waymo was making a lot, but just look at all the startups coming up. Look at ChatGPT, how many subscriptions are happening there, and all of that.
Maybe we have actually crossed over, and a couple of years from now, it’s going to look very, very different.
Apoorv Agarwal
Yeah, the Y-axis is tangible, felt autonomy. Don’t pick the objective—how do I feel about it?
Olivier Godement
Exactly. It vibes more than revenue. But revenue is a good one. We should probably redo that with revenue.
Sherwin Wu
There’s a second thing I wanted to mention as well, which is the scaffolding and the environment in which these things operate. I remember in the early days of self-driving, a lot of the researchers around self-driving were saying that the roads themselves would have to change to accommodate self-driving. There might be sensors everywhere so that the self-driving cars could interact with them, which, in retrospect, I think was overkill.
But I do think self-driving cars have a good amount of scaffolding in the world for them to operate in—not completely unlimited. You have roads. Roads exist, and they’re pretty standardized. You have stoplights. People generally operate in pretty normal ways, and there are all these traffic laws that you can learn.
Whereas AI agents are just dropped in the middle of nowhere, and they kind of have to feel around for things. Going off of what Olivier just said, my hunch is that some of the enterprise deployments that don’t actually work out likely don’t have the scaffolding or infrastructure for these agents to interact with, either.
A lot of the really successful deployments that we’ve made—a lot of what our FDEs end up doing with some of these customers—is creating almost like a platform or some type of scaffolding, building connectors, and organizing the data so that the models have something they can interact with in a more standardized way. My sense is that self-driving cars have actually had this to some degree, with roads, over the course of their deployment.
But I think it’s still very early in the AI-agent space. I would not be surprised if a lot of enterprises and companies just don’t really have the scaffolding ready. So if you drop an AI agent in there, it doesn’t really know what to do, and its impact will be limited.
Once this scaffolding gets built out across some of these companies, I think deployment will also speed up. But, again, to our point earlier, I think there’s no slowdown. Things are still moving very fast.
Bill Gurley
That’s great. Well, you know, I’ve thought about autonomy as a three-part structure. You’ve got perception, you’ve got the reasoning—the brain—and then you’ve got the scaffolding, the last mile of making things work.
7. GPT-5: Release, Benchmarks vs Behavior
Maybe we can dive into the second part, which is the reasoning—the juice that you guys are building with GPT-5 most recently. Huge endeavor. Congrats. It’s the first time you guys have launched a full system, not a model or a set of models, but a full system.
Talk about that—the full arc of that development. What was your focus? Honestly, the benchmarks all seem so saturated. Clearly, it was more than just benchmarks that you were focused on. What is a North Star? Tell us about GPT-5, soup to nuts.
Olivier Godement
It’s been the labor of love of many people for a long time. To your point, I think GPT-5 is amazingly intelligent. You look at the benchmarks, like [SWE-bench?] and the like, and it’s going pretty high.
But to me, equally important and impactful was the craft—the style, the tone, and the behavior of the model. So: capabilities, intelligence, and behavior of the model.
On the behavior of the model, I think it’s the first large-model release for which we’ve worked so closely with a bunch of customers, month after month, to better understand the concrete blockers. Often, it’s not about having a model that is way more intelligent; it’s about having a model that better follows instructions and is more likely to say no when it doesn’t know something.
That super-close customer feedback loop on GPT-5 was pretty impressive to see. I think all the love that GPT-5 has been getting in the past couple of weeks means people are starting to feel that from the builders, essentially. Once you see it, it’s really hard to come back to a model that is extremely intelligent but is essentially an exquisite academic.
Apoorv Agarwal
Are there trade-offs that you made as you were going through it? What were the hardest trade-offs you made as you were building GPT-5?
Sherwin Wu
I actually think a very clear trade-off, which I honestly think we’re still iterating on, is the trade-off between the reasoning tokens—how long it thinks—and performance.
This is something we’ve been working on with our customers since the launch of the reasoning models. These models are so smart, especially if you give them all this thinking time. The feedback I’ve been seeing around GPT-5 Pro has been pretty crazy, too. There are these unsolved problems that none of the other models could handle; you throw them to GPT-5 Pro, and it just one-shots them. It’s pretty crazy.
But the trade-off here is that you’re waiting for 10 minutes. That’s quite a long time. These things just get so smart with more inference time, but for the product builder on the API side, for some of these business use cases, it’s pretty tough to manage that trade-off.
For us, it’s been difficult to figure out where we want to fall on that spectrum. We’ve had to make some trade-offs on how much the model thinks versus how intelligent it should get. As a product builder, there’s a real latency trade-off that you have to deal with. Your user might not be happy waiting 10 minutes for the best answer in the world. They might be more okay with a substandard answer and no wait at all.
Brad Gerstner
Yeah, I mean, even between GPT-5 and GPT-5 Thinking, I have to toggle it now because sometimes I’m so impatient I just want it ASAP. I think there’s an ability to skip, right?
Sherwin Wu
Yeah, that’s right.
Brad Gerstner
And GPT-5 is like, “I’m impatient. I just want a simpler answer.”
Sherwin Wu
That’s right, that’s right.
Brad Gerstner
Well, 4 weeks in, GPT-5—how’s the feedback?
Sherwin Wu
I think feedback has been very positive, especially on the platform side, which has been really great to see. A lot of the things that Olivier mentioned have been coming up in feedback from customers.
The model is extremely good at coding and extremely good at reasoning through different tasks. Especially for coding use cases, when it thinks for a while, it’ll usually solve problems that no other models can solve. I think that’s been a big positive point of feedback.
Brad Gerstner
Yeah, yeah, yeah.
8. GPT-5 Feedback: Instruction Following, Hallucinations, Code Quality
Sherwin Wu
I think there’s an eval that showed that hallucinations basically went to zero for a lot of this. It’s not perfect—there’s still a lot of work to be done—but because of the reasoning in there, it just makes the model more likely to say no and less likely to hallucinate answers. That’s been something that people have really liked as well.
Another bit of feedback has been around instruction-following. It’s really good at instruction-following.
This almost bleeds into the constructive feedback that we're working on. It's so good at instruction following that people need to tweak their prompts; it's almost too literal. That's an interesting trade-off because when you ask developers what they want, they want the model to follow instructions, of course. But once you have a model that is extremely literal, that forces you to express extremely clearly what you want; otherwise, the model may go sideways.
That feedback was interesting. It's almost like the monkey's paw: developers and platform customers ask for better instruction following, and we're like, “Yes, we'll give you really good instruction following.” But it follows instructions almost to a T, so it's obviously something that the team is working through.
A good example of this, by the way, is that some customers would have these prompts. I remember when we were testing GPT-5, one piece of negative feedback we got was that the model was too concise. We were like, “What's going on? Why is the model so concise?” Then we realized it was because they were using their old prompts from other models. With the other models, they have to really beg the model to be concise.
There are 10 lines of “Be concise. Really be concise. Also, keep your answer short.” It turns out that when you give that to GPT-5, it's like, “Oh my gosh, this person really wants it to be concise.” The response would be one sentence, which is too terse. Just by removing the extra prompts around being concise, the model behaved in a much better way and much closer to what they actually end up wanting.
Brad Gerstner
Yeah, it turns out writing the right prompt is still important.
Sherwin Wu
Yes. Prompt engineering is still very, very important. On constructive feedback for GPT-5, there's actually been a good amount as well, which we're all working through. One of the things that I'm really excited for the next snapshot to fix is code quality and small code paradigms or idioms that it might use. I think there was feedback around the types of code and the patterns that it was using, which I think we're working through as well.
The other bit of feedback, which I think we've already made good progress on internally, is around the trade-off between reasoning tokens, thinking, latency, and intelligence. Especially for simpler problems, you don't usually need a lot of thinking. The thinking should ideally be a little bit more dynamic. Of course, we're always trying to squeeze as much reasoning and performance into as few reasoning tokens as possible. So I'd imagine that kind of going down as well.
Brad Gerstner
Yeah. Well, huge congrats. I know it's a work in motion for a bunch of our companies. They've had incredible outcomes with GPT-5. One of them is Expel, a cybersecurity business; it's a huge—
Sherwin Wu
Yeah, I saw the chart from that. It was pretty crazy.
Brad Gerstner
Huge, huge upgrade from whatever they were using prior to that. I think they're going to need a new eval soon.
Sherwin Wu
That's right. They're going to need a new eval.
9. Multimodality: Text, Voice, and Video
Brad Gerstner
It's all about evals. On the multimodality side, obviously you guys announced the real-time API last week. I saw T-Mobile was one of the featured customers on there. Talk about that. Obviously, the text models are leading the pack, but then we've got audio and video. Talk about the progress on the multimodal models. When should we expect to have the next big unlock, and what would that look like?
Sherwin Wu
It's a good question. The teams have been making amazing progress on multimodality: voice, image, and video, frankly. The last-generation models have been unlocking quite a few cool use cases. One piece of feedback that we've received is that, because text was leading the pack so much in intelligence, people felt like, in voice, the model was somewhat less intelligent. Until you actually see it, it does feel weird to have a better answer on text versus voice. That's pretty much the focus that we have at the moment.
I think we filled part of that gap, but not the full gap, for sure. Catching up with text would be one. A second one, which is absolutely fascinating, is that the model is excellent at easy, casual conversation—talking to your coach or your therapist. We basically had to teach the model to speak better in actual, economically valuable work setups.
To give an example, the model has to be able to understand what an SSN is and how to spell an SSN. If one digit is fuzzy, it actually has to repeat it rather than guess. There are lots of inflections like that that someone has in their voice that we are currently teaching the model. That's ongoing work with our customers. Until we actually confront the model with actual customer-support calls at actual scale, it's really hard to get a feel for those gaps. That's a top priority as well.
10. Audio: Realtime API vs Stitched Audio
Apoorv Agarwal
This is completely off script, but an interesting question that comes up in voice models, particularly the real-time API, is that previously people were taking a speech input, converting that to text, then having some layer of intelligence. Then you would have a text-to-speech model that would play it back. It would be a stitch of these 3 parts. The real-time API—you guys have integrated all of that. How does it happen? A lot of the logic is written in text. A lot of the Boolean logic or any function calling is written in text. How does it work with the real-time API?
Olivier Godement
That's an excellent question. The reason why we built the real-time API is that we saw a couple of issues with the stitched model. We call it a stitched-together model: speech-to-text, thinking, and text-to-speech. We saw essentially a couple of issues. One: slowness, with more hops, essentially. Two: loss of signal. With a stitched model, the speech-to-text model is less intelligent.
Brad Gerstner
Yeah, you'd lose the emotion.
Sherwin Wu
Exactly—pauses. When you're doing actual voice, like phone calls, those signals are so important. One of the challenges we have is what you mentioned: it means a slightly different architecture for text versus voice. That's something we're actively working on. But I think it was the right call to start with: let's make the voice experience sound natural to a point where you feel comfortable putting it in production, and then work backward to unify the orchestration logic across modalities.
A lot of customers still stitch these together. It's what worked in the last generation. But what we're interested in seeing is more and more customers moving toward the real-time approach because of how natural it sounds and how much lower the latency is, especially as we uplevel the intelligence of the model.
Bill Gurley
But also, even taking a step back, I will say it's pretty mind-blowing to me that it works. I think it's mind-blowing that these LLMs work at all: you just train them on a bunch of text, and they're autoregressively coming up with the next token, and it sounds super intelligent. That's mind-blowing in and of itself.
But I think it's actually even more mind-blowing that this speech-to-speech setup actually works correctly because you're literally taking the audio bits from someone speaking, streaming them, putting them into the model, and then it's generating audio bits back. To me, it's actually crazy that this works at all, let alone the fact that it can understand accents, tone, pauses, and things like that, and then also be intelligent enough to handle a support call or something like that. If you've gone from text-in, text-out to voice-in, voice-out, that's pretty crazy.
11. Model Customization & Reinforcement Fine-Tuning (RFT)
Apoorv Agarwal
We have a bunch of companies in our portfolio that are using these models: Parloa on the customer-support side, LiveKit on the infrastructure side. There are a bunch of use cases we're starting to see that a speech-to-speech model could address. A lot of the harder ones are still running on what you're calling the “stitched model.” But I hope the day is not far when it's all on the real-time API.
Sherwin Wu
It's going to happen at some point.
Brad Gerstner
Right, right, right. And actually, maybe that’s a good segue into talking about model customization because I suspect that you have such a wide variety of enterprise customers. I think you mentioned hundreds of customers, or maybe more. Each of them has a different use case, a different problem set, a different goal and envelope of parameters that they’re working in—maybe latency, maybe power, maybe others. How do you handle that? Talk about what OpenAI offers enterprises that need a customized version of a great model to make it great for them.
Olivier Godement
Yeah, model customization has actually been something that we’ve invested very deeply in on the API platform since the very beginning. Even before ChatGPT, we had a supervised fine-tuning API available, and people were using it to great effect. The most exciting thing around model customization, I think, obviously resonates quite well with customers because they want to bring in their own custom data and create their own custom version of o3, o4-mini, or even GPT-5, suited to their own needs. It’s very attractive, but the most recent development, which I think is very exciting, has been the introduction of reinforcement fine-tuning. It was something we announced late last year, I think during the 12 Days of OpenAI. We’ve since GA’d it, and we’re continuing to iterate on it.
Brad Gerstner
What is it? Break it down for us.
Olivier Godement
It’s called reinforcement fine-tuning. It’s actually funny—I think we made up the term reinforcement fine-tuning. It wasn’t a real thing until we announced it.
Brad Gerstner
It’s stuck now. I see it all the time. I remember we were discussing it, and I was like, “I don’t know about RFT.”
Olivier Godement
You’re not kidding. You’re not kidding. So, reinforcement fine-tuning introduces reinforcement learning into the fine-tuning process. The original fine-tuning API does something called supervised fine-tuning—we call it SFT. It does not use reinforcement learning; it uses supervised learning. Usually, that means you need a bunch of data, a bunch of prompt-completion pairs. You need to really supervise and tell the model exactly how it should be acting, and then, when you train it on our fine-tuning API, it moves the model closer in that direction.
Reinforcement fine-tuning introduces RL, or reinforcement learning, into the loop. It’s way more complex, way more finicky, but an order of magnitude more powerful. That’s what has really resonated with a lot of our customers. With RFT, the discussion is less about creating a custom model that’s specific to your own use case. You can use your own data and turn the crank on RL to create a best-in-class model for your particular use case. That’s the main difference here.
With RFT, the data set looks a little bit different. Instead of prompt-completion pairs, you really need a set of tasks that are very gradable. You need a grader that’s very objective that you can use here as well. That’s something we’ve invested a lot in over the last year, and we’ve seen a number of customers get really good results on it. We’ve talked about a couple of them across different verticals.
Rogo is a startup in the financial services space. They have a very sophisticated AI team—I think they hired some folks from DeepMind to run their AI program. They’ve been using RFT to get best-in-class results on parsing through financial documents, answering questions about them, and doing tasks around that as well. There’s another startup called Accordance that’s doing this in the tax space. I think they’ve been targeting an eval called TaxBench, which looks at CPA-style tasks as well. Because they’re able to turn it into a very gradable setup, they’re able to turn the RFT crank and get, I think, SOTA results on TaxBench just using our RFT product.
It has shifted the discussion away from just customizing something for your own use case to really leveraging your own data to create a best-in-class, maybe best-in-the-world, model for something that you care about for your business.
Apoorv Agarwal
Yeah, I feel like the base models are getting so good at instruction-following that, for behavior steering, you don’t need to fine-tune at that point. You can describe what you want, and the model is pretty good at it. But pushing the frontier on actual capabilities, my hunch is that RFT will pretty much become the norm. If you’re actually pushing intelligence in your field to a pretty high point, at some point you need to do RL, essentially, with custom environments.
Fascinating. And even going back to the point earlier around top-down versus bottom-up for some of these enterprises, a lot of the data that you end up needing for RFT requires very intricate knowledge about the exact task that you’re doing and understanding how to grade it. A lot of that actually comes from the bottom up. I know a lot of these startups will work with experts in their field to try to get the right tasks and the right feedback to craft some of these data sets.
12. Rapid Fire: Long/Short Picks
Without further ado, we’re going to jump into my favorite section, which is a rapid-fire question. We had a lot of great friends of ours send in some questions for you guys. We’ll start with Altimeter’s favorite game, which is a long-short game. Pick a business, an idea, or a startup that you’re long, and the same short that you would bet against because there’s more hype than reality. Whoever’s ready to go first: long, short.
Sherwin Wu
My long is actually not in the AI space, so this is going to be slightly different.
Brad Gerstner
Wow. Here we go.
Olivier Godement
My short is, though, in the AI space. I’m extremely long esports. By esports, I mean the entire professional gaming industry that’s emerging around video games. It’s very near and dear to my heart. I play a lot of video games, and so I watch a lot of this. Obviously, I’m pretty in the weeds on it. But I actually think there’s incredible untapped potential in esports and incredible growth to be had in this area.
Concretely, what I mean is a really big one: League of Legends. All of the games that Riot Games puts out have their own professional leagues. They have professional tournaments, believe it or not. They rent out stadiums now. But I just think that if you look at what the youth and younger kids are looking at, and where their time is going, it’s predominantly going toward these things. They spend a lot of time on video games.
Brad Gerstner
They watch more esports than soccer or basketball?
Olivier Godement
Yeah, yeah, yeah. A growing number of them do, too. I’ve actually been to some of these events, and it’s very interesting.
Brad Gerstner
He’s very committed to his long.
Olivier Godement
Yeah, yeah, yeah. I’m extremely long on this.
Brad Gerstner
And so they’re booking out stadiums for people to go watch esports.
Olivier Godement
Yeah, yeah, yeah. I literally went to Oracle Arena, the old Warriors’ stadium, to watch one of these, I think, before COVID.
Brad Gerstner
Before COVID? Wow, that’s 5 years ago.
Olivier Godement
6 years ago. So I’ve been following this for a while, and I actually think it had a really big moment during COVID. Everyone was playing video games. I think it’s kind of come back down, so I think it’s undervalued. No one’s really appreciating it now, but it has all the elements to really, really take off.
The youth are doing it. The other thing I’d say is it’s huge in Asia—absolutely massive in Asia. It’s absolutely big in Korea and China as well. We rented out Oracle Arena, I think, or the event I went to was in Oracle Arena. My sense is that in Asia they rent out entire stadiums, like soccer stadiums, and the players are treated like celebrities. Korean culture is really making its way into the U.S. as well, and I think that’s another tailwind for this whole thing. Anyway, esports is something you should keep an eye on because there’s a lot of room for growth.
Brad Gerstner
Very unexpected. Good to hear. Short?
Olivier Godement
My short is a little spicy: I’m short on the entire category of tooling around AI products. This encapsulates a lot of different things. It’s kind of cheating because some of these, I think, are starting to play out already.
But I think 2 years ago, it was maybe eval products, frameworks, or vector stores. I'm pretty short on those. I think nowadays there's a lot of additional excitement around other tooling for AI models. RL environments, I think, are really big right now as well. Unfortunately, I'm very short on those. I don't really see a lot of potential there. I see a lot of potential in reinforcement learning and applying it, but I think the startup space around RL environments is really tough.
The main thing is, 1, it's just a very competitive space. There's a lot of people operating in it. And then, 2, if the last 2 years have shown us anything, the space is evolving so quickly, and it's so difficult to try and adapt and understand what the exact stack is that will really carry through to the next generation of models. I think that just makes it very difficult when you're in the tooling space because today's really hot framework or really hot tool might just not get used in the next generation of models.
I've been noticing the same pattern, which is that the teams that build breakout startups in AI are extremely pragmatic. They are not super intellectual about the perfect world, et cetera. And it's funny because I feel like our generation basically started in tech at a very stable moment, where technology had been building up for years and years with SaaS, cloud, et cetera.
So we were, in a way, raised in that very stable moment where it makes sense at that point to design very good abstractions and tooling because you have a sense of where it's going. But it's so different today. You don't have a way to know what's going to happen in the next 1 or 2 years, so it's almost impossible to define the perfect tooling platform.
Brad Gerstner
Right. Right. Right. Well, that's—there's a lot of that going around right now. Yes. Spicy. A lot of homework there. Olivier, over to you, sir.
Olivier Godement
Long-short. I've been thinking a lot about education for the past month in the context of kids. I'm pretty short on any education that basically emphasizes human memorization at this point. And I say that having mostly been through that education myself, but I learned so much about history facts and legal things. Some of it does shape your way of thinking; a lot of it, frankly, is just knowledge tokens, essentially. And those knowledge tokens, it turns out, other AI models are pretty good at. So I'm quite short on that.
Brad Gerstner
You will need memory when strategy is bionic. You can just think about it straight into your head.
Olivier Godement
Exactly. Exactly. What am I long at? Frankly, I think healthcare is probably the industry that will benefit the most from AI in the next 1 or 2 years. I would say more. I think all the ingredients are here for a perfect storm: a huge amount of structured data—it's basically the heart of pharma companies—and the models are excellent at digesting and processing that kind of data.
There's also a huge amount of admin-heavy, document-heavy culture, but at the same time, companies that are very technical and very R&D-friendly, companies for which technology is, in a way, at the heart of what they do. And so, yeah, I'm pretty bullish on that.
Brad Gerstner
This is life sciences? So you mean life sciences research organizations that are producing drugs. Gotcha.
Olivier Godement
Exactly. It's almost like, over the last 20 or 30 years, these pharma or biotech companies have basically—if you look at the work that they're doing, only a small amount of it is actual research. So much of it ends up being admin and documents and things like that. And that area is just so ripe for something to happen with AI. I think that's what we're seeing with Amgen and some of these other customers.
Brad Gerstner
Exactly.
Olivier Godement
And it's also not what they want to do. I think it's good that we have some regulations there, obviously, but it just means that they have reams and reams of things to go through. And so when you have a technology that's able to really help bring down the cost of something like that, I think it'll just tear right through it.
And I think once governments and institutions realize that, if you step back, it is probably one of the biggest bottlenecks to human progress, right? You step back over the past decade—how many breakthrough drugs have there been? Not that many. How different would life be if you doubled that rate? Essentially. So once you realize what is at stake, then my hunch is that we're going to see quite a bit of momentum in that space.
Brad Gerstner
Wow. All right. Lots of homework there as well. Yeah. Next one: favorite underrated AI tool other than ChatGPT, maybe?
Olivier Godement
I love Granola.
Bill Gurley
Oh man, you stole mine. You stole my answer.
Olivier Godement
I use Granola so much.
Brad Gerstner
Two votes for Granola. Hey, what about ChatGPT Record?
Olivier Godement
I like ChatGPT as well, but there are some features of Granola that I think are really done well. The whole integration with your Google Calendar is excellent. And the quality of the transcription and the summary is pretty good.
Brad Gerstner
Do you just have it on? Because I know your calendar is back-to-back. Do you just have Granola on?
Olivier Godement
The funny thing is that I don't use Granola internally. I use Granola for my personal life mostly.
Brad Gerstner
I see. Yeah, I see. On dates. I'm joking.
Bill Gurley
I was going to say, yeah, Granola is actually going to be mine. So, 2 votes for Granola.
Olivier Godement
I was going to say the easy answer for me is Codex. As a software engineer, it's just gotten so good recently. Codex CLI, especially with GPT-5. Especially for me, I tend to be less time-sensitive about the iteration loop with coding, so leaning into GPT-5 on Codex, I think, has been really interesting.
Brad Gerstner
What about Codex has changed? Because Codex has also been through a journey. Codex has been around for a bit. I remember it launched more than a year ago. What's changed about Codex?
Olivier Godement
Yeah, I was actually going to say Codex CLI has been around for a bit. I feel like it's been less than a year for Codex.
Brad Gerstner
I feel like it's been less than a year for Codex. The time dilation is so crazy. It feels like it's been around for a year with GPT-5. That demo feels like ages ago, and it didn't even come out yet.
Olivier Godement
Probably because it hasn't happened yet.
Brad Gerstner
I think it was a naming thing, okay, but anyway—
Olivier Godement
Oh, there was a Codex model. That's what I'm thinking about.
Brad Gerstner
There was a Codex model. Also, I think the GitHub thing was called Codex.
Olivier Godement
That's right. Yes, yes. I'm talking about our coding product within ChatGPT, which is the Codex Cloud offering and then also Codex CLI. So, actually, maybe if I were to narrow my answer a little bit more, it's Codex CLI, which I've really, really liked.
I like the local environment setup. The thing that's actually made it really useful in the last, I'd say, month or so is, 1, I think the team has done a really good job of getting rid of all the paper cuts—the small product-polish and paper-cut things. It kind of feels like a joy to use now. It feels more reactive.
And then the second thing, honestly, is GPT-5. I just think GPT-5 really allows the product to shine. At the end of the day, this is a product that really is dependent on the underlying model. When you have to iterate and go back and forth with the model 4 or 5 times to get it right, to get it to do the change that you want, versus having it think a little bit longer and one-shot and do exactly what you want, you get this weird, bionic feeling where you're like, “I feel so mind-melded with the model right now, and it perfectly understands what I'm doing.”
Getting that kind of dopamine hit and feedback loop constantly with Codex has made it an indispensable thing that I really, really like.
Sherwin Wu
The other thing I’d say Codex is really good for, for me, is personal projects. I also use it to help me understand codebases. As an engineering manager, I’m not as in the weeds on the actual code, and so you’re able to use Codex to really understand what’s happening with the codebase, have it ask questions and answer them, and really catch up to speed on things as well. Even the non-coding use cases are really useful with Codex CLI.
Brad Gerstner
Fascinating. Sam had this tweet about Codex usage ripping, I think, yesterday. So I wonder what’s going on there, but you’re not alone.
Sherwin Wu
Yeah, I think I’m not alone. Just judging from the Twitter feedback, I think people are really realizing how great of a combination Codex CLI and GPT-5 are.
Bill Gurley
Yeah, I know that team is undergoing a lot of scaling challenges, but the system hasn’t gone down for me, so props to them. But we are in a GPU crunch, so we’ll see how long that goes.
Brad Gerstner
Awesome. All right, the next one: Will there be more software engineers in 10 years or less? There are about 40–50 million full-time, professional software engineers.
Sherwin Wu
That’s what you mean—full-time, actual jobs?
Brad Gerstner
Yeah, because it’s a hard one. I think without a doubt there’s going to be a lot more software engineering going on.
Sherwin Wu
Yes, of course. There’s actually a really great post that was shared, I think, in our internal Slack. It was a Reddit post recently, and I actually think that highlights this. It was a really touching story.
It was a Reddit post about someone who has a brother who’s nonverbal. They have to take care of him. They tried all these types of things to help the brother interact with the world and use computers, but vision tracking didn’t work because I think his vision wasn’t good. All the tools didn’t work, and then this brother ended up using ChatGPT.
I don’t think he used Codex, but he used ChatGPT and basically taught himself how to create a set of tools that were tailor-made for his nonverbal brother—a custom software application just for them. Because of that, he now has this custom setup that was written by his brother and allows him to browse the internet. I think the video was of him watching The Simpsons or something like that, which was really touching.
I think that’s actually what we’ll see a lot more of. This guy’s not a professional software engineer. His title isn’t software engineer, but he did a lot of software engineering—probably pretty good, and definitely good enough for his brother to use. The amount of code, the amount of building that’ll happen, is just going to go through an incredible transformation.
I’m not sure what that means for software engineers like myself. Maybe there’s—of course, more Sherwin.
Brad Gerstner
Of course, more Sherwin.
Sherwin Wu
More of me specifically.
Brad Gerstner
We need more of you.
Sherwin Wu
That’s right. But definitely, there’ll be a lot more software engineering at a lot of companies.
Apoorv Agarwal
I buy that completely. I completely buy the thesis that there is a massive software shortage in the world. We’ve been accepting it for the past 20 years. But the goal of software was never to be that super-rigid, super-hard-to-build artifact. It was to be customized and malleable.
So I expect that we’ll see more of a reconfiguration of people’s jobs and skill sets, where way more people code. I expect that product managers are going to code more and more, for instance. You made your PMs code recently, if I remember right.
Olivier Godement
Oh, yeah, we did that. It was really fun. We started essentially not doing PRDs—product requirements documents. Classic PM thing: You write 5 pages, “My product does that,” et cetera. And PMs have basically been coding prototypes. One is pretty fast with GPT-5 and Codex—just a couple of hours, I think.
Bill Gurley
Fricking fast.
Olivier Godement
And second, it sort of conveys so much more information than a document. You get a feel, essentially, for the feature: Is it right or not?
Brad Gerstner
Yeah, instead of writing English, you can actually now write the actual thing you want. That’s amazing. Advice for high school students who are just starting out their careers?
Olivier Godement
My advice is—I don’t know. Maybe it’s evergreen: Prioritize critical thinking above anything else. If you go into a field that requires extremely high critical-thinking skills—I don’t know, math, physics, or maybe philosophy is in that bucket—you will be fine regardless.
If you go into a field that turns down that thing, and again, it gets back to memorization and pattern matching, I think you will probably be less future-proof.
Brad Gerstner
What’s a good way to sharpen critical thinking?
Olivier Godement
Use ChatGPT and have it test you. That’s a tricky test. Having a world-class tutor who essentially knows how to put the bar about 20% above what you can do all the time is actually probably a really good way to do it.
Brad Gerstner
Nice. Anything from you, sir?
Sherwin Wu
Mine is—I think we’re actually in such an interesting, unique time period. So maybe this is more general advice, not just for high school students, but for the younger generation, even college students.
I think the advice would be: Don’t underestimate how much of an advantage you have relative to the rest of the world right now because of how AI-native you might be, or how versed in these tools you are. My hunch is that high schoolers and college students, when they come into the workplace, are going to have a huge leg up in how to use AI tools and how to actually transform the workplace.
My push for some of the younger high school students is, first, just really immerse yourself in this thing. Second, really take advantage of the fact that you’re in a unique time where no one else in the workforce really understands these tools as deeply, probably, as you do.
A good example of this is that we had our first intern class recently at OpenAI—a lot of software interns. Some of them were just the most incredible Cursor power users I’ve ever seen. They were so productive. I was shocked, by the way. I was like, “Yeah, I know we can get good interns, but I don’t know if they’d be this good.”
I think part of it is they’ve grown up using these tools, for better or worse, in college. But I think the meta-level point is they’re so AI-native. Even Olivier and I are kind of AI-native—we work at OpenAI—but we haven’t been steeped in this and grown up in this.
The advice here would just be: Leverage that. Don’t be afraid to go in and spread this knowledge and take advantage of it in the workplace, because it is a pretty big advantage for them.
I can’t remember who said this to us at Palantir, but every intern class was just getting faster and smarter, like laptops getting smarter every generation. You sure didn’t peak in 2013, when I was an intern.
Bill Gurley
That’s right. There’s a weird spike. That’s summer 2013.
Brad Gerstner
Well, lots happened here. A lot’s happened since you guys joined OpenAI, right? It’s been 3 years and almost 3 years. In your OpenAI journey, what has been the rose moment—your favorite moment; the bud moment, where you’re most excited about something but there’s still opportunity ahead; and the thorn, the toughest moment of your 3-year journey?
Olivier Godement
The thorn is easy for me: What we call the blip, which was the board coup. That was a really tough moment. It’s funny because, after the fact, it actually reunited the company quite a bit. OpenAI had a pretty strong culture before, but there was a feeling of camaraderie that was even stronger. But it was tough on the day.
Bill Gurley
It’s very rare to see that kind of antifragility. Most organizations, after something like that, break apart, but I feel like OpenAI got stronger. OpenAI came back.
Olivier Godement
It’s a good point. I feel it made OpenAI stronger for real, essentially, when they look at it after the fact. When they look at other news, like departures or whatever—bad news, essentially—I feel the company has built a thicker skin and an ability to recover way quicker.
Bill Gurley
I think that’s definitely right. Part of it, too, I think, is also just the culture. I also think this is why it was such a low point for a lot of people.
Olivier Godement
So many people at OpenAI care so deeply about what we're doing, which is why they work so hard. You just care a lot about the work. It almost feels like your life's work. It's a very audacious mission and thing that you're doing, which is why I think the blip was so tough on a lot of people, but also what I think helped bring people back together and how we were able to hold together and get that thick skin as well.
I have a separate worst moment, which was the big outage that we had in December of last year. You remember.
Brad Gerstner
I do.
Olivier Godement
It was a multi-hour outage. It really highlighted to us how essential—almost like a utility—the API was. I think we had a 3- or 4-hour outage sometime in November or December last year. It was really brutal, a pure shitshow. No one could hit ChatGPT. No one could hit the APIs. It was really rough.
That was just really tough from a customer-trust perspective. I remember we talked to a lot of our customers to postmortem with them what happened and our plan moving forward. Thankfully, we haven't had anything close to that since then. I've actually been really happy with all the investments we've made in reliability over the last 6 months. But in that moment, I think it was really tough.
13. Highlights and Lowlights @ OpenAI
On the happy side, on the roses, I think I have 2 of them. The first one would be that GPT-5 was really good. The sprint up to GPT-5 really showed the best of OpenAI: cutting-edge science and research, extreme customer focus, and extreme infrastructure and inference talent. The fact that we were able to ship such a big model and scale it to many, many, many tokens per minute almost immediately, I think, speaks to it.
Brad Gerstner
With no outages.
Olivier Godement
Yeah, really good reliability. I can remember when we shipped GPT-4 Turbo, like a year ago, a year and a half ago, we were terrified by the insane-scale traffic. I feel we've really gotten much better at shipping those massive updates.
The second happy moment for me would be the first Dev Day. It felt like a coming of age for OpenAI. We were embracing that we have a huge community of developers. We were going to ship models and products. I remember seeing all my favorite people, OpenAI or not, nerding out on what they were building and what was coming up next. It felt really like a special moment in time.
That was actually going to be mine as well, so I'll just piggyback off of that: the very first Dev Day, November 2023. I remember it. Obviously, a lot of good things have happened since then, but for me, it was a very memorable moment.
One, it was actually quite a rush up to Dev Day. We shipped a lot, so our team was just really, really sprinting. It was this high-stress environment going up. To add to that, of course, because we're OpenAI, we did a live demo during Sam's keynote of all the stuff that we shipped. I remember being in the back of the audience, sitting with the team and waiting for the demo to happen. Once it finished, we all just let out a huge sigh of relief. We were like, “Oh my God, thank you.”
For me, the most memorable thing was right after Dev Day. All the demos worked well, all the talks worked well, and we had the after-party. Then I was just in a Waymo, driving home at night with the music playing. It was such a great end to Dev Day. That was what I remember. That was my rose for the last few years.
Brad Gerstner
Love it. That's awesome. I assume you guys are, but please tell me if you're AGI-pilled, yes or no. And if so, what was the moment that got you there? What was your aha moment? When did you feel the AGI?
Olivier Godement
I think I'm AGI-pilled. You're definitely AGI-pilled? I am? I've had a couple of them.
The first one was the realization in 2023 that I would never need to code manually ever again. I'm not the best coder; I chose my job for a reason. But realizing that what I thought was a given—that we humans would have to write basically machine language forever—is actually not a given, and that the paradigm shift is huge.
The second feel-the-AGI moment for me was maybe the progress on voice and multimodality. Text, at some point, you get used to it: the machine can write pretty good text. Voice makes it real. But once you start actually talking to something that understands your tone, understands my accent in French, it felt like a moment—machines are going beyond cold, mechanical, deterministic logic to something much more emotional and tangible.
Yeah, that's a great one. Mine are—I do think I am AGI-pilled. I probably gradually became AGI-pilled over the last couple of years. I think there are 2, and for me, I actually get more shocked from the text models. I know the multimodal ones are really great as well. For me, I think they line up with 2 general breakthroughs.
The first one was right when I joined the company in September 2022. It was pre-ChatGPT, 2 months before, about the time GPT-4 already existed internally. I think we were trying to figure out how to deploy it. I think Nick Turley talked about this a lot early on with ChatGPT. But it was the first time I talked to GPT-4, and it was like going from nothing to GPT-4. It was the most mind-blowing experience for me.
For the rest of the world, maybe going from nothing to GPT-3.5 in ChatGPT was the big one, and then going from GPT-3.5 to GPT-4. But for me, and I think for a lot of other people who joined around that time, going from—not nothing, but what was publicly available at the time—to GPT-4 was just incredible. I remember asking it and throwing so many things at it. I was like, “There's no way this thing is going to be able to give an intelligible answer.” And it just knocked it out of the park. It was absolutely incredible. GPT-4 was insane.
I remember GPT-4 came out when I was interviewing with OpenAI, and I was still like, “Should I join?” And then I was like, “Okay. I mean, there is no way I can work on anything else at that point.”
That's true. Yeah, GPT-4 was just crazy. And then the other breakthrough was the reasoning paradigm. I actually think the purest representation of that for me was Deep Research.
Asking it to really look up things that I didn't think it would be able to know, and seeing it think through all of it, be really persistent with the search, get really detailed with the write-up, and all of that—that was pretty crazy. I don't remember the exact query that I threw at it, but I just remember that the feel-the-AGI moments for me are when I'll throw something at the model that I think, “There's no way this thing will be able to get,” and then it just knocks it out of the park. That is kind of the feel-the-AGI moment. I definitely had that with Deep Research, with some of the things I was asking.
Brad Gerstner
Well, this has been great. Thank you so much, folks. You guys are building the future. You guys are inspiring us every day, and I appreciate the conversation.
Olivier Godement
Yeah, thank you so much.
Thank you. Thanks for having us.