Shawn Wang
Okay, we're here in the studio with Akshay from OpenAI. Welcome.
Akshay Kothari
Thank you.
Shawn Wang
And with our trusty co-host, Vibhu. You recently launched ChatGPT Work. You lead Core Product Engineering. It's been a long journey into all this. I find it very interesting that you started with no-code or low-code, with Walrus and Airtable. To some extent, ChatGPT Work is kind of like the super app of super apps. Here is the ultimate no-code: You just write a prompt.
1. The No Code Thesis
Akshay Kothari
Yeah, it's funny how things come full circle. I started my career working in consumer fintech, but after that, there was this hypothesis that the things we were able to do with code as engineers—if we could bring that to many more people in a more accessible way—would be truly magical.
We were working on a startup before LLMs, before vision LLMs, on how to do automated testing with AI. It was just kind of janky back then, but we were doing what we could. Then I worked at Airtable for a while on the same thesis: If we can bring a database, or the primitives behind a database, to people, that would be really useful to them.
Once LLMs came onto the scene, it became clear that this was the missing piece—the missing technology required to bring the magic of code to everyone without them having to know what's going on under the hood. I think this launch, and a lot of the stuff that we've been up to, is the manifestation of that.
Vibhu
How was stuff when you joined? You joined OpenAI in 2023. Now we've got so much more stuff: ChatGPT, the Codex app, ChatGPT for Work. Have things changed?
2. OpenAI Still Feels Startup
Akshay Kothari
Actually, I think the more interesting thing is how things haven't changed. I joined when it was around 500 people. One thing I was worried about was that I was looking for something more early-stage, and I wondered whether it was going to feel startup enough. I joined, and I was like, "This feels even more startup-y than I could ever imagine." That really hasn't changed even till now.
I think the level of bottoms-up ambition, and the ability of anyone to do anything or have an idea and ship it, is really cool. On the mission side, what was really compelling to me was this mission of bringing frontier intelligence to everyone—building AGI and then bringing it to everyone.
I think we acknowledged back then that this vision was not going to be a linear progression. We're probably going to try different products and have different things that succeed and don't. But the vision has stayed the same, and the mission has stayed the same. We're starting to see the pieces fall together, and that's really cool.
Shawn Wang
You worked on Enterprise. A lot of people never touch ChatGPT—ChatGPT Enterprise, though. What is something that you learned from there that you're bringing into your work now?
3. Enterprise Has No Single Use Case
Akshay Kothari
I think there's no one-size-fits-all solution in Enterprise. I remember in the early days of ChatGPT Enterprise, when we talked to customers, everyone was so excited to bring AI into their enterprise. This was a year after ChatGPT was released, and there were all these teams being stood up, like AI deployment teams with enormous budgets.
If you asked anyone what they were excited about solving, at first you'd get the baseline answers: "Yeah, we have all this context and data and all this stuff." But if you asked them what discrete use case they wanted AI to enable in their workplace, you got such a different variance—an explosion of different types of answers.
Using these models and products, you have this box, and you can say anything to it, which is the magic. On the flip side, it also means that you don't know what to do with it. In Enterprise, I think a big part of that is actually meeting the users where they are: What use case were they trying to solve, and how can they use AI to gain leverage there?
Shawn Wang
Do you meaningfully differentiate that from forward-deployed engineering?
Akshay Kothari
I think there's the go-to-market side of it—
Shawn Wang
Yeah.
Akshay Kothari
—and then there's the product side of it. You need someone on the product side. However good we get at the FDE motion, I think at the end of the day, if we have a user who's looking at their computer or looking at their phone, it's our job in the product to enable them and show them where to go. We're really excited about that.
Vibhu
Do you think there have been changes over the past 3 years in adoption? There have been step-function changes. You have reasoning models and whatnot. Is it still the same problem—that Enterprise has a black box and doesn't know what to do with it—or have things changed?
Akshay Kothari
We're seeing now that there's this huge uptake. Everyone's extremely excited about it. It feels like millions, hundreds of millions of people are using ChatGPT. They understand generally how to work with AI.
But every time a new capability gets unlocked—now we're seeing it with agents—there's probably still a contingent of early adopters who truly get it. They're like, "You can do anything. You just have to make sure the right context is there, that it's connected to the right tools, and that you're supervising it. Anything is possible."
But then there's this 10X or 100X bigger market where they don't yet get that, or they don't yet see that. I think that's the next stage here. To answer your question, I think the adoption is there and growing fast, but the opportunity is far, far bigger than that. That's where we want to play, especially with ChatGPT Work.
Shawn Wang
Well, let's skip ahead to ChatGPT Work. It was only announced about a month ago. What was the decision process that led into it? There was this overall merging of the super app. Is that what we're officially calling it? You deprecated the browser as well. I guess, just summarize your last couple of months of working on this thing.
4. ChatGPT Work Grows From Codex
Akshay Kothari
It feels like forever now, but I guess it's only been a few months. I think maybe the one impetus that is most salient is when we re-released Codex, or even internally had Codex. It was really surprising to us. We recently put out some stats on this: There was a real inflection of adoption among non-developers at OpenAI.
Through this product development process, I would go to UX research sessions to talk to people internally. The thing that stuck out to me was that, first, you go talk to strategic finance or marketing or whatever, and they're all using Codex for their use cases. That part is cool, but the thing that really stuck out to me was how proud people were that they were using Codex.
Shawn Wang
It's like, "I'm not supposed to be using it, but I am."
Akshay Kothari
It was that.
Shawn Wang
Yeah.
Akshay Kothari
It was that they were early to this new thing, but it was also that they felt like they had a superpower. What we recognized then was that the power of Codex, the power of agents, was already available to this massive distribution base of people who have come to know and love ChatGPT.
How do we show that to them? How do we bring it to them? That's a hard product problem, and it's a tricky thing. There are many ways you can go about it. That's why we called it the Merge and the Super App over time, and ultimately launched it in ChatGPT Work. How do we do that? It came from that initial realization that the power was not only for developers, much earlier than probably even we thought. It could be extended to everyone.
Shawn Wang
How do you see the products differently? Who is it for? Codex started out as a CLI, then an app. Now there's a merge of ChatGPT, Codex, and ChatGPT Work. Is it the opening for the average user, for Enterprise, or for work? How do you position it?
Akshay Kothari
I think we want to position it for doing work-related things, for lack of a better word. Productivity is actually the pillar that I support. That's the name of the team. The reason we call it productivity and not Enterprise, work, or something like that is because there's also personal productivity. I think ChatGPT Work is—
I've seen people do things in their personal lives that you wouldn't technically classify as work, but these agents are super capable of. One recent example that someone posted about on our Slack was someone who had a missed package. They didn't receive it, and then they got a picture of it from Amazon or whoever the courier was. They asked ChatGPT Work to find out where that package was.
The agent is extremely tenacious. It took the image, looked at a bunch of listings around their neighborhood, and figured out exactly which apartment complex the package was in. It gave them some information. I think there are all these things that are work-y or productivity-related, and I think that's what we want the product to be.
You asked about Codex. We think Codex is a durable brand, but we have a principle: We don't want a user to get stuck in a tab or an experience where they don't get the power of the product. Basically, everything that you can do in the Codex portion of the product on desktop, you can do in ChatGPT Work, and vice versa.
But we made some opinionated product decisions about how much of the Git state, if you're in a Git repo, we want to expose to the end user. How much do we want to make the experience of seeing the agents thinking diff-forward, so that you get exposed to the diffs out of the box? On the safety side, how do we want to think about sandboxing and making sure that we have the right defaults in one state versus the other?
There are some opinions behind that, but we don't want the user to need to choose which experience they're in.
Shawn Wang
That is a good goal for AGI, right? People don't want to choose which version of AGI they want. They just want the AGI to decide for them.
Akshay Kothari
Yeah.
Shawn Wang
Can I get an answer? It's not super clear to me. Is the Codex harness and the ChatGPT Work harness the same? Is it just UI affordances, or are there actually prompt-level or even deeper differences?
5. One Harness Powers Both Products
Akshay Kothari
The harness is the same. The harness is shared. In both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plugins, computer use, or artifacts. You get that power regardless of which experience you're in.
On the UX side, we have opinionated takes about what the UX should be and how it should behave when you're in Codex mode, and some stuff around the sandbox, like I mentioned. But the underlying harness and capabilities should be the same.
Shawn Wang
I'm just curious. Maybe we can—Is there a query that we can run that would look different in the 2 modes?
Akshay Kothari
Yeah. Try to create a retirement calculator spreadsheet, or something, in both modes. In Codex mode, you might have to be in a repo for this, but you'll see the diffs of the sheet that it's creating and the file edits. But in Work, you won't be able to see that.
Shawn Wang
I think that's super clear. The other thing I wanted to dive into was your productivity team. What else is there? First of all, what are the top-level teams other than productivity? Isn't productivity everything?
Akshay Kothari
So—
Shawn Wang
Science.
Akshay Kothari
We have a team focused—
Shawn Wang
Yeah.
Akshay Kothari
—on ChatGPT.
Shawn Wang
Yeah.
Akshay Kothari
The core chat experience for consumers, which is not, I think, all productivity. People are using ChatGPT every day for search, to figure out how to write messages to loved ones, to think about how to learn a new topic, et cetera. There's so much more inside it, like creating images. There's so much more in Chat that the hundreds of millions of users are using, and obviously that warrants a very dedicated effort.
There are teams focused on enterprise, infrastructure, API, and stuff like that.
Shawn Wang
I will bring it up.
Akshay Kothari
Okay.
Shawn Wang
Yeah. So I have them both running. This is Work. There's a Codex version here. I picked the “Five Little Dicks” song, so this will take a while.
Akshay Kothari
Uh-huh.
Shawn Wang
I think we'll just keep it in the background and, as they finish, we'll look into some of the differences.
Akshay Kothari
Yeah. But immediately, I think if you flip back to the Codex version—
Shawn Wang
Yeah.
Akshay Kothari
—you'll see that—
Shawn Wang
It assumes Git. Yeah.
Akshay Kothari
Exactly. The dynamic island assumes that you're in a Git repo. You might miss some stuff because some of it is in the actual chain of thought, but with those changes and how we display that, yeah.
Shawn Wang
Is there an unintuitive—Is there a thing that you wanted to ship, and then you got feedback, and you were like, “No, let's not do it?” What's the thinking behind that?
Akshay Kothari
In ChatGPT Work?
Shawn Wang
Yeah.
Akshay Kothari
I think one direction we could have gone with this was keeping the experiences completely separate. Why—
Shawn Wang
Different apps.
Akshay Kothari
Exactly, like different apps, or even in the same app, completely different experiences. Why merge it all? You know, Codex obviously people love. Why bring these products together?
I think the intuition here is that all of our jobs are changing dramatically with AI. Every few months, I feel like I wake up and I'm doing a completely different thing than I was doing a few months ago. My hypothesis here is—or I should say our hypothesis—that part of what we're building with this technology is giving people leverage.
Maybe it's the more mundane parts of your job, or parts that, if you were able to automate, you'd be able to share more ideas faster or whatever you're able to do now. Because of that, that might actually blur the lines between someone who's only writing code, creating strategy docs, planning events, helping with marketing, doing podcasts, or whatever.
These things are going to get blurred over time. Trying to draw a hard boundary based on who you are is going to be tough. We should enable users to choose, but we shouldn't box them in.
A lot of the work that went into this, like keeping the primitives the same—for example, plugins are unified across this product, ChatGPT, and the cloud—was because of that. It's this thesis that eventually things are going to come together, and we don't want to box anyone in. We want to be prescriptive about when to be in either experience, but we don't want to box anyone in.
Shawn Wang
I wonder if there are users who are very tuned to the old ChatGPT harness, which is effectively now replaced by the Codex harness. I can't imagine what that was, but maybe they're more on the conversational side. Can you compare and contrast the 2 harnesses? Only you've seen it.
Akshay Kothari
Shawn Wang
Yeah. I mean, I think the existing ChatGPT harness still exists today. It exists in this app.
Akshay Kothari
The classic—
Alessio Fanelli
You just start a new chat, and you don't go under Work, right?
Akshay Kothari
Yeah. If you start—
Alessio Fanelli
So—
Akshay Kothari
—a new chat and go to Chat, then—
Alessio Fanelli
Yeah.
Akshay Kothari
—you’re talking to ChatGPT with the Instant model.
Alessio Fanelli
Oh, we can technically do another.
Akshay Kothari
Yeah.
Alessio Fanelli
But I guess it's on Instant.
Shawn Wang
Yeah. So this one's not going to code, or it's going to—
Alessio Fanelli
Oh—
Shawn Wang
—be inline. It's inline in a sandbox.
Akshay Kothari
It'll actually—
Alessio Fanelli
Oh, that's cool.
Akshay Kothari
—we try to push you to go to Work if you're creating—
Shawn Wang
Yeah—
Akshay Kothari
—a spreadsheet.
Shawn Wang
And this is a router decision?
Akshay Kothari
Sorry?
Shawn Wang
Is this a router decision?
Akshay Kothari
This is the decision that the model is making. It sees that you're trying to do something that would be better served in Work mode.
Shawn Wang
Right.
Akshay Kothari
But I think your question was: What are the advantages of the ChatGPT chat harness?
Shawn Wang
It's more broadly that I want to do an oral history of harness engineering. The ChatGPT harness lasted us from, let's call it, the o1 era until now, and now it's effectively being replaced by the Codex harness. They're overlapping somewhat, but I'm curious what changed, if anything.
Akshay Kothari
Mm-hmm.
Shawn Wang
My perspective on this is that there's sort of a constant process of divergence, convergence, divergence, convergence. In ChatGPT, many of the use cases I was talking about before—search or learning—I think we're really optimizing for latency, personality, and different things. The reason people love ChatGPT is that we've been optimizing for those things and working on them for so long.
Akshay Kothari
With Codex, what we learned was that if you give the agent access to this infinitely flexible environment as a computer, it can do really powerful things. When we think about knowledge work, which mode should we choose? It felt more natural to us to bring that to this computer environment and abstract some of the details of the computer away from users who might not be used to it, while giving them that same power.
Ultimately, I think we want the power in all places. We want to meet people where they are. I’m sure there’ll be work down the road to get things to be equivalently capable in all scenarios. It’s just a question of what we’ve historically been focusing on in the product and what we’re focusing on now.
Alessio Fanelli
I think alongside that, outside of just the harness and when to use Codex, ChatGPT, or Work, there are also the new models you’ve released. Any guidance there? People love to min-max what to use: only use Tera on high reasoning, versus, for this, you want to use Sol here and ignore all these—
Akshay Kothari
There are 32 options.
Alessio Fanelli
Yeah, yeah. But that being said, for people who are exploring productivity stuff and trying things for work, who don’t have a breakdown of what all this is, what’s the advice?
Akshay Kothari
Before the advice, I think the first thing is that none of this would be possible without these models. I think you asked earlier what the inspiration for Work was, and early on I mentioned what we were seeing with Codex, but that was also because the models were getting infinitely more capable. That’s happening again. I think it’s another step-function jump now.
To answer the question on advice, we want the default to be the best possible. We want to be opinionated about the default, so we’ve chosen a default that we think is going to be the best for everyone. For power users, we have options under the hood. One could argue that there might be too many right now, and we’re working on simplifying it. You can extend the reasoning level and change between the different model classes if you need to, but the default should be the best for most use cases.
My advice to most people would be to stick to that. If you reach a situation in which you think you want to try a different configuration, and you’re not seeing either the efficiency on the cost side or the quality on the intelligence side, then you can change the defaults and see if you can get something better. But we think the default should be good enough.
Shawn Wang
I’m just going to run something by you, since you have way more experience than me. I’ve recently been doing Sol Lite with Goal, with the idea that Goal basically augments the reasoning effort, but with more terminations and turns. Is that a good way to think about it, as opposed to Sol Ultra or Sol Extra High?
Akshay Kothari
Yeah, it’s hard to say because—
Shawn Wang
Yeah. It’s like an interaction effect.
Akshay Kothari
Exactly. There’s a preference for you as an individual in how you like to collaborate with the models. How many of those terminations, as you call them, do you want, where you can steer or make sure that it’s doing the right thing?
I think people should generally try whatever works for them. Using Ultra or multi-agent setups is best for tasks that are either incredibly complicated, like open explorations, or very parallelizable. Using Goal, I think, is best for tasks where you know you’ll be able to make consistent progress in a way that’s verifiable over time.
For most tasks, they actually don’t fall into either of those buckets, at least when they’re starting. That’s why I think the best first step is trying it with the default configuration and then seeing where you want to go from there.
swyx
Right. You guys worked on a slider, which is actually super helpful for reducing the amount of panic.
Logan Kilpatrick
Yeah. Yeah.
Alessio Fanelli
It’s nice on mobile, at least. There’s a nice slider there.
swyx
It’s nicer.
I haven’t tried it.
swyx
You have the advanced view there, but if you click Advanced view—yeah, yeah.
Ooh.
swyx
Just a nice slider. Very pretty, very colorful.
Logan Kilpatrick
The idea here was to reduce it to 1 dimension, even though there are multiple dimensions, right? Try to project it onto a single dimension for the user—something that represents speed and efficiency on one side, and quality and thoroughness on the other side.
swyx
Yeah.
Logan Kilpatrick
Yeah.
swyx
I’m just puzzled that it uses Sol so much, like the lower—
Alessio Fanelli
No, no. I think the slider, if I’m not mistaken, is—
swyx
Terra.
Alessio Fanelli
Oh, it is.
swyx
Yeah. See?
Nice. Nice.
swyx
So they preset Terra to only be the light one.
Alessio Fanelli
I see.
swyx
I think a lot of people actually would—more people should use Terra. One reason is that Sol keeps running out of capacity.
I’m the reason, you know. Here’s 10 minutes of our retirement calculator.
swyx
There you go.
Oh, that’s the Excel thing we’re displaying. This is Work, and then Codex is still cooking, so we’ll get back into it. I think it’ll be interesting to actually see the thought process and the reasoning. Also, I guess this is 8 minutes on Work. Codex is still cooking.
swyx
Yeah. By the way, do you know Gabriel Chua? He’s part of the OpenAI Singapore team. He showed me this, and I was pretty shocked that it looks like Excel. It edits Excel files. You never paid for an Excel license, right? But somehow this is kind of workable, and it’s agentic Excel.
6. Artifacts Expand Knowledge Work
Logan Kilpatrick
One of the big pushes that we made for this launch was artifacts, right? Both on the model side—I think if you compare this with 5.5 and 5.4 before that, you’ll see that there have been pretty dramatic improvements in the quality of these artifacts—and then also on the product side.
Alessio Fanelli
The UX side is also crazy. Hosted sites and whatnot, no longer needing to host your own little webpage.
swyx
Oh, I have a story about that.
Yeah.
swyx
I can do a separate thing. I’ll need to take the visuals here, but we’ll cut to that later. Was there co-training, I guess, because you’re making this big move, and you launched GPT-5.6 on the same day as ChatGPT Work? Was there influence between the model training teams and the harness teams, or did the launch dates just happen to line up on the same day?
Logan Kilpatrick
I think we collaborate heavily with the research teams, and I think that’s one of the most magical parts of the job—the most fun parts of the job. Using artifacts as an example, a lot of what you’re seeing underneath the hood comes from the work that went into making sure we had the right infrastructure to train the models to get better at this, and then, on the product side, to have the right experience for users to be able to collaborate with the model on an artifact like this.
In fact, this whole viewer—the intuition here is that it’s not necessarily that you wouldn’t need an Excel license. This is stage 1, right? This is probably not what you meant when you were making a retirement calculator.
Alessio Fanelli
Yeah, you can iterate very easily.
Logan Kilpatrick
You want to iterate, and when you see it, if this thing is high fidelity to what you would actually see—or what your coworkers would see if you were to send this to Sean—that makes it so much easier and makes you trust the product in terms of iteration.
Alessio Fanelli
When you say coworkers would see, do you see multiplayer, multi-team collaboration with artifacts? Are there any things you guys think about there?
swyx
You can already share it, right?
Yeah.
Logan Kilpatrick
It’s something that we’re actively thinking about. One thing that we’ve noticed internally, without talking too much about the roadmap, is that there are many times when someone will ping me about something, and I will ask ChatGPT Work the question, and then I’ll ping them back the answer. And then I’ll be thinking—
Alessio Fanelli
The simplest would be the 3 of us all on one hosted—
Logan Kilpatrick
Exactly. And I’ll think about whether I was required in this loop, or whether it was just a rephrasing of what they were asking or pulling from certain context, or whatever.
But when I gave them back the answer, that process was also lossy, right? I gave them just my interpretation of what ChatGPT Work cooked up. But underneath the hood, there's so much context in the rollout and stuff that could be interesting.
swyx
So the answer was to preemptively respond to every inbound request?
No, it was literally just what I do sometimes as my job.
swyx
I know. You copy-paste—
Yeah.
swyx
—and then you're just a message-forwarding service—
Logan Kilpatrick
Yeah. Yeah, exactly.
swyx
—from AI to AI.
Alessio Fanelli
But I think it's interesting, right? It helps people understand the capabilities of what you can ask and delegate.
Logan Kilpatrick
Yeah.
Alessio Fanelli
Oftentimes, people don't realize until they try or someone shows you, and then you're like, “Oh, okay, okay, I see now.”
Logan Kilpatrick
Yeah.
swyx
I think there's also a light security issue where, basically, you're the permissions layer. Yes, I could query everything that you query and get an automated response, but maybe I'm not supposed to see it.
Logan Kilpatrick
Yeah.
swyx
There's no way I would know, because I'm not supposed to know what I don't know.
Logan Kilpatrick
Especially with ChatGPT Work, we're asking you to connect your plugins, and it's pulling from your local files and stuff like that. The amount of context that the agent has access to is deeply personal, and that's something I think we need to preserve. So that'll definitely be a challenge.
swyx
There's Excel, there's PowerPoint, there's Docs—the grand trio of work. What other formats of work do you think about? Obviously, you worked on Airtable. Is there a future where there's OpenAI Airtable? What does that look like if you ever ended up doing it?
It's a really good question. One that you didn't bring up was Sites, and I think that—
swyx
Sites.
Logan Kilpatrick
—that was a core part of this launch. There's one side of Sites that I think people commonly talk about, especially on Twitter or X, as this sort of prototyping tool. We saw that happen with this launch, even. The model slider that you guys were referencing earlier was developed almost fully in a Site.
The collaboration between design, engineering, and product on that was on a Site where we played with the affordance and figured out how it feels and all of that.
But the other aspect that's a little bit less talked about is Sites as an artifact for knowledge work. I was actually talking to someone the other day who's on our corporate finance team, and we were mentioning how now, when they have these reports that they're working on as a team month to month, historically those things were in slide decks and spreadsheets, and now they're just in Sites. Sites is the mechanism that they collaborate across the team.
The reason is because it's somewhat higher bandwidth. These tools like PowerPoint and Excel are infinitely flexible, but at some point you reach the boundary of either, as a human, you may not know how to use some feature, or the product itself doesn't support it. But with a Site, you can do anything. You ask for anything and you can get that. Once people see that magic, I think it's been really valuable.
swyx
Yeah. Let me show you my case study. This involves all the hot topics, including ChatGPT Work, but also 5.6 token billionaires and token maxing and Sites and autoresearch. I’m a fan of this game called Strata. It’s basically a little board game that you play with physical blocks. They come on top of it like that.
Over the weekend, I took 30 photos and just threw them into ChatGPT. 1.7 billion tokens later, out comes this Site with a fully playable thing, with 3D block placement and everything. Because it requires physical blocks, I needed friends to train on it so they can get better, so I can play against them. But I could also do things like train an AI on it, and that gets into autoresearch.
That's your autoresearch.
swyx
So you want to train your own AIs and then make sure they self-play against each other. I need to set both AIs. This is AI versus AI, and they're going to self-play. Obviously, the AIs start out bad, and then you want to define a loss function and get good. I wasn't going to supervise all this. I was down in San Mateo attending a conference.
What I ended up doing was autoresearching this and creating benchmarks, and there were just way too many parameters for me to read. So I started asking it for a Site, and it created this lab panel.
You should be able to go in the sidebar to Sites, at the top of the left sidebar.
swyx
This one? Oh, on the left?
Yeah. Just scroll all the way to the top.
swyx
Oh, it says Sites.
Yeah.
swyx
Oh, there you go. Yeah.
Logan Kilpatrick
Ooh.
swyx
So it creates the Sites. I don't think this is exactly what I wanted, but let me show you what it popped up, right? I think as a research artifact, it is very important to communicate exactly what is being done. It outputs this thing, which I eventually started publishing.
So I moved it off of Sites because I wanted more database and infrastructure than Sites afforded me. But this is a research output that you can start to mess with and try to think about what hyperparameters you're tuning for training AIs. I was trying to make scaling laws and everything and doing all sorts of game optimization stuff.
The fact that you can just throw this up as a research artifact means I no longer need to read ChatGPT output. I read Site output. But then there's also a huge sprawl. Look at how long this thing is. There are so many numbers. It's pretty overwhelming. So then I have to start pruning it from there.
But it's an interesting transition from Markdown, effectively, to putting out a whole functional Site.
Logan Kilpatrick
Yeah.
swyx
I don't know if any of that triggers any stories for you about how it's run internally. Am I doing this right?
Logan Kilpatrick
Yeah. I think this is a workflow we're seeing all different types of teams use, where the canonical artifact that was previously a deck or something is now becoming a Site. And with a Site—because it's just HTML—it's infinitely flexible.
If you want to give more prominence to a certain thing that in a slide deck would feel like it was buried, you can do that. You can have it be the hero image, right? And so I think people are starting to see that. There's obviously more work to be done to make these things much easier to collaborate on.
You mentioned that they're very long and verbose and could be broken up. I'm sure that there's still something to do there. But I think we're starting to see that there is this aspect of this being a really interesting format for people to use that's much more flexible than what they ever had before.
swyx
They're super long. Yeah.
Akshay Kothari
I think Markdown just isn't that optimal for people to read, right? You might as well just write an HTML website and... I don't know. I think you can do a lot with customizing this, right? You have your skills that explain what you want. I noticed they're quite verbose. I don't need a lot of this information.
swyx
It's very verbose.
Akshay Kothari
And then the nice thing about having a Site side by side is you just iterate on what you want and what you don't, right?
swyx
I think your job also becomes kind of meta. You're not designing the products; you're designing a product to make products, and I'm curious how you manage that.
I think one thing that we've been thinking a lot about, when we look at the UX, is how we can balance simplicity with capability. If we're designing a product, as you said, that's made to build other things, right? You can build so many different things. But we can't put that all in front of you because you'll get overwhelmed.
Alessio Fanelli
Yes.
Logan Kilpatrick
And so we had a similar problem, or similar challenges, even with ChatGPT. But especially now, when there's so much that can be done, I think the balance that we're constantly trying to strike is: How can we give the user enough of a UI surface where they can be expressive, they can tell the agent what they need, they can verify that it's using the right tools, it's pulling from the right sources, et cetera, but then it gets out of the way.
And then how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is going to be about how they discover the next use case and the next one after that if they really want to be superpowered by the AI.
Alessio Fanelli
Yeah. It's interesting. Everyone also just has a different way to do it, right? I made a similar version of this same game. I didn't take any pictures of the board or the rules of the game. I threw in a goal. Eighteen minutes and 53 seconds later, and a lot of tokens later, I've got a similar version. Obviously, not with all the autoresearch and whatnot, but—
swyx
You have to do all the latest trends.
And, yeah, I did it with Codex, not Work, but it's interesting, right?
swyx
Yeah. This is obviously GPT Image generating the profile avatars. It's very good for game design. A lot of game designers were really into GPT Image for assets.
I will say the broader takeaway probably is that the reason we do this is more to test the tools, right? This was also a test for GPT-5.1 coming out. I had done the game on 5.5, right? The ability for me to no longer need to feed it the rules—it's a pretty niche game, and it couldn't figure out how to do this on its own.
swyx
Oh, yeah.
Logan Kilpatrick
It's auto-discovery. This is why I'm also very keen—
Alessio Fanelli
Mm, yeah.
Logan Kilpatrick
—on testing the GPT-5.1 capability.
Alessio Fanelli
But, you know, as work comes out and as new things come out, these are just our side quests to test things, right?
Logan Kilpatrick
Yeah. It's some kind of private eval, I guess.
Alessio Fanelli
Yeah.
Logan Kilpatrick
That is not this private.
Alessio Fanelli
But it's also valuable because now you can send this to your friends. I learned about this game through seeing this.
Logan Kilpatrick
It's a hard game. He's very good. It's good when no one is competing with you. But, yes, it's a classic RL problem of self-play, bootstrapping your game AI. You see how easily work becomes personal and personal becomes work—
Alessio Fanelli
Mm-hmm.
Logan Kilpatrick
—because the thing I do for personal use actually directly informs the people I work with because I showed it to them. They were like, “Oh, you can do that with GPT?”
Alessio Fanelli
Yeah.
Logan Kilpatrick
Which, I imagine, is the growth strategy.
Alessio Fanelli
Yeah. The show-not-tell is a big piece that I think we've still not fully cracked: showing people all the things that they can do with the product versus trying to teach that to them through articles or onboarding or whatever.
Logan Kilpatrick
Yeah.
Alessio Fanelli
So, meeting them in the moment.
Logan Kilpatrick
It's a career risk for me because I used to be in developer relations, right, where your job is to show. Then you're like, “What do you mean? You don't need...” Actually, your job is to tell. And then the product people are like, “Well, we don't need you if our product is intuitive enough.”
Alessio Fanelli
Mm-hmm.
Logan Kilpatrick
Yeah. That's the magic of the models. You can tailor the telling or the showing to specifically what the user needs: what they care about, what they've done in the past, and exactly where they are on the adoption journey. I think that's going to be a super-big opportunity. It seems easier and easier now to tailor custom showing, right? People have different use cases.
As much as you said you don't want to segment different people into different buckets, it's also not that hard for people who are in different categories. But the question, I guess, is that you said your team is more broadly on—what was the term you used? Productivity?
Alessio Fanelli
Productivity.
Logan Kilpatrick
Yeah, productivity. So, how—
Alessio Fanelli
Which is now work, basically.
Logan Kilpatrick
Is it work? Is there another distribution that we're not hitting? Is there a group of people that will have something different from ChatGPT, Codex, or Work? Is there more that the mass market isn't targeting?
I see it as a sequencing. The vision is to bring useful agents to everyone. We started with developers. Developers historically are early adopters who are willing to put up with more friction, set things up, et cetera. That's where Codex started.
I think the next opportunity is what we call general knowledge work—all the other functions around developers. When you go from developers to this segment, there are inherent challenges, obviously, with this show-not-tell thing that we're talking about: making the product more understandable and bringing in new capabilities that matter more for this cohort than they matter for developers, things like artifacts and computer use, et cetera.
And then I think the same learnings—similarly to how we took the learnings from developers and brought them to general knowledge work—the next stage will be taking the learnings from general knowledge work and bringing them to everyone, no matter what they're doing in their lives.
We're already seeing that a little bit. This game example that you have is something that's on the border between fun and personal life and your professional life. I use ChatGPT Work full-time at home for everything, for whatever I'm doing. I used it the other day to come up with a meal plan and save that in the computer environment that it has, something that I can continue going back to.
Is everyone doing that yet? Probably not, because the thing says “Work” on it, but eventually we want to get people there.
Alessio Fanelli
ChatGPT Life.
Logan Kilpatrick
Yeah, exactly. ChatGPT Cooking.
But I think there's a lot of opportunity there. I see it as—we built the foundation in software engineering, and we're going to take the same learnings from software engineering to knowledge work, and knowledge work to everyone.
Alessio Fanelli
Do you have any power-user advice? I feel like there's a group of people that will live in it, use it for everything, and stay on it 24/7.
Logan Kilpatrick
Yeah.
Alessio Fanelli
And then there's a bit of a gap between that crew and people who use it for work, use it occasionally, or sometimes type questions. Any advice, learnings, recommendations, or takeaways that you've found help bridge that gap?
Logan Kilpatrick
One thing I've seen is that it really helps to broaden your imagination of what's possible. This has been a learning even for me. The technology has progressed so fast that something that even 3 months ago was, “No way the models can do this,” is now, “Wow, it actually can.”
swyx
Give an example.
Akshay Kothari
We're going through our review cycle internally, and people have always talked about this as something the models are good at. There's a cliché of, “No one wants to be writing reviews, and we just use AI to do it.”
In all seriousness, before, it was just slop, basically. I think it was helpful, but not super productive. Now I've found that the model can do a much better job than me, especially in this environment of pulling context on what people are up to, the things that they've caught, and highlighting wins they may have had that I might not even have seen.
It has access to everything, right? The code, the things that they've done, reviews, Slack, everything. And so it's incredibly powerful in that domain. Just 6 months ago, the last time we did this cycle, I tried using it, but it was not at all helpful. This time, it's been incredibly helpful.
I think continuing to push the frontier of imagination of what's possible, even if you tried something before, is maybe my biggest piece of advice. The other thing is that the more you put in, especially in this environment where the model has access to everything on your computer or in ChatGPT Work, the more you can create artifacts over time and save them in your library, and the model will continue having access to those.
The more information you give it about whatever domain you're in, whether it's your life or your work, the more valuable it becomes. It will become valuable in ways that might surprise you. It might pull from context in a way that may be proactive and that you might not even have thought about, but it needs to have access to those tools or that context first.
Shawn Wang
I just want to talk about the review stuff because that's a very sensitive thing. You're a founder, you've managed people, and you've hired people. As a manager myself, I'm very reticent to put out any LLM-generated things, especially when it comes to people, because it feels like you don't care.
Nick Turley
Mm-hmm.
Shawn Wang
Presumably at OpenAI, people are obviously more open to being evaluated—basically rated by GPT.
But are there any unofficial rules around this? What's the etiquette?
Nick Turley
I think the etiquette is that I would never write something solely via AI and present it as a review for someone. What I was talking about is more like gathering context.
Shawn Wang
Yeah.
Nick Turley
That's the place where it's incredibly helpful.
Shawn Wang
So it's just search.
Nick Turley
Yeah, exactly.
Shawn Wang
It's agentic search. Yeah.
Nick Turley
It's like agentic search, but you can tailor and steer it much more capably than you could before. The thing is, it's all—there's sort of a flywheel happening, right? Because of Codex and ChatGPT, more people are able to do so much more now than ever before. And if you're able to do so much more, it's easy to miss things as well. So I think we need to use these same tools to keep up with all—
Shawn Wang
Yeah.
Nick Turley
—the impact that people are having and understand where we can be helpful.
Shawn Wang
I think the thing is, obviously I run a small company, so it's easy to search. But at the scale of OpenAI, with the amount of messages that you guys put in Slack, do you think that it misses things?
Nick Turley
Probably, but I think that I also miss things.
Shawn Wang
It doesn't matter, right?
Alessio Fanelli
I think sometimes it's—
Shawn Wang
It needs to be human-level.
Nick Turley
It's all relative, right? Yeah.
Alessio Fanelli
Sometimes it's nice when it finds things you wouldn't, right? Right now, my Codex system prompts are set up in such a way that every project I have has a separate NOTES.md, and it just writes learnings to there. Then the global one can pull from all these.
Sometimes it'll be like, "Oh, there's this project you did four months ago. Here's a note that we had," and it randomly pulls it back into context. I would never have thought about it.
Nick Turley
Exactly.
Alessio Fanelli
I'm like, "Okay, this is quite superhuman," right? It's stuff that would save hours on chunking through things or finding something that's already been done. As much as it might miss stuff, I would too, but it's very useful when it finds stuff. And I have a very non-super-engineered solution to this. It's just Markdown files that get pulled whenever they want.
Nick Turley
Yeah. I actually have a funny anecdote about this. Recently, gearing up to this launch, the team had been really cooking on it for a couple of months. Over that time, there was so much conversation and chatter going on in Slack, Docs, and elsewhere. One of the members of the team set up a scheduled-task automation to look at everything that's going on, come up with the best memes, and then post them in one of our shared channels.
There are 2 cool things about this. The first is that I think the models are, over time, actually starting to become—
Shawn Wang
Funny?
Nick Turley
—funny.
Shawn Wang
Nice.
Nick Turley
Whereas a year ago, that was not at all the case. The second is that it was what you were saying: they find things in surprising ways that you may not have thought of and create connections that you may not have thought of. That really helps with meme generation, because then you can see something that genuinely surprises you and is funny in that way.
Obviously, that's not the most productive use of the technology, but it does uncover this capability that's emerging, which is to find information that you otherwise would not know of.
Shawn Wang
Talking about the launch, I think I've pretty much said this is the most successful launch in a long time. I think it's even more successful, personally, than 5.0, and they're announcing 10 million users. Does it feel different? You've been through a lot of launches.
Nick Turley
I think it feels like a culmination. I think 2 things: one, it feels like a culmination of this vision and mission that we've been on for a long time. Like I said, we saw the magic of Codex internally, and then we were extremely excited to bring this to many more people and to see it working—to see us reach the distribution numbers that you mentioned. I think that's huge and super exciting.
The flip side of that is that there's so much more to do, too. That's also really exciting. ChatGPT as a whole—the product that almost everyone equates with AI and loves—has hundreds of millions of users. So 10 million is really cool, but we need to get this to everyone. We need everyone to feel this magic. That's the next step from here.
But yeah, I'm extremely pumped about how it's going so far and the opportunities.
Shawn Wang
Awesome. I did want to also—because I've been tracking the number closely, it transitioned at some point from just Codex users to Codex plus ChatGPT Work, obviously because they're the same harness. The whole point is that you can't count them separately.
Do you have roughly 1 billion ChatGPT users? Why did it just jump to 1 billion right away? Isn't that the default on ChatGPT, or no?
Nick Turley
We don't default you into ChatGPT Work if you're on ChatGPT—
Shawn Wang
If you're free. Yeah.
Nick Turley
It's also only available to paid users right now. I think there's a process of educating users about the value of this product, having them try it, learning from their feedback, and making it better over time.
The goal is to get as many of the people who love ChatGPT today to feel the power of ChatGPT Work, but I think it'll be a journey.
Shawn Wang
Yeah. And Codex will still be alive as a brand for the foreseeable future.
Nick Turley
Yeah.
Shawn Wang
We'll just toggle between them as needed for UI stuff.
Nick Turley
Yeah. I think it's an even stronger point than that. We fully intend to treat developers—developers have been a core market for us for so long—and there's so much more that we can do to make Codex great specifically for software development. We'll continue to do that. This doesn't take away from that at all.
If anything, it should increase the utility of something like Codex, because now you can move seamlessly between writing a diff, creating an artifact, or doing a search over your calendar.
Shawn Wang
Yeah.
Nick Turley
I do wonder how much this terminology leaks to the nontechnical user. Do they have to learn to say "artifact" if they want an artifact?
It's funny: we call it "artifacts" internally because that's what the teams call it.
Shawn Wang
It's actually safe.
Nick Turley
But externally, no one says that. No one calls it an artifact. I think people often describe things using whatever they're used to, right? So if ChatGPT Work is good at creating slides, they'll say ChatGPT Work is good at creating slides, and that's actually what we want.
7. OpenClaw Brings Agents Home
Shawn Wang
Another big thing—it's July 2026—was OpenClaw. I think that's a lot of people's first time really using an agent for personal stuff, but also crossing over to work in essentially the same way.
As far as I understand, OpenClaw is still independent. Did you go through your own OpenClaw moments? Were there any lessons you took from OpenClaw to Codex or back, whatever?
Nick Turley
I think there's a lot of inspiration. I did go through my own OpenClaw moment.
My wife and I set up an OpenClaw to try to manage everything in our house. Not that there's a ton, but it was actually quite useful. We gave it a calendar, and it started creating events for us and stuff. At some point, the laptop that we were running it on died, and then we got a chance to pick it back up.
But there was a lot of inspiration there. In ChatGPT Work, on web and mobile, you get access to this persistent computer environment where you can store files, and those files stay around between sessions. The idea is to be able to enable use cases like this.
One of the members of our team actually uses ChatGPT Work for what they used OpenClaw for before, and it has completely transitioned. It's workout planning and meal tracking, which again is a work-y thing, right? It's not necessarily work, but it's in the personal productivity space.
It has all the same primitives: scheduled tasks, the ability to store files on a file system, and the ability to reference those things over time. So you start to see the same types of use cases emerge, which has been really cool.
Shawn Wang
Is there a point at which ChatGPT Work completely replaces OpenClaw? Obviously, they're independent.
Nick Turley
Yeah, I'm not close to it, so I can't speak to the OpenClaw roadmap.
Shawn Wang
Yeah.
Nick Turley
But I don't think so. I think there's always a need for this incredible open-source technology that that team has built, and I think we can draw inspiration from it in the product.
ChatGPT, I think many more people have heard about and use ChatGPT than have used OpenClaw. If we can take the magic from OpenClaw and bring it to them, I think that'll be a success.
One thing on the ChatGPT Work side that we feel strongly about is that the core experience is that you come to this product and have a conversation—start a session, whatever you want to call it—with this agent, and the magic of the product is that you can do anything in that moment. We would like to create a product where you don't have to click a button or go to a different place, and you can get whatever functionality exists in your finance app or any other product in this one place.
That's the goal. We want an extensible system with plugins where you can connect to the tools that you need in order to accomplish a financial task. Or, if you're doing science work, we have an ability to extend the system such that you can write the tool and it performs well. There'll always be products that we support that are best-in-class at those things, but we want as much of the magic as possible in that core experience.
Shawn Wang
Yeah. Do you think that you can do everything you used to do with Wealthfront in ChatGPT Finance?
Nick Turley
I actually tried it. ChatGPT doesn't yet custody cash and assets for me, so that part, no, not yet. But there was a whole component of retirement planning, financial planning, budgeting, and stuff that we were looking into when I was there. With the finance plugin, that's all possible with ChatGPT today. So I feel like at least that component's replaced for me.
Shawn Wang
I haven't really plugged it in yet. I'm somewhat scared to look at the answer. That's honestly the same reason for health and finances. I'm like, “I don't know.”
Nick Turley
It's really good. It's really cool how, you know, we were talking about the agentic search aspect a little bit earlier, but in conventional UX, the more power you want to give to a user, the more knobs, bells, and whistles you need to add. For these finance and budgeting apps, there's always a bunch of different filters, search bars, and stuff like that.
But now, with the right connectivity to the right data, you can have whatever you want. You can ask any question you want and get the answer, and I think that's super powerful.
Alessio Fanelli
I think it's also nice to just have it centralized in one space, right? You have different health apps. I have one for a smart scale, a watch, and all these different things. It's just nice to centrally co-locate it.
Shawn Wang
Which is part of the whole thing of OpenClaw, right? That it would have a personal OS, which presumably ChatGPT wants to become. I do think that just relying on just-in-time pulling of data, let's say through MCP, CLI, API, or whatever you do, still isn't enough.
I come from a bit of a data engineering background. You still want a data warehouse or some kind of caching or semantic layer. Do you feel that, or do you already have that?
Nick Turley
I can't speak to all the details on how everything works, but I think it depends on the access pattern, right? If you want an answer immediately, then yes, it's very difficult to do that if you need to pull from all of these sources.
But a lot of the use cases that we want to enable in ChatGPT Work aren't necessarily something that you need immediately. It's more like a task that you want the agent to go and do, and that's going to take a certain amount of time. With things like programmatic tool calling and subagents, some of that work is also parallelizable.
And so it's possible—I think it's very possible—that the ceiling on what can be done with MCPs and calling out to these third-party services has been raised substantially.
Shawn Wang
Yeah.
Nick Turley
So we're really excited about that.
Shawn Wang
You mentioned subagents. I gotta double-click on that. Ultra is a new mode. You have special affordances in ChatGPT itself to show off the agents. Can't really do much with them, to be honest. Just watch.
What have been your experiences? Are there any design issues that you would call out to other builders building with subagents?
Nick Turley
I think it sort of goes back to the balance that I was raising earlier about showing builders the power of the tool, but also creating enough of an abstraction not to overwhelm them. With subagents, the thing that we wanted to show is that you can take a task that has many parallel tracks or is complicated in a way that subagents can handle, and this product is for you.
The model can accomplish those goals or try to accomplish those goals. That's the point of showing them in the product, and that's where we've gone with the design. There's another iteration of this where you can see exactly what they're doing and things like that, which I think could converge on being overwhelming with information. This is the deliberate trade-off that we made for now.
Shawn Wang
I mean, you do display quite a lot of transcripts.
Nick Turley
Right. Right.
Alessio Fanelli
I think it's hidden by default, though, right?
Nick Turley
No, it's hidden by default. Yeah.
Alessio Fanelli
Some people could want more. I'm one of those people who will basically throw a lot of stuff at goal, and pretty much every goal I'll tell it to use subagents. Seems redundant, right? But every time I'm like, “Okay, use subagents where possible.” I have a lot of friends who recommend and do the same.
Whereas I'll sometimes talk to people who are like, “Okay, this is where I want you to use subagents for this subtask,” and I'm sure they would appreciate seeing how they're being used. For me, it's primarily 2 things. One is net time efficiency, so span it out across subagents. Two is probably cost, right?
Nick Turley
Right.
Alessio Fanelli
Don't use a big, expensive model. Offload to a lot of smaller, cheaper models. Some people want that level of control.
If you have repetition in what you're doing, say I want something built where I want it to consistently do this every day, I might want to go in and fine-tune subagents here, subagents there.
Shawn Wang
Yeah.
Alessio Fanelli
You can see both, but I think, if I'm not mistaken, it's hidden by default. There's a dropdown where you can get more information about what's been done.
Shawn Wang
Oh, you can change the model that they use?
Alessio Fanelli
I know Anthropic offers this in Cloud Code. You can tell Fable to use Sonnet or Opus to use Sonnet as a subagent. It's a pretty trivial thing. You tell it to spin out subagents with Sonnet, and you know it's cheaper and faster. I would assume if it's not there, it could be built there. But I think there's a side of—
Shawn Wang
It's too many toggles.
Nick Turley
Mm.
Alessio Fanelli
It's not a toggle, actually. You just tell it in chat.
Nick Turley
You're prompting it. Yeah.
Alessio Fanelli
The way I do it is to prompt it, right? I think this is something that gets abstracted unless it's something you built for repetition, right?
If I'm building something for podcast prep—research into people, doing very deep, extensive research—I might want to configure it to use a cheaper, faster model just for web search. I can see a world in which you want both. I think the default is actually pretty good right now, where it's hidden, but you can drop down and get more information about what's been done.
Nick Turley
Mm-hmm.
Alessio Fanelli
I know people talked a lot about it on 5.6's launch. This thing loves to use a lot of subagents and causes the ChatGPT app to just crash because it's so processor-heavy. But—
Shawn Wang
For what it's worth, that's not my experience. I haven't had a crash from subagents.
Alessio Fanelli
I haven't either. We both have big laptops. But I know people brought it up. It was a topic of discussion that we didn't see the same, but it is another vibe eval, right? People are like, “Okay, the amount of subagents Sol is wanting is crazy.” And I'm like, “I think this is okay. I think it's good.” But it's just stuff people bring up.
Nick Turley
I think when we launched the product, too, we weren't as opinionated about who Ultra is for and when they should be using it. Since then, we've made some changes to require you to turn it on and find it in the advanced settings, because that's who it's for.
It's for power users who understand what's going to happen, because, depending on your use case, it can also use more of your limits.
Alessio Fanelli
Yes.
Nick Turley
So that's where I think a lot of the feedback was coming from.
Alessio Fanelli
It's okay. Reset the limits. Always reset the limits.
8. Memory Becomes The Next Frontier
Shawn Wang
Well, today we're resetting because of this. I want to change topics to one last piece of the harness: memory. A lot of people are commenting on memory recently.
ChatGPT’s new memory system used to suck; it’s not very good. And then this guy also basically said the same thing, and Samir, who you presumably work with—
Nick Turley
Mm-hmm.
Shawn Wang
—talking about memory. What can you say there?
Nick Turley
I think that Samir and the team, and the research teams, have made a ton of updates and improvements over time. When I talk to friends and family members about what they love about ChatGPT, the fact that it knows them—that they feel like their ChatGPT is their ChatGPT—I think probably comes up—
Shawn Wang
Yeah.
Nick Turley
—number one. In ChatGPT Work in the cloud, by default, all conversations will inherit from your ChatGPT memory, so they’ll know context about you, and they’ll also be able to write back to this memory.
Shawn Wang
With a small text write? Like, you tell me when you’re writing, right? Is it—
Nick Turley
No, it’s part of the same Memory V3 system that we launched.
Shawn Wang
Okay.
Nick Turley
Yeah. Dreaming V3. Yeah.
Nick Turley
I think that’s been really powerful because going from ChatGPT to ChatGPT Work feels like an extension of what I’ve already been doing with the product for many years.
Shawn Wang
Yeah.
Nick Turley
So that’s been awesome, and it’s awesome to see that people are recognizing—
Shawn Wang
Yeah.
Nick Turley
—the improvements here.
Shawn Wang
Is it basically a retrieval problem? Are you retrieving the right things? Are you over-focusing on the wrong things? Is there more of a false-positive or false-negative problem, if that makes sense? What’s the bigger problem?
Adam Fry
I don’t work on memory directly—
Shawn Wang
Okay.
Adam Fry
—so it’s hard to say what the bigger problem is with certainty, but I think you’re right. There are 2 sides of it. It’s making sure it knows things about you, but then also having the EQ to bring those things up proactively at the right moments, or surprising you in ways that are positive, not negative.
Shawn Wang
Yeah.
Adam Fry
So I think it’s a very challenging problem, but something that I think we feel is a huge opportunity to get right, which is why we’ve made big investments in it.
Alessio Fanelli
How do you see the side of—when you’re building ChatGPT for work, different from the regular chat app and different from Codex, managing memory across different projects, collaboration, and whatnot? How do you see what’s separate from the harness? So if I have 4 threads on 1 project—
Adam Fry
Mm-hmm.
Alessio Fanelli
—any learnings on how to build memory systems there? For background, to steer it a bit, when you do chat-style applications, I’d say you have a lot of one-offs, right?
Adam Fry
Yeah.
Alessio Fanelli
When you switch to work, it might be something you’re doing for a month, something you do a lot, right? Now, as I add more sessions, there’s a lot more than just single-threaded, right?
Adam Fry
Sure.
Alessio Fanelli
And there might be memory there.
Adam Fry
I think first I’d challenge that the depth of the memory, or the value of it, is fundamentally different across chat and work. It is true that there are a lot of shorter sessions on chat, but I think the ChatGPT product has had a ton of longevity—as long as this technology has been around—and people use it for work-y, productivity-related things already today.
And so I think we found that there’s a lot of value. I found this in my personal usage: all these one-offs add up over time into something quite durable and quite a good representation of who I am. I know from time to time something will go viral on X about ChatGPT telling you everything it knows about you, and people are always surprised by how deep that is.
Alessio Fanelli
The “roast me,” you know?
Adam Fry
Exactly. So I think that’s all to say that there’s a lot of depth in the existing ChatGPT product, and that’s why I think we think it’s valuable to bring into the work product. But the other reason I brought that up is because I think hopefully we can use some of the same fundamental primitives and systems to extend memory here as well, and I know this is something that the team that focuses on this is working through right now.
Shawn Wang
I wanted to bring up 1 element of memory that I honestly don’t really use much, and I’m curious if you do: Chronicle, which is up on screen right now. It’s kind of a super memory, or what is it?
Adam Fry
I think the idea is that it can learn from how you’re using your computer, and it’s another input source into memory. I think it’s experimental right now and isn’t on by default, but I’d recommend that you try it.
I think it’s quite interesting because it goes back to a conversation we were having earlier. You were asking, “Can ChatGPT miss things?” Does it, on Slack, when it’s searching, miss things because there’s such a volume of stuff? You can ask the same question about everything that you’re doing on your computer: Is it going to know everything that you’re doing? Is it going to capture the intent and stuff like that?
Probably not, but it probably will find things that you might not know about. If it can surface those to you at relevant times, in proactive ways, when you’re doing tasks, I’ve found at least that it can be quite helpful, so it’s worth trying.
Shawn Wang
So mostly for insights and the longer term?
Adam Fry
Yeah, exactly. Insights, and it builds context that can make you more productive on certain tasks. But it’s hard to describe without feeling it.
Alessio Fanelli
I will say you can feel it pretty well. The idea of what they’re saying here is, “Just check through my memories or check through my logs and add skills.” Pretty underrated, right?
Adam Fry
Yeah.
Shawn Wang
Yeah, but that’s automations. You can repeat that using a cron job.
Adam Fry
Checking through your memories and creating skills?
Shawn Wang
Yeah.
Adam Fry
But I think the creation—
Shawn Wang
I see.
Adam Fry
—of memories from Chronicle itself is what’s different.
Shawn Wang
Okay.
Adam Fry
It’s like you have much deeper memories because you have Chronicle on.
Shawn Wang
It’s there. I don’t use it much, but maybe I just need more examples. I imagine you guys use a lot of it internally, so I’m always fishing for use cases.
Adam Fry
Yeah. I would just try turning it on and—
Shawn Wang
It just auto-works? Like it—
Adam Fry
Yeah, and seeing where—
Shawn Wang
Right.
Adam Fry
—where it might start helping you. I think you’d be surprised.
Shawn Wang
Yeah. Amazing. I think that was about it in terms of the overall coverage of ChatGPT Work. I think there’s been a lot of good progress and discussion on building and all these things. There are a lot of ex-founders in the community, in OpenAI as well. Do you think that things have changed a lot? I guess my overall reflection is on building pre-AI and post-AI.
9. AI Rewrites How Teams Build
Adam Fry
I think things have changed a ton. It’s super exciting to see how quickly you can go from an idea to something real today.
Shawn Wang
Yeah.
Adam Fry
Whereas even before, 5 or 10 years ago, it was fast if you were scrappy and willing to build the minimal viable thing. But now the extent of what you can build is much, much broader. And I think what we’ve seen internally building is that this gives you an opportunity to validate much more quickly, to talk to users, to talk to internal doctors, et cetera, and make sure you’re on the right track.
That loop has become more closed than ever before, and that’s a win for product development. I think it’s a win for consumers and users too, because ideally that means they’re getting much better products out the gate.
Alessio Fanelli
Does it mean your teams are smaller?
Adam Fry
I think there’s much more to do now. People can accomplish more individually or in a small team than they could before—things that would have required more people in the past. But at the same time, there’s also more to do, so I think the teams are much more ambitious.
Alessio Fanelli
Have you seen any changes in the scope of roles and in building teams? How did we used to have teams a few years ago, versus what do ideal teams look like now?
Adam Fry
I think we’ve seen a blurring of the lines between the typical product development functions—between EM, PM, engineer, designer, and so on.
Alessio Fanelli
Yeah, I want to bring up this quote: “There will be only 4 jobs left in tech.” There’s AI slop cannon—the people who just burn a bunch of tokens. And then there is SRE, people who are more responsible. There are grown-ups who sell things, and then there are hot people.
Adam Fry
This is an interesting take. I think my suspicion is that there’s everything...
Everyone will be T-shaped in a way, and AI will enable everyone to become a generalist.
Shawn Wang
Yeah.
Adam Fry
Things that I never would have been able to come up with a design for before—I don't have the visual taste required—but I can iterate on something with the help of AI. People will have a specialty, and that's, I guess, the straight line in the T, or the upward line in the T. You can have a specialty that you're interested in, and with the help of AI, you can go deeper and become better at it over time, but then you'll also be a generalist. With that foundation, what you can accomplish is almost limitless.
Shawn Wang
What are you bottlenecked by in terms of specialties? Do you need more designers? Do you need more slop cannons? Do you need more hot people?
Adam Fry
I think the bottleneck becomes ideas and taste, I guess. Because anyone can build now, I think it really is the era of bottoms-up ambition. Because there's so much to be built, you're always going to be bottlenecked by the amount of ideas and the amount of things that you're doing at any given time.
Shawn Wang
Do you think models help solve that?
Adam Fry
Models?
Shawn Wang
Yeah. I mean, I have a front-end design skill where they give me 4 drastically different examples of what this looks like. Sure, it burns a lot of tokens, but then I'll mostly just condense them down: “Okay, I like this part. I like this part. Let's draw these together.” It's like, yeah, I had a vision, but I don't know.
Adam Fry
The one automation that I would love to work, but doesn't, is: bring me new ideas. Somehow, LLMs are just not it. One interesting part about ideas is that they're not in a vacuum. They usually come from somewhere, and in product development, they're coming from talking to users, reacting to friction that you're seeing or feedback, building on some foundation that you already had planned out before, whatever.
I think that's where there will always be value in these generalists that we talked about: closing that loop and then coming up with those ideas that are grounded in that feedback, talking to users, whatever it is.
Shawn Wang
Cool. You lead the productivity team. How do you define productivity?
Adam Fry
I think our mission is to make it possible for people to do things that they weren't able to do before. Right now, we're thinking about it from the perspective of knowledge work. When I look at knowledge work, I think about how people are no longer siloed by their roles. They're no longer siloed by the background or training that they have. No matter what function you're in, you can suddenly build things.
You can suddenly get access to data that you otherwise might not be able to interpret. I think that extends to your personal life, where we want to give you leverage at the end of the day. We want the models and the product to be able to give you leverage so that you can create time for yourself to do the things that you love.
Shawn Wang
Does that also translate to a way to measure productivity? How do you measure leverage?
Adam Fry
The end is—
Shawn Wang
How do you measure leverage?
Adam Fry
I think we haven't figured this out yet. Part of the reason is that it's so diverse. Everyone has different goals, and really, the true measurement is their ability to achieve that goal. Did we help you, or did we not?
Shawn Wang
Yeah.
Adam Fry
It's very difficult without knowing what that goal is up front and also tailoring it for every individual.
Shawn Wang
And the thumbs-up and thumbs-down from ChatGPT doesn't give you anything, right?
Adam Fry
Right. I mean—
Shawn Wang
Oh, yeah.
Adam Fry
You don't know if they're thumbs-downing the content of the answer, the vibe of it—
Shawn Wang
Yeah.
Adam Fry
—or whether or not it helped them with their goal.
I think that's difficult, but it's something that we will need to figure out, and the industry at large will need to figure out, because that's how we measure success.
Shawn Wang
Do you think productivity and how you measure it has changed? Basically, you said there's a lot more work that can be done, a lot more scope. Has it changed?
Adam Fry
I think it was always true that what you really wanted to measure was whether your team, the individual, or you personally were able to hit the goal, or were closer to hitting whatever your goal was. But I think previously we used proxies for this, like code commits or—
Shawn Wang
Lines of code.
Adam Fry
Lines of code, or whatever.
Shawn Wang
Story points.
Adam Fry
Yeah, exactly. Story points.
Shawn Wang
They're coming back, by the way.
Adam Fry
Maybe. I mean, but that is part of the change. With AI now, I think those proxies are starting to fall apart. The number of tokens you use or the number of pull requests you make are no longer as hypercorrelated with whether your team is able to hit the goal or is on track to hit its goals. I think we'll need to come up with new measurements.
Shawn Wang
For the managers listening, give them one thing to try.
Adam Fry
I think what's important for me is at-bats. Are we as a team building the muscle to have not just quantity of at-bats, but quality? Are we able to go all the way from generating an idea, building it out, getting the feedback, reacting to that feedback, actually validating or invalidating the hypothesis, and going on to the next idea? Are we able to do that really efficiently?
That goes to the actual code that's being written, the designs that are being made, or the specs that are being written, but also the culture of the team. Do we have the humility, and are we able to go through that process many times and stay motivated and excited throughout that? That's the thing that I think is important now, especially when we're on the frontier of this technology. There's so much to build and so much to do. That's probably the most important thing that we look at.
Shawn Wang
Any traps people fall into around measuring productivity or what their teams work on? I feel like there's a lot of, “Okay, we added a lot of LLMs, and we have dashboards for this and that,” but not much has changed, right?
Adam Fry
That is the trap, yes.
Shawn Wang
The broader source of the question is for the managers and teams building: how should they approach this?
Adam Fry
I think maybe the trap is conflating motion and progress. Motion is much easier now than ever before because of the tooling that we have. But progress requires you to be very prescriptive and deliberate about what you're actually trying to achieve. It goes back to our question of measurement, right?
We were talking about whether we, OpenAI, can figure out how to measure productivity for our users. That's a very hard problem because of the diversity. But as a team, you should have a really prescriptive and deliberate view on what progress looks like for you and for your team. If you don't have that, then it's very easy to conflate these 2 things.
Shawn Wang
I think at-bats is a really great thing. I'm really glad. I like the discussion between motion and progress. I think that's a quote that we're going to feature in the write-up. You've been very generous with your time. Thank you so much, and congrats on 10 million.
Adam Fry
Yeah, thank you for having me.
Shawn Wang
The next one at 100 in 2 months.
Adam Fry
Sure.
Shawn Wang
2 weeks. Thank you.