[BidClub_]
Latent Space · · 69 分钟

工作的未来:AI 通才、想法与品味——Akshay Nathan,OpenAI

swyxVibhuAkshay Nathan

YouTube
TL;DR
  • ChatGPT Work 将开发者代理这一切入口转化为拥有1000万用户的产品,但更大规模的分发红利仍在前方。 Nick Turley 称,这次发布是 Codex 在工程之外实现内部普及的“收束”,同时强调 Work 仅面向付费用户,且不是 ChatGPT 的默认产品;相比 ChatGPT 数亿用户的规模,“我们需要让所有人都用上它”。

  • OpenAI 正围绕同一个代理 harness 收敛,同时保留面向编程和通用工作的不同界面。 Codex 与 Work 共享插件、计算机使用、artifacts 等能力;Codex 突出 Git 状态、diff 和面向开发者的沙盒,Work 则将这些细节抽象掉。产品判断很明确:“我们应该让用户可以选择,但不应把他们困在某个选择里。”

  • 企业需求已经不再是约束,真正的难题是把广泛热情转化为具体工作流。 Akshay 看到企业组建拥有“巨额预算”的 AI 部署团队,但由于一个什么都能接收的盒子也会让用户不知道该做什么,各家公司和各职能的用例向四面八方爆发。即使已有数亿人熟悉 AI,他仍认为市场中还有一个尚未理解代理的“10倍或100倍大的市场”。

  • Artifacts 和 Sites 正将 ChatGPT Work 推到应用层之上,由模型为每项任务生成合适的界面。 Logan Kilpatrick 表示,OpenAI 的模型滑杆几乎完全在一个 Site 内开发;公司内部的企业财务团队也已将定期报告从演示文稿和电子表格迁移到 Sites。他的前提是,传统工具最终会撞上边界,而“有了 Site,基本什么都能做”。

  • 持久化上下文可能成为 Work 的复利优势,但权限和信任是对应的负债。 Nick 表示,Memory V3 会在 ChatGPT 与 Work 之间传递个人上下文;Adam Fry 称 Chronicle 还会从计算机活动中获取信息;Logan 则说,插件和本地文件带来的上下文“极度私人化”。Shawn Wang 指出,当 AI 生成的答案包含请求者无权查看的信息时,目前由人充当权限层。

  • 推广路径将先覆盖开发者,再覆盖通用知识工作者,最终进入每个人的私人生活。 Logan 形容 Codex 从能容忍摩擦的早期用户开始,Work 再加入 artifacts、计算机使用和引导式探索;持久化文件、定时任务、金融、膳食规划和家庭协同,正指向一个个人操作层。“你会看到,工作多么容易变成私人生活,私人生活又多么容易变成工作。”

  • AI 提升组织吞吐量的幅度,足以让传统生产力代理指标变得具有误导性。 Adam 预计,未来会出现能够跨职能泛化、同时保留一项专长的 T 型员工,想法和品味会比执行更稀缺。他衡量管理质量的标准,是完整“击球”——想法、构建、反馈、验证——的质量和频率,而不是 tokens、pull requests 或 story points:“陷阱在于把动作和进展混为一谈。”

摘要 · 为研究而整理的核心内容

1. LLMs 补上了 Akshay 长期无代码构想中缺失的一环

  • Shawn Wang 将 Akshay 从低代码产品走到 ChatGPT Work 的路径描述为一个完整闭环。Akshay 表示认同:数据库和自动化测试曾试图把工程原语开放出来,但 LLMs 提供了“让所有人都能获得代码魔法、又不必知道底层发生了什么所需的缺失技术”。

  • Akshay 于2023年加入 OpenAI 时,公司约有500人。他原本担心这里不会有早期创业公司的氛围,后来却发现,这里“比我能想象的任何地方都更像一家创业公司”:自下而上的野心,以及几乎任何人都能把想法做成已发布产品的自主权。

  • 不变的是构建 AGI 并将前沿智能带给所有人的使命。Akshay 从未期待产品线性演进:OpenAI 会尝试多种产品,有些成功,有些失败;但在他看来,过去彼此分离的组件正在“开始拼到一起”。

2. 企业采用暴露了“空盒子悖论”

  • Akshay 在 Enterprise 业务上的核心经验是,“不存在一刀切的解决方案”。早期客户组建了拥有巨额预算、围绕专有数据展开的 AI 部署团队,但当他们被要求提出一个明确的职场用例时,各家公司、各职能给出的答案却向不同方向“爆炸式分化”。

  • 同一个开放式界面同时制造了魔力和采用难题:“你面前有一个盒子,可以对它说任何事”,但用户因此不知道该做什么。产品要实现采用,就必须从每个人已有的任务切入,明确展示 AI 能在哪个环节形成杠杆。

  • Shawn Wang 问,这是否只是前置部署工程(FDE)。Akshay 将商业拓展动作与产品义务区分开来:无论 FDE 多么有效,盯着手机或电脑的人最终仍需要产品本身引导他找到适用的工作流。

3. 非开发者使用 Codex,催生了 Work 及其增长漏斗

  • ChatGPT Work 的催化剂,是 Codex 在非开发者中的采用出现了一个意外拐点。战略财务和市场团队员工不只是使用它;他们为自己成为早期用户感到自豪,并觉得自己获得了“一项超能力”。这种情绪反应表明,代理能力早已远远超出软件开发者的范围。

  • OpenAI 面临的分发问题,变成了如何把这股力量带给那些已经了解并喜爱 ChatGPT 的人。Akshay 将 Merge、Super App 以及最终发布 Work 描述为对这一产品问题的回答——这不是发现代理未来或许能服务非编码者,而是意识到它们“远远早于我们可能预想的时间”就已经服务于非编码者。

  • Shawn Wang 称这次1000万用户规模的发布是 OpenAI 很长一段时间以来最成功的一次,个人排名甚至高于5.0。Nick Turley 称其为一个阶段的收束,但随即重新校准了分母:Work 仅面向付费用户,也不是 ChatGPT 的默认产品,而 ChatGPT 已拥有数亿用户。教育、试用和反馈,仍是从1000万走向“所有人”的必经之路。

4. 一个 harness 支撑不同层级的抽象

  • Akshay 直接给出了架构答案:“harness 是一样的。harness 是共享的。”面向知识工作的改进——插件、计算机使用和 artifacts——会同时流向 Codex 和 Work;两者的用户体验和沙盒默认设置,则反映了对用户需要检查什么的不同假设。

  • Codex 假设用户采用面向代码仓库的工作流:它展示 Git 状态、文件修改、diff,以及更多代理的实现过程。Work 同样可以创建退休计算器电子表格,却不必展示这些机制。OpenAI 希望提供有明确立场的入口,但不希望用户因为选错标签页就失去某项能力。

  • Shawn Wang 以“harness 工程口述史”来描述 Codex 取代经典 ChatGPT harness 的过程;Akshay 回应称,经典 Chat 仍然存在。Shawn 认为它针对延迟、个性、搜索、学习和对话进行了优化;Akshay 则将其与 Codex 的用法对比,后者把“无限灵活的环境当作一台计算机”,用于更长时间、更多工具调用的工作。

  • 更深层的产品押注是,AI 模糊角色边界的速度会快于静态分割的跟进速度。编程、战略、活动、市场和播客工作越来越共享同一套原语,因此 Akshay 预计不同体验之间会反复出现“分化、收敛”,而能力会逐步出现在用户已经身处的地方。

5. 默认设置吸收了维度过多的模型选择问题

  • Alessio Fanelli 问普通用户该如何在 Sol、Terra、不同推理级别,以及 Akshay 戏称的“32个选项”之间做选择。Akshay 的建议是使用默认设置:OpenAI 的目标是让它成为大多数任务的最佳配置,同时简化高级用户控制项,而不是要求所有人对模型类别做极限优化。

  • Shawn Wang 提议用 Sol Lite 加 Goal,替代 Sol Ultra 或 Extra High,以一次深度运行换取更多终止点和调整方向的机会。Akshay 的回答仍然是有条件的:Goal 适合进展一致且可验证的工作;Ultra 或多代理设置则适合高度复杂、开放式或可并行的任务。大多数任务起初并不适合这两种极端。

  • 可视化滑杆将多个变量压缩成一条易懂的轴线,从速度和效率延伸到质量和彻底程度。设计上保留了底层更深的控制项,同时把默认决策简化成非专业用户也能有把握理解的选择。

6. Artifacts 和 Sites 正成为知识工作的原生输出

  • 退休计算器演示生成了一个类似电子表格的 artifact,用户无需离开 Work 就能检查并迭代。Logan Kilpatrick 强调,这只是“第一阶段”,并不意味着 Excel 已经过时:高保真预览让用户更容易核验同事将收到的内容,从而提升迭代过程中的信任。

  • Logan 表示,artifact 质量从5.4到5.5再到5.6出现了大幅提升。被问及5.6与 Work 同时发布是否意味着二者经过联合训练时,他没有做出这一具体判断,而是描述了研究与产品团队之间的大量协作:训练基础设施改善了 artifact 生成,产品团队则围绕它构建了查看器和协作工作流。

  • Sites 将这一构想从电子表格和幻灯片进一步拓展出去。OpenAI 的模型滑杆几乎完全在一个 Site 内由设计、工程和产品团队共同开发;公司内部的企业财务团队现在也用 Sites 取代月度演示文稿和电子表格,因为 HTML 可以根据任务的重要性和交互方式,呈现每一项信息。

  • Swyx 提供了一个极致案例:约30张桌游 Strata 的照片被输入 ChatGPT,经过17亿 tokens 后,他得到了一套可玩的3D实现、自我对弈的 AI、基准测试、 scaling law 实验和研究仪表盘。他最终因为数据库基础设施的需要超出 Sites,且输出多到必须删减,转而“不再读 Markdown,而是开始读 Site 输出”。

7. 界面必须展示能力,但不能把整台机器摊给用户看

  • 打造一个能打造其他产品的产品,构成了 Logan 面临的核心用户体验矛盾:“我们如何在简单性和能力之间取得平衡?”用户需要足够的界面来表达意图、核验工具和来源、纠正代理;但如果暴露每一种可能的操作,产品就会变得无法使用。

  • 最有效的探索机制是展示,而不是讲解。一个典型例子是包裹被放错位置:Work 检查配送照片,搜索附近的房产信息,并找到了包裹可能被放置的公寓楼。这种方式让能力变得具体,是入门文章无法做到的。

  • Logan 将分发描述为一个过程:开发者能够容忍设置和摩擦;通用知识工作者需要 artifacts、计算机使用和更多引导;再从这一群体的经验中学习,把代理带入日常生活。他已经在家中使用 Work,但也承认,只要产品的名字还叫“work”,很多消费者就不会使用它。

8. 更多上下文提升召回,却让协作与信任变得复杂

  • Logan 描述了一种反复出现的协作失败:同事给他发消息,他向 Work 询问答案,然后把自己的理解发回去。他可能加入了有用的框架或权限,但信息在转述过程中有损耗,因为接收者无法检查底层上下文和推导过程。共享 artifact 会话或许能保留更多推理过程,但讨论没有对路线图作出承诺。

  • Swyx 指出了更棘手的多人协作问题:Logan 可能能看到请求者不应看到的信息,使他事实上成为权限层。随着插件和本地文件加入,Logan 也承认代理的上下文变得“极度私人化”;守住这条边界是一个根本性挑战,而不是附带的分享功能问题。

  • Akshay 描述了 AI 在绩效评估中的角色变化。6个月前,他认为 AI 生成的评审辅助大多是“垃圾”;现在,Work 可以搜索代码、评审和 Slack,找出他遗漏的贡献。Nick Turley 补充说,他绝不会提交一份完全由 AI 撰写的评审——有价值的功能是可引导的、代理式的上下文搜集。

  • 更多 AI 输出也让人工监督变得更难,形成一个需要代理来追踪代理所带来影响的飞轮。有个团队安排 Work 挖掘发布讨论中的梗;除了提升幽默感,它还发现了远处的信息和意外的关联——这正是持久笔记数月后仍有价值的那种偶然检索。

9. Work 正在组装个人操作系统的原语,但不主张覆盖所有垂直领域

  • Nick Turley 和妻子曾尝试用 OpenClaw 管理家庭事务和共享日历,直到那台笔记本电脑坏掉。Work 现在提供了类似的原语:一台持久运行的计算机、跨会话保留的文件、定时任务,以及跨时间引用信息;一位同事已经把锻炼计划和膳食追踪迁移到其中。

  • Nick 不认为 Work 必须消灭专业产品或开源系统。目标是建立一个可扩展的对话核心,由插件提供金融、科学或其他领域的工具,而不迫使用户进入不同界面。最优秀的垂直产品可以继续存在,但其中更多“魔法”会从一个会话中被调用。

  • 金融体现了边界。ChatGPT 不托管 Nick 的现金或资产,但它的金融插件已经替他完成退休规划、预算和相关分析。只要连接方式正确,用户就可以直接提出真实问题,而不必在传统强大软件逐渐堆积出来的筛选器、搜索框和配置控制项之间导航。

  • Swyx 认为,即时 MCP 或 API 检索仍需要仓库、缓存或语义层。Nick 对此按访问模式作了保留:即时回答会让联邦式访问变得困难,但许多 Work 任务可以异步运行;程序化工具调用和子代理可以并行检索,大幅抬高第三方数据访问的上限。

10. 子代理和记忆越强大,越应该让复杂性退到幕后

  • Nick 表示,Ultra 让多代理执行变得足够可见,能够传达系统支持复杂并行工作,同时默认隐藏详细 transcript。Alessio Fanelli 认为,专业用户或许希望通过 prompt 指定更便宜的子代理模型,或设定重复任务的拆解方式;Shawn Wang 则反驳说,显式控制有变成“太多开关”的风险。

  • Nick 表示,OpenAI 最初没有明确界定 Ultra 的目标用户。现在它被放在高级用户主动选择的入口之后,因为它可能消耗更多使用额度,面向的是理解这种取舍的人。当前的抽象是有意为之,但他也承认,未来设计可能会展示更多子代理活动。

  • Memory V3 会把上下文从 ChatGPT 带入云端 Work 会话,也可以将信息写回,让 Work 感觉像是多年互动的延伸。Adam Fry 认为,难点不只是检索:系统还需要具备在恰当时刻调用个人知识的“情商”,带来积极的惊喜,而不是冒犯感。

  • Alessio 认为,工作项目需要的记忆深度从根本上高于一次性聊天;Adam 则质疑这一前提。短对话会不断累积,最终形成对一个人的持久表征。Chronicle 又通过计算机活动扩展了这一表征:它仍处于实验阶段且需要用户主动选择,可能错过真实意图,但当之后的任务让某条被遗忘的信息变得有用时,它可以主动将其呈现出来。

11. AI 将瓶颈从生产转向判断

  • Adam Fry 表示,产品团队现在可以比5或10年前快得多地从想法走到逼真的原型、用户反馈和迭代。这并不一定意味着团队规模缩小:每个人完成的事情更多,但 Adam 看到“要做的事情更多了”,因此前沿在团队能力提升的同时,也扩大了团队的野心。

  • 传统产品角色正在模糊。Adam 预计,AI 会把 T 型贡献者变成横向广泛的通才,同时保留其深度专长;没有接受过正规设计训练的人现在也能进行视觉迭代,但专业知识和品味仍决定了哪条方向值得继续。

  • 当被问及模型能否解决剩余的想法瓶颈时,Adam 坦率地指出局限:“我最希望能自动化、但目前还做不到的一件事,就是给我带来新想法。”有价值的想法来自用户访谈、观察到的摩擦、反馈和既有计划,而不是在真空中生成一个听起来可信的答案。

  • 因此,生产力衡量必须回到目标。Tokens、pull requests、commits、代码行数和 story points 作为代理指标的效力都在减弱;即便是点赞,也无法区分答案质量、整体感觉和目标是否真正完成。Adam 更看重完整且高质量的“击球”,并警告管理者:大量 AI 生成的动作,只有在期望结果被明确规定且经过审慎设计时,才会变成进展。

Shawn Wang

Okay, we're here in the studio with Akshay from OpenAI. Welcome.

Akshay Kothari

Thank you.

Shawn Wang

And with our trusty co-host, Vibhu. You recently launched ChatGPT Work. You lead Core Product Engineering. It's been a long journey into all this. I find it very interesting that you started with no-code or low-code, with Walrus and Airtable. To some extent, ChatGPT Work is kind of like the super app of super apps. Here is the ultimate no-code: You just write a prompt.

1. The No Code Thesis

Akshay Kothari

Yeah, it's funny how things come full circle. I started my career working in consumer fintech, but after that, there was this hypothesis that the things we were able to do with code as engineers—if we could bring that to many more people in a more accessible way—would be truly magical.

We were working on a startup before LLMs, before vision LLMs, on how to do automated testing with AI. It was just kind of janky back then, but we were doing what we could. Then I worked at Airtable for a while on the same thesis: If we can bring a database, or the primitives behind a database, to people, that would be really useful to them.

Once LLMs came onto the scene, it became clear that this was the missing piece—the missing technology required to bring the magic of code to everyone without them having to know what's going on under the hood. I think this launch, and a lot of the stuff that we've been up to, is the manifestation of that.

Vibhu

How was stuff when you joined? You joined OpenAI in 2023. Now we've got so much more stuff: ChatGPT, the Codex app, ChatGPT for Work. Have things changed?

2. OpenAI Still Feels Startup

Akshay Kothari

Actually, I think the more interesting thing is how things haven't changed. I joined when it was around 500 people. One thing I was worried about was that I was looking for something more early-stage, and I wondered whether it was going to feel startup enough. I joined, and I was like, "This feels even more startup-y than I could ever imagine." That really hasn't changed even till now.

I think the level of bottoms-up ambition, and the ability of anyone to do anything or have an idea and ship it, is really cool. On the mission side, what was really compelling to me was this mission of bringing frontier intelligence to everyone—building AGI and then bringing it to everyone.

I think we acknowledged back then that this vision was not going to be a linear progression. We're probably going to try different products and have different things that succeed and don't. But the vision has stayed the same, and the mission has stayed the same. We're starting to see the pieces fall together, and that's really cool.

Shawn Wang

You worked on Enterprise. A lot of people never touch ChatGPT—ChatGPT Enterprise, though. What is something that you learned from there that you're bringing into your work now?

3. Enterprise Has No Single Use Case

Akshay Kothari

I think there's no one-size-fits-all solution in Enterprise. I remember in the early days of ChatGPT Enterprise, when we talked to customers, everyone was so excited to bring AI into their enterprise. This was a year after ChatGPT was released, and there were all these teams being stood up, like AI deployment teams with enormous budgets.

If you asked anyone what they were excited about solving, at first you'd get the baseline answers: "Yeah, we have all this context and data and all this stuff." But if you asked them what discrete use case they wanted AI to enable in their workplace, you got such a different variance—an explosion of different types of answers.

Using these models and products, you have this box, and you can say anything to it, which is the magic. On the flip side, it also means that you don't know what to do with it. In Enterprise, I think a big part of that is actually meeting the users where they are: What use case were they trying to solve, and how can they use AI to gain leverage there?

Shawn Wang

Do you meaningfully differentiate that from forward-deployed engineering?

Akshay Kothari

I think there's the go-to-market side of it—

Shawn Wang

Yeah.

Akshay Kothari

—and then there's the product side of it. You need someone on the product side. However good we get at the FDE motion, I think at the end of the day, if we have a user who's looking at their computer or looking at their phone, it's our job in the product to enable them and show them where to go. We're really excited about that.

Vibhu

Do you think there have been changes over the past 3 years in adoption? There have been step-function changes. You have reasoning models and whatnot. Is it still the same problem—that Enterprise has a black box and doesn't know what to do with it—or have things changed?

Akshay Kothari

We're seeing now that there's this huge uptake. Everyone's extremely excited about it. It feels like millions, hundreds of millions of people are using ChatGPT. They understand generally how to work with AI.

But every time a new capability gets unlocked—now we're seeing it with agents—there's probably still a contingent of early adopters who truly get it. They're like, "You can do anything. You just have to make sure the right context is there, that it's connected to the right tools, and that you're supervising it. Anything is possible."

But then there's this 10X or 100X bigger market where they don't yet get that, or they don't yet see that. I think that's the next stage here. To answer your question, I think the adoption is there and growing fast, but the opportunity is far, far bigger than that. That's where we want to play, especially with ChatGPT Work.

Shawn Wang

Well, let's skip ahead to ChatGPT Work. It was only announced about a month ago. What was the decision process that led into it? There was this overall merging of the super app. Is that what we're officially calling it? You deprecated the browser as well. I guess, just summarize your last couple of months of working on this thing.

4. ChatGPT Work Grows From Codex

Akshay Kothari

It feels like forever now, but I guess it's only been a few months. I think maybe the one impetus that is most salient is when we re-released Codex, or even internally had Codex. It was really surprising to us. We recently put out some stats on this: There was a real inflection of adoption among non-developers at OpenAI.

Through this product development process, I would go to UX research sessions to talk to people internally. The thing that stuck out to me was that, first, you go talk to strategic finance or marketing or whatever, and they're all using Codex for their use cases. That part is cool, but the thing that really stuck out to me was how proud people were that they were using Codex.

Shawn Wang

It's like, "I'm not supposed to be using it, but I am."

Akshay Kothari

It was that.

Shawn Wang

Yeah.

Akshay Kothari

It was that they were early to this new thing, but it was also that they felt like they had a superpower. What we recognized then was that the power of Codex, the power of agents, was already available to this massive distribution base of people who have come to know and love ChatGPT.

How do we show that to them? How do we bring it to them? That's a hard product problem, and it's a tricky thing. There are many ways you can go about it. That's why we called it the Merge and the Super App over time, and ultimately launched it in ChatGPT Work. How do we do that? It came from that initial realization that the power was not only for developers, much earlier than probably even we thought. It could be extended to everyone.

Shawn Wang

How do you see the products differently? Who is it for? Codex started out as a CLI, then an app. Now there's a merge of ChatGPT, Codex, and ChatGPT Work. Is it the opening for the average user, for Enterprise, or for work? How do you position it?

Akshay Kothari

I think we want to position it for doing work-related things, for lack of a better word. Productivity is actually the pillar that I support. That's the name of the team. The reason we call it productivity and not Enterprise, work, or something like that is because there's also personal productivity. I think ChatGPT Work is—

I've seen people do things in their personal lives that you wouldn't technically classify as work, but these agents are super capable of. One recent example that someone posted about on our Slack was someone who had a missed package. They didn't receive it, and then they got a picture of it from Amazon or whoever the courier was. They asked ChatGPT Work to find out where that package was.

The agent is extremely tenacious. It took the image, looked at a bunch of listings around their neighborhood, and figured out exactly which apartment complex the package was in. It gave them some information. I think there are all these things that are work-y or productivity-related, and I think that's what we want the product to be.

You asked about Codex. We think Codex is a durable brand, but we have a principle: We don't want a user to get stuck in a tab or an experience where they don't get the power of the product. Basically, everything that you can do in the Codex portion of the product on desktop, you can do in ChatGPT Work, and vice versa.

But we made some opinionated product decisions about how much of the Git state, if you're in a Git repo, we want to expose to the end user. How much do we want to make the experience of seeing the agents thinking diff-forward, so that you get exposed to the diffs out of the box? On the safety side, how do we want to think about sandboxing and making sure that we have the right defaults in one state versus the other?

There are some opinions behind that, but we don't want the user to need to choose which experience they're in.

Shawn Wang

That is a good goal for AGI, right? People don't want to choose which version of AGI they want. They just want the AGI to decide for them.

Akshay Kothari

Yeah.

Shawn Wang

Can I get an answer? It's not super clear to me. Is the Codex harness and the ChatGPT Work harness the same? Is it just UI affordances, or are there actually prompt-level or even deeper differences?

5. One Harness Powers Both Products

Akshay Kothari

The harness is the same. The harness is shared. In both of the products, we made improvements to the harness to make it good for knowledge work, especially as it relates to plugins, computer use, or artifacts. You get that power regardless of which experience you're in.

On the UX side, we have opinionated takes about what the UX should be and how it should behave when you're in Codex mode, and some stuff around the sandbox, like I mentioned. But the underlying harness and capabilities should be the same.

Shawn Wang

I'm just curious. Maybe we can—Is there a query that we can run that would look different in the 2 modes?

Akshay Kothari

Yeah. Try to create a retirement calculator spreadsheet, or something, in both modes. In Codex mode, you might have to be in a repo for this, but you'll see the diffs of the sheet that it's creating and the file edits. But in Work, you won't be able to see that.

Shawn Wang

I think that's super clear. The other thing I wanted to dive into was your productivity team. What else is there? First of all, what are the top-level teams other than productivity? Isn't productivity everything?

Akshay Kothari

So—

Shawn Wang

Science.

Akshay Kothari

We have a team focused—

Shawn Wang

Yeah.

Akshay Kothari

—on ChatGPT.

Shawn Wang

Yeah.

Akshay Kothari

The core chat experience for consumers, which is not, I think, all productivity. People are using ChatGPT every day for search, to figure out how to write messages to loved ones, to think about how to learn a new topic, et cetera. There's so much more inside it, like creating images. There's so much more in Chat that the hundreds of millions of users are using, and obviously that warrants a very dedicated effort.

There are teams focused on enterprise, infrastructure, API, and stuff like that.

Shawn Wang

I will bring it up.

Akshay Kothari

Okay.

Shawn Wang

Yeah. So I have them both running. This is Work. There's a Codex version here. I picked the “Five Little Dicks” song, so this will take a while.

Akshay Kothari

Uh-huh.

Shawn Wang

I think we'll just keep it in the background and, as they finish, we'll look into some of the differences.

Akshay Kothari

Yeah. But immediately, I think if you flip back to the Codex version—

Shawn Wang

Yeah.

Akshay Kothari

—you'll see that—

Shawn Wang

It assumes Git. Yeah.

Akshay Kothari

Exactly. The dynamic island assumes that you're in a Git repo. You might miss some stuff because some of it is in the actual chain of thought, but with those changes and how we display that, yeah.

Shawn Wang

Is there an unintuitive—Is there a thing that you wanted to ship, and then you got feedback, and you were like, “No, let's not do it?” What's the thinking behind that?

Akshay Kothari

In ChatGPT Work?

Shawn Wang

Yeah.

Akshay Kothari

I think one direction we could have gone with this was keeping the experiences completely separate. Why—

Shawn Wang

Different apps.

Akshay Kothari

Exactly, like different apps, or even in the same app, completely different experiences. Why merge it all? You know, Codex obviously people love. Why bring these products together?

I think the intuition here is that all of our jobs are changing dramatically with AI. Every few months, I feel like I wake up and I'm doing a completely different thing than I was doing a few months ago. My hypothesis here is—or I should say our hypothesis—that part of what we're building with this technology is giving people leverage.

Maybe it's the more mundane parts of your job, or parts that, if you were able to automate, you'd be able to share more ideas faster or whatever you're able to do now. Because of that, that might actually blur the lines between someone who's only writing code, creating strategy docs, planning events, helping with marketing, doing podcasts, or whatever.

These things are going to get blurred over time. Trying to draw a hard boundary based on who you are is going to be tough. We should enable users to choose, but we shouldn't box them in.

A lot of the work that went into this, like keeping the primitives the same—for example, plugins are unified across this product, ChatGPT, and the cloud—was because of that. It's this thesis that eventually things are going to come together, and we don't want to box anyone in. We want to be prescriptive about when to be in either experience, but we don't want to box anyone in.

Shawn Wang

I wonder if there are users who are very tuned to the old ChatGPT harness, which is effectively now replaced by the Codex harness. I can't imagine what that was, but maybe they're more on the conversational side. Can you compare and contrast the 2 harnesses? Only you've seen it.

Akshay Kothari

Shawn Wang

Yeah. I mean, I think the existing ChatGPT harness still exists today. It exists in this app.

Akshay Kothari

The classic—

Alessio Fanelli

You just start a new chat, and you don't go under Work, right?

Akshay Kothari

Yeah. If you start—

Alessio Fanelli

So—

Akshay Kothari

—a new chat and go to Chat, then—

Alessio Fanelli

Yeah.

Akshay Kothari

—you’re talking to ChatGPT with the Instant model.

Alessio Fanelli

Oh, we can technically do another.

Akshay Kothari

Yeah.

Alessio Fanelli

But I guess it's on Instant.

Shawn Wang

Yeah. So this one's not going to code, or it's going to—

Alessio Fanelli

Oh—

Shawn Wang

—be inline. It's inline in a sandbox.

Akshay Kothari

It'll actually—

Alessio Fanelli

Oh, that's cool.

Akshay Kothari

—we try to push you to go to Work if you're creating—

Shawn Wang

Yeah—

Akshay Kothari

—a spreadsheet.

Shawn Wang

And this is a router decision?

Akshay Kothari

Sorry?

Shawn Wang

Is this a router decision?

Akshay Kothari

This is the decision that the model is making. It sees that you're trying to do something that would be better served in Work mode.

Shawn Wang

Right.

Akshay Kothari

But I think your question was: What are the advantages of the ChatGPT chat harness?

Shawn Wang

It's more broadly that I want to do an oral history of harness engineering. The ChatGPT harness lasted us from, let's call it, the o1 era until now, and now it's effectively being replaced by the Codex harness. They're overlapping somewhat, but I'm curious what changed, if anything.

Akshay Kothari

Mm-hmm.

Shawn Wang

My perspective on this is that there's sort of a constant process of divergence, convergence, divergence, convergence. In ChatGPT, many of the use cases I was talking about before—search or learning—I think we're really optimizing for latency, personality, and different things. The reason people love ChatGPT is that we've been optimizing for those things and working on them for so long.

Akshay Kothari

With Codex, what we learned was that if you give the agent access to this infinitely flexible environment as a computer, it can do really powerful things. When we think about knowledge work, which mode should we choose? It felt more natural to us to bring that to this computer environment and abstract some of the details of the computer away from users who might not be used to it, while giving them that same power.

Ultimately, I think we want the power in all places. We want to meet people where they are. I’m sure there’ll be work down the road to get things to be equivalently capable in all scenarios. It’s just a question of what we’ve historically been focusing on in the product and what we’re focusing on now.

Alessio Fanelli

I think alongside that, outside of just the harness and when to use Codex, ChatGPT, or Work, there are also the new models you’ve released. Any guidance there? People love to min-max what to use: only use Tera on high reasoning, versus, for this, you want to use Sol here and ignore all these—

Akshay Kothari

There are 32 options.

Alessio Fanelli

Yeah, yeah. But that being said, for people who are exploring productivity stuff and trying things for work, who don’t have a breakdown of what all this is, what’s the advice?

Akshay Kothari

Before the advice, I think the first thing is that none of this would be possible without these models. I think you asked earlier what the inspiration for Work was, and early on I mentioned what we were seeing with Codex, but that was also because the models were getting infinitely more capable. That’s happening again. I think it’s another step-function jump now.

To answer the question on advice, we want the default to be the best possible. We want to be opinionated about the default, so we’ve chosen a default that we think is going to be the best for everyone. For power users, we have options under the hood. One could argue that there might be too many right now, and we’re working on simplifying it. You can extend the reasoning level and change between the different model classes if you need to, but the default should be the best for most use cases.

My advice to most people would be to stick to that. If you reach a situation in which you think you want to try a different configuration, and you’re not seeing either the efficiency on the cost side or the quality on the intelligence side, then you can change the defaults and see if you can get something better. But we think the default should be good enough.

Shawn Wang

I’m just going to run something by you, since you have way more experience than me. I’ve recently been doing Sol Lite with Goal, with the idea that Goal basically augments the reasoning effort, but with more terminations and turns. Is that a good way to think about it, as opposed to Sol Ultra or Sol Extra High?

Akshay Kothari

Yeah, it’s hard to say because—

Shawn Wang

Yeah. It’s like an interaction effect.

Akshay Kothari

Exactly. There’s a preference for you as an individual in how you like to collaborate with the models. How many of those terminations, as you call them, do you want, where you can steer or make sure that it’s doing the right thing?

I think people should generally try whatever works for them. Using Ultra or multi-agent setups is best for tasks that are either incredibly complicated, like open explorations, or very parallelizable. Using Goal, I think, is best for tasks where you know you’ll be able to make consistent progress in a way that’s verifiable over time.

For most tasks, they actually don’t fall into either of those buckets, at least when they’re starting. That’s why I think the best first step is trying it with the default configuration and then seeing where you want to go from there.

swyx

Right. You guys worked on a slider, which is actually super helpful for reducing the amount of panic.

Logan Kilpatrick

Yeah. Yeah.

Alessio Fanelli

It’s nice on mobile, at least. There’s a nice slider there.

swyx

It’s nicer.

I haven’t tried it.

swyx

You have the advanced view there, but if you click Advanced view—yeah, yeah.

Ooh.

swyx

Just a nice slider. Very pretty, very colorful.

Logan Kilpatrick

The idea here was to reduce it to 1 dimension, even though there are multiple dimensions, right? Try to project it onto a single dimension for the user—something that represents speed and efficiency on one side, and quality and thoroughness on the other side.

swyx

Yeah.

Logan Kilpatrick

Yeah.

swyx

I’m just puzzled that it uses Sol so much, like the lower—

Alessio Fanelli

No, no. I think the slider, if I’m not mistaken, is—

swyx

Terra.

Alessio Fanelli

Oh, it is.

swyx

Yeah. See?

Nice. Nice.

swyx

So they preset Terra to only be the light one.

Alessio Fanelli

I see.

swyx

I think a lot of people actually would—more people should use Terra. One reason is that Sol keeps running out of capacity.

I’m the reason, you know. Here’s 10 minutes of our retirement calculator.

swyx

There you go.

Oh, that’s the Excel thing we’re displaying. This is Work, and then Codex is still cooking, so we’ll get back into it. I think it’ll be interesting to actually see the thought process and the reasoning. Also, I guess this is 8 minutes on Work. Codex is still cooking.

swyx

Yeah. By the way, do you know Gabriel Chua? He’s part of the OpenAI Singapore team. He showed me this, and I was pretty shocked that it looks like Excel. It edits Excel files. You never paid for an Excel license, right? But somehow this is kind of workable, and it’s agentic Excel.

6. Artifacts Expand Knowledge Work

Logan Kilpatrick

One of the big pushes that we made for this launch was artifacts, right? Both on the model side—I think if you compare this with 5.5 and 5.4 before that, you’ll see that there have been pretty dramatic improvements in the quality of these artifacts—and then also on the product side.

Alessio Fanelli

The UX side is also crazy. Hosted sites and whatnot, no longer needing to host your own little webpage.

swyx

Oh, I have a story about that.

Yeah.

swyx

I can do a separate thing. I’ll need to take the visuals here, but we’ll cut to that later. Was there co-training, I guess, because you’re making this big move, and you launched GPT-5.6 on the same day as ChatGPT Work? Was there influence between the model training teams and the harness teams, or did the launch dates just happen to line up on the same day?

Logan Kilpatrick

I think we collaborate heavily with the research teams, and I think that’s one of the most magical parts of the job—the most fun parts of the job. Using artifacts as an example, a lot of what you’re seeing underneath the hood comes from the work that went into making sure we had the right infrastructure to train the models to get better at this, and then, on the product side, to have the right experience for users to be able to collaborate with the model on an artifact like this.

In fact, this whole viewer—the intuition here is that it’s not necessarily that you wouldn’t need an Excel license. This is stage 1, right? This is probably not what you meant when you were making a retirement calculator.

Alessio Fanelli

Yeah, you can iterate very easily.

Logan Kilpatrick

You want to iterate, and when you see it, if this thing is high fidelity to what you would actually see—or what your coworkers would see if you were to send this to Sean—that makes it so much easier and makes you trust the product in terms of iteration.

Alessio Fanelli

When you say coworkers would see, do you see multiplayer, multi-team collaboration with artifacts? Are there any things you guys think about there?

swyx

You can already share it, right?

Yeah.

Logan Kilpatrick

It’s something that we’re actively thinking about. One thing that we’ve noticed internally, without talking too much about the roadmap, is that there are many times when someone will ping me about something, and I will ask ChatGPT Work the question, and then I’ll ping them back the answer. And then I’ll be thinking—

Alessio Fanelli

The simplest would be the 3 of us all on one hosted—

Logan Kilpatrick

Exactly. And I’ll think about whether I was required in this loop, or whether it was just a rephrasing of what they were asking or pulling from certain context, or whatever.

But when I gave them back the answer, that process was also lossy, right? I gave them just my interpretation of what ChatGPT Work cooked up. But underneath the hood, there's so much context in the rollout and stuff that could be interesting.

swyx

So the answer was to preemptively respond to every inbound request?

No, it was literally just what I do sometimes as my job.

swyx

I know. You copy-paste—

Yeah.

swyx

—and then you're just a message-forwarding service—

Logan Kilpatrick

Yeah. Yeah, exactly.

swyx

—from AI to AI.

Alessio Fanelli

But I think it's interesting, right? It helps people understand the capabilities of what you can ask and delegate.

Logan Kilpatrick

Yeah.

Alessio Fanelli

Oftentimes, people don't realize until they try or someone shows you, and then you're like, “Oh, okay, okay, I see now.”

Logan Kilpatrick

Yeah.

swyx

I think there's also a light security issue where, basically, you're the permissions layer. Yes, I could query everything that you query and get an automated response, but maybe I'm not supposed to see it.

Logan Kilpatrick

Yeah.

swyx

There's no way I would know, because I'm not supposed to know what I don't know.

Logan Kilpatrick

Especially with ChatGPT Work, we're asking you to connect your plugins, and it's pulling from your local files and stuff like that. The amount of context that the agent has access to is deeply personal, and that's something I think we need to preserve. So that'll definitely be a challenge.

swyx

There's Excel, there's PowerPoint, there's Docs—the grand trio of work. What other formats of work do you think about? Obviously, you worked on Airtable. Is there a future where there's OpenAI Airtable? What does that look like if you ever ended up doing it?

It's a really good question. One that you didn't bring up was Sites, and I think that—

swyx

Sites.

Logan Kilpatrick

—that was a core part of this launch. There's one side of Sites that I think people commonly talk about, especially on Twitter or X, as this sort of prototyping tool. We saw that happen with this launch, even. The model slider that you guys were referencing earlier was developed almost fully in a Site.

The collaboration between design, engineering, and product on that was on a Site where we played with the affordance and figured out how it feels and all of that.

But the other aspect that's a little bit less talked about is Sites as an artifact for knowledge work. I was actually talking to someone the other day who's on our corporate finance team, and we were mentioning how now, when they have these reports that they're working on as a team month to month, historically those things were in slide decks and spreadsheets, and now they're just in Sites. Sites is the mechanism that they collaborate across the team.

The reason is because it's somewhat higher bandwidth. These tools like PowerPoint and Excel are infinitely flexible, but at some point you reach the boundary of either, as a human, you may not know how to use some feature, or the product itself doesn't support it. But with a Site, you can do anything. You ask for anything and you can get that. Once people see that magic, I think it's been really valuable.

swyx

Yeah. Let me show you my case study. This involves all the hot topics, including ChatGPT Work, but also 5.6 token billionaires and token maxing and Sites and autoresearch. I’m a fan of this game called Strata. It’s basically a little board game that you play with physical blocks. They come on top of it like that.

Over the weekend, I took 30 photos and just threw them into ChatGPT. 1.7 billion tokens later, out comes this Site with a fully playable thing, with 3D block placement and everything. Because it requires physical blocks, I needed friends to train on it so they can get better, so I can play against them. But I could also do things like train an AI on it, and that gets into autoresearch.

That's your autoresearch.

swyx

So you want to train your own AIs and then make sure they self-play against each other. I need to set both AIs. This is AI versus AI, and they're going to self-play. Obviously, the AIs start out bad, and then you want to define a loss function and get good. I wasn't going to supervise all this. I was down in San Mateo attending a conference.

What I ended up doing was autoresearching this and creating benchmarks, and there were just way too many parameters for me to read. So I started asking it for a Site, and it created this lab panel.

You should be able to go in the sidebar to Sites, at the top of the left sidebar.

swyx

This one? Oh, on the left?

Yeah. Just scroll all the way to the top.

swyx

Oh, it says Sites.

Yeah.

swyx

Oh, there you go. Yeah.

Logan Kilpatrick

Ooh.

swyx

So it creates the Sites. I don't think this is exactly what I wanted, but let me show you what it popped up, right? I think as a research artifact, it is very important to communicate exactly what is being done. It outputs this thing, which I eventually started publishing.

So I moved it off of Sites because I wanted more database and infrastructure than Sites afforded me. But this is a research output that you can start to mess with and try to think about what hyperparameters you're tuning for training AIs. I was trying to make scaling laws and everything and doing all sorts of game optimization stuff.

The fact that you can just throw this up as a research artifact means I no longer need to read ChatGPT output. I read Site output. But then there's also a huge sprawl. Look at how long this thing is. There are so many numbers. It's pretty overwhelming. So then I have to start pruning it from there.

But it's an interesting transition from Markdown, effectively, to putting out a whole functional Site.

Logan Kilpatrick

Yeah.

swyx

I don't know if any of that triggers any stories for you about how it's run internally. Am I doing this right?

Logan Kilpatrick

Yeah. I think this is a workflow we're seeing all different types of teams use, where the canonical artifact that was previously a deck or something is now becoming a Site. And with a Site—because it's just HTML—it's infinitely flexible.

If you want to give more prominence to a certain thing that in a slide deck would feel like it was buried, you can do that. You can have it be the hero image, right? And so I think people are starting to see that. There's obviously more work to be done to make these things much easier to collaborate on.

You mentioned that they're very long and verbose and could be broken up. I'm sure that there's still something to do there. But I think we're starting to see that there is this aspect of this being a really interesting format for people to use that's much more flexible than what they ever had before.

swyx

They're super long. Yeah.

Akshay Kothari

I think Markdown just isn't that optimal for people to read, right? You might as well just write an HTML website and... I don't know. I think you can do a lot with customizing this, right? You have your skills that explain what you want. I noticed they're quite verbose. I don't need a lot of this information.

swyx

It's very verbose.

Akshay Kothari

And then the nice thing about having a Site side by side is you just iterate on what you want and what you don't, right?

swyx

I think your job also becomes kind of meta. You're not designing the products; you're designing a product to make products, and I'm curious how you manage that.

I think one thing that we've been thinking a lot about, when we look at the UX, is how we can balance simplicity with capability. If we're designing a product, as you said, that's made to build other things, right? You can build so many different things. But we can't put that all in front of you because you'll get overwhelmed.

Alessio Fanelli

Yes.

Logan Kilpatrick

And so we had a similar problem, or similar challenges, even with ChatGPT. But especially now, when there's so much that can be done, I think the balance that we're constantly trying to strike is: How can we give the user enough of a UI surface where they can be expressive, they can tell the agent what they need, they can verify that it's using the right tools, it's pulling from the right sources, et cetera, but then it gets out of the way.

And then how can we build the right system such that we can show them instead of telling them what can be done? Because so much of this is going to be about how they discover the next use case and the next one after that if they really want to be superpowered by the AI.

Alessio Fanelli

Yeah. It's interesting. Everyone also just has a different way to do it, right? I made a similar version of this same game. I didn't take any pictures of the board or the rules of the game. I threw in a goal. Eighteen minutes and 53 seconds later, and a lot of tokens later, I've got a similar version. Obviously, not with all the autoresearch and whatnot, but—

swyx

You have to do all the latest trends.

And, yeah, I did it with Codex, not Work, but it's interesting, right?

swyx

Yeah. This is obviously GPT Image generating the profile avatars. It's very good for game design. A lot of game designers were really into GPT Image for assets.

I will say the broader takeaway probably is that the reason we do this is more to test the tools, right? This was also a test for GPT-5.1 coming out. I had done the game on 5.5, right? The ability for me to no longer need to feed it the rules—it's a pretty niche game, and it couldn't figure out how to do this on its own.

swyx

Oh, yeah.

Logan Kilpatrick

It's auto-discovery. This is why I'm also very keen—

Alessio Fanelli

Mm, yeah.

Logan Kilpatrick

—on testing the GPT-5.1 capability.

Alessio Fanelli

But, you know, as work comes out and as new things come out, these are just our side quests to test things, right?

Logan Kilpatrick

Yeah. It's some kind of private eval, I guess.

Alessio Fanelli

Yeah.

Logan Kilpatrick

That is not this private.

Alessio Fanelli

But it's also valuable because now you can send this to your friends. I learned about this game through seeing this.

Logan Kilpatrick

It's a hard game. He's very good. It's good when no one is competing with you. But, yes, it's a classic RL problem of self-play, bootstrapping your game AI. You see how easily work becomes personal and personal becomes work—

Alessio Fanelli

Mm-hmm.

Logan Kilpatrick

—because the thing I do for personal use actually directly informs the people I work with because I showed it to them. They were like, “Oh, you can do that with GPT?”

Alessio Fanelli

Yeah.

Logan Kilpatrick

Which, I imagine, is the growth strategy.

Alessio Fanelli

Yeah. The show-not-tell is a big piece that I think we've still not fully cracked: showing people all the things that they can do with the product versus trying to teach that to them through articles or onboarding or whatever.

Logan Kilpatrick

Yeah.

Alessio Fanelli

So, meeting them in the moment.

Logan Kilpatrick

It's a career risk for me because I used to be in developer relations, right, where your job is to show. Then you're like, “What do you mean? You don't need...” Actually, your job is to tell. And then the product people are like, “Well, we don't need you if our product is intuitive enough.”

Alessio Fanelli

Mm-hmm.

Logan Kilpatrick

Yeah. That's the magic of the models. You can tailor the telling or the showing to specifically what the user needs: what they care about, what they've done in the past, and exactly where they are on the adoption journey. I think that's going to be a super-big opportunity. It seems easier and easier now to tailor custom showing, right? People have different use cases.

As much as you said you don't want to segment different people into different buckets, it's also not that hard for people who are in different categories. But the question, I guess, is that you said your team is more broadly on—what was the term you used? Productivity?

Alessio Fanelli

Productivity.

Logan Kilpatrick

Yeah, productivity. So, how—

Alessio Fanelli

Which is now work, basically.

Logan Kilpatrick

Is it work? Is there another distribution that we're not hitting? Is there a group of people that will have something different from ChatGPT, Codex, or Work? Is there more that the mass market isn't targeting?

I see it as a sequencing. The vision is to bring useful agents to everyone. We started with developers. Developers historically are early adopters who are willing to put up with more friction, set things up, et cetera. That's where Codex started.

I think the next opportunity is what we call general knowledge work—all the other functions around developers. When you go from developers to this segment, there are inherent challenges, obviously, with this show-not-tell thing that we're talking about: making the product more understandable and bringing in new capabilities that matter more for this cohort than they matter for developers, things like artifacts and computer use, et cetera.

And then I think the same learnings—similarly to how we took the learnings from developers and brought them to general knowledge work—the next stage will be taking the learnings from general knowledge work and bringing them to everyone, no matter what they're doing in their lives.

We're already seeing that a little bit. This game example that you have is something that's on the border between fun and personal life and your professional life. I use ChatGPT Work full-time at home for everything, for whatever I'm doing. I used it the other day to come up with a meal plan and save that in the computer environment that it has, something that I can continue going back to.

Is everyone doing that yet? Probably not, because the thing says “Work” on it, but eventually we want to get people there.

Alessio Fanelli

ChatGPT Life.

Logan Kilpatrick

Yeah, exactly. ChatGPT Cooking.

But I think there's a lot of opportunity there. I see it as—we built the foundation in software engineering, and we're going to take the same learnings from software engineering to knowledge work, and knowledge work to everyone.

Alessio Fanelli

Do you have any power-user advice? I feel like there's a group of people that will live in it, use it for everything, and stay on it 24/7.

Logan Kilpatrick

Yeah.

Alessio Fanelli

And then there's a bit of a gap between that crew and people who use it for work, use it occasionally, or sometimes type questions. Any advice, learnings, recommendations, or takeaways that you've found help bridge that gap?

Logan Kilpatrick

One thing I've seen is that it really helps to broaden your imagination of what's possible. This has been a learning even for me. The technology has progressed so fast that something that even 3 months ago was, “No way the models can do this,” is now, “Wow, it actually can.”

swyx

Give an example.

Akshay Kothari

We're going through our review cycle internally, and people have always talked about this as something the models are good at. There's a cliché of, “No one wants to be writing reviews, and we just use AI to do it.”

In all seriousness, before, it was just slop, basically. I think it was helpful, but not super productive. Now I've found that the model can do a much better job than me, especially in this environment of pulling context on what people are up to, the things that they've caught, and highlighting wins they may have had that I might not even have seen.

It has access to everything, right? The code, the things that they've done, reviews, Slack, everything. And so it's incredibly powerful in that domain. Just 6 months ago, the last time we did this cycle, I tried using it, but it was not at all helpful. This time, it's been incredibly helpful.

I think continuing to push the frontier of imagination of what's possible, even if you tried something before, is maybe my biggest piece of advice. The other thing is that the more you put in, especially in this environment where the model has access to everything on your computer or in ChatGPT Work, the more you can create artifacts over time and save them in your library, and the model will continue having access to those.

The more information you give it about whatever domain you're in, whether it's your life or your work, the more valuable it becomes. It will become valuable in ways that might surprise you. It might pull from context in a way that may be proactive and that you might not even have thought about, but it needs to have access to those tools or that context first.

Shawn Wang

I just want to talk about the review stuff because that's a very sensitive thing. You're a founder, you've managed people, and you've hired people. As a manager myself, I'm very reticent to put out any LLM-generated things, especially when it comes to people, because it feels like you don't care.

Nick Turley

Mm-hmm.

Shawn Wang

Presumably at OpenAI, people are obviously more open to being evaluated—basically rated by GPT.

But are there any unofficial rules around this? What's the etiquette?

Nick Turley

I think the etiquette is that I would never write something solely via AI and present it as a review for someone. What I was talking about is more like gathering context.

Shawn Wang

Yeah.

Nick Turley

That's the place where it's incredibly helpful.

Shawn Wang

So it's just search.

Nick Turley

Yeah, exactly.

Shawn Wang

It's agentic search. Yeah.

Nick Turley

It's like agentic search, but you can tailor and steer it much more capably than you could before. The thing is, it's all—there's sort of a flywheel happening, right? Because of Codex and ChatGPT, more people are able to do so much more now than ever before. And if you're able to do so much more, it's easy to miss things as well. So I think we need to use these same tools to keep up with all—

Shawn Wang

Yeah.

Nick Turley

—the impact that people are having and understand where we can be helpful.

Shawn Wang

I think the thing is, obviously I run a small company, so it's easy to search. But at the scale of OpenAI, with the amount of messages that you guys put in Slack, do you think that it misses things?

Nick Turley

Probably, but I think that I also miss things.

Shawn Wang

It doesn't matter, right?

Alessio Fanelli

I think sometimes it's—

Shawn Wang

It needs to be human-level.

Nick Turley

It's all relative, right? Yeah.

Alessio Fanelli

Sometimes it's nice when it finds things you wouldn't, right? Right now, my Codex system prompts are set up in such a way that every project I have has a separate NOTES.md, and it just writes learnings to there. Then the global one can pull from all these.

Sometimes it'll be like, "Oh, there's this project you did four months ago. Here's a note that we had," and it randomly pulls it back into context. I would never have thought about it.

Nick Turley

Exactly.

Alessio Fanelli

I'm like, "Okay, this is quite superhuman," right? It's stuff that would save hours on chunking through things or finding something that's already been done. As much as it might miss stuff, I would too, but it's very useful when it finds stuff. And I have a very non-super-engineered solution to this. It's just Markdown files that get pulled whenever they want.

Nick Turley

Yeah. I actually have a funny anecdote about this. Recently, gearing up to this launch, the team had been really cooking on it for a couple of months. Over that time, there was so much conversation and chatter going on in Slack, Docs, and elsewhere. One of the members of the team set up a scheduled-task automation to look at everything that's going on, come up with the best memes, and then post them in one of our shared channels.

There are 2 cool things about this. The first is that I think the models are, over time, actually starting to become—

Shawn Wang

Funny?

Nick Turley

—funny.

Shawn Wang

Nice.

Nick Turley

Whereas a year ago, that was not at all the case. The second is that it was what you were saying: they find things in surprising ways that you may not have thought of and create connections that you may not have thought of. That really helps with meme generation, because then you can see something that genuinely surprises you and is funny in that way.

Obviously, that's not the most productive use of the technology, but it does uncover this capability that's emerging, which is to find information that you otherwise would not know of.

Shawn Wang

Talking about the launch, I think I've pretty much said this is the most successful launch in a long time. I think it's even more successful, personally, than 5.0, and they're announcing 10 million users. Does it feel different? You've been through a lot of launches.

Nick Turley

I think it feels like a culmination. I think 2 things: one, it feels like a culmination of this vision and mission that we've been on for a long time. Like I said, we saw the magic of Codex internally, and then we were extremely excited to bring this to many more people and to see it working—to see us reach the distribution numbers that you mentioned. I think that's huge and super exciting.

The flip side of that is that there's so much more to do, too. That's also really exciting. ChatGPT as a whole—the product that almost everyone equates with AI and loves—has hundreds of millions of users. So 10 million is really cool, but we need to get this to everyone. We need everyone to feel this magic. That's the next step from here.

But yeah, I'm extremely pumped about how it's going so far and the opportunities.

Shawn Wang

Awesome. I did want to also—because I've been tracking the number closely, it transitioned at some point from just Codex users to Codex plus ChatGPT Work, obviously because they're the same harness. The whole point is that you can't count them separately.

Do you have roughly 1 billion ChatGPT users? Why did it just jump to 1 billion right away? Isn't that the default on ChatGPT, or no?

Nick Turley

We don't default you into ChatGPT Work if you're on ChatGPT—

Shawn Wang

If you're free. Yeah.

Nick Turley

It's also only available to paid users right now. I think there's a process of educating users about the value of this product, having them try it, learning from their feedback, and making it better over time.

The goal is to get as many of the people who love ChatGPT today to feel the power of ChatGPT Work, but I think it'll be a journey.

Shawn Wang

Yeah. And Codex will still be alive as a brand for the foreseeable future.

Nick Turley

Yeah.

Shawn Wang

We'll just toggle between them as needed for UI stuff.

Nick Turley

Yeah. I think it's an even stronger point than that. We fully intend to treat developers—developers have been a core market for us for so long—and there's so much more that we can do to make Codex great specifically for software development. We'll continue to do that. This doesn't take away from that at all.

If anything, it should increase the utility of something like Codex, because now you can move seamlessly between writing a diff, creating an artifact, or doing a search over your calendar.

Shawn Wang

Yeah.

Nick Turley

I do wonder how much this terminology leaks to the nontechnical user. Do they have to learn to say "artifact" if they want an artifact?

It's funny: we call it "artifacts" internally because that's what the teams call it.

Shawn Wang

It's actually safe.

Nick Turley

But externally, no one says that. No one calls it an artifact. I think people often describe things using whatever they're used to, right? So if ChatGPT Work is good at creating slides, they'll say ChatGPT Work is good at creating slides, and that's actually what we want.

7. OpenClaw Brings Agents Home

Shawn Wang

Another big thing—it's July 2026—was OpenClaw. I think that's a lot of people's first time really using an agent for personal stuff, but also crossing over to work in essentially the same way.

As far as I understand, OpenClaw is still independent. Did you go through your own OpenClaw moments? Were there any lessons you took from OpenClaw to Codex or back, whatever?

Nick Turley

I think there's a lot of inspiration. I did go through my own OpenClaw moment.

My wife and I set up an OpenClaw to try to manage everything in our house. Not that there's a ton, but it was actually quite useful. We gave it a calendar, and it started creating events for us and stuff. At some point, the laptop that we were running it on died, and then we got a chance to pick it back up.

But there was a lot of inspiration there. In ChatGPT Work, on web and mobile, you get access to this persistent computer environment where you can store files, and those files stay around between sessions. The idea is to be able to enable use cases like this.

One of the members of our team actually uses ChatGPT Work for what they used OpenClaw for before, and it has completely transitioned. It's workout planning and meal tracking, which again is a work-y thing, right? It's not necessarily work, but it's in the personal productivity space.

It has all the same primitives: scheduled tasks, the ability to store files on a file system, and the ability to reference those things over time. So you start to see the same types of use cases emerge, which has been really cool.

Shawn Wang

Is there a point at which ChatGPT Work completely replaces OpenClaw? Obviously, they're independent.

Nick Turley

Yeah, I'm not close to it, so I can't speak to the OpenClaw roadmap.

Shawn Wang

Yeah.

Nick Turley

But I don't think so. I think there's always a need for this incredible open-source technology that that team has built, and I think we can draw inspiration from it in the product.

ChatGPT, I think many more people have heard about and use ChatGPT than have used OpenClaw. If we can take the magic from OpenClaw and bring it to them, I think that'll be a success.

One thing on the ChatGPT Work side that we feel strongly about is that the core experience is that you come to this product and have a conversation—start a session, whatever you want to call it—with this agent, and the magic of the product is that you can do anything in that moment. We would like to create a product where you don't have to click a button or go to a different place, and you can get whatever functionality exists in your finance app or any other product in this one place.

That's the goal. We want an extensible system with plugins where you can connect to the tools that you need in order to accomplish a financial task. Or, if you're doing science work, we have an ability to extend the system such that you can write the tool and it performs well. There'll always be products that we support that are best-in-class at those things, but we want as much of the magic as possible in that core experience.

Shawn Wang

Yeah. Do you think that you can do everything you used to do with Wealthfront in ChatGPT Finance?

Nick Turley

I actually tried it. ChatGPT doesn't yet custody cash and assets for me, so that part, no, not yet. But there was a whole component of retirement planning, financial planning, budgeting, and stuff that we were looking into when I was there. With the finance plugin, that's all possible with ChatGPT today. So I feel like at least that component's replaced for me.

Shawn Wang

I haven't really plugged it in yet. I'm somewhat scared to look at the answer. That's honestly the same reason for health and finances. I'm like, “I don't know.”

Nick Turley

It's really good. It's really cool how, you know, we were talking about the agentic search aspect a little bit earlier, but in conventional UX, the more power you want to give to a user, the more knobs, bells, and whistles you need to add. For these finance and budgeting apps, there's always a bunch of different filters, search bars, and stuff like that.

But now, with the right connectivity to the right data, you can have whatever you want. You can ask any question you want and get the answer, and I think that's super powerful.

Alessio Fanelli

I think it's also nice to just have it centralized in one space, right? You have different health apps. I have one for a smart scale, a watch, and all these different things. It's just nice to centrally co-locate it.

Shawn Wang

Which is part of the whole thing of OpenClaw, right? That it would have a personal OS, which presumably ChatGPT wants to become. I do think that just relying on just-in-time pulling of data, let's say through MCP, CLI, API, or whatever you do, still isn't enough.

I come from a bit of a data engineering background. You still want a data warehouse or some kind of caching or semantic layer. Do you feel that, or do you already have that?

Nick Turley

I can't speak to all the details on how everything works, but I think it depends on the access pattern, right? If you want an answer immediately, then yes, it's very difficult to do that if you need to pull from all of these sources.

But a lot of the use cases that we want to enable in ChatGPT Work aren't necessarily something that you need immediately. It's more like a task that you want the agent to go and do, and that's going to take a certain amount of time. With things like programmatic tool calling and subagents, some of that work is also parallelizable.

And so it's possible—I think it's very possible—that the ceiling on what can be done with MCPs and calling out to these third-party services has been raised substantially.

Shawn Wang

Yeah.

Nick Turley

So we're really excited about that.

Shawn Wang

You mentioned subagents. I gotta double-click on that. Ultra is a new mode. You have special affordances in ChatGPT itself to show off the agents. Can't really do much with them, to be honest. Just watch.

What have been your experiences? Are there any design issues that you would call out to other builders building with subagents?

Nick Turley

I think it sort of goes back to the balance that I was raising earlier about showing builders the power of the tool, but also creating enough of an abstraction not to overwhelm them. With subagents, the thing that we wanted to show is that you can take a task that has many parallel tracks or is complicated in a way that subagents can handle, and this product is for you.

The model can accomplish those goals or try to accomplish those goals. That's the point of showing them in the product, and that's where we've gone with the design. There's another iteration of this where you can see exactly what they're doing and things like that, which I think could converge on being overwhelming with information. This is the deliberate trade-off that we made for now.

Shawn Wang

I mean, you do display quite a lot of transcripts.

Nick Turley

Right. Right.

Alessio Fanelli

I think it's hidden by default, though, right?

Nick Turley

No, it's hidden by default. Yeah.

Alessio Fanelli

Some people could want more. I'm one of those people who will basically throw a lot of stuff at goal, and pretty much every goal I'll tell it to use subagents. Seems redundant, right? But every time I'm like, “Okay, use subagents where possible.” I have a lot of friends who recommend and do the same.

Whereas I'll sometimes talk to people who are like, “Okay, this is where I want you to use subagents for this subtask,” and I'm sure they would appreciate seeing how they're being used. For me, it's primarily 2 things. One is net time efficiency, so span it out across subagents. Two is probably cost, right?

Nick Turley

Right.

Alessio Fanelli

Don't use a big, expensive model. Offload to a lot of smaller, cheaper models. Some people want that level of control.

If you have repetition in what you're doing, say I want something built where I want it to consistently do this every day, I might want to go in and fine-tune subagents here, subagents there.

Shawn Wang

Yeah.

Alessio Fanelli

You can see both, but I think, if I'm not mistaken, it's hidden by default. There's a dropdown where you can get more information about what's been done.

Shawn Wang

Oh, you can change the model that they use?

Alessio Fanelli

I know Anthropic offers this in Cloud Code. You can tell Fable to use Sonnet or Opus to use Sonnet as a subagent. It's a pretty trivial thing. You tell it to spin out subagents with Sonnet, and you know it's cheaper and faster. I would assume if it's not there, it could be built there. But I think there's a side of—

Shawn Wang

It's too many toggles.

Nick Turley

Mm.

Alessio Fanelli

It's not a toggle, actually. You just tell it in chat.

Nick Turley

You're prompting it. Yeah.

Alessio Fanelli

The way I do it is to prompt it, right? I think this is something that gets abstracted unless it's something you built for repetition, right?

If I'm building something for podcast prep—research into people, doing very deep, extensive research—I might want to configure it to use a cheaper, faster model just for web search. I can see a world in which you want both. I think the default is actually pretty good right now, where it's hidden, but you can drop down and get more information about what's been done.

Nick Turley

Mm-hmm.

Alessio Fanelli

I know people talked a lot about it on 5.6's launch. This thing loves to use a lot of subagents and causes the ChatGPT app to just crash because it's so processor-heavy. But—

Shawn Wang

For what it's worth, that's not my experience. I haven't had a crash from subagents.

Alessio Fanelli

I haven't either. We both have big laptops. But I know people brought it up. It was a topic of discussion that we didn't see the same, but it is another vibe eval, right? People are like, “Okay, the amount of subagents Sol is wanting is crazy.” And I'm like, “I think this is okay. I think it's good.” But it's just stuff people bring up.

Nick Turley

I think when we launched the product, too, we weren't as opinionated about who Ultra is for and when they should be using it. Since then, we've made some changes to require you to turn it on and find it in the advanced settings, because that's who it's for.

It's for power users who understand what's going to happen, because, depending on your use case, it can also use more of your limits.

Alessio Fanelli

Yes.

Nick Turley

So that's where I think a lot of the feedback was coming from.

Alessio Fanelli

It's okay. Reset the limits. Always reset the limits.

8. Memory Becomes The Next Frontier

Shawn Wang

Well, today we're resetting because of this. I want to change topics to one last piece of the harness: memory. A lot of people are commenting on memory recently.

ChatGPT’s new memory system used to suck; it’s not very good. And then this guy also basically said the same thing, and Samir, who you presumably work with—

Nick Turley

Mm-hmm.

Shawn Wang

—talking about memory. What can you say there?

Nick Turley

I think that Samir and the team, and the research teams, have made a ton of updates and improvements over time. When I talk to friends and family members about what they love about ChatGPT, the fact that it knows them—that they feel like their ChatGPT is their ChatGPT—I think probably comes up—

Shawn Wang

Yeah.

Nick Turley

—number one. In ChatGPT Work in the cloud, by default, all conversations will inherit from your ChatGPT memory, so they’ll know context about you, and they’ll also be able to write back to this memory.

Shawn Wang

With a small text write? Like, you tell me when you’re writing, right? Is it—

Nick Turley

No, it’s part of the same Memory V3 system that we launched.

Shawn Wang

Okay.

Nick Turley

Yeah. Dreaming V3. Yeah.

Nick Turley

I think that’s been really powerful because going from ChatGPT to ChatGPT Work feels like an extension of what I’ve already been doing with the product for many years.

Shawn Wang

Yeah.

Nick Turley

So that’s been awesome, and it’s awesome to see that people are recognizing—

Shawn Wang

Yeah.

Nick Turley

—the improvements here.

Shawn Wang

Is it basically a retrieval problem? Are you retrieving the right things? Are you over-focusing on the wrong things? Is there more of a false-positive or false-negative problem, if that makes sense? What’s the bigger problem?

Adam Fry

I don’t work on memory directly—

Shawn Wang

Okay.

Adam Fry

—so it’s hard to say what the bigger problem is with certainty, but I think you’re right. There are 2 sides of it. It’s making sure it knows things about you, but then also having the EQ to bring those things up proactively at the right moments, or surprising you in ways that are positive, not negative.

Shawn Wang

Yeah.

Adam Fry

So I think it’s a very challenging problem, but something that I think we feel is a huge opportunity to get right, which is why we’ve made big investments in it.

Alessio Fanelli

How do you see the side of—when you’re building ChatGPT for work, different from the regular chat app and different from Codex, managing memory across different projects, collaboration, and whatnot? How do you see what’s separate from the harness? So if I have 4 threads on 1 project—

Adam Fry

Mm-hmm.

Alessio Fanelli

—any learnings on how to build memory systems there? For background, to steer it a bit, when you do chat-style applications, I’d say you have a lot of one-offs, right?

Adam Fry

Yeah.

Alessio Fanelli

When you switch to work, it might be something you’re doing for a month, something you do a lot, right? Now, as I add more sessions, there’s a lot more than just single-threaded, right?

Adam Fry

Sure.

Alessio Fanelli

And there might be memory there.

Adam Fry

I think first I’d challenge that the depth of the memory, or the value of it, is fundamentally different across chat and work. It is true that there are a lot of shorter sessions on chat, but I think the ChatGPT product has had a ton of longevity—as long as this technology has been around—and people use it for work-y, productivity-related things already today.

And so I think we found that there’s a lot of value. I found this in my personal usage: all these one-offs add up over time into something quite durable and quite a good representation of who I am. I know from time to time something will go viral on X about ChatGPT telling you everything it knows about you, and people are always surprised by how deep that is.

Alessio Fanelli

The “roast me,” you know?

Adam Fry

Exactly. So I think that’s all to say that there’s a lot of depth in the existing ChatGPT product, and that’s why I think we think it’s valuable to bring into the work product. But the other reason I brought that up is because I think hopefully we can use some of the same fundamental primitives and systems to extend memory here as well, and I know this is something that the team that focuses on this is working through right now.

Shawn Wang

I wanted to bring up 1 element of memory that I honestly don’t really use much, and I’m curious if you do: Chronicle, which is up on screen right now. It’s kind of a super memory, or what is it?

Adam Fry

I think the idea is that it can learn from how you’re using your computer, and it’s another input source into memory. I think it’s experimental right now and isn’t on by default, but I’d recommend that you try it.

I think it’s quite interesting because it goes back to a conversation we were having earlier. You were asking, “Can ChatGPT miss things?” Does it, on Slack, when it’s searching, miss things because there’s such a volume of stuff? You can ask the same question about everything that you’re doing on your computer: Is it going to know everything that you’re doing? Is it going to capture the intent and stuff like that?

Probably not, but it probably will find things that you might not know about. If it can surface those to you at relevant times, in proactive ways, when you’re doing tasks, I’ve found at least that it can be quite helpful, so it’s worth trying.

Shawn Wang

So mostly for insights and the longer term?

Adam Fry

Yeah, exactly. Insights, and it builds context that can make you more productive on certain tasks. But it’s hard to describe without feeling it.

Alessio Fanelli

I will say you can feel it pretty well. The idea of what they’re saying here is, “Just check through my memories or check through my logs and add skills.” Pretty underrated, right?

Adam Fry

Yeah.

Shawn Wang

Yeah, but that’s automations. You can repeat that using a cron job.

Adam Fry

Checking through your memories and creating skills?

Shawn Wang

Yeah.

Adam Fry

But I think the creation—

Shawn Wang

I see.

Adam Fry

—of memories from Chronicle itself is what’s different.

Shawn Wang

Okay.

Adam Fry

It’s like you have much deeper memories because you have Chronicle on.

Shawn Wang

It’s there. I don’t use it much, but maybe I just need more examples. I imagine you guys use a lot of it internally, so I’m always fishing for use cases.

Adam Fry

Yeah. I would just try turning it on and—

Shawn Wang

It just auto-works? Like it—

Adam Fry

Yeah, and seeing where—

Shawn Wang

Right.

Adam Fry

—where it might start helping you. I think you’d be surprised.

Shawn Wang

Yeah. Amazing. I think that was about it in terms of the overall coverage of ChatGPT Work. I think there’s been a lot of good progress and discussion on building and all these things. There are a lot of ex-founders in the community, in OpenAI as well. Do you think that things have changed a lot? I guess my overall reflection is on building pre-AI and post-AI.

9. AI Rewrites How Teams Build

Adam Fry

I think things have changed a ton. It’s super exciting to see how quickly you can go from an idea to something real today.

Shawn Wang

Yeah.

Adam Fry

Whereas even before, 5 or 10 years ago, it was fast if you were scrappy and willing to build the minimal viable thing. But now the extent of what you can build is much, much broader. And I think what we’ve seen internally building is that this gives you an opportunity to validate much more quickly, to talk to users, to talk to internal doctors, et cetera, and make sure you’re on the right track.

That loop has become more closed than ever before, and that’s a win for product development. I think it’s a win for consumers and users too, because ideally that means they’re getting much better products out the gate.

Alessio Fanelli

Does it mean your teams are smaller?

Adam Fry

I think there’s much more to do now. People can accomplish more individually or in a small team than they could before—things that would have required more people in the past. But at the same time, there’s also more to do, so I think the teams are much more ambitious.

Alessio Fanelli

Have you seen any changes in the scope of roles and in building teams? How did we used to have teams a few years ago, versus what do ideal teams look like now?

Adam Fry

I think we’ve seen a blurring of the lines between the typical product development functions—between EM, PM, engineer, designer, and so on.

Alessio Fanelli

Yeah, I want to bring up this quote: “There will be only 4 jobs left in tech.” There’s AI slop cannon—the people who just burn a bunch of tokens. And then there is SRE, people who are more responsible. There are grown-ups who sell things, and then there are hot people.

Adam Fry

This is an interesting take. I think my suspicion is that there’s everything...

Everyone will be T-shaped in a way, and AI will enable everyone to become a generalist.

Shawn Wang

Yeah.

Adam Fry

Things that I never would have been able to come up with a design for before—I don't have the visual taste required—but I can iterate on something with the help of AI. People will have a specialty, and that's, I guess, the straight line in the T, or the upward line in the T. You can have a specialty that you're interested in, and with the help of AI, you can go deeper and become better at it over time, but then you'll also be a generalist. With that foundation, what you can accomplish is almost limitless.

Shawn Wang

What are you bottlenecked by in terms of specialties? Do you need more designers? Do you need more slop cannons? Do you need more hot people?

Adam Fry

I think the bottleneck becomes ideas and taste, I guess. Because anyone can build now, I think it really is the era of bottoms-up ambition. Because there's so much to be built, you're always going to be bottlenecked by the amount of ideas and the amount of things that you're doing at any given time.

Shawn Wang

Do you think models help solve that?

Adam Fry

Models?

Shawn Wang

Yeah. I mean, I have a front-end design skill where they give me 4 drastically different examples of what this looks like. Sure, it burns a lot of tokens, but then I'll mostly just condense them down: “Okay, I like this part. I like this part. Let's draw these together.” It's like, yeah, I had a vision, but I don't know.

Adam Fry

The one automation that I would love to work, but doesn't, is: bring me new ideas. Somehow, LLMs are just not it. One interesting part about ideas is that they're not in a vacuum. They usually come from somewhere, and in product development, they're coming from talking to users, reacting to friction that you're seeing or feedback, building on some foundation that you already had planned out before, whatever.

I think that's where there will always be value in these generalists that we talked about: closing that loop and then coming up with those ideas that are grounded in that feedback, talking to users, whatever it is.

Shawn Wang

Cool. You lead the productivity team. How do you define productivity?

Adam Fry

I think our mission is to make it possible for people to do things that they weren't able to do before. Right now, we're thinking about it from the perspective of knowledge work. When I look at knowledge work, I think about how people are no longer siloed by their roles. They're no longer siloed by the background or training that they have. No matter what function you're in, you can suddenly build things.

You can suddenly get access to data that you otherwise might not be able to interpret. I think that extends to your personal life, where we want to give you leverage at the end of the day. We want the models and the product to be able to give you leverage so that you can create time for yourself to do the things that you love.

Shawn Wang

Does that also translate to a way to measure productivity? How do you measure leverage?

Adam Fry

The end is—

Shawn Wang

How do you measure leverage?

Adam Fry

I think we haven't figured this out yet. Part of the reason is that it's so diverse. Everyone has different goals, and really, the true measurement is their ability to achieve that goal. Did we help you, or did we not?

Shawn Wang

Yeah.

Adam Fry

It's very difficult without knowing what that goal is up front and also tailoring it for every individual.

Shawn Wang

And the thumbs-up and thumbs-down from ChatGPT doesn't give you anything, right?

Adam Fry

Right. I mean—

Shawn Wang

Oh, yeah.

Adam Fry

You don't know if they're thumbs-downing the content of the answer, the vibe of it—

Shawn Wang

Yeah.

Adam Fry

—or whether or not it helped them with their goal.

I think that's difficult, but it's something that we will need to figure out, and the industry at large will need to figure out, because that's how we measure success.

Shawn Wang

Do you think productivity and how you measure it has changed? Basically, you said there's a lot more work that can be done, a lot more scope. Has it changed?

Adam Fry

I think it was always true that what you really wanted to measure was whether your team, the individual, or you personally were able to hit the goal, or were closer to hitting whatever your goal was. But I think previously we used proxies for this, like code commits or—

Shawn Wang

Lines of code.

Adam Fry

Lines of code, or whatever.

Shawn Wang

Story points.

Adam Fry

Yeah, exactly. Story points.

Shawn Wang

They're coming back, by the way.

Adam Fry

Maybe. I mean, but that is part of the change. With AI now, I think those proxies are starting to fall apart. The number of tokens you use or the number of pull requests you make are no longer as hypercorrelated with whether your team is able to hit the goal or is on track to hit its goals. I think we'll need to come up with new measurements.

Shawn Wang

For the managers listening, give them one thing to try.

Adam Fry

I think what's important for me is at-bats. Are we as a team building the muscle to have not just quantity of at-bats, but quality? Are we able to go all the way from generating an idea, building it out, getting the feedback, reacting to that feedback, actually validating or invalidating the hypothesis, and going on to the next idea? Are we able to do that really efficiently?

That goes to the actual code that's being written, the designs that are being made, or the specs that are being written, but also the culture of the team. Do we have the humility, and are we able to go through that process many times and stay motivated and excited throughout that? That's the thing that I think is important now, especially when we're on the frontier of this technology. There's so much to build and so much to do. That's probably the most important thing that we look at.

Shawn Wang

Any traps people fall into around measuring productivity or what their teams work on? I feel like there's a lot of, “Okay, we added a lot of LLMs, and we have dashboards for this and that,” but not much has changed, right?

Adam Fry

That is the trap, yes.

Shawn Wang

The broader source of the question is for the managers and teams building: how should they approach this?

Adam Fry

I think maybe the trap is conflating motion and progress. Motion is much easier now than ever before because of the tooling that we have. But progress requires you to be very prescriptive and deliberate about what you're actually trying to achieve. It goes back to our question of measurement, right?

We were talking about whether we, OpenAI, can figure out how to measure productivity for our users. That's a very hard problem because of the diversity. But as a team, you should have a really prescriptive and deliberate view on what progress looks like for you and for your team. If you don't have that, then it's very easy to conflate these 2 things.

Shawn Wang

I think at-bats is a really great thing. I'm really glad. I like the discussion between motion and progress. I think that's a quote that we're going to feature in the write-up. You've been very generous with your time. Thank you so much, and congrats on 10 million.

Adam Fry

Yeah, thank you for having me.

Shawn Wang

The next one at 100 in 2 months.

Adam Fry

Sure.

Shawn Wang

2 weeks. Thank you.