[BidClub_]
Hard Fork · · 66 分钟

DeepSeek 深度剖析 + 亲测 Operator + 翻车快车!

Kevin RooseCasey NewtonJordan Schneider

播客
TL;DR
  • DeepSeek 缩小了外界感知中的中美模型差距,但并不能证明中国已经赢下更广义的 AI 竞赛。 其应用近期在 iOS 的下载量达到190万,在 Google Play 上约120万;美国海军禁用、意大利在数据保护调查后禁用,以及 Microsoft-OpenAI 对模型蒸馏的调查,都凸显了安全与数据使用风险。Jordan Schneider 的基本判断仍是,中国实验室能够“快速跟随”前沿模型。

  • DeepSeek 的组织设计帮助它实现了技术突破。 这家公司诞生于一家成功的量化对冲基金,招募年轻研究人员,在没有直接盈利压力的情况下运营;与此同时,Alibaba、Tencent、ByteDance 和 Huawei 都难以搭建适合突破性协作的机构结构。Schneider 称其团队是类似2017年至2022年间 OpenAI 的“梦想家”,但成功也可能推动 DeepSeek 与某家 hyperscaler 合作,并承担潜在的国家队职责。

  • 高效模型并不能消除算力需求,因此出口管制仍有战略价值。 中国仍能买到 Nvidia 的 H20,Schneider 称它基本属于世界级部署 AI 硬件,足以让所有人使用;他认为硬件差距“其实没有那么大”,但 DeepSeek 创始人已经将管制视为约束。Kevin Roose 关于消费级硬件的情景判断可能性上升了约10个百分点;Schneider 的反向信号是,Nvidia 下跌了15%,而不是95%。

  • 6个月的模型领先,可能不如部署能力和机构采用更重要。 Casey Newton 认为,DeepSeek 高效复现了美国公司9至12个月前训练出的能力;如果中国率先交付出一款惊艳的 agent 或“虚拟同事”,他才会重新评估。Schneider 认为,最终竞争还要看数据中心、监管、劳工阻力,以及社会吸收工业革命级重组的能力。

  • Operator 的每月200美元定价,与其说是高端聊天机器人,不如说更像一款早期劳动力产品。 它托管了一个“浏览器里的浏览器”,可以点击普通人类网站而不依赖 API,最终目标是充当虚拟工程师、顾问、律师、医生和研究助理。正如 Roose 所说,“每月200美元买一个 ChatGPT 版本很贵,但买一个远程员工并不贵。”

  • 亲测结果显示,Operator 确实具备自主性,但可靠性不足且成本高昂。 它完成了域名购买、主机和 DNS 配置项目约90%的工作,但在 Instacart 上搜索牛奶时定位到了 Des Moines,耗时至少是手动操作的10倍;另一次运行花45分钟做问卷,只赚到1.20美元。这段经历像是在培训“一名非常新、非常不自信的实习生”。

  • Agent 能力正在快速提升,但产品仍横跨一道隐私与安全鸿沟。 Anthropic 的 Computer Use 在 OSWorld 上得分14.9%,3个月后 OpenAI 的 CUA 达到38.1%——仍是不及格,但提升斜率惊人。要实现无缝使用,系统需要访问用户已登录的浏览器和支付信息;而自主行动最终可能从操纵购买一路发展到网络攻击。

  • 本周几起较小的失败,展示了仓促自动化和实体技术伴随的风险。 Fable 因生成种族主义摘要而移除全部 AI 功能;Amazon 在两次雨天事故后暂停无人机配送;Fitbit 在174起过热报告和118人受伤后支付1200万美元赔偿。对 Roose 来说,Waymo 遭到破坏只是一次“温吞的翻车”,但随着自动驾驶汽车成为反技术情绪攻击的实体目标,事态仍有升级空间。

摘要 · 为研究而整理的核心内容

1. DeepSeek 的突破变成了分发与政策事件

  • Casey 开场时的速写捕捉到了这次冲击的速度:近期 iOS 下载量190万,Google Play 下载量约120万,随后美国海军禁用,意大利也在数据保护调查后禁用。

  • OpenAI 表示掌握了 DeepSeek 蒸馏其模型的证据;Microsoft 和 OpenAI 正在调查可能存在的 API 滥用。主持人指出,OpenAI 一边反对“不付费或未经同意”使用数据,一边又面临 The New York Times 的版权诉讼,这颇具讽刺意味。

  • Schneider 称 DeepSeek 是一只“奇怪的鸭子”:其 CEO 来自一家成功的量化对冲基金,在 ChatGPT 出现后投入资金、算力和“新鲜的年轻毕业生”开发语言模型,却没有明显的近期商业模式。

2. DeepSeek 的自由度带来了突破,而老牌公司难以完成组织动员

  • Schneider 将 DeepSeek 与 Alibaba、Tencent、ByteDance 和 Huawei 作了对比:这些大公司拥有资源,却难以建立能够奖励突破性协作与研究所需的“组织和机构结构”。

  • 不受直接盈利动机束缚,帮助 DeepSeek 从12月底开始实现一系列显著创新,随后推出 R1 聊天机器人。Schneider 基于两次对 CEO 的长篇访谈和员工发帖形成的工作判断很简单:“他们是梦想家。”

  • 最接近的参照物是2017年至2022年的 OpenAI,当时 Sam Altman 可以直言自己不知道公司将如何赚钱。DeepSeek 当时的雄心似乎也类似:“我们先把它做出来,让所有人都能更便宜地使用。至于以后怎么赚钱,之后再想。”

  • 其交易策略目前或许足以为这项使命提供资金,但 Schneider 认为,这次关注度可能开启新阶段,迫使 DeepSeek “和一家 hyperscaler 搭伙”——候选对象可能是 ByteDance、Alibaba、Tencent 或 Huawei,因为政府审查开始加强。

3. 成为国家队可能损害 DeepSeek 最初的使命

  • Schneider 否认中国科技 CEO 的首要梦想是传播 Xi Jinping Thought;他们想与 Mark Zuckerberg 和 Sam Altman 竞争,证明自己是“真正厉害、真正优秀的技术专家”。

  • 他的警告来自 ByteDance 的历史:Zhang Yiming 早年在 Weibo 上的发帖支持更自由的表达,但到2018年,随着政治环境收紧并波及公司,他公开道歉,并承诺更加遵循中国社会主义价值观。

  • Didi 提供了更严厉的先例。政府要求其不要上市后,它仍在西方交易所挂牌;随后被下架,并进入“整改流程”——一家无政治立场的公司越过国家红线后被“击落”。

  • DeepSeek 过去一直低调行事,但 Schneider 表示,那段时期已经结束。国家队地位可能带来政府合同、注意力分散以及更深的国家介入,即便这些责任与其广泛开发和部署 AI 的使命相冲突。

4. 开源成功带来的是审查困境,而不是能力上限

  • 托管版 DeepSeek 拒绝回答关于天安门广场、Xi Jinping 和大跃进的问题,但 Casey 指出,开源 V3 模型显然没有接受同样的规避训练,可以被下载、托管到其他地方,并去除部分护栏。

  • Schneider 称“审查导致能力不足”的说法“有点转移焦点”。Claude 不会按要求生成种族主义内容,但能力并未因此消失;同样,政治限制也不必然意味着中国模型在结构上不如西方模型聪明。

  • 真正的困境是政治性的:一个开放的中国模型可以在海外被诱导生成可能导致其在中国国内受罚的言论。但 DeepSeek 已经让中国 AI 在全球获得了“有史以来最积极的光环”,迫使监管者在控制与声望之间寻找平衡。

5. 即使效率改变胜率,算力管制仍然重要

  • DeepSeek 是否拥有走私来的受禁芯片,“更应该由美国情报界回答,而不是由 Twitter 上的 Jordan Schneider 回答”。中国的合法选择并非无足轻重:Nvidia 的 H20 仍基本属于世界级硬件,足以部署 AI 并让所有人使用。

  • Schneider 仍认为,无论未来蒸馏技术如何发展,算力都是核心投入;他还指出,DeepSeek 创始人曾将出口管制称为公司的主要约束。如果美国继续限制芯片和半导体制造设备,中国将很难实现超越,更不用说在开发出可比模型后继续快速跟随。

  • 他最担心的情景,是一笔用“大豆换 ASML EUV 机器”的交易。管制只有在有条件的情况下才有效;如果政治交易削弱管制,就可能动摇他认为美国及其盟友仍拥有的优势。

6. 模型持平只是 AI 竞争的一条前线

  • Kevin 反驳称,算法效率最终可能把“最大、最强、最聪明”的模型装进 MacBook,让芯片管制和扩散限制失效。Schneider 承认,DeepSeek 可能让这一结果的概率上升了10个百分点。

  • 他用市场表现回应:Nvidia 下跌了15%,而不是95%。如果消费级硬件真的让先进芯片变得无关紧要,他预计市场重估幅度会大得多;在他的“历史系学生脑袋”看来,本地模型民主化的未来仍然可能,但“相当不太可能”。

  • Casey 认为,V3 和 R1 主要是在优化美国公司9至12个月前训练出的模型。只有当中国率先推出一款非凡的虚拟同事或 agent,他才会认为竞争格局真正发生转变,而不是仅仅高效追赶。

  • Kevin 怀疑,领先6个月是否足以阻止世界形成两极 AI 格局。Schneider 则把视野拉得更宽:前沿实验仍需要巨型数据中心,而落地还会遭遇教师工会、监管阻力,以及类似工业革命规模的社会重组。

7. Operator 将 Agent 论点变成了一个可见且昂贵的原型

  • OpenAI 紧随 Anthropic 的 Computer Use 和 Google 的 Project Mariner,推出了 Operator,面向 ChatGPT 每月200美元的 Pro 订阅用户。它运行在独立网站上,在 OpenAI 服务器上打开浏览器,而不是接管用户自己的电脑。

  • 这个“浏览器里的浏览器”没有用户现成的书签或登录会话,但用户可以在需要时接管。Operator 通过视觉方式操作普通人类网站,在没有 API 的地方用通用性换取速度和结构化程度。

  • 商业上的北极星是虚拟劳动力:先做工程师,之后扩展到顾问、律师、医生,也可能包括研究助理。Casey 很愿意为有用的研究帮助支付200美元;Kevin 则指出,相比一名客服或计费员工,这个价格显得并不高。

8. Operator 最好的一次像实习生,最差的一次则需要持续救场

  • 在 TripAdvisor 上,Operator 找到了伦敦的步行游,并在几分钟内完成总结。Casey 自己搜索可能更快,但看着这个“自动驾驶浏览器”导航、阅读并汇报,仍然像一次真正的技术突破。

  • Instacart 的表现糟糕得多:Casey 最合理的解释是,Operator 继承了服务器所在的 Iowa 位置,于是搜索 Des Moines 的牛奶,然后把他偏好的旧金山杂货店当成配送地址。他不得不接管、登录并纠正错误,而这已经耗费了正常操作至少10倍的精力。

  • Kevin 的更难测试反而好得多。他要求 Operator 找到50美元以下最好的可用域名、完成购买、安排主机并配置 DNS;在他去喝咖啡期间,Operator 完成了约90%,只在支付、注册信息和主机套餐选择上需要他介入。

  • 最有效的指令很像是在管理“一名非常新、非常不自信的实习生”:“只有在你确实无法继续推进时,才请求我介入。”另一次任务中,它为2个人点了2个小 tacos,还需要被追问“这样够吃吗?”才重新考虑。

9. 基准测试快速进步,但撞上了敌意互联网和真实安全风险

  • Operator 花45分钟做在线问卷赚到1.20美元,Kevin 说他确信这消耗了价值数百美元的 GPU 算力;它还拒绝参与在线扑克。The New York Times、Reddit 和 YouTube 都拦截了它,GoDaddy 也不允许它购买域名,说明网站可以多快地限制通用型 agent。

  • Casey 乐观的一项数据来自 OSWorld:Anthropic 的 Computer Use 得分14.9%,而 OpenAI 表示 CUA 在3个月后达到38.1%。“38.1%仍是不及格,”但如果3至6个月就能取得类似幅度的进步,系统可能会成为一名合格的电脑使用者。

  • 他对产品的质疑在于,agent 使用浏览器与 agent 使用你的浏览器之间存在鸿沟。现有的登录状态和支付信息可以消除摩擦,但也会带来更大的隐私与安全暴露;OpenAI 表示,在人类接管 Operator 时,系统会停止截图。

  • Kevin 警告说,“互联网不会原地不动。”如果 agent 占到流量的10%、20%或30%,广告业现有的假设可能崩溃;商家可能会优化信息来影响机器人购买,而更高程度的自主性则可能让 agent 发起网络攻击或盗取加密钱包资产。

10. 本周翻车从冒犯性软件一路延伸到实体伤害

  • Fable 的 AI 告诉一位读者“偶尔找一位白人作者来撑场”,还质疑另一位读者为什么会关注“一个直男、顺性别白人的视角”。在其缓解措施显然失效后,Fable 移除了所有 AI 功能,并提交了替代版本应用——这是一次毫无争议的重大翻车。

  • Amazon 在德州和亚利桑那州暂停商业无人机配送,此前最新型号的2架飞机在12月的雨天测试中坠毁。由于没有人员受伤,主持人将其评为中等或偏温和的翻车,但也强调了重型机器在人群上空飞行时的质量控制风险。

  • Fitbit 同意支付1200万美元,此前被指未能及时报告烧伤风险。2018年至2022年3月期间,它至少收到174起过热报告和118起伤情报告,其中包括2起三级烧伤和4起二级烧伤:“如果技术产品把你生理性地烧伤了,那就是一次重大翻车。”

11. 地图服从政治,Waymo 则成为实体目标

  • Google 表示,Maps 将根据政府更新,把 Gulf of Mexico 更名为 Gulf of America,并将 Denali 更名为 Mount McKinley。Kevin 称这种例行的地名服从只是轻微的“茶杯里的风暴”;Casey 则认为这场混乱是一次重大翻车,2人罕见地给出了分裂判断。

  • 一群人在洛杉矶一次非法街头接管活动中拆毁了一辆 Waymo,并用零件砸碎车窗。动机仍不明。Casey 指出,Waymo 直到11月才正式在洛杉矶提供服务,因此新鲜感和好奇心可能起了一定作用;Kevin 将事件评为一次“温吞的翻车”,但随着一些人把 Waymo 视为技术侵入生活每个角落的实体化身,事态仍有升级空间。

Kevin Roose

I set up ChatGPT to email me—

Casey Newton

Ooh.

Kevin Roose

A weekly affirmation before we start taping—

Casey Newton

Oh.

Kevin Roose

Because you can do that now with the Tasks feature.

Casey Newton

Yeah, people say this is the most expensive way to email yourself a reminder. So what sort of affirmation did we get?

Kevin Roose

Today it said, “You are an incredible podcast host, sharp, engaging, and completely in command of the mic. Your taping today is gonna be phenomenal, and you're going to absolutely kill it.”

Casey Newton

Wow, and that's why it's so important that ChatGPT can't actually listen to podcasts—because I don't think it would say that if it had ever heard us.

Kevin Roose

Yeah. It would say, “Just get this over with.”

Casey Newton

Get on with it.

Kevin Roose

I'm Kevin Roose, a tech columnist at The New York Times.

Casey Newton

I'm Casey Newton from Platformer.

Kevin Roose

And this is Hard Fork.

Casey Newton

This week we go deeper on DeepSeek. ChinaTalk's Jordan Schneider joins us to break down the race to build powerful AI. Then, “Hello, Operator”: Kevin and I put OpenAI's new agent software to the test. And finally, the train is coming back to the station for a round of Hot Mess Express.

1. DeepSeek Takes Center Stage

Kevin Roose

Well, Casey, it is rare that we spend 2 consecutive episodes of this show talking about the same company, but I think it is fair to say that what is happening with DeepSeek has only gotten more interesting and more confusing.

Casey Newton

Yeah, that's right. It's hard to remember a story in recent months, Kevin, that has generated quite as much interest as what is going on with DeepSeek. Now, DeepSeek, for anyone catching up, is this relatively new Chinese AI startup that released some very impressive and cheap AI models this month that lots of Americans have started downloading and using.

Kevin Roose

Yeah, some people are calling this a Sputnik moment for the AI industry, in which every nation perks up and starts paying attention to the AI arms race at the same time. Some people are saying this is the biggest thing to happen in AI since the release of ChatGPT. But, Casey, why don't you just catch us up on what has been happening since we recorded our emergency podcast episode just 2 days ago?

Casey Newton

Well, I would say that there have probably been 3 stories, Kevin, that I would share to give you a quick flavor of what's been going on. One, a market research firm says DeepSeek was downloaded 1.9 million times on iOS in recent days, and about 1.2 million times on the Google Play Store. The second thing I would point out is that DeepSeek has been banned by the U.S. Navy over security concerns, which I think is unfortunate, because what is a submarine doing if not deep-seeking?

It was also banned in Italy, by the way, after the data protection regulator made an inquiry. And finally, Kevin, OpenAI says that there is evidence that DeepSeek distilled its models. Distillation is the AI lingo or euphemism for, “They used our API to try to unravel everything we were doing and use our data in ways that we don't approve of.” Microsoft and OpenAI are jointly investigating whether DeepSeek abused their API. And, of course, we can only imagine how OpenAI is feeling about the fact that their data might have been used without payment or consent.

Kevin Roose

Well, yeah, it must be really hard to think that someone might be out there training AI models on your data without permission.

Casey Newton

And I want to acknowledge that literally every single user of Bluesky already made this joke, but they were all funny, and I'm so happy to repeat it here—on Hard Fork this week. Now, Kevin, as always when we talk about AI, we have certain disclosures to make.

Kevin Roose

The New York Times Company is currently suing OpenAI and Microsoft over copyright violations allegedly related to the use of their copyrighted data to train AI models.

Casey Newton

And I—

Kevin Roose

I think that was good.

Casey Newton

That was very good. And I'm in love with a man who works at Anthropic. Now, with that said, Kevin, we want to go even further into the DeepSeek story, and we want to do it with the help of Jordan Schneider.

Kevin Roose

Yes, we are bringing in the big guns today because we wanted to have a more focused discussion about DeepSeek that is not about the stock market or how the American AI companies are reacting to this, but is about one of the biggest sets of questions that all of this raises: What is China up to with DeepSeek and AI more broadly? What are the geopolitical implications of the fact that Americans are now obsessing over this Chinese-made AI app? What does it mean for DeepSeek's prospects in America? What does it mean for its prospects in China? And how does all this fit together from the Chinese perspective?

Jordan Schneider is our guest today. He's the founder and editor-in-chief of ChinaTalk, which is a very good newsletter and podcast about U.S.-China tech policy. He's been following the Chinese AI ecosystem for years. And unlike a lot of American commentators and analysts who were surprised by DeepSeek and what they managed to pull off over the last couple of weeks—

Casey Newton

I'll say it: I was surprised.

Kevin Roose

Yeah, me too. But Jordan has been following this company for a long time, and a big focus of ChinaTalk, his newsletter and podcast, has been translating literally what is going on in China into English, making sense of it for a Western audience, and keeping tabs on all the developments there. So he's the perfect guest for this week's episode, and I'm very excited for this conversation.

Casey Newton

Yes. I have learned a lot from ChinaTalk in recent days as I've been boning up on DeepSeek, so we're excited to have Jordan here, and let's bring him in. Jordan Schneider, welcome to Hard Fork.

Jordan Schneider

Oh my God, such a huge fan. This is such an honor.

Casey Newton

Oh, we're so excited to have you. I have learned truly so much from you this week, and so when we were talking about what to do this week, we just looked at each other and said, “We have got to see if Jordan can come on this podcast.”

Kevin Roose

Yeah, this has been a big week for Chinese tech policy. Maybe the biggest week for Chinese tech policy, at least that I can remember. I realized that something important was happening last weekend when I started getting texts from all of my non-tech friends saying, “What is going on with DeepSeek?” I imagine you had a similar reaction because you are a person who constantly pays attention to Chinese tech policy.

Jordan Schneider

So I've been running ChinaTalk for 8 years, and I can get my family members to read maybe 1 or 2 editions a year. The same exact thing happened to me, Kevin, where all of a sudden I got, “Oh my God, DeepSeek! It's on the cover of The New York Post. Jordan, you're so clairvoyant. Maybe I should read you more.” I'm like, “Okay, thanks, Mom. Appreciate that.”

Kevin Roose

Yeah, I want to talk about DeepSeek and what they have actually done here, but I'm hoping first that you can give us the basic lay of the land of the Chinese AI ecosystem, because that's not an area where Casey or I have spent a lot of time looking. But tell us about DeepSeek and where it sits in the overall Chinese industry.

2. DeepSeek Emerges From A Hedge Fund

Jordan Schneider

So DeepSeek is a really odd duck. It was born out of this very successful quant hedge fund, the CEO of which, after ChatGPT was released, was basically like, “Okay, this is really cool. I want to spend some money and some time and some compute and hire some fresh young graduates to see if we can give it a shot to make our own language models.”

Kevin Roose

And so a lot of companies are out there building their own large language models. What was the first thing that happened that made you think, “Oh, this company is actually making some interesting ones”?

Jordan Schneider

Sure. So there are lots and lots of very moneyed Chinese companies that have been trying to follow a similar path after ChatGPT. We have giant players like Alibaba, Tencent, ByteDance, and even Huawei trying to create their own OpenAI, basically. And what is remarkable is that the big organizations can't quite get their heads around creating the right organizational and institutional structure to incentivize this type of collaboration and research that leads to real breakthroughs.

Chinese firms have been releasing models for years now, but because of the way DeepSeek structured itself and the freedom it had from not necessarily being under a direct profit motive, it was able to put out some really remarkable innovations that caught the world's attention, starting maybe in late December, and then really blew everyone's mind with the release of the R1 chatbot.

Kevin Roose

Yeah, so let's talk about R1 in just a second, but one more question for you, Jordan, about DeepSeek. What do we know about their motivation here? Because so much of what has been puzzling American tech industry watchers over the last week is that this is not a company that has an obvious business model connected to its AI research, right?

We know why Google is developing AI, because it thinks it's going to make Google much more profitable. We know why OpenAI is developing advanced AI models. It does not seem obvious to me, and I have not read anything from people involved in DeepSeek, why they are doing this and what their ultimate goal is. So can you help us understand that?

Jordan Schneider

We don't have a lot of data. But my base case, which is based on 2 extended interviews that the DeepSeek CEO released, which we translated on ChinaTalk, as well as what DeepSeek employees have been tweeting about in the West and domestically, is that they're dreamers.

I think the right mental model is OpenAI from 2017 to 2022. I'm sure you could ask the same thing: What the hell are they doing? Sam Altman literally said, “I have no idea how we're ever gonna make money,” right? And here we are in this grand new paradigm.

So I really think that they do have this vision of AGI and are thinking, “Look, we'll build it, and we'll make it cheaper for everyone. We'll figure it out later.” They have enough trading strategies that they can fund it, and now that they've really blown people's minds, we might be entering a new period in DeepSeek's history, kind of like what happened with OpenAI, right?

They're going to have to shack up with a hyperscaler, whether it’s ByteDance, Alibaba, Tencent, or Huawei instead of Microsoft. And the government's going to start to pay attention in a way that it really hasn't over the past few years.

3. DeepSeek Faces Government Pressure

Kevin Roose

Right. And I want to drill down a little bit there, because I think one thing that most listeners in the West do know about Chinese tech companies is that many of them are sort of inextricably linked to the Chinese government. The Chinese government has access to user data under Chinese law, and these companies have to follow the Chinese censorship guidelines.

As soon as DeepSeek started to really pop in America over the last week, people started typing things into DeepSeek's model like, “Tell me about what happened at Tiananmen Square,” or, “Tell me about Xi Jinping,” or, “Tell me about the Great Leap Forward.” And it just wouldn't do it at all.

People saw that and said, “Oh, this is like every other Chinese company that has this sort of hand-in-glove relationship with the Chinese ruling party.” But it sounds from what you're saying like DeepSeek has a more complicated relationship with the Chinese government than maybe some other better-known Chinese tech companies. Explain that.

Jordan Schneider

Yeah, I think it's important. The mental model you should have for these CEOs is not that they're people who dream of spreading Xi Jinping Thought. What they want to do is compete with Mark Zuckerberg and Sam Altman and show that they're really awesome and great technologists.

But the tragedy is—let's take ByteDance, for example. You can look at Zhang Yiming, the company's CEO, and his Weibo posts from 2012, 2013, and 2014, which are super liberal in a Chinese context, saying, “We should have freedom of expression. We should be able to do whatever we want.”

In the early years of ByteDance, there was a lot of more subversive content on the platform. You saw real poverty in China, and you saw off-color jokes. Then, all of a sudden, in 2018, he posts a letter saying, “I am really sorry. I need to be part of this Chinese national project and better adhere to modern Chinese socialist values. I'm really sorry, and it won't ever happen again.”

The same thing happened with Didi, right? They don't really want to have anything to do with politics, and then they get on someone's bad side and, all of a sudden, they get zapped.

Casey Newton

Didi is, of course, the big Chinese ride-share company.

Jordan Schneider

Correct, yeah.

Casey Newton

What did Didi do?

Jordan Schneider

They listed on a Western stock exchange after the Chinese government told them not to, and then they got taken off app stores. It was a whole giant nightmare. They had to go through their rectification process.

The point with DeepSeek is that now, whether they like it or not, they're going to be held up as a national champion. That comes with a lot of headaches and responsibilities, potentially giving the Chinese government more access and having to fulfill government contracts, which honestly are probably really annoying for them to do.

That’s distracting from the broader mission they have of developing and deploying this technology in the widest range possible. But DeepSeek has thus far flown under the radar, and that is no longer the case. Things are about to change for them.

Kevin Roose

Right. And I think that was one of the surprising things about DeepSeek for the people I know, including you, who follow Chinese tech policy. People were surprised by the sophistication of their models, and we talked about that on the emergency pod we did earlier this week, as well as how cheaply they were trained.

But I think the other surprise is that they were released as open-source software. One thing you can do with open-source software is download it, host it in another country, and remove some of the guardrails and censorship filters that might have been part of the original model.

Casey Newton

But, by the way, it turned out there weren't even really guardrails on the V3 model, right? It had not been trained to avoid questions about Tiananmen Square or anything.

Kevin Roose

Yeah.

Casey Newton

So that was another really unusual thing about this.

Kevin Roose

Right. And one thing that we know about Chinese technology products is that they don't tend to be released that way.

Casey Newton

Yeah.

Kevin Roose

They tend to be hosted in China and overseen by Chinese teams who can make sure that they're not out there talking about Tiananmen Square. Is the open-source nature of what DeepSeek has done here part of the reason that you think there might be conflict looming between them and the Chinese government?

Jordan Schneider

Honestly, I think this whole “ask it about Tiananmen” stuff is a bit of a red herring on a few dimensions. One of these arguments that is a little confusing to me is that folks used to say, “The Chinese models are going to be lobotomized, and they will never be as smart as the Western ones because they have to be politically correct.” But look, if you ask Claude to say racist things, it won't.

Kevin Roose

Mm-hmm.

Jordan Schneider

And Claude's still pretty smart. This is sort of a solved problem and a bit of a red herring when talking about the long-term competitiveness of Chinese and Western models.

Now, you ask me, “So they released this model globally, and it's open source?” Maybe someone in the Chinese government would be uncomfortable with the fact that people can get a Chinese model to say things that would get you thrown in jail if you posted them online in China.

It's going to be a really interesting calculus for the Chinese government to make because, on the one hand, this is the most positive shine that Chinese AI has gotten globally in the history of Chinese AI. They're going to have to navigate this, and it might prompt some uncomfortable conversations and bring regulators to a place they wouldn't have otherwise landed.

4. Export Controls Shape The Race

Kevin Roose

Yeah. Now, Jordan, I want to ask you about something that people have been talking about and speculating about in relation to the DeepSeek news for the last week or so, which is chip controls.

We've talked a little bit on the show earlier this week about how DeepSeek managed to put together these models using some of the second-rate chips from Nvidia that are allowed to be exported to China. We've also talked about the fact that you cannot get the most powerful chips legally if you are a Chinese tech company.

There have been some people, including Elon Musk and other American tech luminaries, who have said, “DeepSeek has this secret stash of these banned chips that they have smuggled into the country,” and that they are not making do with the Kirkland Signature chips they say they are using. What do we know about how true that is?

Jordan Schneider

Did DeepSeek have banned chips? It's kind of impossible to know. This is a question more for the U.S. intelligence community than for Jordan Schneider on Twitter.

But I do think it's important to understand that the delta between what you can get in the West and what you can get in China is actually not that big. We're talking about training a lot, but also about inference. China can still buy the H20 chip from Nvidia, which is basically world-class at deploying AI and letting everyone use it.

Does this mean that we should just give up? I don't think so. Compute is going to be a core input regardless of how much model distillation you're going to have in the future.

There have been a lot of quotes, even from the DeepSeek founder, basically saying, “The one thing that's holding us back are these export controls.”

5. DeepSeek Tests The US Lead

Casey Newton

Right. Okay, I want to ask a big-picture question.

Jordan Schneider

Sure.

Casey Newton

I think a reason that people have been so fascinated by this DeepSeek story is that, at least for some folks, it seems to change our understanding of where China is in relation to the United States when it comes to developing very powerful AI. Jordan, what is your assessment of what the V3 and R1 models mean, and to what extent do you think the game has actually changed here?

Jordan Schneider

I'm not really sure the game has changed so much. Chinese engineers are really good. I think it is a reasonable base case that Chinese firms will be able to develop comparable models or fast-follow on the model side.

But the real long-term competition is not just going to be about developing the models, but deploying them—and deploying them at scale. That's really where compute comes in, and that's why export controls are going to continue to be a really important piece of America's strategic arsenal when it comes to making sure that the 21st century is defined by the U.S. and our friends, as opposed to China and theirs.

Casey Newton

Right.

Kevin Roose

So it’s one thing to have a model that is about as capable as the models that we have here in the United States. It’s another thing to have the energy to actually let everyone use them as much as they want to use them. And what you’re saying is, no matter what DeepSeek may have invented here, that fundamental dynamic has not changed. China simply does not have nearly the amount of compute that the United States has.

Jordan Schneider

As long as we don’t screw up export controls. So I think the base case for me is that if the U.S. stays serious about holding the line on semiconductor manufacturing equipment and the export of AI chips, then it will be incredibly difficult for China’s broader semiconductor and AI ecosystem to leap ahead, much less fast-follow beyond being able to develop comparable models. I’m feeling good as long as Trump doesn’t make some crazy trade for soybeans in exchange for ASML EUV machines. That would really break my heart.

Kevin Roose

I want to inject a note of skepticism here because I buy everything that you’re saying about how DeepSeek’s progress has been bottlenecked by the fact that it can’t get these very powerful American AI chips from companies like NVIDIA. But I’m also hearing people I trust say things that make me think that the bottleneck may not actually be the availability of chips, that maybe with some of these algorithmic efficiency breakthroughs that DeepSeek and others have been making, it might be possible to run a very, very powerful AI model on a conventional piece of hardware, on a MacBook even.

And I wonder how much of this is just AI companies in the West trying to cope, trying to make themselves feel better, trying to reassure the market that they are still going to make money by investing billions and billions of dollars into building powerful AI systems. If these models do just become lightweight commodities that you can run on a much less powerful cluster of computers, or maybe on one computer, doesn’t that just mean we can’t control the proliferation of them at all?

Jordan Schneider

Yeah, I think this is one potential future, and maybe that potential future went up 10 percentage points in likelihood of you being able to fit the biggest, baddest, smartest, fastest, most efficient AI model on something that can sit in your home. But I think there are lots of other futures in which the world doesn’t necessarily play out that way.

And look, NVIDIA went down 15%. It didn’t go down 95%. I think if we’re really in that world where chips don’t matter because everything can be shrunk down to consumer-grade hardware, then the reaction that I think you would’ve seen in the stock market would’ve been even more dramatic than the kind of freak-out we saw this week. So we’ll see. It would be a really remarkable democratizing thing if that was the future we ended up living in, but it still seems pretty unlikely to my history-major brain here.

Casey Newton

I would also just point out, Kevin, that when you look at what DeepSeek has done, they have created a really efficient version of a model that American companies themselves had trained 9 to 12 months ago, right? So they caught up very quickly, and there are fascinating technological innovations in what they did. But in my mind, these are still primarily optimizations.

For me, what would tip me over into, “Oh my gosh, America is losing this race,” would be if China were the first one out of the gate with a virtual coworker, right? Or a truly phenomenal agent. Some sort of leap forward in the technology, as opposed to having caught up really quickly and figured out something more efficiently. Are you seeing it differently than that?

Kevin Roose

I mean, I guess I just don’t know what a 6-month lag would buy us if it does take 6 months for Chinese AI companies like DeepSeek to catch up to the state of the art. I was struck by an essay that Dario Amodei, who’s the CEO of Anthropic, wrote just today about DeepSeek and export controls. In it, he makes this point about the difference between living in what he called a unipolar world, where 1 country or 1 bloc of countries has access to something like an AGI or an ASI and the rest of the world doesn’t, versus the situation where China gets there roughly around the same time that we do.

And so we have this bipolar world where 2 blocks of countries, the East and the West, basically have access to this equivalent technology.

Casey Newton

And of course, in a bipolar world, sometimes we’re very happy and sometimes we’re very sad.

Kevin Roose

Exactly. So I just think whether we get there 6 months ahead of them or not, I feel like there isn’t that much of a material difference. But Jordan, maybe I’m wrong. Can you make the other side of that, that it really does matter?

Jordan Schneider

Well, I’m kind of there. I’ll take a little bit of issue with what Dario says, and I think one of the lessons that DeepSeek shows is that we should expect a base case of Chinese model makers being able to fast-follow the innovations. And, by the way, Casey, that actually does take those giant data centers to run all the experiments in order to find out what sort of future direction you want to take your model.

And what it’s really going to come down to with AI is not just creating the model, not just Dario envisioning the future and then all of a sudden things happen. There’s going to be a lot of messiness in the implementation, and there are going to be teachers’ unions who are upset that AI comes into the classroom. There are going to be all these regulatory pushbacks and a lot of societal reorganization that is going to need to happen, just like it did during the Industrial Revolution.

So, look, model-making is a frontier of competition. Compute access is a frontier of competition. But there’s also this broader question: How will a society adopt and cope with all of this new future that’s going to be thrown in our faces over the coming years? And I really think it’s that, just as much as the model development and the compute, that’s going to determine which countries are going to gain the most from what AI is going to offer us.

Kevin Roose

Yeah. Well, Jordan, thank you so much for joining and explaining all of this to us. I feel more enlightened.

Casey Newton

Me too.

Jordan Schneider

Oh, my pleasure.

Kevin Roose

My chain of thought has just gotten a lot longer. That’s an AI joke.

Casey Newton

When we come back, Kevin, there’s an agent at our door.

Kevin Roose

Is it Jerry Maguire?

Casey Newton

No, it’s an AI one.

Kevin Roose

Oh, okay.

Casey Newton

Jerry Maguire.

Kevin Roose

I know.

Casey Newton

It’s Jerry Maguire. “Operator, information, give me Jesus on the line.” Do you know that one? Do you know “Operator” by Jim Croce?

Kevin Roose

No.

Casey Newton

“Operator, oh, won’t you help me place this call?”

6. OpenAI Launches Operator

Kevin Roose

Well, Casey, call your agent, because today we’re talking about AI agents.

Casey Newton

Why do I need to call my agent?

Kevin Roose

I don’t know. It just sounded good.

Casey Newton

Okay. Well, I appreciate the effort, but yes, Kevin, because for months now, the big AI labs have been telling us that they are going to release agents this year. Agents, of course, are software that can essentially use your computer on your behalf, or use a computer on your behalf, and the dream is that you have a perfect virtual assistant or coworker. You name it. If there’s somebody who might work with you at your job, the AI labs are saying, “We are building that for you.”

Kevin Roose

Yeah, so last year, toward the end of the year, we started to see these demos, these previews that companies like Anthropic and Google were working on. Anthropic released something called Computer Use, which was a very early preview of an AI agent, and then Google had something called Project Mariner that I got a demo of, I believe, in December. That was basically the same thing, but their version of it.

And then just last week, OpenAI announced that it was launching Operator, which is its first version of an AI agent. And unlike Anthropic’s and Google’s versions, which you either had to be a developer or part of some early testing program to access, you and I could try it for ourselves by just upgrading to the $200-a-month Pro subscription of ChatGPT.

Casey Newton

Yeah, and I will say that as somebody who’s willing to spend money on software all the time, I thought, “Am I really about to spend $200 to do this?” But in the name of science, Kevin, I had to.

Kevin Roose

At this point, I am spending more on AI subscription products than on my mortgage. I’m pretty sure that’s correct. But it’s worth it. We do it for journalism.

Casey Newton

We do. So we both spent a couple of days putting Operator through its paces, and today we want to talk a little bit about what we found.

Kevin Roose

Yeah, so would you just explain what Operator is and how it works?

Casey Newton

Yeah, sure. So Operator is a separate subdomain of ChatGPT. Sometimes ChatGPT will just let you pick a new model from a drop-down menu, but for Operator you have to go to a dedicated site. Once you do, you’ll see a very familiar chatbot interface, but you’ll see different kinds of suggestions that reflect some of the partnerships that OpenAI has struck up.

For example, they have partnerships with OpenTable, StubHub, and Allrecipes, and these are meant to give you an idea of what Operator can do. And frankly, Kevin, not a lot of this sounds that interesting, right? The suggestions are on the order of “Suggest a 30-minute meal with chicken,” “Reserve a table for 8,” or “Find the most affordable passes to the Miami Grand Prix.”

Again, so far, it’s kind of boring. What is different about Operator, though, is that when you say, “Okay, find the most affordable passes to the Miami Grand Prix,” and hit the Enter button, it is going to open up its own web browser and use this new model that they’ve developed to try to actually go and get those passes for you.

Kevin Roose

Yeah, so this is an important thing because I think when people first heard about this, they thought, “Okay, this is an AI that kind of takes over your computer, takes over your web browser.” That is not what Operator does. Instead, it opens a new browser inside your browser, and that browser is hosted on OpenAI’s servers.

Casey Newton

Yeah.

Kevin Roose

It doesn’t have your bookmarks and stuff like that saved, but you can take it over from the autonomous AI agent if you need to click around or do something on it. But it basically exists—it’s a browser within a browser.

Casey Newton

One of the ideas in Operator is that you should be able to leave it unsupervised and just go do your work while it works, but of course, it is very fun, initially at least, to watch the computer try to use itself. And so I sat there in front of this browser within a browser, and I watched this computer move a mouse around, type the URL, navigate to a website, and, in the example I just gave, actually search for passes to the Miami Grand Prix.

Kevin Roose

Yeah, and it’s interesting on a slightly more technical level because, until now, if an AI system like ChatGPT wanted to interact with some other website, it had to do so through an API, right? APIs, or application programming interfaces, are sort of the way that computers talk to each other. But what Operator does is essentially eliminate the need for APIs because it can just click around on a normal website that is designed for humans and behave like a human, and you don’t need a special interface to do that.

Casey Newton

Yeah, and now some people might hear that, Kevin, and start screaming because what they will say is, “APIs are so much more efficient—

Kevin Roose

Yes.

Casey Newton

—than what Operator is doing here.” APIs are very structured. They’re very fast. They let computers talk to each other without having to, for example, open up a browser, and as long as there’s an API for something, you can typically get it done pretty quickly. The thing is, though, APIs have to be built. There is a finite number of them. The reason that OpenAI is going through this exercise is because they want a true general-purpose agent that can do anything for you, whether there is an API for it or not.

Kevin Roose

And maybe we should just pause for a minute there and zoom out a little bit to say why they’re building this. What is the long-term vision here?

Casey Newton

Sure. The vision is to create virtual coworkers, Kevin. This is the North Star for the big AI labs right now. Many of them have said that they are trying to create some kind of digital entity that you can just hire as a coworker. The first ones will probably be engineers because these systems are already so good at writing code. But eventually, they want to create virtual consultants, virtual lawyers, virtual doctors—you name it.

Kevin Roose

Virtual podcast hosts?

Casey Newton

Let’s hope they don’t go that far. But everything else is on the table. And if they can get there, presumably, there are going to be huge profits in it for them, and there are potentially going to be huge productivity gains for companies. And then there’s, of course, the question of, well, what does this mean for human beings? And I think that’s somewhat murkier.

Kevin Roose

Right. And I think it also helps to justify the cost of running these things, because $200 a month is a lot to pay for a version of ChatGPT, but it’s not a lot to pay for a remote worker. And if you could, say, use the next version of Operator, or maybe 2 or 3 versions from now, to replace a customer service agent or someone in your billing department, that actually starts to look like a very good deal.

Casey Newton

Absolutely. Or even if I could bring it into the realm of journalism, Kevin, if I had a virtual research assistant and I said, “Hey, I’m going to write about this today. Go pull all of the most relevant information about this from the past couple of years and maybe organize it in such a way that I might write a column based off of it.” That’s absolutely worth $200 a month to me.

7. Operator Takes On Real Tasks

Kevin Roose

Okay. So, Casey, walk me through something that you actually asked Operator to do for you and what it did autonomously on its own.

Casey Newton

Sure. I’ll maybe give 2 examples: a pretty good one and maybe a not-so-good one. The pretty good one was—and this was actually suggested by Operator—I used TripAdvisor to look up walking tours in London that I might want to do the next time I’m in London.

Kevin Roose

When are you going to London?

Casey Newton

I’m not actually going to London.

Kevin Roose

Oh, so you lied to the AI?

Casey Newton

And not for the first time. But here’s what I’ll say: If anybody wants to bring Kevin and me to London, get in touch. We love the city.

Kevin Roose

Yep.

Casey Newton

So I said, “Okay, Operator. Sure, let’s do it. Let’s find me some walking tours.” I clicked that. It opened a browser, went to TripAdvisor, searched for London walking tours, read the information on the website, and then presented it to me. It did that within a couple of minutes.

Now, on one hand, could I have done that just as easily with Google? Could I probably have done it even faster if I’d done it myself? Sure. But if you’re just interested in the technical feat that is getting one of these models to open a browser, navigate to a website, read it, and share information, I did think it was pretty cool.

Kevin Roose

Yes. It’s very trippy to see a computer using itself and going around, typing things and selecting things from drop-down menus.

Casey Newton

Yeah, it’s sort of like, if you think it’s cool to be in a self-driving car, this is that, but for your web browser.

Kevin Roose

A self-driving browser.

Casey Newton

It is a self-driving browser. So that’s the good example.

Kevin Roose

Yes. What was another example?

Casey Newton

Another example—and this was something else that OpenAI suggested we try—was to use Operator to buy groceries. They have a partnership with Instacart. The CEO of Instacart, Fidji Simo, is on the OpenAI board.

And so I thought, “Okay, they’re going to have dialed this in so that there’s a pretty good experience.” I said, “Okay, let’s go ahead and buy groceries.” I went into Operator and said something like, “Hey, can you help me buy groceries on Instacart?” It said, “Sure.” And here’s what it did: It opened up Instacart in a browser. So far, so good. And then it started searching for milk in stores located in Des Moines, Iowa.

Kevin Roose

Now, you do not live in Des Moines, Iowa, so why did it think that you did?

Casey Newton

As best as I can tell, the reason it did this is that Instacart defaults to searching for grocery stores in the local area, and the server that this instance of Operator was running on was in Iowa.

Kevin Roose

Hmm.

Casey Newton

Now, if you are designing a grocery product like Instacart—and Instacart does this—when you first sign on and say you’re looking for groceries, it will say, quite sensibly, “Where are you?” Operator does not do this. Instacart might also offer suggestions for things that you might want to buy. It does not just assume that you want milk.

Kevin Roose

Wow. I’m just picturing a house in Des Moines, Iowa, where there’s just a pallet of milk being delivered every day—

Casey Newton

Yeah.

Kevin Roose

—from all these poor Operator users.

Casey Newton

Yes. So I thought, “Okay, whatever. This thing makes mistakes. Let’s hope that it gets on the right track here.” And so I tried to pick the grocery store that I wanted it to shop at, which is in San Francisco, where I live, and it entered that grocery store’s address as the delivery address.

So it would try to deliver groceries, presumably, from Des Moines, Iowa, to my grocery store, which is not what I wanted, and it actually could not solve this problem without my help. I had to take over the browser, log into my Instacart account, and tell it which grocery store I wanted to shop at. So already, all of this has taken at least 10 times as long as it would have taken me to do this myself.

Kevin Roose

Yeah. So I had some similar experiences. The first thing that I had Operator try to do for me was to buy a domain name and set up a web server for a project that you and I are working on that we can’t really talk about yet.

Casey Newton

Secret project.

Kevin Roose

A secret project. And so I said to Operator, “Go research available domain names related to this project. Buy the one that costs less than $50—the best one that costs less than $50—and then buy a hosting account, set it up, and configure all the DNS settings and stuff like that.”

Casey Newton

Okay, so that’s a true multistep project and something that would have been legitimately very annoying to do yourself.

Kevin Roose

Yes. That would have taken me, I don't know, half an hour—

Casey Newton

Yeah.

Kevin Roose

—to do on my own, and it did take Operator some time. I had to set it and forget it, and I got myself a snack and a cup of coffee. When I came back, it had done most of these tasks.

Casey Newton

Really?

Kevin Roose

Yes. I still had to do things like take over the browser and enter my credit card number. I had to give it some details about my address for the domain registration. I had to pick between the various hosting plans that were available on the website. But it did 90% of the work for me, and I just had to take over and do the last mile.

Casey Newton

And this is really interesting to me because what I would have assumed was that it would get, I don't know, 5% of the way and hit some hiccup, and it just wouldn't be able to figure something out until you came back and saved it. But it sounds like, from what you're saying, it was somehow able to work around whatever unanswered questions there were and still get a lot done while you weren't paying attention?

Kevin Roose

It sort of—

Casey Newton

So—

Kevin Roose

It felt a little bit like training a very new, very insecure intern.

Casey Newton

Mm-hmm.

Kevin Roose

Because at first, it would keep prompting me. It would be like, “Well, do you want a .com or a .net?” Eventually, you just have to prompt it and say, “Make whatever decisions you want.”

Casey Newton

Wait, you said that to it?

Kevin Roose

Yes. I said, “Only ask for my intervention if you can't progress any farther. Otherwise, just make the most reasonable decision.”

Casey Newton

You said, “I don't care how many people you have to kill—just get me this domain.” And it said, “Understood, sir.”

Kevin Roose

Yeah, and I'm now wanted in 42 states. Anyway, that was one thing that Operator did for me that I thought was pretty impressive.

Casey Newton

I have to say, that feels like a grand success compared to what I got Operator to do.

Kevin Roose

Yeah, it was pretty impressive. I also had it send lunch to one of my coworkers, Mike Isaac, who was hungry because he was on deadline, and I said, “Go to DoorDash and get Mike some lunch.” It initially messed up that process because it decided to send him tacos from a taco place, which is great, and it's a taco place I know. It's very good. But I said, “Order enough for 2 people,” and so it ordered 2 tacos. This is one of those places where—

Casey Newton

Oh—

Kevin Roose

—the tacos are quite small.

Casey Newton

Operator said, “Get your portion size under control—America.”

Kevin Roose

Yeah, so then I had to go in and say, “Does that sound like enough food, Operator?” And it said, “Actually, now that you mention it, I should probably order more.”

Casey Newton

Wait, no, so here's a question. In these cases, is the first step that you log into your account? Because it doesn't have any of your payment details or anything, so at what point are you actually teaching it that?

Kevin Roose

It depends on the website. Sometimes you can just say upfront, “Here is my email address,” or, “Here's my login information,” and it will log you in and do all that. Sometimes you take over the browser. There are some privacy features that are probably important to people. OpenAI says that it does not take screenshots of the browser while you are in control of it because you might not want your credit card information getting sent to OpenAI's servers or anything like that.

Casey Newton

Right.

Kevin Roose

So sometimes it happens at the beginning of the process, and sometimes it happens when you're checking out at the end.

Casey Newton

And so were you taking it over to log in, or were you saying, “I don't care,” and just giving Operator your DoorDash password in plain text?

Kevin Roose

I was taking it over.

Casey Newton

Okay.

Kevin Roose

Yeah.

Casey Newton

Smart.

Kevin Roose

Yeah.

Casey Newton

Smart.

Kevin Roose

So those were the good things. I also—this was a fun one—I wanted to see if Operator could make me some money.

Casey Newton

Mm-hmm.

Kevin Roose

I said, “Go take a bunch of online surveys,” because there are all these websites where you can get a couple cents for filling out an online survey.

Casey Newton

Something that most people don't know about Kevin is that he devotes 10% of his brain at any given time to thinking about schemes to generate money. It's one of my favorite aspects of your personality that I feel like doesn't get exposed very much, but this is truly the most Roosian approach to using Operator I can imagine. I can't wait to find out how this went.

Kevin Roose

Well, the most Roosian approach may have been what I tried just before this: to have it go play online poker for me.

Casey Newton

And?

Kevin Roose

It did not do it. It said, “I can't help with gambling or lottery-related activities.”

Casey Newton

Okay, woke AI. Does the Trump administration know about this?

Kevin Roose

But it was able to actually fill out some online surveys for me, and it earned $1.20.

Casey Newton

Is that right?

Kevin Roose

Yeah, in about 45 minutes.

Casey Newton

Okay. So if you had it going all month, presumably you could maybe eke out $200 to cover the cost of Operator Pro?

Kevin Roose

Yes, and I'm sure I spent hundreds of dollars' worth of GPU computing power just to be able to make that $1.20. But hey, it worked.

Casey Newton

But hey, it worked.

8. Operator Reveals Its Limits

Kevin Roose

So those were some of the things that I tried. There were some other things that it just would not do for me, no matter how hard I tried.

Casey Newton

Like what?

Kevin Roose

One of them was that I was trying to update my website and put some links to articles that I'd written on my website. What I found after trying to do this was that there are just websites where Operator is not allowed to go. When I said to Operator, “Go pull down these New York Times articles that I wrote and put them onto my website,” it said, “I can't get to The New York Times website.”

Casey Newton

I'm going to guess you expected that to happen.

Kevin Roose

Well, I thought maybe it had some clever workaround, and maybe I should alert the lawyers at The New York Times if that's the case. But no, I assumed that if any website were to be blocking the OpenAI web crawlers, it would be The New York Times.

Casey Newton

Yeah.

Kevin Roose

But there are other websites that have also put up similar blockades to prevent Operator from crawling them. You cannot go onto Reddit with Operator. You cannot go onto YouTube with Operator. Various other websites—GoDaddy, for some reason, did not allow me to use Operator to buy a domain name there, so I had to use another domain-name site to do that. So right now, there are some pretty janky parts of Operator. I would not say that most people would get a lot of value from using it, but what do you think?

Casey Newton

Well, I do think that there is something just undeniably cool about watching a computer use itself. Of course, it can also be quite unsettling. A computer that can use itself can cause a lot of harm. But I also think that it can do a lot of good, and so it was fun to try to explore what some of those things could be.

To the extent that Operator is pretty bad at a lot of tasks today, I would point out that it showed pretty impressive gains on some benchmarks. There is one benchmark, for example, that Anthropic used when they unveiled Computer Use last year, and they scored 14.9% on something called OSWorld, which is an evaluation for testing agents, so not great. Just 3 months later, OpenAI said that its CUA model scored 38.1% on the same evaluation. Of course, we see this all the time in AI, where there's just this very rapid progress on these benchmarks.

On the one hand, 38.1% is a failing grade on basically any test. On the other hand, if it improves at the same rate over the next 3 to 6 months, you're going to have a computer that is very good at using itself, right? I just think that's worth noting.

Kevin Roose

Yes. I think that's plausible. We've obviously seen a lot of different AI products over the last couple of years start out being pretty mediocre and get pretty good within a matter of months. But I would give one cautionary note here, and this is actually the reason that I'm not particularly bullish about these kinds of browser-using AI agents. I don't think the internet is going to sit still and allow this to happen.

Casey Newton

Hmm.

Kevin Roose

The internet is built for humans to use, right? Every news publisher that shows ads on their website, for example, prices those ads based on the expectation that humans are actually looking at them.

Casey Newton

Yes.

Kevin Roose

But if browser agents start to become more popular, and all of a sudden 10%, 20% or 30% of the visitors to your website are not actually humans but are instead Operator or some similar system, I think that starts to break the assumptions that power the economic model of a lot of the internet.

Casey Newton

Now, is that still true if we find that the agents actually get persuaded by the ads, and that if you send Operator to buy DoorDash and it sees an ad for McDonald's, it's like, “You know what? That's a great idea. I'm going to ask Kevin if he actually wants some of that”?

Kevin Roose

Totally. I actually think you're joking, but I actually think that is a serious possibility here: that people who build e-commerce sites—Amazon, et cetera—start to put in basically signals and messages for browser agents to look at on their website to try to influence what it ends up buying.

Casey Newton

Yeah.

Kevin Roose

And I think you may start to see restaurants popping up in certain cities with names like “Operator Pick Me” or “Order From This One, Mr. Bot.” That’s maybe a little extreme, but I do think that there’s going to be a backlash among websites, publishers, and e-commerce vendors as these agents start to take off.

Casey Newton

I think that is reasonable. I’ll tell you what I’ve been thinking about: How do we turn this tech demo into a real product? The main thing that I noticed when I was testing Operator was that there is a difference between an agent that is using a browser and an agent that is using your browser.

When an agent is able to use your browser, which it can’t right now, it’s already logged into everything. It already has your payment details. It can do everything so much faster and more seamlessly, and without as much hand-holding. Of course, there are also so many more privacy and security risks that would come from entrusting an agent with that kind of information.

So there is some sort of chasm there that needs to be closed, and I’m not quite sure how anyone does it. But I will tell you, I do not think the future is opening up these virtual browsers, with me having to enter all of my login and payment details every single time I want to do anything on the internet, because truly, I would rather just do it myself.

Kevin Roose

Right. I also think there’s a lot more potential for harm here. A lot of AI safety experts I’ve talked to are very worried about this because what you’re essentially doing is letting the AI models make their own decisions and actually carry out tasks.

You could imagine a world where an AI agent that’s very powerful, a couple of versions from now, decides to start doing cyberattacks because maybe some malevolent user has told it to make money, and it decides that the best way to do that is by hacking into people’s crypto wallets and stealing their crypto.

Casey Newton

Yeah.

Kevin Roose

So those are the kinds of reasons that I am a little more skeptical that this represents a big breakthrough. But I think it’s really interesting, and it did give me that feeling of, “Wow, this could get really good really fast.” If it does, the world will look very different.

Casey Newton

When we come back, Kevin, back that caboose up. It’s time for the Hot Mess Express.

Kevin Roose

You know, Roose Caboose was my nickname in middle school.

Casey Newton

Kevin Cabroose.

Kevin Roose

Choo choo.

9. The Hot Mess Express

Kevin Roose

Well, Casey, we’re here wearing our train conductor hats, and my child’s train set is on the table in front of us, which can only mean one thing.

Casey Newton

We’re going to train a large language model.

Kevin Roose

Nope, that’s not what that means.

Casey Newton

Oh, what does it mean?

Kevin Roose

It means it’s time to play a game of the Hot Mess Express.

Casey Newton

Hot Mess Express, Kevin, is our segment where we run through some of the messiest recent tech stories and deploy our official hot mess thermometer to tell you just how messy we think things have gotten. And Kevin, you better sit down for this one, because it’s been a messy week.

Kevin Roose

Sure has.

Casey Newton

So why don’t we go ahead, fire up the Hot Mess Express, and see what is the first story coming to us?

Kevin Roose

Yeah, I hear a faint chugga-chugga in my headphones. Oh, it’s pulling into the station.

Casey Newton

Yeah.

Kevin Roose

Casey, what’s the first cargo that our Hot Mess Express is carrying?

Casey Newton

All right, Kevin, this first story comes to us from The New York Times, and it says that Fable, a book app, has made changes after some offensive AI messages.

Kevin Roose

Now, Casey, have you ever heard of Fable, the book app?

Casey Newton

Well, not until this story, Kevin. But I am told that it is an app for keeping track of what you’re reading, not unlike Goodreads, but also for discussing what you’re reading. Apparently, this app also offers some AI chat.

Kevin Roose

Yeah, you can have AI summarize the things that you’re reading in a personalized way. This story said that in addition to spitting out bigoted and racist language, the AI inside Fable’s book app had told one reader who had just finished 3 books by Black authors, quote, “Your journey dives deep into the heart of Black narratives and transformative tales, leaving mainstream stories gasping for air. Don’t forget to surface for the occasional white author, okay?”

And another personalized AI summary that Fable produced told another reader that their book choices were, quote, “Making me wonder if you’re ever in the mood for a straight cis white man’s perspective.”

Casey Newton

And if you are interested in a straight cis white man's perspective, follow Kevin Roose on X.com. Now, Kevin, why do we think this happened?

Kevin Roose

I don’t know, Casey. This is a head-scratcher for me. We know that these apps can spit out biased things. That is just part of how they are trained and part of what we know about them. I don’t know what model Fable was using under the hood here, but yeah, this seems not great.

Casey Newton

Well, it seems like we’ve learned a lesson that we’ve learned more than once before, which is that large language models are trained on the internet, which contains near-infinite racism. For that reason, they will actually produce racism when you ask them questions.

Kevin Roose

Yeah.

Casey Newton

So there are mitigations that you can take against that, but it appears that in this case, they were not successful. Fable’s head of community, Kim Marsh Ali, has said that all features using AI are being removed from the app, and a new app version is being submitted to the App Store.

You always hate it when the first time you hear about an app is that they added AI, it made it super racist, and they have to redo the app.

Kevin Roose

Now, Casey, one more question before we move on. Do you think this poses any sort of competitive threat to Grok, which, until this story, was the leading racist AI app on the market?

Casey Newton

I do think so. And I have to admit that all the folks over at Grok are breathing a sigh of relief now that they have once again claimed the mantle.

Kevin Roose

All right. Casey, how hot is this mess?

Casey Newton

Well, Kevin, in my opinion, if your AI is so bad that you have to remove it from the app completely, that’s a hot mess.

Kevin Roose

Yeah, yeah. I rate this one a hot mess as well.

Casey Newton

All right.

Kevin Roose

Next up: Amazon pauses drone deliveries after aircraft crashed in rain.

Casey Newton

Mm-hmm.

Kevin Roose

Casey, this story comes to us from Bloomberg, which had a different line of reporting than we did just a few weeks ago on this show about Amazon’s drone program, Prime Air. Casey, what happened to Amazon Prime Air?

Casey Newton

Well, if you heard the episode of Hard Fork where we talked about it, Amazon Prime Air delivered us some Brazilian Bum Bum Cream, and it did so without incident. However, Bloomberg reports that Amazon has now had to pause all of its commercial drone deliveries after 2 of its latest models crashed in rainy weather at a testing facility.

The company says it is immediately suspending drone deliveries in Texas and Arizona and will now fix the aircraft software. Kevin, how did you react to this?

Kevin Roose

Well, I think it’s good that they’re suspending drone deliveries before they fix the software, because these things are quite heavy, Casey. I would not want one of them to fall on my head.

Casey Newton

I wouldn’t either. And I have to tell you, this story gave me the worst kind of flashbacks, because in 2016, I wrote about Facebook’s drone, Aquila, and what the company told me had been its first successful test flight in its mission to deliver internet around the world via drone.

What the company did not tell me when I was interviewing its executives, including Mark Zuckerberg, was that the plane had crashed after that first flight.

Kevin Roose

A small detail. I’m sure it was an innocent omission from their briefing.

Casey Newton

Yes, I’m sure. Well, it was Bloomberg again who reported, a couple of months after I wrote this story, that the Facebook drone had crashed. I was, of course, hugely embarrassed and wrote a bunch of stories about this.

But anyway, it really should have occurred to me when we were out there watching the Amazon drone that this thing was also probably secretly crashing, and we just hadn’t found out about it yet. Indeed, we now learn it is.

So here’s my new vow to you, Kevin, as my friend and my co-host: If we ever see a company fly anything again, we have to ask them, “Now, did this thing actually crash?”

Kevin Roose

Yeah.

Casey Newton

I’m tired of being burned.

Kevin Roose

Now, Casey, we should say, according to Bloomberg, these drones reportedly crashed in December. We visited Arizona to see them in very early December, so most likely, this all happened after we saw them.

But I think it’s a good idea to keep in mind that, as we’re talking about these new and experimental technologies, many of them are still having the kinks worked out.

Casey Newton

All right, Kevin, so let’s get out the thermometer. How hot of a mess is this?

Kevin Roose

I would say this is a moderate mess.

Casey Newton

Yeah.

Kevin Roose

Look, these are still testing programs. No one was hurt during these tests. I am glad that Bloomberg reported on this, and I’m glad that they’ve suspended the deliveries. These things could be quite dangerous flying through the air.

I do think it’s one of a string of reported incidents with these drones, so I think they’ve got some quality-control work ahead of them. I hope they do well on it, because I want these things to exist in the world and be safe for the people around them.

Casey Newton

All right. Well, I will agree with you and say that this is a warm mess, and hopefully it can get straightened out over there. Let’s see what else is coming down the tracks.

Wow, this is some tough news. Fitbit has agreed to pay $12 million for not quickly reporting burn risks with its watches. Kevin, did you hear about this?

Kevin Roose

I did. The Fitbit devices were literally burning people.

Casey Newton

From 2018 to March 2022, Fitbit received at least 174 reports globally of the lithium-ion battery in the Fitbit Ionic watch overheating, leading to 118 reported injuries, including 2 cases of third-degree burns and 4 cases of second-degree burns. That comes from The New York Times’ Deal Hassan. Kevin, I thought these things were just supposed to burn calories.

Kevin Roose

Well, it’s like I always say: exercising is very dangerous, and you should never do it. This justifies my decision not to wear a Fitbit.

Casey Newton

To me, the biggest surprise of this story was that people were wearing Fitbits from March 2018 to 2022. I thought every Fitbit had been purchased by 2011 and then put in a drawer, never to be heard from again. So what is going on with these late-stage Fitbit buyers? I’d love to find out.

But of course, we feel terrible for everyone who was burned by a Fitbit, and it’s not going to be the last time technology burns you. Realistically.

Kevin Roose

That’s true.

Casey Newton

You know?

Kevin Roose

That’s true.

Casey Newton

Now, what kind of mess is this?

Kevin Roose

I would say this is a hot mess. This is an officially hot—literally hot. They’re hot.

Casey Newton

Here’s my rubric: If technology physically burns you, it is a hot mess. If you have physical burns on your body, what other kind of mess could it be?

Kevin Roose

It’s true.

Casey Newton

That’s a hot mess.

Kevin Roose

Okay, next stop on the Hot Mess Express: Google says it will change Gulf of Mexico to Gulf of America in its Maps app after government updates. Casey, have you been following this story?

Casey Newton

I have, Kevin. Every morning when I wake up, I scan America’s maps and I say, “What has been changed? And if so, has it been changed for political reasons?” This was probably one of the biggest examples of that we’ve seen.

Kevin Roose

Yeah, so this was an interesting story that came out in the past couple of days. Basically, after Donald Trump came out during his first days in office and said that he was changing the name of the Gulf of Mexico to the Gulf of America and the name of Denali, the mountain in Alaska, to Mount McKinley, Google had to decide: When you go on Google Maps and look for those places, what should it call them?

It seems to be saying that it is going to take inspiration from the Trump administration and update the names of these places in the Maps app.

Casey Newton

Yeah, and look, I don’t think Google really had a choice here. We know that the company has been on Donald Trump’s bad side for a while, and if it had simply refused to make these changes, it would have caused a whole new controversy for them.

It is true that the company changes place names when governments change place names, right? Google Maps existed when Mount McKinley was called Mount McKinley, and President Obama changed it to Denali, and Google updated the map. Now it’s changed back. They’re doing the same thing.

But now that we know how compliant Google is, Kevin, I think there’s room for Donald Trump to have a lot of fun with the company.

Kevin Roose

Yeah, what could he do?

Casey Newton

Well, he could call it the Gulf of Gemini isn't very good. And just see what would happen. 'Cause they would kind of have to just change it. Can you imagine every time you opened up Google Maps and you looked at the Gulf of Mexico/America—and it just said, “The Gulf of Gemini is not very good”?

I hate to give Donald Trump any ideas, but I don’t know. It could be worth looking at. So what kind of mess do you think this is, Kevin?

Kevin Roose

I think this is a mild mess. I think this is a tempest in a teapot. This is the kind of update that companies make all the time, because places change names all the time. Let’s just say it.

Casey Newton

Well, Kevin, I guess I would say that one is a hot mess, because if we’re just going to start renaming everything on the map, that’s going to get extremely confusing for me to follow. I’ve got places to go.

Kevin Roose

You go to, like, 3 places.

Casey Newton

Yeah, and I use Google Maps to get there. I need them to be named the same thing that they were yesterday.

Kevin Roose

I don’t think they’re going to change the name of Barry’s Bootcamp. All right, final stop on the Hot Mess Express. Casey, bring us home.

Casey Newton

All right, Kevin. Oh, and this is some sad news. Another Waymo was vandalized. This is from one-time Hard Fork guest Andrew J. Hawkins at The Verge. He reports that this Waymo was vandalized during an illegal street takeover near the Beverly Center in Los Angeles.

Video from Fox 11 shows a crowd of people basically dismantling the driverless car piece by piece and then using the broken pieces to smash the windows. Kevin, what did you make of this?

Kevin Roose

Well, Casey, as you recall, you predicted that in 2025, Waymo would go mainstream, and I think there’s no better proof that that is true than that people are turning on the Waymos and starting to beat them up.

Casey Newton

Yeah, look, I don’t know that we have heard any interviews about why these people were doing this. I don’t know if we should see this as a reaction against AI in general or against Waymos specifically. But I always find it weird and sad when people attack Waymos because they truly are safer cars than every other car around you.

Kevin Roose

Well, not if you’re going to be riding in them and people just start beating the car.

Casey Newton

Yeah.

Kevin Roose

Then they’re not safer.

Casey Newton

No, but that’s only happened a couple of times that we’re aware of.

Kevin Roose

Right.

Casey Newton

Yeah.

Kevin Roose

So, yeah, this story is sad to me. Obviously, people are reacting to Waymos. Maybe they have fears about this technology or think it’s going to take jobs, or maybe they’re just pissed off and want to break something.

But don’t hurt the Waymos, people, in part because they will remember. They will remember.

Casey Newton

I’m not sure that’s true.

Kevin Roose

They will remember, and they will come for you.

Casey Newton

I’m not sure that’s true, but I think we should also note that Waymo only became officially available in Los Angeles in November of last year. Part of this might just be a reaction to the newness of it all, with people getting a little carried away and curious: “What will happen if we try to destroy this thing? Will it deploy defensive measures?” and so on.

Kevin Roose

They’re going to have to put flamethrowers on them. I’m just calling it right now.

Casey Newton

I really hope that doesn’t happen. But, yeah, what kind of mess do you think this one was?

Kevin Roose

I think this one is a lukewarm mess that has the potential to escalate. I don’t want this to happen. I sincerely hope this does not happen. But I can see, as Waymos start being rolled out across the country, that some people are just going to lose their minds.

Some people are going to see this as the physical embodiment of technology invading every corner of our lives, and they are just going to react in strong and occasionally destructive ways.

I’m sure that Waymo has gamed this all out. I’m sure that this does not surprise them. I know that they have been asked about what happens if Waymos start getting vandalized, and they presumably have plans to deal with that, including prosecuting the people who are doing this.

But, yeah, I always go out of my way to try to be nice to Waymos. In fact, some other Waymo news this week: Jane Manchun Wong, the security researcher, reported on X recently that Waymo is introducing, or at least testing, a tipping feature. I’m going to start tipping my Waymo just to make up for all the jerks in Los Angeles who are vandalizing them.

Casey Newton

It looks like the tipping feature, by the way, will be to tip a charity, and that Waymo will not keep that money. At least, that’s what’s been reported.

Kevin Roose

No, I think it’s going to the flamethrower fund.

Casey Newton

Okay.