AI 安全强势回归 + 与 Nikesh Arora 共谈 Mythos 乱象 + Hot Mess Express
Claude Mythos 将危险的网络能力从假设变成现实,迫使华盛顿迅速掉头重做安全政策。 传闻中的行政命令将建立类似 Biden 时期的模型发布前审查机制,而 Trump 重返白宫首日就取消了这套机制,此前共和党曾批评相关测试反创新。Casey Newton 的结论是:本届政府的世界观“经不起现实的检验”。
政府仍缺乏连贯的模型开放策略,芯片、承包商、中国和盟友网络安全都面临政策风险。 五角大楼一边推动将 Anthropic 认定为供应链风险,一边安装 Mythos 扫描漏洞;Trump 正在研究向中国开放 Nvidia 芯片,但政府尚未决定是否向中国开放 Mythos。结果是,本届政府“同时在安装和卸载 Anthropic”。
网络防御的时间尺度已从以天计的响应问题,变成以分钟计的基础设施竞赛。 Palo Alto Networks 发现26个关键漏洞利用,涉及75项问题,而通常基线不到5个;Mozilla 4月发布了423项修复,远高于2025年每月约22项的平均水平。Nikesh Arora 表示,传统防御“是按天设计的”,企业必须彻底改造系统,做到“用 AI 对抗 AI”。
Mythos 的威力来自持续算力对多个弱点的串联,但它既不是自动运行的,也并非不会出错。 Arora 称其误报率约30%,只有在 Palo Alto 提供代码用途、预期行为,以及过去5年1万次攻击产生的威胁情报后,模型表现才有所提升。Mythos 与 GPT-5.5 Cyber 找到的问题并不相同,而 Mythos 高度消耗算力的“ultra mode”则让持续试错和“漏洞串联”更加有效。
短期经济账更偏向攻击者,而大规模修复周期需要有规模的网络安全厂商和系统集成商参与。 防守方必须100%正确,攻击者只需要找到1个可用漏洞;防火墙可以提供“临时脚手架”,但开源依赖和无人管理的终端仍然难以及时修补。Arora 预计,企业将经历一轮持续3-6个月的漏洞积压“清理”,IBM、PwC、Deloitte 和 Accenture 等公司将参与其中。
最暴露的,是业务依赖技术、但核心能力在其他领域的企业。 相比资源充足的金融机构,Arora 更担心医院、小企业、工业运营商和诊所——这些“95%在做别的事”的公司往往没有工程师。他以 Change Healthcare 遭遇的攻击为例,说明一次事故如何让整个医生生态停摆。消费电子邮件和电信服务商也需要加强前置拦截,因为 AI 很快会让钓鱼攻击变得明显更逼真。
AI agents 恰恰因为足够有用、能够持有凭证并自主行动,扩大了攻击面。 Arora 称 OpenClaw“从安全角度看很可怕”,因为它可以被授予跨账户行动所需的权限和凭证;他把 OpenClaw 安装在隔离环境中,无法访问日历或邮箱,因此“完全没用”。能力需要访问权限、访问权限又制造风险,这一取舍将决定企业如何部署 agent。
Arora 预计,AI 带来的生产率提升会先扩大工程产出,而不是先消灭工程需求。 功能积压已经排到6-12个月之后,因此30%-60%的效率提升可以用于开发更多功能;已宣布的7%、15%或20%裁员,也可能是在为具备新技能的人腾出空间。他更广泛的判断是,企业将经历“持续10年的转型”,财务等职能的效率提升将为 token 和新增 AI 算力买单。
1. Mythos 颠覆华盛顿“让它先跑”(let them cook)的共识
Kevin Roose 开场提供的证据,是一份传闻中的行政命令:成立 AI 工作组,并可能要求前沿模型在发布前接受政府审查。它与被 Biden 撤销的框架高度相似;此前,类似测试曾被斥为“共产主义”、反创新,甚至会让美国输掉与中国的 AI 竞赛。
Casey 对因果关系的解释非常直接:Mythos 似乎能在大量程序中发现新颖且可利用的弱点,因此本届政府的放任立场“经不起现实的检验”。当官员看到一个模型如果被广泛发布、可能造成“巨大的伤害”时,现实问题就变成政府如何阻止这种伤害。
这场机构争斗,一边是 CAISI——更名后的 U.S. AI Safety Institute——另一边是 NSA 等情报机构;同时也是前 AI 沙皇 David Sacks 的“让它先跑”理念,与新近愿意把模型能力视为国家安全威胁的共和党官员之间的对撞。
Casey 认为,Trump 的立场一直比共和党整体意见更窄:共和党和民主党都对 AI 持强烈怀疑态度,共和党州议员早已争相监管 AI。Mythos 只是让联邦层面奉行“油门踩到底、不踩刹车”(all gas, no brakes)的派系不得不面对现实账单。
2. 安全监管仍有执行缺口,也存在审查风险
CAISI 专门招聘研究人员来评估日益危险的模型,其中包括一些即便对服务本届政府有所顾虑、仍选择加入的人。Casey 认为,机制上的缺口仍然存在:如果评估人员认定某个模型过于危险,但开发商以“商业需要”为由坚持发布,政府该怎么办?
Kevin 提出的反驳值得保留:监管可能演变成政治胁迫,就像社交媒体监管曾经出现的情况。Casey 同意,发布前审查可能构成事先限制,官员可能会禁止某个模型,“不是因为它真的危险,而只是因为它看起来很 woke、很 gay”;最终可能需要通过诉讼解决。
两人暂时的权衡仍偏向干预。Casey 目前宁愿听到这样的话:“那个疯狂的网络模型,不要把它给所有人。”Kevin 则对政府终于承认这项技术可能需要采取行动表示欢迎:“能拿到一点小胜利,我就拿一点。”
3. 芯片、模型与盟友撞在一起,暴露出混乱的对华策略
Trump 与 Jensen Huang、Elon Musk、Tim Cook 以及 Meta 的 Dina Powell McCormick 一同前往中国,据报道,与 Xi Jinping 的 AI 会谈也在议程之中。Kevin 认为其中存在内在矛盾:一边向中国出售构建 Mythos 级系统所需的 Nvidia 硬件,一边试图拒绝向中国提供今天的 Mythos。
五角大楼内部也存在同样的矛盾。它正在为将 Anthropic 认定为供应链风险进行辩护;这一认定源于 Anthropic 拒绝在合同中允许任何“合法使用”。但与此同时,在 Anthropic 技术据称正被移除的这段时期,五角大楼又在部署 Mythos。
一名中国智库代表在新加坡接触 Anthropic,寻求模型访问权限;德国则提出建立自己的 CAISI 类机构,并要求获得最先进模型。Casey 倾向于加强西方合作:修复互联网范围内的漏洞可能需要“我们能获得的所有帮助”。毕竟,美国实际上是在告诉盟友自己已经赢了,盟友要么“喜欢”,要么学会接受现实。
Kevin 表示,一旦模型能够找到零日漏洞,且军方和情报机构开始围绕模型改变行为,“AI 只是普通技术”的立场就无法维持。他更悲观的预测是,矛盾将持续下去,直到“某个重大事件”迫使官员真正警醒。
4. 网络攻陷已从按天计缩短至按分钟计
Arora 以7年的对比界定了转折点:过去,攻击者进入系统后还需要几天才能窃取组织的“核心资产”;有了 AI,这段时间已压缩到几分钟。为人类响应速度设计的防御体系,现在需要在同样的时间尺度内自动发现并采取行动。
Palo Alto 披露了26个关键漏洞利用,涉及75项问题,而正常基线不到5个,约为通常发现速度的5-7倍。Arora 将这次覆盖所有产品、由数百名工程师参与的行动称为一次“彻底清理”,目标是处理长期积累的“技术债或漏洞债”。
这一模式并不只出现在 Palo Alto。Mozilla 4月推送了423项安全修复,而2025年平均每月约22项;Google 的威胁情报团队发现了首个利用其认为由 AI 开发的零日漏洞的攻击者;Canvas 攻击则迫使公司停运,并就被盗数据展开谈判。
Arora 提醒,不应假设7倍增幅会永久持续;这次集中审计应能清理 Palo Alto 相当一部分积压。但每家企业现在都必须判断,自己的旧代码中有多少存在类似弱点,而开源组件通常会比专有软件更慢得到修复。
5. Mythos 是极度依赖上下文的能力倍增器,不是魔法扫描器
Arora 对 Mythos 的第一印象没有那么惊艳,因为它会标记太多问题:约30%的发现属于误报,每一项都需要人工验证。随着工程师解释代码原本要实现的功能,以及正常行为应是什么样,Mythos 的实用性才逐渐提升。
Palo Alto 随后提供了自己的专有威胁数据库,其中包含过去5年约1万次攻击所采用的技术,并要求模型判断已知方法能否适用于新的上下文。Arora 形容,这是把“过去全部的人类训练”交给模型,用来构建未来的防御体系。
Mythos 和 GPT-5.5 Cyber 发现了不同的弱点,这表明它们的训练数据和落地方式带来了互补覆盖,而不是由某一个模型给出唯一且终局的答案。对 Arora 来说,这种差异意味着“还有很多东西会被找出来”。
这些系统识别的不只是有缺陷的代码,也包括配置错误。Arora 举的最清晰例子是:一个产品控制面板为了方便远程操作而暴露在互联网上、没有关闭。“如果我能找到它,其他人也能找到它。”
6. 补丁周期跟不上自动化攻击
Casey 以 Palo Alto 的发现挑战传统90天负责任披露窗口:AI 辅助攻击者可能在25分钟内取得初始访问权限并外传数据。Arora 同意窗口会进一步缩短,但“究竟会缩短多少”仍没有定论。
SaaS 是相对容易的场景:软件可以集中调查、修补和部署,Palo Alto 已在2-3周内完成。笔记本电脑、服务器、交换机和路由器则需要组织自行行动,因此 Arora 预计,随着积压被清理,企业将经历3-6个月异常密集的更新周期。
IBM、PwC、Deloitte 和 Accenture 等系统集成商正在调动修复资源。如果无法立即修补,Palo Alto 可以把已知的脆弱路径编码进边界防火墙特征中,搭建“临时脚手架”,在企业修复后端代码的同时阻断攻击。
攻击者与防守者的竞争仍具有结构性不对称:“我们必须100%正确,坏人只需要正确1次。”即使拦下4个漏洞、仍有1个漏洞被成功利用,防守方得到的结果依然是“零”。因此 Arora 认为,当前攻击者从相当的模型能力中获取了更多价值。
7. 限制 Mythos 访问权限只能为防守方争取时间,无法带来永久优势
Arora 肯定 Anthropic 和 OpenAI 在防守方仍有时间反应时,试图负责任地展示“可能做到什么”。他说,两家公司“大体上做对了大部分事情”;它们在发布过程中各自搞砸了一些环节,但不存在简单的模型分发政策。
Mythos 的差异化能力在于消耗算力的“ultra mode”,它可以持续远长于已发布模型通常使用的“flash mode”。这种持续性让模型能够尝试多种技术并串联成功步骤,因此算力成本和风险更多取决于持续搜索,而不只是基础能力。
Arora 因此倾向于先给企业时间修复系统,再让同等能力普遍开放。如果攻击者获得 Mythos 级工具,勒索软件和国家级经济破坏仍会是熟悉的结果;改变的是攻击的“速度和规模”,而不是其基本性质。
4-6周的评估窗口让防守方得以研究模型行为,并在 Arora 所说的“AI 攻击海啸”到来前构建 AI 驱动的传感器。真正的竞赛在于,这些防护和补丁能否赶在国家、开源行动者或第三方复现相关能力之前到位。
8. 资源不足的行业与消费者构成薄弱层
Arora 最担心的是那些业务“95%在做别的事”、只有5%与技术相关的组织:医院、小企业、工业制造商、基础设施运营商和诊所。金融机构拥有更深的工程能力;而 Change Healthcare 这样的上游事故,就可能让一家医生诊所陷入瘫痪。
消费者得到的保护弱于企业,因为他们缺少有效的统一守门人。企业防御系统可以在一个客户处观察到钓鱼邮件发送者,再在其他地方将其拦截;个人邮箱和电信服务商则需要针对冒充攻击加强控制,这类攻击对“在做 AI”的公司而言本应很容易识别。
Kevin 提出的消费者常规防护方案是使用强密码和多因素认证,Arora 则强调服务商层面的控制,并敦促用户安装软件更新。Kevin 每天收到假冒 X 的密码重置邮件,正说明了即将到来的问题:他预计,6个月或1年内,同样的诱饵会变得逼真得多。
9. 自主 agent 把有用的权限变成安全负债
Arora 称 OpenClaw“从安全角度看很可怕”,因为它可以获得凭证和权限,随后跨用户账户行动。一次晚餐上,一名热心用户展示了名为 Zara 的 agent,旁边的人则反应道:“我靠,这简直是安全噩梦。”
他自己的 OpenClaw 运行在一台隔离设备上,与日历和邮箱断开连接,因此“完全没用”。机制很简单:没有访问权限的 agent 无法预订会议或回复邮件,而拥有访问权限的 agent 则可以代表用户采取行动。
工程师们同时感到兴奋、过劳和恐惧。在 Palo Alto 9,000多名技术人员中,Arora 看到了各种情绪:眼下的工具很有前景,但它们在未来2-3年的影响仍然极不确定。
10. AI 会先扩大积压工作,再压缩技术就业
Arora 否定了30%、40%、50%或60%的生产率提升必然意味着工程师减少这一推论:“我需要更多工程师。”产品路线图已经排到6-12个月之后,原因是团队缺乏产能,因此第一轮效率提升应投入长期搁置的功能和测试。
宣布裁员7%、15%或20%的公司,可能是在“重塑”组织,而不是永久缩小规模,从而腾出空间招聘具备新技能的人。整个企业中,财务、人力资源和其他职能的效率提升,可能是支付 token 费用的资金来源。
他对宏观趋势的概括是“一场转型意愿的海啸”,以及持续10年的商业转型。CFO 或 HR 负责人并不关心抽象的 AI;他们要的是更高效的运营,无论具体形式是自动化评估、面试,还是内部工作流。
Kevin 认为“AI 测评官”会带来糟糕的候选人体验,但 Arora 认为,它评估领域技能的能力可能强于对话。对于求职者声称自己熟悉 AI,他偏好的测试方式是看实际产出:“给我看看。”一个根据菜谱生成购物清单的 agent,并不能证明真正的能力。
11. Hot Mess Express 展示了采用潮与信任危机的碰撞
Venmo 开始测试新用户默认仅对好友可见,试图修补一项长期存在的隐私问题:平台公开交易记录曾帮助记者识别涉及 Joe Biden、J.D. Vance 和 Matt Gaetz 的账户或付款。主持人认为这是迟来的清理工作:它“过去真是一团乱麻”,如今这条容易被调查追踪的路径终于要被关闭。
据报道,Amazon 员工生成了不必要的 Meshclaw 活动,以增加 token 消耗,让自己在管理者面前看起来表现更好。Casey 引用了 Goodhart 定律:“当一个指标变成目标,它就不再是一个好的指标。”他将结果戏称为“hot mesh”。
University of Central Florida 的艺术与人文学科毕业生对 AI 是“下一次工业革命”的说法报以嘘声;Kevin 看到的学生中,约80%会说:“我讨厌这个。”Casey 为现场观众辩护,而 Kevin 则要求任何曾在学业中使用 ChatGPT 的人,在起哄前先披露这段经历。
Grindr 的 Madonna 广告播放了“嗨,Grindr,我是 Mother”,即使用户手机音量已关闭也可能继续外放,导致用户在家人身边被迫出柜;Casey 称这是危险的混乱,不只是营销失误。其他消息包括:Dua Lipa 就 Samsung 包装问题索赔1500万美元;eBay 将 GameStop 那份资金明显不足的550亿美元报价斥为“既不可信,也没有吸引力”;证词还显示,Elon Musk 曾考虑将 OpenAI 的控制权交给自己的孩子。
Casey, will you record my audiobook for me?
Yes, I would love to, actually.
Okay, thanks. I got the briefing yesterday on what this would entail for me.
Mm-hmm.
They want 36 hours in the studio to record this audiobook.
That's—wait, hold on. 8, 16, 24. That's over 4 days' worth. That's 4 and a half days of recording. That's almost a full week. Oh my God.
Yeah. I know. But apparently people have a connection to us because of our voices—
Yeah.
—so they didn't want me using an AI clone to do it.
It makes it—you know what? I really think there would be a case that I should do this because it would force me to read your book. You know what I mean? Then I really can't get out of it. I'm on the hook to read this thing for real. So that might be the best way to do it.
You can insert your little snotty wisecracks if you want. Mystery Science Theater it.
Yeah, a little extra commentary on the side. Like, “Oh, I see we're using that transition again. Hmm.” “Oh boy, he really ended this whole thing with ‘time will tell.’” “I would have suggested a different direction. Was this book edited?”
No, wait, now I actually want you to do it. I'm Kevin Roose, the tech columnist at The New York Times.
I'm Casey Newton from Platformer.
And this is Hard Fork.
This week, is AI safety back? The Trump administration seems to be changing its tune. Then, Palo Alto Networks CEO Nikesh Arora joins us to discuss what's real and what's hype in the freak-out over Claude Mythos. And finally, the train has returned to the station. It's the Hot Mess Express.
Buckle up.
People don't typically buckle a seat belt on a train.
This is a very safe train.
All right.
1. AI Safety Is Back
Well, the big news this week is that President Trump headed to China with a cohort of American business executives to have a series of meetings about Chinese trade policy, AI, and other things with Xi Jinping and other leading Chinese officials.
Now, is it true that when they walked off the plane, a bunch of H100s fell out of the leg of Jensen Huang's pants?
I haven't heard that confirmed—
Okay.
—but I'll look into it.
Thank you.
I want to talk about this less through the lens of President Trump and United States trade policy than through this larger shift that I think we've both observed over the past week or so, which is that after several years of dismissing AI safety and doomer fear-mongering about AI, the Trump administration—or at least parts of the Trump administration—seems to be getting quite scared about what's happening.
Yes, and while this is something that I think was honestly inevitable, it still has been jarring to see it happen because it seems like this administration has really turned on a dime when it comes to this subject.
Yeah. So let's talk about what's been going on and some of the data points that support the idea that the Trump administration is changing its AI posture, or at least has several different AI postures that it's considering. But first, let's do our AI disclosures. I work for The New York Times, which is suing OpenAI, Microsoft, and Perplexity.
And my fiancé works at Anthropic.
2. Trump Reconsiders AI Oversight
So first, there was this executive order—or rumored executive order—that my colleagues at The New York Times reported on last week. This would be a new executive order to create an AI working group that would bring together tech executives and government officials to potentially come up with new ways of overseeing or regulating AI. One of the potential plans being discussed is a formal government review process for new AI models before they're released. So this is still ongoing. We still don't know exactly what the executive order will or won't include, but we are expecting more news on that.
Yes, and the reason that is notable, Kevin, is that on President Trump's first day in office in his second term, he canceled President Biden's executive order on AI, which, among other things, included a very similar kind of review process for new frontier AI models. The Biden people were very confident that we would one day get models that could be used to commit great harm, and so they wanted to get a handle on that before those models were released. And when Biden did that, many Republicans were saying, “This is anti-innovation. You're going to make us lose to China.” Well, well, well, now the shoe is on the other foot and they're saying, “Hey, slow down. Don't release those things quite so fast.”
It's so remarkable how fast the Overton window has shifted on this idea. I mean, as you just said, during the Biden administration and during the SB 1047 fight here in California over this proposed AI bill, people in tech, on the tech right, and among the more libertarian crowd were incensed about the idea that the government might ask them—
Yes.
—to do pre-release testing of their models and then submit the results to the government. They called this communist. They were implying that this would be the end of free enterprise as we know it, and now, just a couple of years later, they are reportedly considering doing something similar. So what do you think happened here?
Well, I think that basically, to use a phrase you sometimes like to use, the Trump administration's view of AI just did not survive contact with reality, right? In a word, what has changed here is Mythos, the model that Anthropic has now released as a preview to a very small group that now includes many federal agencies. This model is apparently very good at finding novel vulnerabilities in code that can be used to create exploits, and that appears to be true across many, many, many programs. And so the administration, I think, took a look at this and the serious people over there said, “Look, whatever your views may be about free trade and the threat of losing to China, we have a model right now that, if it were just unleashed on the public, could create vast amounts of harm.” And I think, to their credit, the Trump administration said, “Okay, then what would be a policy to prevent harm from happening?”
3. Mythos Triggers A Turf War
Yes, Mythos is the proximate cause here for a lot of this, but I think it's also worth talking about the various factions within the Trump administration that appear to be battling over control of this new AI regulatory push. There appears to be a turf war breaking out between the Center for AI Standards and Innovation, or CAISI—
Shout-out to Casey.
—which was formerly known as the U.S. AI Safety Institute. This was a group within the Commerce Department that was set up under the Biden administration. The Trump administration came in and basically didn't like that they considered this group a bunch of doomers. So they made some changes, including changing the name. But this is a group of AI researchers and safety experts who work in the Commerce Department and want to be involved in vetting new models.
And there's just something so funny about these people coming in and saying, “AI safety is such a stupid idea that we have to remove ‘safety’ from the name of this institute,” and then, one year later, being like, “Well, AI safety is really going to be a focus for us from now on.”
Yeah. So there are some people who believe that the vetting of frontier models should take place within the intelligence community, including the NSA and various other organizations. So there's some turf war there. There's also just this interesting kind of posture war over whether the let-it-rip approach to AI development—or, as former AI czar David Sacks put it, the “let them cook” philosophy of laissez-faire regulation—should prevail, or this more hawkish, safety-oriented faction within the Republican Party that does see these models as a big threat and wants to take steps to reel them in.
Right. So do we know at this point who seems to be winning that battle? And do you think it matters to the average person which side gains the upper hand?
I do. I think there's obviously going to be some back-and-forth. We'll see, when this executive order comes out, what they do about the testing requirements and where they locate that—whether it's, “We're going to let the NSA do this,” or, “We're going to let CAISI do this.” I think that all might matter a little bit, but I think the general posture of the administration changing from “AI safety is ridiculous, and these doomers are using hyped-up fears to enact regulatory capture” is very different from what we are seeing now, which is, “Oh, wait, these models are very powerful, and we don't want our adversaries to get access to them.”
But we should also say it is entirely confused and incoherent right now at the level of the federal government, because on one level, you have President Trump inviting Jensen Huang of NVIDIA onto Air Force One to fly with him to China to try to make a deal to presumably open up the export of NVIDIA's most powerful AI chips to China, while at the same time, you have other high-ranking government officials saying, “We need to institute some kind of safety regime because these models are potentially very dangerous.”
Yes, and nowhere is that schism more apparent than in the Pentagon, Kevin, where, on the one hand, the Pentagon has designated Anthropic as a supply-chain risk because it refused to amend its contract to enable any, quote, “lawful use of its technology,” as we talked about on the show for a few months.
That designation, the Pentagon is still arguing for in court, but at the same time, we learned that this week, during the period when the Pentagon is supposed to be unwinding all of Anthropic’s technology from the Pentagon, it is also implementing Mythos and using it to try to scan for vulnerabilities.
It’s truly wild.
I want to be in the meeting where the person who has to remove Anthropic from the Pentagon sits down with the person who’s installing Anthropic into the Pentagon and just hears what those talks are like.
Yeah. So aside from the obvious incoherence and maybe hypocrisy of these conflicting positions, which side do you think is going to come out on top here?
Well, obviously, I’m always going to side with CAISI. You know, CAISI is a great agency. Great people over there. And honestly, they were just set up to do this exact thing, right? When it was established under President Biden, the idea was, these models are getting better. Pretty soon, they’re going to be dangerous. We need to have a way of evaluating them before they’re released.
And frankly, they’ve just hired a lot of people who I think ordinarily might not work in a Trump administration but felt like, “This is so important that I’m going to swallow hard and go over there and try to serve my country by protecting us from the worst things that AI can do.” So to me, it seems like they would be very well set up to do this kind of work.
Where I think we still have an obvious gap, though, Kevin, is that it’s not entirely clear to me what is supposed to happen when a company like Anthropic comes up with a model that is too dangerous to release in the view of something like CAISI but wants to release it anyway. And I assume we are just going to get there. Sometime within the next 6 months, one of these companies is going to say, “Yeah, it’s risky, but we think it’s fine to put out there. We have business imperatives. We’re going to talk ourselves into it.” Then what happens?
Yeah. I mean, it’s also just so clearly unfortunate that the issue of AI safety has become polarized in the way that it did over the past couple of years, that caring about safety and talking about safety became vaguely woke-coded. And people in the Trump administration thought it was a bunch of hysterical liberals using fears of AI to get heavy-handed regulation into place.
I don’t think that was ever true, but I think it has become especially untrue now when you have very senior people in the Republican Party talking about how we need to restrain these systems. It’s frustrating because I think you and I both saw that this technology is real. It’s going to get even more powerful than it is, and at that point, it’s not going to matter whether you’re a Republican or a Democrat. You do not want this stuff falling into the hands of our adversaries.
True, but I think the Trump administration was always out on a limb here in a really weird way. We have talked a lot recently about what the surveys show when it comes to public opinion of AI in America. Republicans and Democrats are largely aligned in being deeply skeptical of it and even outright hating it, and that’s why you see so many Republican state legislators trying to pass laws to rein in AI, right?
You did not have to convince Republican state legislatures that AI was dangerous and needed to be regulated. They were racing to do it, and the Trump administration has had to put a lot of energy into trying to pass a moratorium so that it can preserve its all-gas, no-brakes approach to AI.
So what I think happened here was that there was basically a minority of Republicans that happened to be running the country who said, “Let the labs do whatever they want,” and then Mythos comes out, the bill comes due, and they sort of have their pants down. They have to change their tune.
Yeah.
Just to throw a lot of metaphors in there.
Yeah. Pants, tunes.
Yeah. I’ll come up with more. Don’t worry.
We’re getting there.
Yeah.
One other thing here is that you are starting to see the issue of catastrophic or existential risk floating up and percolating on the right. This is something that people like Bernie Sanders have now been talking about on the left for a couple of months. But on Tuesday of this week, Ted Cruz was talking about catastrophic risk and the need to protect against it.
I just think that the improvement of these models and the fact that they are so clearly useful for dangerous things like cyberattacks is going to scramble some of the usual partisan allegiances here.
Yeah, I mean, look, the idea that a large language model might eventually get so good that it could break into your computer and wreak havoc—that was not a liberal view. That was just a view grounded in an observation of the rate of improvement in the model.
In truth, I am glad that they are reversing course on this and doing it before we’ve had a massive catastrophe. Maybe an asterisk there, though, which is that I truly feel like every single day for the past week I’ve seen news of a major cyberattack, and increasingly we’re getting word that these may have had AI systems involved in identifying these vulnerabilities.
Yeah, so there may be a catastrophe unfolding under our noses; we just don’t know about it yet.
Yeah. Stay tuned for next week’s episode.
4. China Wants Mythos
I want to talk a little bit about this China trip and what, if anything, we think that has to do with AI regulations.
Mm-hmm.
There was some reporting in The Wall Street Journal last week that both the U.S. and China have been considering a series of official discussions around AI. We know that AI is on the agenda for President Trump’s meetings with Xi in China this week. And we also know that China has been looking to get access to Mythos.
There was a great story recently in The Times that talked about the fact that a representative from a Chinese think tank approached Anthropic officials at a meeting in Singapore last month to basically lobby them to open this model up to China.
And we want to give them the Hard Fork Chutzpah Award for shooting your shot.
Yeah.
If you work at a Chinese think tank and you think Dario Amodei was about to hand you Mythos, that is truly—I aspire to your level of self-confidence.
Listen, you miss 100% of the shots you don’t take.
It’s true.
So along with Jensen Huang, who finagled a last-minute invite on Air Force One after there were news reports that he was not going to be going on this trip, Elon Musk, Tim Cook, and Dina Powell McCormick from Meta are also on the trip with Trump. What’s going on here, and how would you characterize the blunt rotation among those tech executives?
You know, this is a group of executives that are aligned with the Trump administration, and they have all found, in various ways, that the more time you spend flattering President Trump, the more tax breaks and other forms of relief your company gets.
This is exactly what we talked about expecting Tim Cook to do once he announced that he’d be stepping down as CEO. You’re just kind of a Trump whisperer, and you follow him around and say, “Go, President Trump, and also please give Apple what we want.”
Meta, Apple, and Nvidia have all had huge success with this administration, and now, as their reward, they get to be photographed with the president flying around China.
Yeah. I think I am just very unsure where all of this settles out, because I can imagine Trump wanting to go to China and make a bunch of deals, and obviously Jensen Huang and Nvidia want to be able to sell their chips in China.
So I can see them, on one hand, giving some kind of expanded access to Chinese AI companies to get these American chips, but then I can also see them not wanting China to get access to models like Claude. I just don’t know how that resolves, and I see it as basically inherently contradictory that you want to give China, or sell China, the means to make its own Mythos-caliber models while at the same time trying to block it from getting access to the one that we have today.
This is where it would be helpful to have a coherent strategy, but we don’t, right? It’s like the same administration that is installing and uninstalling Anthropic at the same time is having a similar level of confusion over in China, where it seems like the administration is highly susceptible to blowing wherever the wind is today.
Yeah. I mean, I am generally not all that optimistic about the government’s ability to regulate technology in a way that is timely and relevant. And I hope I’m wrong here, but I think that we will see this sort of incoherence and contradiction until there is some big event that forces everyone to sit up straight.
I mean, my question is, will this be a case of the same AI-safety-minded people who were dismissed for the past couple of years by the Trump administration being proven right again in the future, when it turns out that China did use access to American technology to build Mythos- or better-level models? And will there be any regrets that we paved the way for them to do that?
I don’t think it’s unlikely.
Yeah. It's interesting, though. I had a conversation with a federal official recently, in the last couple of weeks, where this person was basically telling me that AI is just a normal technology, taking the line that we've heard again and again from the people who don't want to regulate this stuff: “This is just the internet. This is just the PC. It's not some special technology that requires special rules.”
That position has just become so untenable, to me at least, when you have models that are out there finding zero-day exploits. Clearly, our military and our intelligence agencies don't think this is a normal technology. They think it's more like a step change that requires them to act in different ways. So I am very curious what happens to the “AI is normal technology” camp inside the Trump administration as the technology continues to grow.
Yeah.
They may change their arguments, or they may not. That's the thing. You just don't know how committed these people are to their view.
You don't.
We should also talk about some of the international reaction to Mythos, because it's not just China that wants into this thing. Germany's Digital Affairs and Cybersecurity Agency has come out this week with a proposal for establishing its own version of something like the U.S. CAISI. They are also demanding access to state-of-the-art models like Mythos.
So it just seems like this model has forced conversations around the world about who should have access to which models. Should the public have access? Should governments have access? Which governments should have access? It just seems like we are in a new era of AI brinkmanship.
For sure. What I hope that we will see in the coming months is more and more cooperation. The whole reason that we had that series of AI action summits over the past few years was to try to get more cooperation among the Western powers with this stuff. And then last year, the U.S. sort of came in and said, “That's over. The U.S. is winning the AI race, and you can like it or learn to live with it,” basically.
Right.
So it's no wonder to me that these other Western powers are seeking access to these models, and I think there's probably honestly a good case that they should get access to these models. When it comes to fixing every vulnerability on the internet, I think we could probably use all the help we can get.
Yeah. I remember that AI action summit. I didn't go to the most recent one in India, but at the one in Paris before that, it was just like, “Oh, we're not going to talk about any of this.” We're not going to talk about the dangers that this technology might create because we're so invested in this sort of accelerationist posture. So, how far we've come—and yet we are still in the very early innings of this.
Does it make you wonder what would've happened if the Trump administration had just been listening to Hard Fork a year ago? Could they have saved themselves some trouble here?
It's possible.
Yeah.
So, Casey, the politics of AI and AI regulation are obviously shifting very quickly. We may learn more this week after these meetings in China, but what is your take on what this latest burst of news signals about AI or AI regulation?
My take is that this is a rare bit of good news when it comes to AI regulation. I am somebody who's been worried about AI safety for a long time, and one of the main reasons I've been worried about it is that our government has seemed to have this feeling of, “Let's just see what happens.” Whereas to me, it seemed pretty obvious what was going to happen.
Now we have arrived at that point. We have a superpowerful model, and to their credit, the Trump administration is saying, “Okay, it seems like we were wrong about how capable these models were going to be. Let's make some changes.”
And do you think there's any way that this turns out to backfire? I'm just remembering people wanting social media to be regulated, and then when the Trump administration started doing things in the realm of social media, it amounted to what you and I would consider sort of censorship—
Yeah.
—or at least wanting to strong-arm the social media companies into doing their bidding. So do you think there's a possibility that something similar happens with AI, where we get the regulation, but it's just the wrong kind? Or this pre-release testing is testing for the wrong kind of thing?
Yes. I am very sympathetic to those who believe that this could amount to a kind of prior restraint on free speech, and that there is a risk that members of the Trump administration will effectively say, “You can't release that model, not because it's actually dangerous, but just because it seems woke and gay.” I think that we need to keep an eye out for that, and potentially someone is going to need to sue over it.
Yes.
But when I look at how I want to balance those things, for the moment, I would rather have an administration saying, “The crazy cyber model—don't give that to everyone.”
Yeah. I think I'm landing at a pretty similar place. I'm a little worried that this regulatory push from the right is going to be confused and maybe too sudden, and there's going to be some sort of overreaction that ends up with something more like the sort of censorship that you mentioned.
But I am glad that after many years of kind of denying that this technology was important, and that it would become as good as the people at the lab said, our government at least appears open to the idea that maybe they need to step in and do something here.
Mm-hmm, mm-hmm.
I'll take the little wins where I can get them.
That's what I'm saying. When's the last time we talked about a win on this show?
Yeah. When we come back, what is Claude Mythos doing to the world of cybersecurity? We'll talk to Palo Alto Networks CEO Nikesh Arora.
Well, Kevin, is it just me, or every time you look at the tech news, do you see some new cyberattack that seems to have befallen some company or another?
Yes. This is my experience of social media over the past 2 weeks. I log in, I see 3 posts from companies about how they've discovered more bugs in a 24-hour window than in the previous 80 years of their company's history. And then everyone's reposting that with just, “It begins,” or, “It is over,” or, “Hide your kids.”
Yes. Just to name a few of those, Mozilla was one of those companies saying that it had pushed 423 security bug fixes in April alone, compared to an average of about 22 per month throughout 2025.
Google announced on Monday that, for the first time ever, its threat intelligence group had identified an attacker using a zero-day exploit that the group believes was developed with AI, so that's kind of a grim milestone.
And then, if you're a student, perhaps you noticed the cyberattack on the learning platform Canvas last week, which forced the site down for several hours. The company behind Canvas, which is called Instructure, had to negotiate a deal with hackers for the return and destruction of the stolen data.
So, on the one hand, there are cyberattacks going on all the time, but it does seem like some new inflection point has been reached. Of course, one reason that people think we might be seeing more of these is AI.
Yes. So we have talked about Claude Mythos Preview, the model that Anthropic did not release widely but released to a select group of companies and open-source maintainers. And today we're actually going to talk to someone who has used Mythos and who has been on the front lines of this frantic sprint to secure the infrastructure of modern life.
Yes. Our guest today is Nikesh Arora. Nikesh is the CEO and chairman of Palo Alto Networks, the largest cybersecurity firm in the world, which supports more than 70,000 customers, including the vast majority of the Fortune 100.
And as you mentioned, Kevin, Palo Alto was among the organizations given early access to Claude Mythos, as well as GPT-5.5 Cyber.
Yes. Nikesh is one of the people I think is best positioned to see the effects that these models are having on cybersecurity because they work so broadly across industries. They're also a big government contractor, so I'm really interested in what he thinks is different about this new class of models.
Yes. And something I appreciate about Nikesh is that, in an industry where there is a lot of hype because, of course, the more scared a cybersecurity executive can make you, the more likely you might be to buy their software, Nikesh is somebody who I think tries to maintain an even-keeled approach here and not ring alarm bells where none are needed.
That said, I do think that he is quite concerned about some of the things that he's seeing.
Well, let's bring him in. Nikesh Arora, welcome to Hard Fork.
Well, thank you for having me.
I want to just start with your account of what it feels like to run a major cybersecurity company right now. Casey and I have talked with people at these companies for many years, usually because something terrible has happened, and I feel like the vibe we get is, “This is the worst, most dangerous time ever in cybersecurity.” What is your subjective experience as someone who's been in this field for a long time?
I'm perhaps a little more relaxed than what you're trying to ascribe to people who come here and tell us it's the worst moment.
Historically, what’s happened is, in the last 7 years, the time from somebody breaching an organization and being able to extract, we’ll say, crown jewels has been measured in days. Unfortunately, with the emergence of AI and the arrival of agentic technologies, that time frame has shrunk down to minutes. When that happens in minutes, your defense systems have to be able to be activated and defend yourself within minutes.
Fundamentally, the cybersecurity infrastructure was designed for days. Some parts of it are making it to seconds—the good parts where you know how to stop them—but we have to basically overhaul the backend infrastructure to make sure it’s AI-ready, so we can fight AI with AI. You’re seeing that. You’re seeing AIs out there. You’re seeing people like Anthropic launch models like Mythos. You’re seeing OpenAI do that with GPT-5.5 Cyber. They’re showing you the art of the possible from a bad-actor perspective.
We have to make sure we move as fast as they do, or faster, perhaps, to try and plug those holes and make the infrastructure better.
Hmm. So your company recently put out a report on some patches that you all—
Yes.
—have made to your own systems.
Yes.
You disclosed 26 critical exploits covering 75 issues, and you said that’s against a typical baseline of under 5.
Yeah.
Meaning that they discovered five to seven times as many in a comparable period.
Five to seven times—
Yeah.
—depending, yeah.
Yeah. Is that pretty standard for the kind of spike you’re seeing in exploits, or discovered exploits, as a result of Mythos and similar models?
5. Mythos Finds Hidden Vulnerabilities
What we’ve discovered with some of the newer models that have come out in the last few weeks, perhaps a month or so, is that AI models are getting very good at coding. As the models start to understand what good code looks like, they also start developing an understanding of what bad code looks like. If you point the model at all the code repositories you have and say, “Okay, now look through all this code and find me bad code,” it will. Unfortunately, humans have been writing bad code for a very long time.
On average, we’ll find about one-fifth or one-seventh of what was found in the last 6 weeks using these models. Of course, remember, we ran a concerted effort to see what the models were going to find. We had hundreds of engineers working on it to make sure we looked under every rock and ran every product through it.
It’s almost like it’s a great cleansing moment, right? We found 7 times the volume that we would’ve normally found in a normal period. It’s not going to happen again, hopefully, because we’ve hopefully cleared out a whole bunch of what we’ll call the tech debt or the vulnerability debt. But I think a lot of organizations will have to go through this moment to understand how much of their code written in the past suffers from these vulnerabilities.
They’ll have to do their own work. They’ll have to make sure that it’s fixed. I think the challenge we’re going to run into is that most companies use a large corpus of open source, and open source doesn’t get patched or remediated as quickly as your own proprietary code can. The other thing we found, very interestingly, with Mythos and other models is that it’s really good at daisy-chaining vulnerabilities, and that’s what needs to be contended with.
I’m trying to get a sense of the scale of this issue, because I feel like within the past few weeks, I’ve heard a lot of stories like the one you just described about your own company. Mozilla has been publishing blog posts about discovering—
That’s exactly—
—hundreds of bugs over a period where maybe previously they only would’ve discovered a couple of dozen. My sense is that as more companies undertake this audit, they’re going to find that they have similar problems.
What’s the time scale that we might expect for these kinds of issues to be fixed? Is there enough time to fix particularly critical infrastructure before our adversaries gain access to similarly capable models?
That’s a great question, Casey. I think that’s what should keep us up at night, right? Not every organization has the resources to fix code that could’ve been written 20 years ago.
The good news is that most cyber defenders have had access to the models. They understand the scale and enormity of the problem to some degree. What we’ve been able to do is enlist the support of many of the system integrators in the world, like IBM, PricewaterhouseCoopers, Deloitte, and Accenture, who are all rallying to make sure they make resources available to many of these customers to patch these things.
But I think we’re in the midst of testing an interesting solution: Once we know the vulnerabilities in an organization, we can write signatures into our perimeter-defense firewalls to say, “If you see somebody trying to go in this direction, we know there’s an unpatched piece of code behind it. Block them.”
So we can create a temporary scaffolding to let organizations have a little bit more time to go fix their vulnerabilities. But it has to be done, and the risk, as you rightly articulated, is that open-source actors, nation-states, or third parties can start building models that are similar to what Anthropic or OpenAI have built. The risk is that they get there faster than the patches have been enabled in many enterprises.
Hmm. Yeah.
I want to understand a little bit more about the defense side of this now that you have access to the Mythos model. There’s been a lot written about it. It’s the subject of much debate at the highest levels of power, and I just want to ask: What is it like to use it? Does it feel different than using Claude Code? If you’ve used another Anthropic product, does it feel kind of the same? What is it like to use Mythos?
6. Mythos Needs More Context
In the beginning, it was not that impactful. When you’re looking for bad code, it’s going to find everything. Remember, 30% of them are false positives.
Mm.
It’s not always going to get the right thing, but unfortunately, we’ve got to test every one of them out to see which is real. What became more and more fascinating was that the more context we gave it, the better it became.
What do you mean?
You show it a piece of code, and it doesn’t know what the code is trying to achieve.
Right.
So you have to give it context, saying, “Well, this code works—”
So you’re not just pointing it and saying, “Go test this firewall—
No, no, no.
—and tell me what you find.” You’re actually giving it some instructions beyond that.
You have to give it context in terms of what the purpose of the code is, what it does, and what normal behavior is supposed to look like. Then you have to give it more context in terms of other threat research.
The models don’t have all the threat research in the world. We sit on hoards of threat data saying, “This is how 10,000 attacks have been conducted in the past 5 years,” which is data we store and hold because we write machine-learning algorithms to protect us from these instances. So we say, “We’re arming you with all the past known techniques that have been used. Can you see if some of those known techniques can be applied in this scenario?”
Effectively, you’re giving it all the human training of the past to make sure that in the future you can build defenses against those techniques.
You’ve mentioned using both Mythos and GPT-5.5 Cyber. I’m curious, in your mind, how comparable those models are. Are they in the same class, or is one different than the other?
The most fascinating part is that they both found different things.
Hmm.
That tells you that, based on their grounding and their training—whatever they’ve been used to train on—one of them was better at certain things, and the other one was better at some other things. It just tells you that there’s still a lot that’s going to get found.
Hmm. One thing that stuck out to me as I was reading some of your blog posts and your postmortems about your experiments with Mythos is, if a cybersecurity company is finding 5 to 7 times more vulnerabilities using this model—
Yes.
—the average bank, the average insurance company—
To say nothing of Kevin’s personal website.
My personal website—I mean, we’re going to be looking at many multiples of that, right?
Yes, yes.
Or is it the case that everything is so centralized and runs through just a few platforms that the average institution is not as screwed as I think they are?
I wouldn’t say the average institution. There’s a lot of work that needs to be done. It’s not just good at finding vulnerabilities. The other thing we also found as part of our testing is that it can even take a look at products you might be using to power your website, which you may have misconfigured.
Mm-hmm.
That’s not a vulnerability. That’s human error in the way you’re using the product, where you’ve left the door open.
For example, many people will take products and say, “Ah, it’s easier if this control plane of this product was accessible from home or from the internet so I could just go access it from wherever I am and manage this thing.” Well, you should not leave control planes of most products in your company exposed to the internet, because if I can find it, other people can find it too.
Right.
Right.
When Mythos was first announced, there were a lot of people who were very skeptical. They said, “Oh, this is just marketing hype,” or, “Anthropic doesn’t have the compute to serve this model,” which is why they’re only releasing it to a select group of companies. A month or so later, do you still hear that kind of thing from people in your industry, that maybe this isn’t the sort of apocalyptic moment that Anthropic and others have said it is?
Yeah, I look at it slightly from a longer-term perspective. I think what the Mythos model showed is what the art of the possible is going to be in the future once we are compute-unconstrained or have better models in the future that are trained better. It gave us a window into what’s coming, I think, which was very useful.
I think that’s a bit of a tough rap toward Mythos: that it did this on purpose. Remember, these companies, whether it’s OpenAI or Anthropic, are working their way through trying to understand how to do this. Both Anthropic and OpenAI want to do it right. They want to do it so that AI is not used in a bad way, at least in this instance. I think they were trying to do the right thing. I think there is no easy solve to this. I give them credit for trying to do the right thing, and I think they partly got most of it right. Some of it they fumbled along the way, but I give credit to both of them for trying to get it done right.
Speaking of how we fix this, for decades cybersecurity has operated using this sort of a 90-day responsible disclosure window—
Yes.
…where I find something, I find a bug, I privately notify you, but in 90 days, I’m going to go public with this, so you better get your act together and fix it. Companies often do take 90 days or longer to implement those bug fixes. I read a blog post this week by a researcher named Himanshu Anand who wrote that, in his opinion, the 90-day responsible disclosure window is dead. I also saw that in your own company’s blog post last week, you guys said that within 25 minutes in an AI-assisted scenario, somebody could get initial access to a system and exfiltrate the data. So do you agree that this 90-day window is dead? And if so, what the heck do we do about it?
Look, I think the principle of the 90-day window is to allow the owners of the product, the piece of software, or the piece of code to have enough time to investigate, fix it, and make sure their customers are secured. I think the 90-day window is going to shrink, as you have rightly articulated. How much does it shrink? It’s still up for debate. How long do we have?
Think about what we just did. We announced this morning that we’ve patched almost 30 critical vulnerabilities. We’ve known about these for 2 or 3 weeks. We’ve had time to test them. We had time to build patches and pretty much deploy everything that’s available from a SaaS software perspective. So the challenge is not the SaaS software. SaaS software you can find, you can fix, you can deploy. It’s not a problem. The challenge is when there’s a laptop sitting in front of you, and I’ve got to go make sure you update your laptop because you’re required to do something with it.
And I can tell you, he will go 6 months without installing the mandatory updates. I’m not even kidding.
Delay.
Yeah, yeah.
Delay. I’m starting to see more of those just in my products, and I’m getting more requests to update system software. Is that Mythos-related? No, seriously, I’m wondering to myself every time I see it, I’m like, “Oh, what did Mythos find now?” So are we starting to see, as consumers, evidence that some of these systems need to be patched more frequently?
As I said, there is going to be a cleansing of the vulnerability backlog that has been built over the years. You will most likely experience, in the next 3 to 6 months, a lot more of it if you’re an enterprise. You’ll experience it in a lot more boxes that you buy. You buy servers, you buy switches, you buy routers. All those things where you have code lying on them will have to be looked at and patched or upgraded over time. So you’re going to see some of that cleansing happen, but hopefully you can power through it and get to the other side.
But it sounds like it is just a good time to install those software updates when you get them.
Yes.
Yeah.
I highly recommend you do that.
Yeah.
7. Attackers Gain The Advantage
One persistent question about these models is whether they favor attackers or defenders. I guess I’m just going to put that question to you: Is this technology better for people who want to break into systems or people who want to safeguard systems? And if you had attackers and defenders with an equal model, who would win?
That’s a great question.
The classic Batman versus Superman.
Remember, it’s an unbalanced fight to start with. We have to be right 100 percent of the time. The bad guys are right once. So it’s an uneven playing field from that perspective.
If the model can find you 5 vulnerabilities and you can exploit 1 of them, it’s a win for them and a loss for us. It doesn’t matter if you protect against the other 4. We don’t get an 80 percent grade for protecting the other 4. We get 0 because it was able to find something to breach us.
So for now, the bad actor is most likely able to use it much better than the good people. That’s not a model constraint or model fault; it’s because the model doesn’t protect. Remember, the sensors protect. The sensors we apply around your perimeter protect. The sensor has to be smart enough to understand what the model is going to find, and that’s why, given the fact that we got this window of 4 to 6 weeks to test and understand them, we’re busy building defense techniques to make sure that as this tsunami of AI-based attacks starts to arrive, we have enough defense capability, which is still powered by AI, to give us the real-time response that we need.
Is there a sector of the economy that you’re most worried about when it comes to cybersecurity and the new capabilities of AI systems?
The challenge always is the companies that use technology where their core business is 95 percent something else, and the 5 percent part is technology. You can take that to mean small businesses. You can take that to mean core industrial manufacturing output-type businesses, where they’re not spending as much time thinking about the technology. They’re busy digging for gold or building infrastructure or something else.
Or hospitals—
Exactly. So those people—
…they use technology. So you’re worried about the non-tech businesses—
Yes.
…that may not have as many resources or as many—
Yeah.
…engineers working on—
And I’m not worried about financial institutions. They have more engineers than I do, so they will rally against it. They’ll put the resources to work, and they’ve been protecting themselves for a very long time. They understand the implications of these things.
So, a poor doctor’s office. Remember that there was a breach at Change Healthcare, I think almost a year ago, maybe slightly more, which caused a whole bunch of the physician ecosystem to come to a halt, and the physicians didn’t know what to do about it.
Mm. Yeah.
Hmm. For the moment, do you breathe a sigh of relief that these models are not generally available, or do you think they could be released and it wouldn’t be that big of a deal?
Well, they have been released, right? Both Claude Opus 4.7 Cyber and OpenAI’s GPT-5.5 have been released with cyber capabilities and guardrails.
But not Mythos.
Not Mythos.
But Mythos has another unique property, which perhaps goes toward your conversation about constraints: Mythos runs in Ultra mode. Ultra mode is a compute-consumptive mode, which allows the model to persist for much longer than the Flash mode that most models are released in. So if you’re—
So what you’re saying is it can just work for a lot longer—
That’s right.
…spend a lot more compute—
That’s right.
…than other models.
So the compute cost is from the persistence, perhaps, not from the capability. The persistence allows the daisy-chaining to happen much more effectively, because it’s trying different techniques, trying to see which one’s most likely to work. So that’s what causes the daisy-chaining to happen in a more effective fashion. That’s why.
So is it a good thing that the average person doesn’t have access to that right now?
I think so. I think every company should have a chance to fix these things in the meantime. But again, I don’t know who the average person is in this case, right? If every company out there is an average person, then they should have access to it—
Because they have to fix their stuff. You mean the average bad person?
Basically, I'm just thinking about all of these cyberattacks that we've seen over the past couple of weeks, and I'm assuming that they do not have access to a Mythos-level model, so I'm asking myself, well, what if they did?
Yeah. If they did, they'll find a way to attack companies much faster.
Yeah.
Right. I don't think the nature of the attacks changes. I don't think the nature of the outcomes changes. Most likely, they will be used to leverage ransomware, perhaps cause economic harm if you're looking at it from a nation-state perspective. So I think the fundamentals of how the bad-actor industry works aren't going to change. What does change is the pace and volume of attacks that are going to be made possible by the availability of these models.
I want to talk a little bit about what, if anything, an average person can do here. I myself am the subject of an ongoing phishing attack.
Somebody must like you.
I mean, I hope so. But basically, almost every day somebody tries to get me to reset my X password from an email address that has nothing to do with x.com. Because I'm looking at my emails on the desktop, that's very easy for me to see, and I'm not fooled. Congratulations.
That's me. I've been trying to steal your bank—
I mean, how could you? But I also believe that within 6 months or a year, one of those emails is going to come in, and it's just going to look way more convincing, right? It's just going to figure out a way to trick me. One of my frustrations with talking about cybersecurity in general is that it tends to leave people with the sense of, “Well, everything's really bad. Sorry. Good luck to you.” Usually, we give people advice like, “Create a strong password” and “Use multifactor authentication.”
That's right.
Is that good enough, or do people need to update the playbook?
Look, I think one of the things that has always frustrated me is that, if you think about it, we have much better cybersecurity solutions in the enterprise world than we do for consumers. For example, if you had a corporate email and all the phishing attacks came to your corporate email, we'd be pretty good at sussing these out, because if we see the X email address you're talking about—which isn't actually X—at 1 customer, we'll block it everywhere else.
Now, the problem with the consumer world is that it doesn't have any such gatekeepers, right? We're effectively the gatekeepers of the enterprise, but the consumer world doesn't have a gatekeeper. The consumer gatekeepers are the email providers. The consumer gatekeepers are the telecom networks that provide our mobile service. If you were getting an attack on your corporate mobile device and we were sitting in front of it, it wouldn't happen. But on our personal devices, we can all get spam, we can all get phished, and we can all have all this stuff happen to us.
I think part of the frustration I have is that there are some consumer companies that need to employ better cyber controls for all of us consumers, which they should be doing, but they're not.
Well, any particular controls come to mind that you'd like to see out there?
Well, think about email, right? Is it hard for the email provider to figure out that this is not an X email address?
Right.
We should. These same guys are building AI, right?
Right. Right.
These guys are building AI that's going to anticipate what we want and do it for us, so somebody just needs to pay attention to it.
For what it's worth, though, this is my paid Google Workspace for my work account. You're absolutely right. It seems like a very simple classifier for Google to make, just to be like, “Hmm, this probably isn't coming from x.com.”
8. AI Changes The Engineering Workforce
How are your engineers feeling about all this? I imagine they're working a lot these days. Are they excited because there's this new set of tools available to them? Are they stressed out because, all of a sudden, their workload just got 5 times bigger? What is the mood?
Yes.
All of it.
All of the above. Look, if you think about it, if you're a technologist, this is a phenomenal time to be doing this, right? There's so much opportunity to learn and to understand. Some people are fearful, asking, “How is this even going to work?” And then you can find, I think, every emotion you can think of in probably every engineering team out there.
We have 9,000-plus technical people. I think it's not just the tool in front of us; I think it's the uncertainty of what this holds in the next 2 or 3 years. People are seeing OpenClaw being deployed. Now, OpenClaw is a scary thing from a security perspective. It's going to take all your permissions, all your credentials, and do all kinds of stuff for you, but it's cool.
The early adopters are doing cool shit. I had dinner with somebody who came to my house. He said, “I got OpenClaw on my phone. It's doing everything. I've given it a name. It's called Zara, and it's doing all the things I'm asking it to do.” And the guy sitting next to me says, “Holy shit, that's a security nightmare.”
Yeah.
You're worried about your xAI, X, you know, posting to change your password. You don't need to change your password. OpenClaw is just going to tweet on your behalf because it's had a moment last night.
Totally.
Right?
Yeah.
Yeah, and for all of my objectionable tweets over the years, I would like to formally say that was my OpenClaw acting autonomously.
There we go. So are you personally running any of this insecure stuff? Are you running OpenClaw? Are you experimenting with this stuff just from a “I need to understand the landscape” perspective?
On a segregated device that has no connection—
Yeah.
—to many of my things, which makes it totally useless, by the way.
Yeah. Yeah.
It's like she can't even book a meeting in my schedule because she doesn't have access to my schedule. It can't respond to an email on my behalf because it doesn't have access to my email. So I'm still using it the old-fashioned way, which is using Gemini in the meantime.
I did do that. I took my earnings script, sent it to Gemini, and said, “What do you think?” 2 quarters ago. It said, “Are you trying to hide something? You're too enthusiastic. You use the words ‘momentum’ and ‘excited’ much more than you normally do.”
Wow.
I was like, “Holy shit.” That's not bad. So I have to tone it down.
Yeah. That's very funny. Is it changing your hiring plans at all?
Yes.
I mean, you employ thousands of cybersecurity engineers—
Yes.
—and researchers.
Yes.
You may need fewer of those people in the future, or…?
No. I need more. I think this is the fallacy out there, right? The fallacy is that organizations are going to get 30%, 40%, 50%, 60% more productive from a development perspective and a testing perspective, so we need fewer people.
The problem is, every technologist that you talk to has a feature-request list that's longer than their arm, and typically, people have product road maps that are 6 to 12 months out. Why is that? Because they don't have enough people, or they cannot serialize something because it takes a lot of effort to get it done.
So I think the first thing that's going to happen is, as we create more capacity, we're going to try to fill the technological backlog and try to make that work. I do understand there are people out there reshaping their technical organizations by creating capacity.
Everybody who's out there saying, “I'm reducing my headcount by 7% or 15% or 20%,” which you're beginning to see recently, I think they're just creating capacity. They're saying, “That capacity allows me to hire more people—
Mm-hmm.
—and make room for people that I need who have the newer skill set.”
Hmm. They're not just spending that salary money on tokens instead.
Look, I think the interesting part is, I was saying this earlier—I was speaking somewhere else—and the part we don't realize is that we're dealing with a tsunami of a desire to transform. I think we're in a decade-long transformation of business ahead of us.
Imagine, you have a new technology. My CFO would never come and say, “I want to use AI to transform my team.” He wants to transform his team and see if he can do it much more efficiently, but he wants AI. My head of HR wants AI because she wants to create an AI interviewer, an AI assessor, instead of having humans do it. So every function wants more AI to deploy.
Now, the question is, where's the money going to come from? It's probably going to come from efficiency in those teams, in those functions, so that's what's going to pay for the tokens.
I have to say, I don't think anyone wants to be interviewed by the AI assessor. That's not a good vibe, you know?
I don't know.
Would you want to be interviewed for a job by an AI?
I think AI is most likely going to be better at assessing my domain skills than a human being.
Really?
Yes. If you're trying to hire a good coder, if you're trying to hire somebody who knows agent AI really well, sitting and talking to them isn't going to get me a better answer if they can sit and code and deploy OpenClaw in front of me.
I've literally done that interview. The guy says, “Well, I'm really conversant with AI.” I'm like, “Really? That's cool.” I'm like, “What have you done for us?” He's like, “Well, I built myself an agent.” I'm like, “Show me.”
“What do you mean?” “You're on Zoom. Show me.” They show you this bizarre, simplistic thing, like, “Oh, I got her to make a shopping list from the recipe I saw.” I'm like, “Dude—”
It's an AI girlfriend. I actually shouldn't show you this.
That could be true.
Yeah.
Yeah.
I said, “Well, now we have an HR problem.”
Well, Nikesh, thanks so much for coming in. Really great to talk to you.
Yeah.
And good luck out there. Fascinating.
Thank you, Kevin.
Please tell Mythos to spare our families in the coming uprising. When we come back, it's time for the Hot Mess Express.
9. The Hot Mess Express
Well, Casey, we've got a train to catch today. The Hot Mess Express is here.
Hot Mess Express.
The Hot Mess Express is, of course, our segment where we take a look at the various calamities befalling people in and around the tech industry and, at the end of discussing them, decide what kind of mess this was.
What's pulling up to the station today?
Well, let's see what's first here on the tracks.
You just love the sound effect.
Our first story today comes from The Verge. Oh, and this is truly the end of an era. Venmo is starting to test a big redesign of its app, and as part of the changes, it will be implementing a major new privacy feature. The onboarding process for new users will set their posts to only be viewable by their friends by default instead of being public.
This is very sad for me because for years now, every time I've opened up Venmo to pay a friend, I've seen a recent transaction from someone I hooked up with once in 2016. The thought that other people aren't going to have that experience makes me really sad.
As a nosy person who loves to gossip, I am sad about this story because it was always fun to see which of your random phone contacts had been paying their fractional share of the rent or paying back for dinner. People put various jokey things on their transactions: “illicit drug deal,” “foreign arms trade,” et cetera. It's just sad that we won't get to experience that.
Yeah. The public-by-default Venmo transactions also gave us many great stories over the years, including Joe Biden's secret Venmo, which was a BuzzFeed story. J.D. Vance had a public Venmo that Wired reported on. Matt Gaetz's Venmo payments were part of a federal inquiry into his payments to women, according to The New York Times.
I guess all of us investigative reporters are going to have to find a new easy way of writing a story, Casey.
Yeah. Now the only baffling security breach from these apps is that Telegram still notifies you when one of your phone contacts joins. I always love to screenshot that and send it to people and then be like, “Crypto or drugs? What is it this week?”
The only 2 possible answers. So what kind of mess is this Venmo mess?
This is unfortunately a cleanup, not a mess.
Yeah.
This used to be a very hot mess, and now, belatedly, it is getting cleaned up.
Fair enough. RIP. Let's see what's else coming down the tracks.
Oh, well, this was interesting, Casey, and ties in closely to something that you've written about recently. Amazon has started to widely deploy its in-house Meshclaw product in recent weeks, which allows employees to create AI agents that can connect to workplace software and carry out tasks on a user's behalf. But some employees are saying that colleagues are using the software to automate additional unnecessary AI activity to increase their consumption of tokens, which will then, of course, make them look better to their bosses.
So, did we see that one coming, or what?
Yeah. Yes. I believe you invoked Goodhart's law about what happens when a measure becomes a target.
When a measure becomes a target, it ceases to be a good measure. That is, of course, Goodhart's law.
Thank you so much for that. Yes, and I imagine that at the famously frugal Amazon, they are loving this era of people just spending a bunch of random tokens to move up the leaderboard.
Here's the thing. I've talked to a lot of Amazon employees over the years. Tokens are the only thing at that company that is free. You want a Diet Coke from the vending machine? Get out your wallet, okay? So these guys finally find something free, and now they're getting in trouble.
Yeah. The good news is they have unlimited tokens. The bad news is they can only use them on Meshclaw.
Yeah, I'm going to say that this is actually a hot mesh.
Yeah.
That's what kind of a mess this is.
Very good.
All right. Next up, Casey, this comes to us from 404 Media, and boy, did I see this clip in about 14 different places over the past week: “Students boo commencement speaker after she calls AI ‘the next industrial revolution.’” You see this one?
Yes.
Yeah. On May 8, commencement speaker Gloria Caulfield, who is the vice president of strategic alliances at Tavistock Group, told graduates of the University of Central Florida's College of Arts and Humanities and Nicholson School of Communication that AI is the next industrial revolution. She was met with thousands of booing graduates, and someone in the crowd, Casey, yelled, “AI sucks.”
What did you make of this commencement moment?
Here's my thing: students are allowed to feel however they want about AI.
Yeah.
But if you boo the commencement speaker for suggesting that AI is a big deal, I want to see your ChatGPT history. If you've used AI to write your exams, to help you with your problem sets, in any way for your academic work, you are not allowed to boo it at commencement. That is my rule.
I don't know. I think these students were fine to boo. Ms. Caulfield was, after all, addressing the College of Arts and Humanities, which I'm guessing is probably not the group of students at the university that is most excited to see AI come into their lives.
So here's the thing I'll say that is sincere. I think people are radically underestimating how mobilized young people are against AI right now. I see this every time I go to a college to talk to students. There's a small group of them who are running OpenClaws and very excited, and 80% of them are like, “I hate this.”
Yeah. So look, if you have to give a commencement speech within the next few months—a highly relatable situation that many of our listeners will be in—now you know.
Yeah.
Careful how you talk about AI.
Yeah.
Okay, Casey, our next story comes to us from the good folks at Variety. Dua Lipa has filed a $15 million lawsuit against Samsung for using her face to sell TVs, and this one is honestly pretty incredible.
Samsung has apparently used Dua Lipa's image on the cardboard packaging of its TVs starting last year. When Ms. Lipa became aware of it, she demanded that the company stop using her image and apparently could not get through to anyone at Samsung. Samsung finally responded on Monday and said this was all the fault of some third-party content partner.
Samsung said, “We have great respect for Ms. Lipa and the intellectual property of all artists,” and they are actively seeking and remain open to a constructive resolution with Ms. Lipa's team. Well, it sounds like a constructive resolution could be taking her face off the packaging and paying her $15 million.
I understand her concern, because the thing people always forget about Samsung products is that they do explode when you least expect them. There was, of course, the famous series of explosions related to their phones. So if I see my face on a Samsung TV, I'm thinking, “I do not want to be the literal face of an exploding piece of hardware.”
Yeah. What kind of mess is this?
This is a true hot mess because the TV could have exploded.
Yeah.
Now, here, you want to read one?
Okay.
Okay.
All right. This next one comes to us from our colleagues at The New York Times: eBay rejects GameStop's $55 billion takeover bid. Last week, GameStop offered $55 billion to eBay in an unsolicited takeover attempt. According to some interviews, they appeared not to have $55 billion, which would put a damper on their plans.
This week, eBay officially said no to the GameStop offer, calling it, quote, “Neither credible nor attractive.”
And—
Which is also what our last iTunes review of this podcast said.
And there you have it.
This one is an interesting story from the world of what I like to call companies that I can't believe still exist. I don't know what's happening on eBay. I don't know what's happening to GameStop. But what I do know is these companies probably don't belong together, Kevin.
Yeah, I find this fascinating because it is just the Internet-brain CEO of GameStop. He's this guy, Ryan Cohen, who rose to prominence during the meme-stock mania of 2020 and 2021. And now you can just do whatever you want. If you're the CEO of a company, you can just say, "We're going to buy a company that's five times bigger than us." How? Shame on you for asking.
I mean, is it unreasonable, given their history, to expect that they could have announced this, and GameStop stock could have gone through the roof, and all of a sudden they would have had $55 billion to buy eBay?
Yeah.
But that didn't happen.
Well, if they had done this deal in typical GameStop fashion, they would have offered about half of what the market value for eBay was—because it's used and probably doesn't even work on your console anymore.
I like jokes that you'll only get if you've returned a video game to GameStop.
Listen, for our younger listeners, there used to be a time when you could walk into GameStop with a box of old video games that you wanted to get rid of, and they would offer you between 50 cents and $1 for each video game.
All right, this is the sort of mess where we're explaining the joke.
Okay.
Okay. So we've got a few more items, Kevin. Shein and Temu are fighting it out in U.K. courts, as Shein has accused Temu of, quote, "astonishing levels of copyright infringement," and Temu accused Shein of waging, quote, "an aggressive and relentless battle using copyright allegations to undermine competition." This comes to us from Bloomberg.
The whole trial revolves around thousands of photographs that Shein says are from its website. According to Shein's lawyers, Temu sold identical clothing items using the same images and is seeking to piggyback off Shein's own investment in building up its supply chain and training and upskilling suppliers. What do you make of this fight?
The fast-fashion brands are fighting.
They're fighting.
There's no one I'm rooting for in this fight. I've never bought an item of clothing from either of them. But it is very funny that two of the brands that have made their entire existence out of ripping off clothing from more established purveyors are now fighting each other about which one's ripping off the other one.
Yeah, truly a situation where—Is there a way they both could lose—
Yeah.
—and learn a hard lesson—
Yes.
—about intellectual property.
Yes.
We're rooting for them. Next up—favorite story of the week, Kevin, and I imagine you heard about this one. People are seriously pissed that Grindr outed them with its latest Madonna ad. Did this happen to you?
No.
Okay. So this issue stems from the fact that Madonna has been doing this big campaign inside Grindr to promote her upcoming album, Confessions on a Dance Floor 2, which is a concept album about a 68-year-old woman who still wants to be at a nightclub after midnight. She's advertising on Grindr, and apparently, over the past week, when you opened up Grindr, even if you had your phone volume turned off, you would hear the sound of Madonna saying loudly, "Hi, Grindr, it's Mother."
No.
First of all, it's Grandmother. Sorry. Second of all, apparently, people who were not out to their families were opening Grindr at the dinner table—which, you're already putting yourself in harm's way there, maybe. But the last thing they expected was to have Madonna saying, "Hey, look at this guy. He's on Grindr right now." So, truly, one of the most misconceived ad campaigns in recent history.
Wow.
Yeah.
That's so wild. It's like if they put U2's Songs of Innocence on your phone, but it just outed you to your family.
Yeah. The song was "You're Gay." That was the song. Here's the thing: This is a dangerous mess. It is not always safe for people to be outed to people in their immediate surroundings.
Yes.
So shame on Grindr. They really should have known better.
Yes. Push notifications should be illegal.
All right, and one more car coming down the train tracks here, Kevin. This is from the Elon-OpenAI trial this week. Sam Altman was on the witness stand Tuesday and testified that, at one point, Elon thought he should run OpenAI. Sam asked him, "Hey, what do you think would happen to the company if you died?" And according to Sam, Elon replied, "I haven't thought about it a ton, but maybe control should pass to my children?"
Question mark, question mark.
Question mark, question mark. So what do you think? Let me just ask it this way: Do you think we would be better off if OpenAI was a hereditary monarchy controlled by the Musk clan?
I do. You know, we always talk about what is the ideal governance structure for AGI.
Yes.
I think we can all agree that it would be best if Elon's 27 children were involved somehow.
Yeah, or they just pick one at random, and one is probably—I don't know—11 years old and rides a skateboard around town. They're like, "All right, kid, you run AGI now. Best of luck." So, yeah, that continues to be a legal mess.
Yeah, the whole trial has been fascinating to me, less because I care about the actual legal issue on trial and more because it has produced all these amazing and incriminating files from the early days of OpenAI, including all of their texts and emails and messy dramas. I live for it.
Yeah, look, it's very hard to run a successful company without a lot of executives saying a bunch of really stupid things and writing them down. We see it over and over again.
Yeah.
So let that be a lesson to us.
Yep. Hot mess.
And that is it for the Hot Mess Express. Thank you to all of this week's passengers, and best of luck with your messes. Try to stay on the right side of the tracks. Hey, before we go, one request. We want to hear what it's like for people who are undergoing major career changes in response to AI. So for example, if you have recently left a computer or desk job to do something more manual, like HVAC installation or tree trimming, we would love to hear how it's going. So anything in that realm, please send us an email. We would love for you to share your story with our audience. Our email, again, of course, is hardfork@nytimes.com. Tell us about your career shift and why you're making the change. Hard Fork is produced by Whitney Jones and Rachel Cohn. We're edited by Viren Pavich. We're fact-checked by Caitlin Love. Today's show was engineered by Chris Wood. Original music by Alicia Meitube, Rowan Nemestcho, and Dan Powell. Video production by Jake Nickell and Chris Schott. You can watch this whole episode on YouTube at youtube.com/hardfork. Special thanks to Paula Schumann, Pui Wing Tam, and Dalia Haddad. You can email us at hardfork@nytimes.com with what you would do with Mythos if you could.