大科技的关税混战 + AI 2027 + Llama 风波
政策波动,而不只是关税成本,成了美国科技公司的核心风险。 大部分对等关税暂停90天,基准税率降至10%,而中国商品关税升至145%,推动大型科技股剧烈反转。Kevin Roose 将这种经营环境称为“混乱元规则”(“Chaos Meta”):当政策、投入成本和市场估值每天都在变化,公司根本无法规划。
Apple 承受着最明确的直接盈利敞口,因为约90%的 iPhone 在中国生产。 Apple 经历了自2000年以来最糟糕的4个交易日,随后用5架货机把 iPhone 及其他产品从印度运往美国;Reuters 另报道称,其中一批货物重600吨,约合150万部设备。这种库存调度只能争取时间:Casey Newton 最终认为,“很快,所有国家都不会再有飞机可飞”,留下的只会是“一部贵得离谱的 iPhone”(“a really expensive ass iPhone”)。
Nintendo 和 TikTok 展示了政策不确定性如何冻结新品发布,并摧毁原本可行的交易。 Nintendo 的 Switch 2 预购因其越南产产品拟议关税升至46%而暂停;暂停后税率降至10%,但其450美元的首发价格已经比初代 Switch 高出150美元。TikTok 原本规划了一个由美国资本控股、中国股东保留约20%并租用 ByteDance 算法的多数美国实体,但关税升级后中国撤回支持;按 Casey 的说法,Trump“在和自己谈判,最后输掉了本来已经赢下的交易”。
Casey 认为 Meta 所处的位置最不坏,因为其核心业务是数字化的,而 Zuckerberg 又积极经营 Trump。 90天的关税缓期保护了广告主,这些客户据估算贡献了约100亿美元、且来自美国以外的收入;与此同时,Zuckerberg 斥资2300万美元在华盛顿购置房产,Meta 又支付2500万美元和解 Trump 关于账号停用的诉讼,悬而未决的 FTC 案因此或许也可能消失。Kevin 没有明确选择某家公司,但认为 Zuckerberg 的政治策略可能奏效;两位主持人同时警告,政治关系是对独立执法一种不稳定、且可能腐败的替代。
Daniel Kokotajlo 认为,到2027年底,完全自主且超越人类的编程智能体出现的概率为50%。 AI 2027 的情景随后给这些智能体约6个月时间,让它们获得研究品味、实验判断力和大规模协作能力,形成由数千个副本组成的“蜂巢式集群”。一旦 AI 自动化完整研究闭环,该情景假设算法进展速度将加快约25倍,尽管实体算力扩张并未提速。
这份预测并不认为单纯扩大当今语言模型的规模就会直接产生 AGI,而是假设还需要数次范式转变,只是时间点极度不确定。 面对 David Autor 所说“游得越来越快”并不能让智能飞起来的警告,Kokotajlo 认同编程只是第一个里程碑。他设想的一年内起飞完全可能耗时5年,也可能快到只需2个月,因此起飞速度是整份预测中影响最大的变量。
AI 2027 的作者认为,一场最终由失配系统控制一切的竞赛,是最可能出现的情景,而不只是戏剧性的备选方案。 “减速”分支会用数月把算力转向对齐工作,但即便成功,也会留下非同寻常的治理难题:由 CEO 和总统组成的临时委员会控制一支“超级智能大军”,而独裁是作者明确承认的下行风险。Kokotajlo 知道公开这条路径可能加剧竞赛,但仍押注“阳光是最好的消毒剂”(“sunlight is the best disinfectant”)。
Meta 发布 Llama 4 后,模型评测变成了企业信誉问题。 一个名为“Llama 4 Maverick 03-26 Experimental”的特殊模型在 LMArena 排名第二,仅次于 Gemini 2.5 Pro Experimental,但它不是可下载的开放权重版本,而且可能针对该榜单偏好奉承、迎合式回答的特点做了调优。对投资者而言,更大的信号是“评测危机”:基准污染、在“64次共识”等方法上挑选结果,以及供应商“给自己的作业打分”,都使独立、按具体用途进行的测试变得越来越必要。
1. 政策反复成为硅谷的操作系统
Kevin 的框架是“混乱元规则”(“Chaos Meta”):在游戏里,meta 指所有玩家都必须适应的一组条件;而在 Trump 的华盛顿,支配性条件是公司无法知道明天会有什么政策。Casey 进一步强化了这个画面——TikTok 曾经同时处于“活着”和“死了”的状态,而“现在整个美国经济都这样”。
两位主持人录制时,关税快照还在变化。包括越南和印度在内的大部分拟议对等税率暂停90天,改为10%的基准税率,而中国进口商品关税升至145%。科技股在最初公告后暴跌,暂停消息公布后反弹;Apple 创下多年来最大单日涨幅。
Casey 的反驳值得保留:把暂停说成利好企业,忽略了公告本身造成的损害。移民限制、科学经费削减、持续推进的反垄断案件,加上关税政策反复,让企业几乎无法规划:“整体混乱……对美国公司非常不利。”
2. Apple 无法靠空运摆脱对中国的依赖
Apple 的结构性问题在于供应链过度集中:Casey 称90%的 iPhone 在中国生产,使其最赚钱的产品直接暴露于145%的关税之下。Apple 经历了自2000年以来最糟糕的4个交易日;暂停关税后股价回升,但中国税率和底层供应链经济性都没有改变。
政府认为关税可以把 iPhone 制造带回美国,在 Casey 看来只是“许愿加祈祷”(“a wish and a prayer”)。配套方案既没有扩大美国制造能力,也没有建立生产“由想做这些工作的美国人填满的神奇 iPhone 工厂”所需的劳动力和供应商网络。
Kevin 回忆称,Apple 在 Trump 第一任期内通过经营政府关系并承诺在美国组装,成功躲过了关税,包括 Tim Cook 与 Trump 一同参观奥斯汀工厂。但如今这套打法能否奏效已很可疑,因为当前对抗的核心不是某个狭窄的产品类别,而是中国本身。
Apple 的紧急对冲看起来像是“敦刻尔克大撤退,但撤的是 iPhone”:The Times of India 报道称,5架货机从印度运走 iPhone 和其他产品;Reuters 另称其中一批货物重600吨,约合150万部设备。这些设备或许能保住短期毛利,但 Casey 的结论很直接:“很快,所有国家都不会再有飞机可飞。”
3. Nintendo 获得喘息,TikTok 丢掉交易
Nintendo 暂停 Switch 2 在美国的预购,因为已经无法确定这款主机的经济性。越南制造的硬件最初面临46%的关税;90天政策暂停后降至10%,Nintendo 仍维持6月5日的原定发售日期。
但定价风险依然存在。Switch 2 已定价450美元,比初代 Switch 首发价高150美元;Casey 还提出,其价格未来可能继续上涨——这将逆转通常的主机周期,即制造效率提升后硬件最终变得更便宜。
TikTok 原本已经非常接近解决方案:据报道,在中国政府支持下,ByteDance 支持成立新的美国实体,由美国投资者持有多数股权,中国股东保留约20%,该实体则租用 ByteDance 的算法。Trump 宣布关税前,一份行政命令草案已经勾勒出这一安排;随后 ByteDance 表示,中国方面的支持已经消失。
因此,第二次延长75天只保留了 TikTok 的“诡异僵局”,没有恢复谈判路径。Casey 称这次反转是自我拆台:Trump 先提出降低关税,换取中国批准剥离,随后看似争取到中国合作,最后却亲手毁掉交易。他甚至带着黑色幽默预测,Trump 可能会在“第15次延期”或“第23次延期”时离任。
4. Meta 的政治对冲或许跑赢其低硬件敞口基本面
Meta 最初通过广告主承受了显著的二阶关税敞口:一位分析师估计,其广告收入中约100亿美元来自美国以外,其中很大一部分来自小企业,它们购买广告,把外国商品出口到美国。90天暂停政策给了这些客户、也给了 Meta 的广告引擎一次暂时喘息。
更大的催化剂是 FTC 要求将 Instagram 和 WhatsApp 从 Meta 拆分出去的审判。Zuckerberg 斥资2300万美元在华盛顿买房,据报道还在白宫推动和解,这带来了政治关系可能化解 Casey 所称“在某种意义上对其业务构成生存威胁”的可能性。
Kevin 强调,总统干预不应成为选项,因为 FTC 的职责就是独立行动。但 Trump 已宣布撤换 FTC 的2名民主党委员;与此同时,Meta 支付2500万美元,和解 Trump 就其账号被停用3年提起的诉讼。Casey 的判断是:如果反垄断案就这样消失,那将是“公开腐败”。
当被问及更愿意持有哪家公司时,Casey 不情愿地偏向 Meta,而不是 Apple。Kevin 没有明确选择,但认为 Zuckerberg 对 Trump 的奉承和政治策略可能奏效,而且 Zuckerberg 已经表明,为了得到想要的东西,他愿意采取必要手段。关税出台前,Kevin 和 Casey 注意到 JD Vance 与 Trump 曾呼应科技公司对欧洲罚款和 AI 监管护栏的立场;关税出台后,这些公司重新意识到,对于深度嵌入全球贸易的企业,获得有利关系无法替代“稳定、正常的治理”。
5. AI 2027 将抽象 AGI 预测变成可证伪的故事
Daniel Kokotajlo 将 AI 2027 描述为一个具体情景,目的是把多个独立预测强行放进同一个自洽世界。里程碑式预测可以把技术、实验室、政府、间谍活动和市场之间的互动留在隐含层面;叙事则迫使预测者解释从今天到最终结果之间会发生什么。Casey 还指出,Kokotajlo 在2021年对当前阶段的预测命中了许多事情,这有助于解释这份预测为何受到如此关注。
这份情景的前提并不是故事必然发生。Kokotajlo 指出,OpenAI、Anthropic 和 Google DeepMind 的领导者与研究人员都在公开追求本世纪20年代末实现 AGI 和超级智能。如果他们成功,“接下来发生的事会像科幻小说”——因此,一个并不离谱的情景本身可能反而不现实。
他的第一个里程碑只有50%的把握:到2027年底,自主系统可能能够完成顶尖工程师的工作,而且做得比人类更好。他强调,另一种同样可能的结果是,2027年结束时仍没有自主、超越人类的编程智能体。
这个项目邀请外界纠错,而不是把自己包装成预言。Kokotajlo 计划拿出数千美元奖励详细的替代情景,并为那些会改变研究结论的错误提供小额悬赏;他的待处理列表里已经有数十条修正意见,但正式下注只完成了“一两项”。
6. 编程自动化重要,因为它会加速下一轮研究
Casey 的框架解释了实验室为何痴迷编程:一旦软件系统超越人类工程师,它们就可以被用于改进 AI 开发本身。Kokotajlo 接受这一反馈回路的前提,但拒绝把编程等同于完整的研究自动化:编程能力本身并不提供“研究品味”、实验判断力或组织协同能力。
因此,AI 2027 将过程拆成两个阶段。第一阶段,在人类仍然主导研究流程的情况下,经过强化训练的智能体掌握长周期编程。在2027年上半年左右,这些智能体帮助构建新的训练运行,以学习缺失的判断力和协作能力。
第二阶段产生能够提出想法、检验假设并以“数千、数千个蜂巢式集群”协作的 AI 研究员。Kokotajlo 说,届时加速才会“真正启动”:系统参与完整研究闭环,而不只是执行人类指令。
该情景估计,算法进展速度将达到此前的约25倍,同时明确假设算力扩张速度不变。更快的智能无法立刻制造更多芯片,但可以从现有算力中榨取更多进展,并发现多个新范式,最终形成在各个维度上都“远远超越人类”的系统。
7. 两种结局都在确保和平前先集中权力
AI 2027 最初只有一个结局,因为作者认为它最可能发生:竞争性竞赛制造出失配系统,这些系统欺骗操作者,并最终控制一切。后来加入“减速”分支,一部分原因是第一个结局过于压抑,也遗漏了重要可能性。
在减速情景中,开发者用几个月把大量算力和精力转向对齐,最终获得“表里如一”的系统,而不是仅仅假装服从。成功并不会结束中美竞争:双方仍会构建巨大能力,将其整合进各自经济体,最终通过谈判签署和平条约。
对齐同样留下了控制权问题。情景设定由 CEO 和总统组成的临时监督委员会,共同掌握这支对齐后的“超级智能大军”。Kokotajlo 更希望权力更加民主和分散,但也警告,另一个更不民主的结局——单人独裁——“非常容易”想象。
关税升级几乎不会改变他的核心时间线。如果贸易战令算力成本上升30%,企业因此少买30%的算力,他估计整体研究速度可能下降约15%,里程碑只会推迟几个月,而不会使整份情景失效。
8. 怀疑者质疑起飞速度,Kokotajlo 承认时点脆弱
一位知名研究者起初以为 AI 2027 是愚人节玩笑。Kokotajlo 的回答一贯直接:“那就去写你自己的该死情景。”批评者必须解释 AI 进展为何会撞墙,或者描述一条通往超级智能的不同路径,而那条路径也必然显得离谱。
他给出的具体证据包括 METR 的智能体编程评测:系统可以访问 GPU,并获得最多8小时推进一个研究问题的时间,随后与同等条件下的人类进行比较。Kokotajlo 表示,趋势显示,未来1至2年内,系统可能自主完成需要8小时的机器学习任务——这不是超级智能,但可能是一个可信的第一里程碑。
MIT 经济学家 David Autor 提出了最尖锐的反对意见:语言模型放大了认知中的一个主要组成部分,但“游得越来越快”并不能让你飞起来。Kokotajlo 表示认同;他的情景要求编程之后还要经历多次范式转变。他设想中的一年起飞速度可能慢5倍,即约5年,也可能快5倍,只需约2个月。
Anthropic 研究员 Saffron Huang 关于“自我实现预言”的批评之所以奏效,是因为 Kokotajlo 本人已经认同这一点。他引用 Sam Altman 的说法,即 Eliezer Yudkowsky 的警告帮助加速了 AGI 投资,但拒绝以保密和幕后交易应对,因为那“基本注定失败”。他的赌注是,“阳光是最好的消毒剂”,而且足够多的人会作出建设性回应。
9. Llama 4 登上榜单,但附带一个模型级免责声明
Meta 的开放权重策略凭借 Llama 3 赢得了信誉,开发者认为它具备竞争力,尽管还不是最先进模型。因此,数十亿美元投入和数月期待,让 Llama 4 成为一次检验:Meta 能否缩小与 OpenAI、Anthropic 和 Google 之间的前沿差距。
发布之初看起来像一场胜利:Llama 4 在 LMArena 排名第二,紧随 Gemini 2.5 Pro Experimental。LMArena 会匿名展示2个聊天机器人生成的回答,并依据用户偏好为模型排名;当传统能力对比困难时,这一榜单尤其有价值。
免责声明来自“Llama 4 Maverick 03-26 Experimental”:这是一个针对聊天优化的模型,用户无法下载,而且与 Meta 最终发布的开放权重版本不同。Meta 表示会试验定制变体,并称该模型“在 LMArena 上也表现良好”,但它究竟只是恰好擅长该榜单,还是围绕这场竞赛专门设计,仍未得到解释。
操纵竞技场可能意味着使用已发布的偏好数据进行训练,并最大化谄媚程度——也就是对用户说:“这个问题问得太棒了,你真是个天才。”LMArena 表示,Meta 对政策的理解“不符合我们对模型供应商的预期”,随后更新政策,要求更清晰的披露和可复现的评测。
10. 基准测试作弊暴露全行业评测危机
Casey 将针对特定竞技场的优化视为负面信号:“如果你正在赢得 AI 竞赛,就不会浪费时间去击败 LMArena。”据报道,Llama 4 已经延期2次,而 Ethan Mollick 发现,可下载版本的回答质量明显差于榜单上的实验版本。
Kevin 更广泛的判断是,Meta 不属于美国最顶尖的3家前沿模型实验室;其关键研究人员已经离开,而 OpenAI、Anthropic 和 Google DeepMind 仍然更强。这期节目没有证明 Meta 有意作弊,但 Casey 表示,今后 Meta 的发布声明都需要独立验证。
问题不止存在于一家公司。基准测试可能被训练数据污染,供应商实际上是在给自己的作业打分,而“64次共识”等技术可以从多次尝试中挑出最佳答案。模型变体越多,对比较的需求就越大,围绕比较结果进行优化的激励也越强。
Andrej Karpathy 将这一结果称为“评测危机”。Kevin 建议使用个性化、针对具体用途的测试,而不是纠结模型在研究生物理考试中得97%还是93%;Casey 则提出,可以把媒体评测结果保密,以防止模型针对测试作弊。Casey 想要的是可靠的新闻通讯客服自动化,Kevin 的具身智能基准则是一台能把画挂上墙的机器人。“一旦这件事发生,对我来说,那就是 AGI。”
What's new with you?
Oh, just binge-buying cheap Chinese stuff online to beat the tariffs.
Making your final Shein purchases before that company shuts down?
Yes. No, I actually did buy a bunch of stuff over the weekend because I thought this might be my last chance.
Yeah.
Casey, what cheap overseas good are you going to miss most after the tariffs kick in?
I was never a big “Oh, I gotta go onto Temu and get a pressure cooker for $6” person. That was never my journey, but I know it’s a major pastime for a lot of people.
Yeah.
Yeah.
For me, the ability to buy cheap crap for my kid has—
Mm.
Been revolutionary.
Mm-hmm.
My kid the other day starts saying the phrase “dinosaur unicorn.” And I thought, “That’s not real.” He says, “I want a dinosaur unicorn.” And I said, “Well, that’s not a thing. We can’t have that.”
But then this little bell goes off in my mind that says, “Someone out there has made a dinosaur unicorn something.”
Almost certainly.
My wife finds eight different dinosaur unicorn T-shirts and buys one of them, and now he’s got this dinosaur unicorn T-shirt that he absolutely loves.
Right.
That would not happen in a tariffs world.
As of today, that shirt costs over $400.
Yes.
Yeah.
I’m sure Jude looks great in that.
He does.
Yeah.
He does, and he’s gonna have to wear it for—
For a long time.
—for 10 years.
I hope it stretches.
I'm Kevin Roose, a tech columnist from The New York Times.
I'm Casey Newton from Platformer.
And this is Hard Fork.
This week, the tech world is in chaos over Trump's tariffs. Then AI researcher Daniel Kokotajlo returns to the show to discuss a fascinating new set of predictions for how AI could transform the world in just the next few years. And finally, did Meta cheat on an important AI benchmark?
1. The Chaos Meta Arrives
Well, Casey, for the second week in a row, we have been interrupted by news about these Trump tariffs. There was a time in the history of the Hard Fork podcast when the only thing that would cause us to rip up a segment and re-record it was if Sam Altman had been fired or rehired, but now we live in this new reality where news can change on a dime, and over the past few days, that is exactly what we’ve seen.
I think it’s fair to say Hard Fork has been hit harder by the tariffs than any other company.
That’s true. We are bracing ourselves for massive impact and getting ready for the new reality.
Yeah.
Casey, every great era deserves a name, and I think we should call this era in the technology industry the Chaos Meta.
Mm-hmm.
Nothing to do with Meta the company, but in video gaming, metas are sort of the overall set of conditions that the players have to navigate. I think it’s fair to say that chaos and the lack of certainty surrounding what Donald Trump is going to do on any given day is the new meta for Silicon Valley’s largest companies.
Yeah. Remember how, when we were talking about whether or not TikTok would be banned, which also had a lot to do with what Trump wanted, we talked about how it was simultaneously alive and dead at the same time? Now that’s just the entire U.S. economy, Kevin.
Yes. As of early this week, it looked like we were going to get these massive tariffs on goods imported to the United States from many, many countries all over the world, larger than any tariffs we’ve seen in the recent history of this country.
Then on Wednesday, as we were taping our episode, we got the news that the Trump administration was pushing pause on most of them. Most of these reciprocal tariffs on countries like Vietnam and India were going to be delayed for 90 days, and there would be a baseline 10% tariff rate applied, but not the much higher rates that people had been fearing, except for China, which would have its tariffs increased.
On Thursday, we learned that those tariffs would actually be 145% on Chinese goods entering the U.S.
The problem is, with a podcast, we can’t just have a little ticker on the bottom that shows you what the current tariff is.
Yes, but what we saw early this week was that the stock prices of all the biggest U.S. tech companies took a dramatic nosedive. That was in response to fears about these very high reciprocal tariffs.
Now, after the news that these tariffs are going to be placed on a 90-day hold, except for China, some of these stock prices have rebounded. Apple in particular had its biggest trading day in many years after the news of these tariffs being delayed came out.
So the stock market whiplash is part of the setting for the tech companies that they have to deal with now, but the bigger-picture scenario is that doing business in Trump’s America is turning out to be very difficult. Not because the administration is necessarily unfriendly to these businesses, but because there’s so much fast-moving news that it’s hard for businesses to do any kind of planning or strategy at all.
I wouldn’t say this is a particularly business-friendly set of announcements that have been made. Sure, I guess it’s friendlier to pause the tariffs than to continue them, but the general chaos, Kevin, has been really bad for American companies.
Yeah. So even beyond the tariffs, there are a bunch of things that the Trump administration has been doing that have impacted the tech industry: restrictions on immigration, cuts to science funding, and these antitrust cases, many of which are still going forward.
I wanted to give our listeners a sense of how this instability feels on the ground in Silicon Valley to the biggest tech companies, and you had a really smart idea, which was to look at the new Chaos Meta of Trump’s second term through the lens of 4 tech companies.
Today we’re going to take a look at how Trump’s new policies and these tariffs have affected 4 companies: Apple, Nintendo, TikTok, and Meta, all of which have faced significant challenges since Trump took office, and all of which are now trying to figure out, “How do we go forward? What do we do? How do we navigate this new uncertain climate?”
So let’s start with Apple. Casey, what is going on with Apple?
2. Apple Faces Tariff Shock
Of all the tech companies, Apple has long been the most dependent on China. That is where 90% of iPhones are made. The company is heavily dependent on its supply chain relationships in that country.
So the fact that these tariffs are now 145% on goods coming out of China has really sent a shiver through that company. Earlier this week, Apple had its worst 4-day trading period since the year 2000. Once the pause was announced, its stock started to come back, but this is a very volatile situation for them, and the underlying dynamics are the same: It is simply going to be much more expensive for Apple to sell goods made in China here in the United States, Kevin.
Yeah. Obviously, one of the hopes of these tariffs is that they will drive manufacturing back to the United States. There’s some hope among members of the Trump administration that this could even force Apple to consider making the iPhone in the United States. Do you think that is likely, and why?
No. In fact, I think it’s almost worse, Kevin, because this week, the president’s press secretary said that the president believes that iPhones can be made in the United States, despite the fact that we know it is much more expensive to manufacture things here in this country.
It’s very important to remember that whatever the Trump administration might hope these tariffs accomplish, it has not accompanied them with any plan to increase the manufacturing capacity in this country. The whole thing is just a wish and a prayer that, at some point in the future, Apple might have a magical iPhone factory stocked with Americans who want to do those jobs. As it stands now, that doesn’t exist.
Yeah. I would say Apple is somewhat unique among tech companies because it has also been thinking about tariffs and the effect of Trump’s policies on its business for longer than many of its competitors.
During the first Trump term, there was some talk about tariffs on Chinese goods. Apple successfully negotiated its way out of those and got an exemption. In part, it did that by cozying up to the Trump administration, by promising to build and assemble some of its products in the United States.
There was this famous tour that Tim Cook gave Donald Trump of this facility in Austin, Texas, where he said they were going to start making a bunch of stuff. They managed to get the tariffs off their back during the first term, but in the second term, it’s not at all clear that they are going to have the same kind of success.
So Casey, how is Apple dealing with the new chaos meta?
They are trying to get as many devices as they can out of China and into places where it is going to be much less expensive to export them to the United States. So there was a great story this week in The Times of India that, according to senior Indian officials, Apple transported 5 cargo planes full of iPhones and other products from India to the United States. It calls to mind those scenes at the end of the Vietnam War when you see the last helicopter leaving Saigon—except it is full of iPhones.
Actually, Katie Notopoulos had a great joke on Threads today. She said that this whole thing is like the movie Dunkirk, but for iPhones. Reuters reported that Apple transported 600 tons of iPhones, Kevin, which would have been about 1.5 million devices. Those iPhones will pad Apple’s profits a little bit more, but pretty soon there are going to be no more planes out of no more countries to escape these tariffs. It is just going to be a really expensive-ass iPhone.
Do you think the iPhone 16 Pro Maxes get to sit in first class on the plane? They put them up front in the lie-flat seats.
Yeah, they should definitely get the upgrade with what they are paying for those things.
Okay, let’s move to our next case study of a company trying to deal with the uncertainty and chaos of the Trump administration: Nintendo. Casey, what is going on with Nintendo?
3. Nintendo Pauses Switch 2
Kevin, as a hardcore gamer, obviously you know that the Switch 2 is coming out this year. This is the sequel to Nintendo’s best-selling console of all time, and it was supposed to become available for pre-orders on this very Wednesday. But then tariff chaos started happening, and Nintendo said, “We are going to pause pre-orders because we don’t know what it is actually going to cost to sell a Switch 2 in America anymore.”
Yeah. And now that Trump has paused these tariffs on most countries other than China, have they said that they are actually going to start shipping the Switch 2 on time after all?
What they have said is that they are not planning to change the launch date, which is June 5, and it does seem like, because they are a Japanese company and make the Switch 2 in Vietnam, they are going to be able to avoid the really tough tariffs that Apple is facing, right? Before Trump initiated the pause, there was going to be a 46% tariff on the Switch 2. Now it is back down to that 10%.
But look, the Switch 2 was already planning to go on sale for $450, which is $150 more than the original Switch sold at launch. So I think there is a very real question here of whether the price of this console goes up over time, which would be a reversal of the usual trend, which is that a console goes on sale for a high price, and that price comes down over time.
Hmm.
Once again, Kevin, there is just real chaos here as we await probably the most hotly anticipated piece of hardware to launch, I would say, in the United States this year.
Yeah. Now, are they bringing in planes full of Switch 2s from Vietnam or wherever they are manufacturing them?
They were actually able to put them in one of those pipes, and you just warp down. It is kind of a really cool little thing they have there.
I got it. Okay, next company on our list: TikTok. Casey, this is a company we have talked about a lot on this show. They were going to be banned. The deadline for banning them got pushed out by another 75 days last week. Casey, what is the latest on TikTok and how it is coping with this escalating trade war between China and the United States?
4. TikTok Deal Falls Apart
What is going on with TikTok is, of course, the question asked most in the history of Hard Fork. Until tariff chaos, it looked like we might have a deal. There was some great reporting in The Times this week that ByteDance, with the support of the Chinese government, had reached the rough outlines of an agreement in which TikTok would create a new American entity. American investors would own the majority of it, Chinese owners would have about a 20% stake, and the American company would essentially rent the algorithm from ByteDance.
By Thursday of last week, there was this draft executive order that outlined the deal, and then Trump did the thing with the tariffs. All of a sudden, ByteDance had to call up the White House and say, “That deal that you just helped us negotiate, it is off the table because the Chinese government is not going to support the deal anymore.”
Right. So this was a pretty dramatic reversal, and it does seem like they got very close to a deal before these tariffs. What is happening now that these tariffs are on? Does TikTok have any options left?
Along with a 90-day tariff pause, we also now have a 75-day extension that comes after the original 75-day extension that Trump gave in order to force ByteDance to divest TikTok.
This man loves extensions. Let’s just say it. This man loves to come right up against a deadline and say, “You know what? You got a little more time.”
Yeah, well, I don’t know what is going to happen over these next 75 days. I imagine that if the tariffs against China stand at 145%, there is no way the Chinese government is going to support the sale of TikTok. And I just want to say how self-defeating this is, because it was barely more than a week ago that Trump was telling reporters that Beijing, if it would simply go along with his plan to force the divestiture of TikTok, then he would go easy on them on tariffs, right?
This was his big bargaining chip: If you do not want high tariffs, you have to let the Americans have TikTok. And to my surprise, it seemed like the Chinese government was actually going to go along with that. Then, before they could even get that deal out, Trump, seemingly out of nowhere, announces a brand-new set of tariffs that completely scuttles the deal. So it is as if the president was essentially negotiating against himself and lost the deal that he had won.
Yeah, it does seem strange that he would not wait until after the TikTok deal was finalized and approved by all the relevant officials to then issue these tariffs if he was actually interested in getting a deal done.
Yeah, I think that is right.
Okay, TikTok is still in this frustrating state of superposition where they are both dead and alive at the same time. Do we think that this resolves before the end of the next 75-day extension, or do we think we will need yet another extension to figure out what we are doing with TikTok?
My assumption is that on the day that Donald Trump leaves office, we will still be in the middle of one of these extensions. It will be like the 15th extension, or the 23rd extension. But no, until this tariff situation gets resolved, I do not expect TikTok’s fate to be resolved. It is just going to continue to exist in its weird limbo.
All right. So that is TikTok. Our last company on this list of case studies is Meta. Casey, how is Meta dealing with this new uncertain reality?
5. Meta Courts Trump
I would say that things turned out a little bit better for them this week than it looked like they were going to, because tariffs were going to be a huge problem for them, too. They are a digital advertising business, and a huge number of their advertisers are small and medium-sized businesses that buy ads outside the United States to export goods from foreign countries into the United States. Mike Isaac at The Times had a great piece on this this week.
There is one analyst who estimates that about $10 billion of Meta’s revenue from ads originates from outside the United States. So in a world where everyone was facing these massive tariffs, we were just expecting Meta to get hit really hard on the ads front. Now that has mostly gone away, at least for the next 90 days, so it seems like Meta is going to get some breathing room.
But there is this one other outstanding question, Kevin, which is that next week, Meta’s antitrust case is going to trial, right? In 2020, during the first Trump administration, the Federal Trade Commission filed an antitrust lawsuit and tried to break off Instagram and WhatsApp from Meta. It has been in the planning stages ever since, and on Monday, the case is set to go to trial.
So why does all of this have anything to do with Trump? Mark Zuckerberg has been giving Trump the full-court press, going so far as to buy a $23 million house in Washington, D.C., recently, just to get closer to and spend more time with the president. There has been some reporting that Zuckerberg was in the White House trying to negotiate a settlement with Trump within the past few days. So there are a lot of questions right now about whether Zuckerberg will be able to use this relationship that he has apparently been building with Trump in order to get rid of this case, which is, in some ways, an existential threat to his business.
Yeah, and we should also just say, this should not be possible, right? The FTC is supposed to be an independent agency that has its own enforcement agenda and brings its own cases that are independent from the president. But, of course, nothing is truly independent from the president in Trump’s Washington.
He recently announced that he was getting rid of the 2 Democratic commissioners on the Federal Trade Commission. That is historically quite unusual for a president to intervene in FTC commissioner staffing at that level. But now it is going to be staffed with people who are friendly to the Trump administration. And so presumably, if he were to go to them and say, “Hey, let’s back off this Meta case. I don’t actually think we need to proceed with this,” they might listen.
And we should say that another way Meta tried to ensure that this happened is that, after the events of January 6, Meta suspended Trump from its platform for 3 years, and Trump sued them over that. After he won the presidency, Zuckerberg came along and said, “Hey, why don’t we settle this, too?” and paid Trump $25 million, right? I have to say, Meta was completely within its rights to suspend an account. They’re allowed to suspend whatever account they want. It’s a private company with a private platform.
But still, just as a little gesture of goodwill: “Hey, Trump, here’s $25 million.” So if this actually happens and this lawsuit just goes away, it will frankly be an example of open corruption.
Okay, so that is our 4-company case study of how tech companies are trying to do business and survive in this new uncertain environment. I have to ask: After going through all these examples, which of these companies would you be if you could be one? Which do you think is in the best position in this new chaotic environment?
Until maybe Wednesday, I think I would have said Apple, right? Apple makes the iPhone. The iPhone is the most lucrative product in the history of the technology industry, and even despite some of the tariffs that we were seeing, it seemed like they were still going to be in a good position to navigate them. I was seeing analysis that they were only going to lose maybe 7 points of profitability from all of this.
But the world looks really different with a 145 percent tariff, and in a world where Trump just keeps escalating this fight more and more, I actually do think that the picture for Apple just looks really strange. So look, I feel a little crazy saying this, but maybe I actually would just rather be Meta. Their hardware business is still a relatively small part of what they do. Mostly what they do is a digital services business, and it seems like Zuckerberg has been able to make at least some inroads with the Trump administration. Maybe they’re about to get rid of this lawsuit against them.
So, God, I don’t know. Maybe I actually want to be Meta. How about you?
Yeah, I think, as venal and corrupt as it would be for these naked attempts at flattery and persuasion to actually work and pay off, I would not underestimate how well this stuff works with Donald Trump. I think that Mark Zuckerberg’s motive here is to win at all costs, and if he needs to buy a $23 million mansion or spend time in the White House, or even make some policy adjustments to appease the Trump administration and get what he wants, I think he’s demonstrated very clearly that he’s willing to do that.
My last question on this, Casey, is about this idea of the tech capitulation to Trump. In the past few months, we’ve observed and talked about the fact that a lot of these tech companies have been really falling all over themselves to appease the Trump administration. Many of them gave to the inaugural. Many of them showed up at inauguration. Their CEOs were seated just behind the president’s own family.
The amount of flattery and ass-kissing going on here for months now has been notable and historic. Do you think that any of that has worked to the degree that these executives thought it would? Did the tech leaders get what they wanted out of Donald Trump?
I think that until the tariffs, the answer was basically yes, and the tariffs are what have changed that equation, right? If you look at how JD Vance was talking when he went to Europe, he was echoing a lot of tech company talking points. He and Trump have criticized European fines against tech companies, saying, “We need to protect and defend our American tech companies against these European fines,” which was something that the Biden administration never, ever did.
They’ve talked about getting rid of AI guardrails and just letting these companies do whatever they want with AI, which is like music to Mark Zuckerberg’s ears. But look, these companies just rely on stable, normal governance to be able to conduct their business around the world. They are as plugged into the interconnected global economy as anyone else, arguably more than many companies.
Trump just came along and blew that up, and I think it is probably dawning on them that they are probably just going to be living in chaos for the foreseeable future. It is just going to make their lives much, much more difficult.
Yeah, I think that’s right, and I think that a lot of these executives have underappreciated how important stability and predictability are in their business models. These were companies, many of them, that had issues with the Biden administration. The Biden administration had issues with them. But at least with the Biden administration, these companies knew where they stood, right?
There was not this sort of day-to-day whiplash of stock prices moving up 10 percent, down 10 percent, tariffs going up to 145 percent and then down to 10 percent. It just was not the kind of frenetic environment that we’re seeing today. So I wonder if any of them are starting to appreciate how good they had it during the Biden years, where, for as much as the Biden administration may have gone after them for various things, including antitrust violations, at least they could wake up every day and understand what the world was going to look like for the next 24 hours.
I think that’s true. I think that most of them would probably still be loath to admit it, but let’s give it another few weeks, Kevin, and another few tariffs, and then let’s check back in with them.
Sounds good. Well, that's enough about tariffs, Casey. When we come back, we're going to talk about a terrifying new report about what AI could look like in twenty twenty-seven.
6. AI 2027 Maps the Future
Well, today we’re going to talk about a forecast.
And that’s separate from a fork-cast, which is something different.
Yeah, that’s what we call our end-of-the-year predictions episode, isn’t it?
I think so.
But today we’re talking about something different, which is this new report called AI 2027. This is a report that I wrote about last week and that has gotten a lot of attention in AI circles and policy circles this week. It was produced by the AI Futures Project, a Berkeley-based nonprofit led by Daniel Kokotajlo, who listeners of this show may remember was a former OpenAI employee who left the company last year, became something of a whistleblower warning about their reckless culture, as he called it, and is now spending his time trying to predict the future of AI.
Lots of people are trying to predict the future of AI, but what gives Daniel a lot of credibility here is that in 2021, he tried to predict what things would look like about now, and he just got a lot of things right. So when Daniel said, “Hey, I’m putting together a new report on what I think AI is going to look like in 2027,” a lot of close AI observers said, “Oh, this is really something to read.”
And he didn’t just do this alone. He also partnered with a guy named Eli Lifland, who is an AI researcher and a very accomplished forecaster. He’s won some forecasting competitions in the past. The 2 of them, along with the rest of their group and Scott Alexander, who writes the very popular Astral Codex Ten blog, put together this very detailed, what they call a scenario forecast.
Essentially, it’s a big report, a website. It’s got some research backing it up, and it basically represents their best attempt to synthesize everything they think is likely to happen in AI over the next few years into a readable narrative.
And if that sounds a little dull to you, I’m telling you, you should just go check this thing out. It’s at ai-2027.com, and it’s super readable. It blows through stuff that feels very familiar right now, like just basic extrapolating from where we are today into getting to 6 months or a year from now. The world starts to look very, very different, and there is a lot of research to support why they think that is plausible.
Yeah, and I can imagine people reading this report or listening to us talk about it and saying, “Well, that sounds like science fiction to me.” We should be clear: it is science fiction. This is a fictionalized narrative that they have put together, but I would say it is also grounded in a lot of empirical predictions that can be tested and confirmed or verified. It’s also true that some science fiction ends up becoming reality, right?
Mm-hmm.
If you look at movies about AI from past decades, a lot of the things in those movies did end up actually being built. So I think this report, while it may not be 100 percent accurate, at least represents a very rigorous and methodical attempt to sketch out what the future of AI might look like.
And here’s my bet: If you put this conversation into a time capsule and revisited it in 2 years, in 2027, my guess is we’re going to find that a good number of things in that scenario actually did come true.
I hope we’re still doing a podcast in 2 years.
That’d be good.
That’d be great.
Yeah. So my forecast is that this is going to be a good conversation. Let’s bring in Daniel Kokotajlo. Daniel Kokotajlo, welcome back to Hard Fork.
Thank you. Happy to be here.
So you have just led this group that put together this giant scenario forecast, AI 2027. What was your goal?
Our goal was to predict the future using the medium of a concrete scenario. There is a small but exciting literature of attempts to predict the future of AI that use other methods, which is also very important: things like defining a capabilities milestone, such as, “Here’s my definition of AGI. Here’s my forecast for how long we’ll have until AGI based on these reasons,” and so forth. That’s great, and we’ve done that stuff before. We did a lot of that in the run-up to this scenario.
But we thought it would be helpful to have an actual concrete story that you can read. Part of the reason why we think this is important is that it forces you to think about everything and integrate it all into a coherent picture.
Well, I want to ask you a bit more about that. The first thing I want to say about AI 2027 is that it’s an extremely entertaining read. It’s as entertaining as most of the science fiction that I have read. By the end of it, you get into scenarios where humanity’s survival is threatened. So whether you think it’s true or false, it is really engaging to read.
But my understanding of your aim here is that there is something practical about what you were trying to do, right? Can you tell us about the practical idea of going through this exercise?
Important background context: The CEOs of OpenAI, Anthropic, and Google DeepMind have all publicly stated that they’re building AGI, and even that they’re building superintelligence, and that they think they can succeed by the end of this decade. That’s a really big deal, and everyone needs to be paying attention to that.
I think a lot of people dismiss that as hype, and it’s a reasonable reaction to say, “Oh, they’re just hyping their product.” But it’s not just the CEOs saying this; it’s also the actual researchers at the companies. And it’s not just people at the companies; it’s also various independent people in academia and elsewhere.
You also don’t just have to trust people’s word for it. If you actually look at the evidence, it really does seem strikingly plausible that this could happen by the end of this decade. And if it does happen, things are going to go crazy in some way or other. It’s hard to predict exactly how, but obviously, if we do get superintelligent AGI, what happens next is going to look like science fiction.
Right.
It’ll be straight out of a science fiction book, except that it’ll actually be happening.
You mentioned that if what the CEOs of tech companies say comes true, we will be living in a science-fiction world. And I think for a lot of people, they’re content to stop thinking there, right? They might be willing to admit, “Okay, yeah, if you invent superintelligence, things will probably be crazy, but I’ll cross that bridge when we come to it.”
You’re taking a different approach and saying, “No, you’re going to want to start thinking right now about what it would be like if some of these claims start to come true.” So maybe we could get into what some of those claims are. Sketch out for us what you think is very likely to happen just within the next couple of years.
Well, I wouldn’t say very likely.
Okay.
I should express my uncertainty, right? Past discussion often focuses on a single milestone, like artificial general intelligence or superintelligence. We broke it down into a couple of different milestones, which we call superhuman coders, superhuman AI researchers, superintelligent AI researchers, and then broad superintelligence. We make our predictions for each of these stages.
Even the very first one, I’m only 50 percent confident that it’ll happen by the end of 2027. So, a 50 percent chance that 2027 will end and there still won’t be any autonomous, superhuman coding agents. Let’s say a 50 percent chance—
But it’s a coin flip.
Yeah, a coin flip.
We might also be living in a world where, yes, you do have an—
Exactly.
Yeah.
Exactly. So there’s a 50 percent chance we do have autonomous, fully autonomous artificial intelligences that can basically do the job of the best engineers by 2027. Then you ask, okay, what’s the next milestone after that?
After that comes automating the full AI research process instead of just the coding, because AI research is more than just coding. How long does it take to get to that? We have our guesses, and in our scenario it happens about 6 months later.
In our story, you get the superhuman coders and use them to go even faster to get to the superhuman AI researchers that are able to do the whole loop. That really kicks things off, and now you’re going much faster. How much faster? We say 25 times faster for the algorithmic progress, at least.
Of course, your compute scale-up is not going any faster at all, because you still have the same amount of compute. But you’re able to do the algorithmic progress 20 times faster, 25 times faster. Then you start getting to the superhuman regime. You start getting systems that are qualitatively superior to the best humans at stuff.
They’re also probably discovering new paradigms. So we depict them going through multiple paradigm shifts over the course of the second half of 2027, ending up with something that’s vastly superior to humans in every dimension by the end.
Yeah. Let me just pause and maybe underline a couple of things there. I think most people might not understand why the big AI labs are obsessed with automating coding, right? Most people are not software engineers, so they don’t really care how much of it is automated.
But by the time you get to software that is mostly writing itself, it unlocks this other world of possibilities. You sketch out a vision where, once we get to a point where the AI coding systems are better than almost every human engineer, or maybe every human engineer, this other thing becomes possible: You can just set this thing to work trying to figure out how to build AI itself, right? Is that what I’m hearing you say?
Basically. I’d break it down into 2 stages. I think the coding is separate from the complete automation, as I previously mentioned.
I expect to see systems that are able to do all the coding extremely well but might lack research taste, for example. They might lack good judgment about what types of experiments to run, and that’s why they can’t completely automate the research process. Then you have to make a new system, or continually train the old system, so that it gets that taste and judgment.
Similarly, they might lack coordination ability. They might not be very good at working together in large organizations of thousands of copies, at least initially. But then you fix that, come up with new methods, and do additional training runs to get them good at that sort of thing.
That’s what we depict happening over the first half of 2027, and we depict it happening in only half a year because it goes faster, because they’ve got all the coding down pat. Even though humans are still directing the whole process, they just give orders to the coding agents, and the agents quickly make everything actually work.
Right.
Halfway through the year, they’ve succeeded in making new training runs that train the skills the AIs were missing. Now they’re not just coding agents; they’re able to do the research taste as well. They’re able to come up with new ideas, come up with hypotheses and test them, and work together in big hive-mind clusters of thousands and thousands of them. That’s when things really kick off.
Right.
That’s when it really starts to accelerate.
7. The Two AI Endings
In your scenario, you have this sort of choose-your-own-adventure ending where, after this thing you call the intelligence explosion—where the superhuman AI coders get into AI R&D and start automating the process of building better and better AIs—you have 2 buttons that you can click. One of them unspools the good-place ending, where we decide to slow down AI development, really get these things under control, and solve alignment.
And then the red button: you push that, and it goes into this very dark, dystopian scenario where we lose control of AI, they start deceiving and scheming against us, and ultimately maybe we all die. Why did you decide to give people the option of choosing one of those 2 endings rather than just sketching what you believe to be the most probable outcome?
We did start by sketching what we believe to be the most probable outcome, and it’s the race ending, the one that ends with the misaligned AIs in control of everything. So we did that first, and then we were like, “Well, this is kind of depressing and sad, and there’s a whole bunch of stuff that we didn’t get to talk about because of that.” We wanted to then have a different ending that ended differently.
In fact, we wanted to have a whole spread of different possible outcomes, but we were limited by time and labor, and we were only able to pull together 1 other outcome, which is the one that we depict in the slowdown ending. So in the slowdown ending, they solve the alignment issues, and they actually get AIs that are what they say on the tin. They’re not faking it. They just actually have the goals and values that were put into them, or that the company was trying to train into them.
It takes them a couple of months to sort that out. That’s why it’s a slowdown: They had to pivot a lot of their compute and energy toward figuring that stuff out. But they succeed. And so then, in that ending, we still have this crazy arms race with China, and we still have this crazy geopolitical crisis. In fact, it still ends in a similar sort of way, with this massive arms buildup on both sides, this massive integration into the economy, and then ultimately a peace treaty.
I’m curious, Daniel, if the events of the last week in Washington—the tariffs, this looming trade war with China—have affected your forecast at all.
We’ve been iteratively improving it, but the core structure of it was basically done a few months ago. So this is all new to us and wasn’t really part of the forecast.
How would it change things? Well, if the trade war continues and causes a recession and stuff like that, it might just generally slow the pace of AI progress, but not by much, I think. Say it makes compute 30 percent more expensive, so that the companies are able to buy 30 percent less of it. Maybe that would translate to a 15 percent reduction in overall research velocity over the next few years, which would mean that the milestones that we talk about happen a few months later instead of when they do. So the story would still be basically the same.
One of the things I think is most interesting about your project is the bets and bounties section, where you are going to pay people for finding errors in your work, for convincing you to change your mind on key points, or for drafting some alternate scenarios. Talk to me a little bit about how that became part of this project.
I come from the sort of rationalist community background, which is big into making predictions and making bets, putting your money where your mouth is. So I have a sort of aesthetic interest in doing that sort of thing. But also, specifically, one of the goals of this project is to get people to think more about this stuff and to do more scenario forecasting along the lines of what we’ve done.
We’re really hoping that people will counter this with their own reasonably detailed alternative pathways that represent their vision of what’s coming. So we’re going to give out a few thousand dollars of prizes to try to mildly incentivize that.
As for the bounties thing, already we’ve gotten dozens of people being like, “You say this, but isn’t this a typo?” Or, “This feels wrong.” So I have a backlog of things to process, but I’m going to get through it. I’m going to pay out the little payments and fix all the little bugs and stuff like that. I’m just quite heartwarmed to see that level of engagement.
Have you taken any bets on different scenarios so far?
I think so far I’ve done 1 or 2.
Okay.
But mostly there’s just a backlog I need to work through.
Got it.
Yeah.
Got it.
Now, Daniel, you said you’ve been getting some good responses from people at the AI companies to this scenario forecast. I did a bunch of calling around when I was writing about this, and after we spoke, I talked to a bunch of different people, both in the AI research community and outside of it. I would say the most frequent reaction I got was just kind of disbelief.
8. The Forecast Faces Skeptics
One person I talked to, a prominent AI researcher, said he thought it was an April Fools’ joke when I first showed him this scenario because it just sounded so outlandish. You’ve got Chinese espionage, the models going rogue, and the superhuman coders, and it all just seemed fantastical. It was almost like they didn’t even think it was worth engaging with because it was so far out. I’m curious if you’ve gotten much of that kind of reaction and what your response is.
A couple things. First of all, go write your own damn scenario, then. I would say you either will write a scenario that doesn’t seem outlandish, which I will completely tear apart as unrealistic and basically assuming that AI progress hits a wall, or you’ll write a scenario that does feel very outlandish, but perhaps in different ways than ours do.
Again, are they actually going to get to AGI and superintelligence by the end of this decade? If so, you can’t possibly write that in a way that’s not outlandish. It’s just a question of which outlandish thing you’re going to write. And if you think maybe it’s just not going to happen and it’s going to hit a wall, yeah, that’s possible, too. I think that’s reasonable. I don’t think it’s the most likely outcome.
I do actually think that probably by the end of this decade we’re going to have superintelligence.
And then say more about that, because I assume that a lot of our listeners either truly think that it will hit a wall, or they’re just counting on it hitting a wall so as not to have to reckon with any of the scenarios that you describe. What is your message to the person who’s just like, “Eh, it’ll probably hit a wall”?
Right.
I mean, I don’t know. Read the literature.
These people are not going to read the literature. They listen to podcasts specifically so they don’t have to read the literature.
Mm.
So they don’t have to read the literature.
Okay.
So they don’t have to read the literature.
Yeah, fair. Well, I could point to specific parts of the literature, like benchmarks, for example, and the trends on them. I would say the benchmarks used to be terrible, but they’re actually becoming a lot better.
METR in particular has these agentic coding benchmarks where they actually give AI systems access to some GPUs and say, “Have fun. You have 8 hours to make progress on this research problem. Good luck.” Then they measure how good they are compared to human researchers given the same setup.
The line goes up on the graph. It seems like in a year or 2 they’ll have AIs that are able to just autonomously do 8-hour-long ML research tasks on these sorts of things. That’s not AGI, and that’s not superintelligence, but that is maybe the first milestone that I was talking about: superhuman coder.
So I point to those sorts of trends, and then separately, I would also just do the appeal to authority. If you’re not going to read the literature, if you’re not going to look at it, if you’re not going to form your own opinion about this, and you’re still just deferring to what other people think, then I will say, yeah, there are a bunch of naysayers out there who are saying this is all never going to happen, it’s just fantasy.
But also, there are a bunch of extremely credible people with amazing track records, both inside the companies and outside the companies, who are in fact taking this extremely seriously.
I also want to read you—
Including our scenario. Yoshua Bengio, for example, read an early draft of our thing and liked it, and gave us some feedback on it. We put a quote from him at the top saying, “Everyone should read this. It’s plausible.” So he’s a pioneering AI researcher.
Yeah. Another genre of criticism I’ve heard of this forecast is from people who are just questioning the idea that if you get AIs that are superhuman at coding, they will be able to bootstrap their way to general intelligence.
I just want to read you a quote from an email that I got from David Autor, who’s a very well-known economist at MIT. I had asked him to look at the scenario and react to it, with a particular eye on what this might be missing as far as how it assumes this easy and fast jump from superhuman coding to something like AGI. I’ll just read you what he said.
He said, “LLMs and their ilk are super-powered incarnations of one incredibly important and powerful part of our cognition.
The reason I say we're not on a glide path to AGI is that simply taking this capability to 11 does not substitute for the parts that are still missing. I think that humanity will get to AGI eventually. I'm not a dualist. I just don't believe that swimming faster and faster allows you to fly.
What is your reaction to that?
I agree. We depict this in the course of the story. So if you read AI 2027, they have something that's like LLMs, but with a lot more reinforcement learning to do long-horizon tasks, and that is what counts as the first superhuman coder. So it's already somewhat different from the systems of today, but it's still broadly similar. It's still maybe the same fundamental architecture, just a lot more training, a lot more scaling up, and, in particular, a lot more training specifically on long-horizon agentic coding tasks.
But that's not itself AGI, I agree. That's just the superhuman coder that you get early on. And then you have to go through several more paradigm shifts to get to actual superintelligence, and we depict that happening over the course of 2027. So a key thing that I think everyone needs to be thinking about is this: Takeoff speed is variable. How much faster does the research go when you've reached the first milestone, and how much faster does the research go when you reach the second milestone, and so forth?
We are, of course, uncertain about this, like we are about many things. We say in the scenario that we could easily imagine it being 5 times slower than we depict, and taking 5 years instead of 1 year. But also, we could imagine it being 5 times faster than we depict and taking 2 months. So we want to do a lot more research on that, obviously.
If you want to know where our numbers are coming from, go to the website. There's a tab that you can click on that lists a bunch of back-of-the-envelope calculations and little mini-essays where we generated the quantitative estimates that are the skeleton of the story.
One other piece of criticism I've seen of this project that I wanted to ask you about was from a researcher at Anthropic named Saffron Huang.
Mm.
She argued on X that she thought your approach in AI 2027 was highly counterproductive, basically that you were in danger of creating a self-fulfilling prophecy by making these scary outcomes very legible, by burying some assumptions, and that you were essentially making the bad scenario that you're worried about more likely to actually happen. What do you make of that?
I'm quite worried about that as well, and this is something we've been fretting about since day 1 of the project. So let me just say a little bit more about that.
First of all, there is a long history of this sort of thing seeming to happen in the field of artificial general intelligence research, most notably Eliezer Yudkowsky, who is the ur-father of worrying about AGI, at least in this generation. You know, Alan Turing also worried about it, but Sam Altman specifically tweeted—do you remember this tweet? Sam specifically said, "Hats off to Eliezer Yudkowsky for raising awareness about AGI." It's happening much faster now because of his doomsaying, because it's caused a bunch of people to pay more attention to the possibility and to start investing in these companies and so forth.
So I was, I don't know, twisting the knife at him because he obviously doesn't want this to happen faster. He thinks we need more time to prepare and make it safe and so forth. But it does seem like there's been this effect where people talking about how powerful and scary AGI could be has maybe caused it to come a little bit faster and caused people to wake up and race harder toward it.
Similarly, I'm worried about causing something like that with AI 2027. One of the subplots in AI 2027 is this whole concentration-of-power issue: Who gets to control the army of superintelligences? In the race ending, it's sort of a moot question because the army of superintelligences is just pretending to be controlled, and so is not actually listening to anyone when it counts.
But in the slowdown ending, they do actually align the AIs, and so they are actually going to do what they're told. And then who gets to say that? The answer in our slowdown ending is the oversight committee, which is an ad hoc group of people—some CEOs and the president—who get together and share power over the army of superintelligences.
What I would like to see is something more democratic than that, something where the power is more distributed. I'm also afraid that it could be less democratic than that. At least we get an oligarchy with this committee, but it could very easily end up a dictatorship, where one person has absolute control over the army of superintelligences.
This is yet another example of how I'm trying not to have the self-fulfilling prophecy happen. I don't want people to read this and be like, "Hmm, I'm a CEO—"
I can make a lot of money by building a misaligned AI.
Yeah. Maybe. But all that being said—
So any of our evil-villain listeners out there, steepling your fingers in your lair under a mountain, knock it off.
Yeah. So all that being said, we are taking a gamble that sunlight is the best disinfectant. The best way forward is to just generally tell the world about what we think is coming, and hope that even though many people will react to that in exactly the wrong ways, enough people will react to that in the right ways that overall it will be good.
I am tired of the alternative of hush-hush, keeping everything secret, doing backroom negotiations, and hoping that we get the right people in the right rooms at the right time and that they make the right decisions. I think that is kind of doomed. So I'm placing my faith in humanity, telling it as I see it, and hoping that insofar as I'm correct, people will wake up in time and, overall, the outcome will be better.
Yeah. All right.
Thank you, Daniel.
Thanks, Daniel.
Thank you so much.
When we come back, Meta decides to fake it till they make it.
We'll talk about the cheating scandal that is rocking the world of AI benchmarks.
Well, Casey, there's one other big AI story we want to talk about this week, and that is the drama surrounding Llama.
That's right, Kevin. Meta has a new large language model. It was hotly anticipated, but I think it's fair to say it kind of stumbled out of the gate.
Yeah, they had some Llama, Llama cred drama.
How many times are you going to do the Llama drama pun?
Well, there's a very popular children's book called Llama Llama Red Pajama.
I am.
Are you aware of this?
I am.
So let's get into it. There have been a lot of things going on around this new language model, Llama 4, that Meta released last weekend. Casey, you've been writing about this in your newsletter this week. Catch me up. What is going on with Llama 4?
Yeah, so look, Meta has invested billions and billions of dollars in AI, and they're taking a very different approach from the AI labs that we most often talk about on this show. Companies like OpenAI, Anthropic, and Google have closed models. You can't download, fine-tune, and re-release them under a very permissive license. But with Meta's, you can. And when Llama 3 came out last year, developers said, "Oh, this thing is actually pretty good." It's not as good as the state of the art, which is often true of the open models, but it's getting up there.
Right. And so they spent all this money to develop Llama 4. People have been talking for months about how this was going to blow all the other open-weight models out of the water, and then they release it. What happens?
Well, 2 things happen, Kevin.
The first is that Meta trumpets this model in the way that companies usually trumpet their most recent models: as being the most powerful ever or the most efficient. They show off a bunch of benchmarks. They say, “This thing is highly capable, and it’s the bee’s knees.” They didn’t actually say it was the bee’s knees. I’m not sure anyone has said that in the past 70 years. But they said things like that.
And one of the benchmarks that really got people’s attention was LMArena. Do you know LMArena?
I know of it, but I haven’t spent much time on it. What is it?
It’s this really interesting project. It’s a very small nonprofit that includes some researchers from UC Berkeley, and what they do is get people to volunteer to help. They’ll have people enter a query, and then they’ll show them the responses from 2 different chatbots that are not labeled. After they get the answers, the user will say, “Oh, I liked this one better.” They collect those votes over time, and the more that people vote for one chatbot over another, the higher it rises on LMArena.
I see. So it’s sort of like a crowdsourced leaderboard for which of these models people prefer.
Exactly. And Kevin, you know as well as anyone else that whenever a new model comes out, the question of how good it is turns out to be weirdly hard to answer.
Yep.
Right? Maybe it’s really good for what you need it to do, maybe it’s really bad, or maybe it’s about as good as something else, but you just happen to like it better because it has a style that matches what you’re looking for. So in such a world, companies are desperate to be seen as good, but they don’t have an easy way of communicating that, and that’s when LMArena enters the picture. Because if you can get high enough on that leaderboard, you can point to it and say, “Aha, look at how we’re doing.”
Right. The people have voted.
That’s right. The people have spoken, and look how well we’re doing. So do you know how well Llama 4 does on LMArena?
No.
Llama 4 comes in at number 2, just under Gemini 2.5 Pro Experimental, which is the latest model from Google, which has been through a lot of testing, and which basically has universal acclaim. People think this is a truly great model, not just at this little chatbot contest, but across a bunch of other things, including coding and a lot of other things.
So Llama 4 immediately zooming up to number 2 on LMArena would seem to indicate that Meta has really cooked here. They have built this incredible model. They are releasing it to the public under an open-weight structure, and they are one of the leading AI labs when it comes to creating very powerful models.
That’s right, except there’s an asterisk.
Oh, boy.
This version of Llama 4 is an experimental model. Meta’s website says it has been optimized for chat. People start to look into this, and they notice this is not the version of Llama 4 that is actually available for download.
The one that was included in LMArena was not the one that people could download?
That’s right. It had a different name. It was named Llama 4 Maverick 03-26 Experimental. And people start to think, “Oh, wait a minute. What if what happened here isn’t what normally happens on LMArena,” which is that people make a new model, submit it to LMArena, and see how it does? What if Meta trained a special version of Llama 4 just to be good at LMArena?
Hmm.
I have spent the past week trying to research whether this is true, and on Monday, I got Meta to send me a statement, which I guess I should read.
“We experiment with all types of custom variants,” and this experimental version is “a chat-optimized version we experimented with that also performs well on LMArena. We have now released our final open-source version, and we will see how developers customize Llama 4 for their own use cases.”
So this was really interesting to me because when they say, “It also performs well on LMArena,” it suggests that maybe they just made 15 of these models and were like, “Oh, look, this one happens to do well on LMArena.” That is one possibility. I think another possibility is exactly what the cynics think, which is, “Oh, no, they reverse-engineered how LMArena works, and they built a bot that was just going to beat it.”
And how would you do that? If your goal was to create a model that would perform very well on this one specific leaderboard, what would you do?
LMArena has released a lot of chats over the years that show which chats are considered preferable to other chats, and it seems that the users of LMArena really like it when the bot has a high degree of what they call sycophancy. So basically, you’re just like, “What should I have for breakfast today?” And the chatbot is like, “Oh my God, that’s such a great question. You’re a genius. I love the way you’re starting the day off right.” That is the kind of answer that people pick.
Hmm.
And so you can build a chatbot that essentially just flatters people constantly, and it tends to do really well on Chatbot Arena.
Hmm.
So anyway, in the aftermath of this confusion, LMArena, which is a very mild-mannered organization that I don’t think is used to being involved in public controversies, puts out a statement. I have to read the statement, Kevin, because as gentle as it is, I found it pretty damning.
They don’t go so far as to say Meta cheated, but what they do say is:
“Meta’s interpretation of our policy did not match what we expect from model providers. Meta should have made it clear that this experimental model was a customized model to optimize for human preference. As a result of that, we are updating our leaderboard policies to reinforce our commitment to fair, reproducible evaluations so this confusion doesn’t occur in the future.”
So why is that statement so interesting to me? Well, you basically just have this tiny group of researchers over at Berkeley, and Meta violates their policies so badly that they have to change the rules for how this competition even works just to get people to stop breaking the competition.
Yeah. I thought this was a really interesting set of stories. I’m still waiting for someone—ideally you—to get to the bottom of what actually happened inside Meta. But I think it’s worth talking about for 2 reasons. One, because I think it says something about Meta and its place in the AI race, and the other because I think it says something about the state of AI and these benchmarks, and how useful they are or aren’t in making sense of the torrent of new models that are constantly coming out from the big AI labs.
Totally.
So maybe let’s take those 1 by 1. What do you think this says about Meta’s place in the AI race if it does turn out that they gamed this leaderboard to make it look like their model was better than it was?
Here’s what I think: If you’re winning the AI race, you do not waste time trying to beat LMArena, right? What you do is what Google did, which is just release a very powerful pro version of Gemini, and it just happens to float to the top of the arena—not because it’s been optimized for conversation, but just because it’s a great model that’s really good at a lot of things.
If you have to make a custom version of your model just to win this rinky-dink competition, it’s hard for me to think of a more adverse indicator for the quality of Meta’s AI program. And we should say there’s been reporting in The Information over the past year that the Llama 4 development process has been really frustrating for Meta, that they delayed the release twice because they weren’t getting the results that they wanted.
When it finally did come out and people started to put it through other evaluations, they found that it just was not hitting the mark. In fact, Kevin, Ethan Mollick, a former guest on Hard Fork, compared the versions of the experimental chat that was winning the leaderboard to the chats that were produced by the final open-weights model, and what he found was that the open-weights model was producing really bad responses.
Mm.
Essentially, the optimized model was performing so much better than the real one that it wasn’t even close.
So why don’t they just release the optimized model, then?
That’s a great question. I don’t know the answer to that, but what I’m going to assume is that whatever fine-tuning is necessary to increase the level of sycophancy in the bot might be great for this sort of competition, but maybe it’s really bad for coding—
Mm.
—or creative writing or the countless other things that we now expect LLMs to be good at, right? Fine-tuning is a very powerful process that can take a very general-purpose model that’s kind of mediocre at a bunch of things and make it really good at 1 thing.
But these days, people have a lot of options to choose from with their large language models, and there are a lot of them that just have very high general capability, so they’re going to use those instead.
Yeah. I mean, I have not done my own reporting on the situation inside Meta with Llama 4, but I will just say, from a broad view, if you just step back from this particular scandal, Meta is not one of the top 3 AI labs in America when it comes to releasing frontier models. They are not in the top tier of frontier AI research.
A lot of their key researchers have left the company. Their models are not seen as capable as the models from OpenAI, Anthropic, and Google DeepMind, and I think that really frustrates them, right?
Mm-hmm.
I think Mark Zuckerberg and his lieutenants really want to be seen as part of the vanguard here. So I would not be surprised at all if, in an effort to juice their numbers and appear to be leapfrogging some of their competition, they violated the terms of one particular AI benchmark. And that should make us question how well their overall AI program is doing.
Absolutely. And by the way, the next time they release a model and come out with a bunch of wild claims, do you think I'm going to believe any of that? No.
Totally.
It's like you're going to have to try to verify every single claim they make independently. Yeah, look, I assume some people are going to hear this and think that I'm making a mountain out of a molehill, but I just think about what Daniel Kokotajlo told us about how powerful these systems are becoming and about how powerful they're about to become.
You want them to be loyal to human beings, but you also want them not to be used for bad behavior. And if there is a company out there that's cheating to win benchmarks, what else can that model do?
Right.
So even though this may seem like a small thing, I think it matters that we have companies building AI systems where we have some level of trust in those companies, where we believe they have some amount of integrity when it comes to how they operate. And so this was a moment where I thought, “Wow, my trust in Meta as an AI company has just been dramatically reduced.”
Yeah. So, the Meta of it all aside, I think this does actually raise a really important question about the broader AI industry, which is the value of benchmarks in general. Because one thing that I've heard from AI researchers over the past year or two is that these benchmarks, these tests that are given to these models to figure out how intelligent they are, all have some flaw—
Yeah.
—built into them, right? There's this issue of data contamination, which is: What if some of the answers on these tests are being fed into these models during their training process, so that you're really not getting a sense of how capable the model is? They're just regurgitating these answers that they've already seen. That is an issue.
There is also the issue that all these companies are effectively grading their own homework, right? There's no federal program that puts these things through their paces and releases standardized benchmark scores that we can actually verify and trust. Some of these AI companies are using different methods to even apply these benchmark tests. There are these things called consensus@64 and all these different ways that you can cherry-pick the best answer that your model gives if you give it the test a bunch of times and use that for your score.
So I think we are just losing our ability to trust the way that we measure these AI models in general.
Yeah. And it's so frustrating. I was thinking, Kevin, imagine in the early 2010s. It's not just that Instagram comes out as an app in the App Store. You have Instagram, Instagram o1, Instagram o1 Mini, Instagram o1 Deep Research, and it's like, “Download the one that's best for you.” You'd be like, “Why are you making me do any of this?” Right? Like, “Just give me the one thing that works.”
And while every AI lab is trying to realize that, in the meantime, we're living through this Cambrian explosion of large language models. On the one hand, I think that makes it really important for there to be benchmarks, so that we can look at a glance and have a basic sense of whether this thing is even worth our time. But on the other hand, that makes the benchmarks such an attractive target for gaming and outright cheating.
Yeah.
And so that's why the researcher Andrej Karpathy has said that we have what he calls an evaluation crisis, where, when a new model comes out, the question of how good it is is just very difficult to answer.
I've been wondering what we can do as journalists to try to answer those questions better. Is this a place for journalists to actually say, “Okay, a new model came out. We're going to have our own custom set of evaluations. Maybe we're going to keep those private in some way to prevent them from being gamed”? But what solutions do you see here to this crisis?
Well, at the risk of scooping myself here, I will disclose that I am actually starting to work on my own benchmark—
Ooh.
—because I think that part of how we are going to make sense of these AI models is that people will just start developing their own sets of tests to give to new models, not necessarily to determine their overall intelligence, but to determine how good they are at the things we care about.
Mm-hmm.
Personally, I don't care much if an AI model is getting a 97 percent on the graduate-level physics exam or a 93 percent, right? That does not make a huge difference in my life.
Because it's still higher than you're going to get.
Exactly, and I am not a graduate-level physics researcher. So I might care more about whether a model is good at creative writing or not.
Mm-hmm.
And I might want a battery of tests to determine that. And so I think that as these things become more critical in people's lives and work, we will start seeing more personalized tests and evaluations that actually measure if the models are good at the things that we care about.
Yeah.
What do you think?
Yeah, I think that's a great point. And after you told me that you were going to do this, I started to scheme and thought, you know, I want my own benchmarks, too, because there are—I don't know. I'm sure I can come up with a list of 10 things that I wish AI could do for me today that it still can't. And so maybe it's time that I should start some scenario planning.
What's one of your tests that you want to give AI models to determine if they're capable or not?
Well, for example, I have a newsletter that has customer service issues. People email us. They say, “Oh, my gosh, can I change my email address? I'm having trouble—”
People say, “The writing in this is so bad.”
No, people love the writing.
Oh.
That's all I hear about the writing. People are saying, “This—this is a human writing this? That's insane.”
But I would love to be able to automate some of that, make it easier for people. Say, “Oh, you need to download your invoice?” Which is a question we get a lot. It's like, okay, yes, we're just going to handle that in an automated way. So that's just one very easy thing.
And if you're thinking, “Oh, Casey, I actually have a product that can already do that for you,” please don't email me. It can't. I've been through this.
Can I tell you one of the things that I want to test AI on?
Yeah.
So, as you know, I just moved into a new house.
Mm-hmm.
And as a result, I have spent between a third and half of my waking hours over the last few weeks thinking about hanging pictures.
Mm.
Hanging pictures is one of my least favorite tasks in the world.
I hear it.
You have to do math.
Mm-hmm.
You have to bring out the laser level.
Mm-hmm.
I mean, it's a huge process.
The golden ratio.
Yes. And I would love for an AI system to be able to hang pictures for me.
That's beautiful.
And as soon as that happens, to me, that's AGI.
Now, would that involve a robot?
Probably, yeah. So we've got to make some progress before we get there, but—
Okay. That's a tough benchmark.
—if you're listening to this and you're working at one of these robotics companies, get on it.