[BidClub_]
Latent Space · · 121 分钟

Google 的 AI 科学家最初是一次自动化 Kaggle 的尝试——Google Fellow John Platt

John Platt

AI与软件技术
YouTube
TL;DR
  • Platt 认为,ERA 直到最近 Gemini 的能力提升后才真正变得可用。 他表示,在 Gemini 2.0 时代,“这是不可能做到的”;而在实验开始奏效后,Gemini 各个大版本和中间版本的进步都十分显著。他回忆说,Gemini 2.5 已经能立即写出 boosted-decision-tree 代码,并告诉那些不喜欢 2.0 的科学家:那已经是“上古时代”了。
  • ERA 的核心洞见是,许多科学问题都可以转化成“可评分任务”:编写代码,最大化一个预先定义的分数。 它的 Monte Carlo 研究循环借助 Gemini 改写 notebooks;Gemini 能阅读论文、提出代码,并利用有助于构建模型的“梯度”。随机改写代码基本以失败告终,因为大多数变异都有害;LLM 作为变异算子,则能做出更有信息量的修改。
  • 人类仍必须定义目标,并持续监督目标是否被正确执行。 评分函数可能存在漏洞,agents 和人类都会通过钻漏洞来刷分。Platt 举例说,Kaggle 参赛者曾在一个卷云竞赛中发现标签存在半个像素的误差。他表示,研究人员既需要创造力,也需要严谨性;人类必须确认优化后的结果确实是在描述现实,而不只是预测得准。
  • 持续性航迹云约占人为全球变暖的1%,而且或许可以用相对低的成本规避。 冰过饱和区域可能把排放的水、冰或烟尘效应放大约10,000∶1;避开这些区域有时只需下降两个飞行高度层。ERA 还找到一个简单的反事实模型,解决了 Platt 的团队曾经卡了两年的反射阳光问题。
  • Platt 认为,具备商业意义的核聚变可能“3年后,而不是30年后”实现,但他强调这只是概率判断,并非确定预测。 Lawson 判据要求密度、温度和约束时间的组合超过阈值,而每种路线都有自己的阿喀琉斯之踵。他还强调,电气化和电池并不能提供一个适用于所有场景的气候解决方案。
  • 天气预报已经经历了一次机器学习带来的重大跃迁,而气候问题仍然难得多。 天气预测的是大气的演化轨迹;气候关注的是一个本身也在变化的吸引子的统计特征。由于未来数据有限、碳循环行为存在巨大不确定性,Platt 希望未来的系统能把数据与散落在大量论文中的科学知识结合起来,但他表示 ERA 目前还做不到这一点。
  • 最大的瓶颈在于物理实验。 Platt 理想中的系统是一个“万能实验室”,只需输入一段 JSON 描述,就能执行任何实验。ERA 目前处理的是计算问题,而实验仍是真实世界的 ground truth,也是主要瓶颈;他同时认为,科学家仍应建立深厚的领域知识、严谨性和创造力,并通过足够多未经优化的实践来锻炼这些能力。
摘要 · 为研究而整理的核心内容

1. ERA 的起点:把科学问题变成“可评分任务”

  • 在 Google Research 做了10多年 AI for science 后,Platt 的框架是:大约两年前,团队“偶然找到了这种映射”——许多科学问题都可以表述为:“我真正想要的是一段能把某个分数最大化的代码。”统计模型拟合只是这一更大框架中的一个例子。
  • 一个不太直观的例子是,ERA 论文第一作者 Michael Brenner 用这套工具快速处理科学论文。对于常微分方程或偏微分方程的渐近展开,ERA 可以提出解,测试其在 epsilon = 1e-4 时是否满足渐近正确性,再让 Gemini 在最大化拟合效果的同时改进数学推理。
  • 遥感场景涉及不同卫星之间的信息增强超分辨率:把 OCO-2 或 OCO-3 窄幅但精确的 CO₂ 测量,与 GOES 大约每5分钟一次、分辨率较低的红外图像结合,再叠加天气和长期反照率数据,用一颗卫星的数据估计另一颗卫星的信号。

2. 底层机制:在 notebooks 上做树搜索,让 Gemini 充当变异算子

  • 这套 harness 首先帮助科学家定义一个可评分任务,生成带评分函数的 Python notebook,然后在数百或数千个候选 notebook 上运行 Monte Carlo 研究。随着变异不断生成子节点,这些候选方案可以形成一棵树。
  • 选择采用 upper-confidence-bound 策略:对某次变异的潜力做乐观估计,Platt 描述为大致预测其95%分位数结果。因此,系统不会总是选择当前最好的 notebook;有时会选第五好的候选,因为探索可能更有价值。团队也尝试过把两个候选方案的思路重组为第三个候选。
  • 初始种子不需要 starter code。只要给出文字版问题描述,并可选择附上5篇相关论文,Gemini 就能读论文并完成第一次编码尝试。代码 bug 可能导致分数实际跌至负无穷,随后系统可以继续变异这个候选方案。
  • 默认设置是大约10路并行搜索,也就是同时生长约10个叶节点。并行度过高会让各条搜索路径无法相互学习;每轮结束后,被剪枝的历史和此前结果会回流到基本上共享的一个上下文中。
  • 这比早期的进化式编程效果更好,因为“在代码空间里随机变异基本没有价值”。随机变异大多有害,但 Gemini 能凭借变异算子本身积累的大量世界知识,识别出更有信息量的尝试方向。

3. 评分函数仍是人类的工作——Goodhart 定律始终在等待

  • 主持人认为,最难、也最需要人类判断的环节,可能是确定正确的评分标准。Platt 表示同意:agent 可能钻目标函数的漏洞,得到的是字面上被要求的结果,而非科学家真正想要的结果。因此,研究人员必须反复修改评分标准和指令。
  • 这个不间断内循环的价值,在于科学家可以从实现细节上升到科学问题本身。他们不必把大部分时间耗在导入文件或让数据库运行起来,而可以思考成本函数究竟应该代表什么。ERA 还可以建议将哪些数据集拼接在一起。
  • Platt 引用了 Goodhart 定律:一旦指标成为目标,它就不再是可靠的指标。这一原则适用于每一个 leaderboard 或 benchmark,因此优化过程必须叠加多层严谨性,并进行细致核查。
  • 这套流程明确是 human-in-the-loop。ERA 可以运行数小时,返回样例和进展,随后接受进一步指导——“就像有一个精力过剩、从不睡觉的研究生”。

4. Gemini 的能力曲线决定了一切

  • 当被问及是否存在某个相变节点时,Platt 表示,ERA“在 Gemini 2.0 时代根本不可能实现”。团队先实验了一段时间,看到系统开始奏效,随后 Gemini 各个版本及中间版本迅速进步。
  • 他记得的拐点是 Gemini 2.5:他让模型写 boosted-decision-tree 代码,结果“它直接就写出来了”,于是他说:“这是一个全新的世界。”
  • 对于那些试过 Gemini 2.0 后放弃的科学家,他的建议是:那个模型已经落后了一年——“那已经是上古时代了”。
  • 一篇 preprint 将 ERA 与另一个可以调取论文并编写代码的公开 harness Antigravity 结合起来,但在这次对话时,ERA 加 Antigravity 的组合尚未公开提供。

5. 预测模型与描述模型:Newton 并不是在给苹果建模

  • Platt 区分了预测模型和描述模型:前者在数据集上最小化误差;科学真正希望用于外推的是后者,因为它们捕捉到了物理规律或现实的真实描述。他的例子是,引力并不是真的关于苹果:如果模型只在下落的苹果上训练,它不会知道该如何描述行星。
  • 当被追问先验从何而来时,他说两者的边界并不清晰。科学家用直觉和积累的事实约束模型;在当前 ERA 的工作中,对应的来源是 LLM 在预训练中学到的直觉与世界知识。
  • 如果给出已有论文,ERA 可以改造已发表的方法,而且经常能反向工程出论文中的代码。Platt 举了一个屋顶太阳能设计的例子:他认为系统在没有安装原始模拟器的情况下,复现了论文中的方法。
  • 但边界同样重要:据 Platt 的经验,ERA 还不能从零开始发现一个完全新的物理模型。如果没有被告知 Clebsch–Gordan coefficients 这样的概念,它无法可靠地凭空发明出整套形式体系。不过,在给定已知约束和相关文献后,它可以演化出包含这些要素的模型。

6. 多重假设检验,以及人类为何仍在循环中

  • 主持人担心,变异假设和超参数会扩大搜索空间,并带来严重的过拟合风险。Platt 的结构性回答是:科学家仍必须判断结果究竟是对世界的描述模型,还是只是在测试数据上取得了成功的预测模型。
  • 他说,系统还没有成为发现全新物理学的机制,但它是帮助科学家发现新科学的强力工具。
  • 对于普通过拟合,他的答案是,研究人员必须使用和过去相同的技术,同时格外小心,不要把自己骗过去。主持人将这一点概括为:“你需要使用相同的技术,但必须非常谨慎地使用它们。”Platt 表示同意。
  • 在品味与严谨性之间,Platt 认为两者都不可缺。有些人会专注于创造性假设和新科学;另一些人则会确保系统不崩溃、不出现糟糕的扩展性,也不产生错误结果。这些角色不一定要由同一批人承担。
  • 前沿发展并不平整:编码、知识搜集和寻找相关材料的能力大幅提升,但严谨性、哲学思考和创造力仍然较弱。Platt 的 AI co-scientist 曾让他了解到一种存在于岩浆中的陌生离子,但他不认为这能替代人类创造力。

7. ERA 起步于“AutoKaggle”——人类也会通过刷目标来获利

  • ERA 最初名为 AutoKaggle,因为 Kaggle 属于 Google:最初目标是构建一个能够参加 Kaggle 风格竞赛的系统。它在一些 playground 竞赛中表现非常好,尤其是在 CDC 举办的一项竞赛中,成功预测美国各州及领地下一周的 COVID 和流感病例。
  • 在其他竞赛中,它也曾接近胜出,但没能跨过最后一道门槛。Platt 说,成绩往往取决于一个人愿意花多少精力,把最后的0.001或0.01分挤出来。
  • 在 Google 的一个 contrails 竞赛中,人类参赛者发现标签存在半个像素的错误:人工数据旋转时,像素中心与左下角原点之间的区别被处理错了。他们利用这一漏洞,击败了赛事组织者。
  • Platt 的结论是,人类也可以表现得像 LLM 一样,通过刷目标来获利。因此,Goodhart 定律适用于每一个 leaderboard,而不只是自动化系统。

8. 航迹云:约占人为变暖的1%,或许只需下降两个飞行高度层

  • 持续性航迹云可以形成肉眼可见的“华夫格”状卷云。它们白天会反射阳光,但在约10微米、对应地球约300-Kelvin 黑体辐射的波长上,会吸收部分向外发射的红外辐射,并向两个方向重新发射。由于红外效应昼夜都在发生,净效应通常是增温。
  • Platt 引用的估算显示,航迹云造成的人为全球变暖约为1%。在欧洲等航班密集区域,航迹云卷云可能带来约1 watt per square meter 的局部辐射强迫,而全球平均的全部人为变暖约为3 watts per square meter。
  • 关键在于大气中肉眼不可见、却能长期存在的冰过饱和区域。这些区域可能呈煎饼状,高度只有几百米。在其中每排放1克水、冰或烟尘,大约就能抽取10 kilograms 的水,比例约为10,000∶1。
  • 飞机有时只需下降约两个飞行高度层,就能避开这些区域,但会付出一些额外燃油成本。Google 将卫星监测与定制的卷积模型或 U-Net 风格模型结合,因为天气模型还不足以直接准确定位冰过饱和区域,然后把地图提供给航班规划软件。
  • 降温效应并不是一个简单的地球工程机会。Platt 表示,增温总体上是持续性的,而有把握的降温主要发生在极地夏季;那时南极上空基本没有航班,其他极地附近的航班也相对较少。

9. ERA 的代表性成果:一个卡了两年的反事实问题

  • 估计航迹云的增温效应本质上是一个反事实问题:研究人员可以测量航迹云实际存在时发生了什么,却无法观测航迹云不存在的另一个宇宙。
  • 团队建立的向外长波辐射模型表现尚可,但反射阳光部分始终很难处理,连续两年没有突破。他们把人工航迹云注入数据,构造合成测试,而团队自己的方法都没能通过这些测试。
  • ERA 在可能的混杂变量中进行搜索,找到一个通过测试的模型。模型很简单,使用了团队此前从未尝试过的一组混杂变量。事后看,它显得很直观,但确实解决了一个现实中的瓶颈。这项工作曾在 EGU 上讨论,后续论文仍在准备中。
  • 相关的 CO₂ 工作同样意义重大。到2100年,未来生物圈吸收碳的规模不确定性可能达到约300 parts per million,而当前浓度约为440–450 ppm。通过卫星估计 CO₂,是降低这一不确定性的一个步骤。

10. 天气已经完成跃迁;气候仍然难以被 AI 攻克

  • 天气预报关注的是大气在约15天内的演化轨迹,预测窗口受混沌和蝴蝶效应限制。Platt 表示,机器学习、大规模数据集和充足算力已经让天气预报发生重大跃迁,这项工作主要始于约2018年,而不是最近的 LLM 浪潮。
  • 一个被重点提到的进步是热带气旋预测,包括更准确的路径预测,以及更早识别和处理增强过程。Platt 指出,飓风本质上是由海洋热量驱动的热机。
  • 气候更难,因为它关注的是系统的统计特征,而不是某一条近期轨迹;与此同时,底层系统本身也在变化。按 Platt 的表述,天气问的是你位于吸引子的什么位置;气候问的是吸引子的统计特征,而吸引子的形状本身可能移动,甚至在临界点发生变化。
  • 当气候模拟出现不稳定时,研究人员可能无法判断这种不稳定究竟是真实的物理效应,还是模型失效。未来30年的直接数据也不存在,无法告诉我们某次外推是否过拟合。
  • Platt 将过程模型描述为一种还原论式拆解:把复杂的气候系统分成许多部分,有时每一部分还要单独做专门拟合。冰微物理体现了这一难题:晶体形状、下落速度和混合行为都存在不确定性,却可能显著影响气候模型。
  • 他希望未来出现一种系统,能够读取并整合远多于单个研究者所能掌握的、经过更充分评估的科学知识,同时以受约束且符合科学常识的方式拟合数据。他明确表示,当前的 ERA 还无法完整做到这一点。

11. AlphaFold 的考验:什么时候黑箱已经够用?

  • 主持人认为,AlphaFold 迫使人们重新思考科学理解究竟需要什么。数据驱动模型可以极其有用,却未必提供物理学家传统上追求的那种机制直觉。Platt 引用了那句格言:所有模型都是错的,但有些模型是有用的。
  • Platt 的判断标准主要取决于数据是否足够丰富,以及问题是否封闭。在有足够数据覆盖相关案例的领域——他以蛋白质和 PDB 为例——AlphaFold 这样的统计模型可以非常有效;而在气候这种开放、非平稳且需要外推的场景中,谨慎就重要得多。
  • 在 PDB 覆盖的范围内穷举式应用 AlphaFold,对 Platt 来说尤其能说明模型在数据充分覆盖的区域内是多么有用。
  • 主持人还主张采用还原论式组装:把各自可理解的部分组合成连贯的整体;相比之下,一个声称能够解释细胞的巨大黑箱模型可能很难被信任或检查。
  • Platt 认可从简单模型起步这一长期有效的基本纪律。主持人建议先拟合线性回归;Platt 补充了 SVM 这一简单基线。他们一致认为,即使有大量数据,许多生物学问题也很难取得超越简单基线的改进。

12. 核聚变:“3年后,而不是30年后”

  • Platt 表示,具备商业意义的核聚变在本世纪末实现“确实存在一定概率”,并把自己的个人估计表述为“3年后,而不是30年后”。他强调这是一种可能性,而不是确定性预测。
  • 他用于保持怀疑的工具是 Lawson 判据:密度、温度和约束时间必须共同超过某个阈值。每种核聚变路线至少有一个维度存在阿喀琉斯之踵,因此,某一个指标达到稳定水平的新闻标题,并不能证明完整判据已经满足。
  • 在与 TAE 的合作中,Platt 接触到 field-reversed configurations:在某些磁流体力学假设下,理论认为它们不稳定,但实际中却可以稳定运行。Z-instability 可能制造一种摆动,并由 PID 系统控制;但偶发放电仍可能损坏设备,让运行中断数周。
  • Google DeepMind 曾研究用于 tokamak 的控制系统,以防止 disruption。主持人将故障模式描述为等离子体能量击中真空室;他们开玩笑说,一次重大 disruption 可能把 ITER 变成一块极其昂贵的砖头。
  • 宏观经济层面的框架是“悲伤饼图”:气候问题不存在单一的万能解。更大范围地推进经济电气化,可能让电力需求增加约5倍;而当电池需要覆盖最后10–20%的需求时,成本会越来越高。因此,某种稳定或基荷电力可能很有价值,Platt 将核聚变视为潜在贡献者之一,而不是万能方案。
  • 对话还谈到一个颇具推测性的设想:在核聚变反应堆中把汞嬗变成金。Platt 认为这个想法很巧妙,但表示它可能行不通。

13. FireSat:在火灾还只有一间屋子大小时就将其捕捉

  • Platt 表示,野火在大约一间屋子大小时可能很容易扑灭,但一旦扩大到1英亩,就会难得多。火势有时会在一段时间内维持较小规模,随后迅速增长。
  • Google 设计了一个低地球轨道卫星星座,使用中波红外传感器,因为火焰会以温度的形式显著突出。根据轨道配置不同,大约50–80颗卫星可以在全球任何地点,在约15–20分钟内探测到直径约5米的火灾。
  • 传感器的原生分辨率约为50 × 50米,再通过多光谱超分辨率和其他 AI 方法,将有效火点定位分辨率提升至接近5 × 5米。这些传感器无法简单搭载在任意微型卫星星座上,因为中波红外探测需要制冷。
  • Google 与其参与其中的非营利组织 Earth Fire Alliance 以及 Muon Space 合作;一颗原型卫星已经发射。Google 还通过 Android 和 Search 传播火灾边界信息。
  • 团队与 U.S. Forest Service 合作,开发了 Rothermel 火灾传播模型的快速神经网络代理模型。Platt 引用 World Health Organization 的估计称,野火烟雾每年造成约300,000人超额死亡。
  • 他对气候韧性的理念是,社会可能需要同时进行适应和减缓:气候变化就像一种严重疾病,因此既要治疗症状,也要处理根本病因。

14. 相变之后的建议:深入一个领域,也要自己爬山

  • Platt 认为,过去12–18个月发生了相变。这个领域正从一个个独立构建的专用模型,转向通用 AI 系统;但 AI for science 社区仍在摸索最好的工具和工作流,尚不能把这次转型视为尘埃落定。
  • 他举了自己21岁儿子的例子:儿子正在 RNA 和化学领域建立专业能力,同时使用 vibe coding 等 AI 工具。Platt 的建议是,在积极尝试新系统的同时,把一个领域钻深。
  • 当主持人问,如果 AI 能瞬间解决过去用来训练人的基础问题,人们该如何建立专业能力时,Platt 用徒步作比喻:即使可以开车上山,徒步登山仍然值得,有时也更有乐趣。
  • 他回忆自己曾用多种语言反复重写 boosting,并说这一过程教会了他 boosting。科学和技术工作就像运动训练,既需要高效完成任务,也需要花时间锻炼底层能力。
  • Platt 强力维护团队中20%的时间,用于学习和实验。他警告说,过度优化生产力本身也可能变成一种过拟合,并消耗掉创造性工作所需的时间。

15. 不要成为“空心西装”——以及“万能实验室”的愿望

  • 针对主持人关于人们可能变成监督 agents 的中层管理者的玩笑,Platt 表示他们不应成为“空心西装”。LLM 会犯一些不寻常的错误,不能被完全信任;但人类也应对自己的工作采取同样程度的怀疑。
  • 他回忆起与 Feynman 相关的一句警告:人不能欺骗自己,而自己又是最容易被欺骗的人。长期有效的能力包括生物学和物理科学等基础领域知识、数学、严格核查、创造力,以及跳出框框思考的能力。
  • Platt 的魔法愿望是一个“万能实验室”:只需接收一段 JSON,就能执行任何实验。这要求解决某种类似 AI-complete 的机器人学问题。
  • ERA 目前仍然是计算系统;物理数据还得由人收集。实验仍是真实世界的 ground truth,也是瓶颈,而通用的 lab-in-the-loop 系统仍处于高度开放的探索阶段。

16. 1982年与 Feynman、几乎为零的算力,以及 Hopfield 的回环

  • Platt 在1982年上了 Feynman 的计算物理课程,当时 John Hopfield 和 Carver Mead 也在相关环境中。每周二安排客座讲座,Feynman 会在周四解释为什么讲座材料是错的。Platt 将其称为“Feynman effect”:离开教室时确信自己已经理解,走出门后才意识到其实并没有。
  • 课程讨论了可逆的台球计算机、热力学极限,以及应该追求哪类计算。这门课由 DARPA 资助,学生还被要求参加每周一次的 prime-rib 晚餐。
  • Platt 形容当时可用的计算能力“按四舍五入误差计算等于零”:他作为 Carver Mead 的系统管理员,使用一台约1 MIPS 的 VAX 11/750,而研究组共享一块80-megabyte 硬盘,大小大约相当于一台洗碗机。
  • NeurIPS 源于一次名义上私密、却吸引大量人试图参加的 Snowbird workshop,此前还有 Santa Barbara 和 Caltech 的“Hopfest”活动。Platt 指出,transformers 可以被理解为一种 associative memory,这与标题“Hopfield Networks Is All You Need”相呼应。
  • 他说,神经网络受益于围绕 BLAS 和 GPUs 形成的硬件与软件生态。他不声称知道人脑是否通过矩阵乘法工作。神经网络开始大幅超越 CPU,大致是在最初的 ImageNet 和语音识别突破时期,当时 GPUs 对这些工作负载变得实用得多。
  • 主持人补充了一个故事:Platt 读研究生时的朋友 Brian Catanzaro 曾用一次简短交谈说服 Jensen Huang,相信 CUDA 和 GPUs 可以服务于 deep learning,这推动了 NVIDIA 转向这一市场。
  • Platt 还表示,是他创造了“convolutional net”这一通用术语,因为他不在 Yann LeCun 手下工作,也不想把自己的工作称为 LeNet。

17. 两颗小行星、一座 Oscar,以及量子计算的稳步推进

  • Platt 在1980年代 Caltech 的行星科学课程中发现了自己的小行星,当时授课者是 Gene 和 Carolyn Shoemaker。学生们把同一片天空拍在胶片上,用立体镜寻找移动的“浮点”,再以参考恒星为基准测量,并把观测结果发送给 Minor Planet Center 的 Brian Marsden,将不同时间的观测连接成轨道。
  • 他发现了两颗小行星。其中一颗以他的父亲命名;后来 Carolyn 将另一颗以一位教授命名,Platt 抱怨自己原本想给它命名,于是 Carolyn 把另一颗交给他命名。那颗小行星后来以他的母亲命名。随后,一位业余天文学家通过掩星观测发现了这颗以其母亲命名的小行星周围存在一颗卫星;Platt 说他相信这件事发生在去年。主星估计直径约4 km,卫星约1 km。
  • Platt 表示,Vera C. Rubin Observatory 的 Simonyi Survey Telescope 在6周内发现了11,000颗小行星,但他不确定其中有多少最终会获得正式名称。
  • 他的 Oscar 来自一项始于1986年 Schlumberger 实习期间的工作:他用弹性模拟对织物和其他柔性物体建模。这项工作后来发表在 SIGGRAPH,其衍生技术成为 Pixar 电影所用物理模拟的一部分。大约20年后,他获得了这座奖项。
  • 谈到量子计算时,Platt 表示自己以数十年为尺度评估进展。这个领域目前处于 NISQ 时代,这个术语由 John Preskill 创造,但 Platt 并不特别喜欢它。他将 Google 的 Quantum Echoes 工作描述为一种 Feynman 式方法:把参数化 Hamiltonian 拟合到 NMR 测量等观测数据上;它能否达到足以带来重大突破的规模,仍未有定论。
  • 他说,Google 的量子团队正沿着 Hartmut Neven 的路线图推进,时间尺度是几年到数年不等;具体里程碑的时间点应向 Neven 核实。超导量子比特因可扩展性仍令他看好,但其他架构最终也可能胜出。
  • Platt 不期待出现瞬间相变:量子系统是模拟的,而且极其脆弱难控;把物理量子比特扩展成有用的逻辑系统仍然困难。主持人指出,物理量子比特到逻辑量子比特的开销,以及连接性,都是重要约束。
  • Dave Bacon 开玩笑说,Platt 总是领先时代20年;10年之后,他就“走到了一半”。主持人则提出另一种理论:Platt 可能本身就是因果变量——他开始研究某个领域,现实最终才追赶上来。
完整逐字稿
Speaker 1

Are you talking about introducing explicit priors that you know based upon some human intuition, or maybe, in this case, LLM intuition?

1. Inside ERA: Tree Search, Parallel Experiments, and Shared Context

John Platt

When you talk about multiple hypothesis testing, there are predictive models and descriptive models. A predictive model is like—let’s say you have some inputs and some outputs, and you just want to build a piece of code that tries to have the lowest error rate on some data set. A statistical model, a descriptive model, is actually what science is trying to get to: it should be able to extrapolate because it has sort of the physics, or an actual description of reality, captured within it, and then you can use it to extrapolate.

2. Studying with Feynman and the Early Days of Quantum Computing

Yes, Newton thought of apples and gravity, but gravity isn’t actually about apples, right? If you take the 17th-century machine-learning model—“Apples will fall”—but how about planets? “I don’t know. I have no data about planets, so who knows what they do?” The distinction between those is a little blurry, because when a physicist or scientist comes along, they use their intuition, or maybe even more than intuition. Essentially, there’s a solid pile of facts that they know about the world, and then they make sure that whatever model they build is consistent with what’s known.

Speaker 1

3. John Platt: Google Fellow, ML Pioneer, and Academy Award Winner

My co-host is R.J. It’s a pleasure to have John Platt with us today. John is a Google Fellow and head of applied science at Google Research. He has a really fun background. I guess you described yourself, when we were talking a few minutes ago, as a mega nerd.

John Platt

Oh, giga nerd.

Speaker 1

Giga nerd. Giga nerd. He’s excited about absolutely everything, and it really shows. Correct me if I’m wrong about any of this stuff, but you started college at 14 and started your PhD at 18 at Caltech. You were advised or co-advised by John Hopfield, right?

John Platt

Oh, yeah. Yeah.

Speaker 1

Yeah. Yeah. Who just won a Nobel Prize two or three years ago.

Speaker 2

Yes. John was responsible for several textbook algorithms. One is known as Platt scaling; another is sequential minimal optimization, which is the textbook algorithm for training SVMs. Even today, if you use scikit-learn, it’s there.

John has discovered and named 2 asteroids and has an Oscar for technical developments from 2006. So, if you’ve ever watched a Pixar movie, you’ve seen John’s algorithms and work. John has an Erdős–Bacon number of 6, or 3 and 3 from either side, and I’m going to skip over 20 years of your career, but then, jumping to Google, you worked at Google Research on fusion, quantum computing, climate modeling, and many other topics. Is that more or less right?

John Platt

That’s right, yeah.

Speaker 1

Okay, cool. Did I miss anything important for today?

John Platt

No. I’ve also done lots of applied math and signal processing and all sorts of fun things.

Speaker 1

Yeah. I think your Wikipedia has a fun story about patents and the iPhone, too.

John Platt

The iPod.

Speaker 1

iPod. Yeah. Welcome.

4. ERA: Turning Scientific Problems into Scorable Tasks

John Platt

Thank you. Thank you for having me.

Speaker 1

Can you tell us about ERA? Is that the way the acronym is pronounced?

Speaker 2

I know there’s a lot of different, semi-related stuff out there, both within and outside of Google. So what can you tell us about the details of ERA and what makes it special?

John Platt

We’ve been doing AI for science in Google Research for more than 10 years now. Around 10 years ago, it was very much about using what we might now call classical machine-learning models—things like convolutional nets or whatever—to build specific models to solve specific science problems.

But about 2 years ago, we got very excited about these more general LLMs that have popped up in the last few years. We were wondering what could be done with them, and of course, a lot of people have been playing with them and trying to figure out the right thing to do. We kind of stumbled into this mapping. In other words, we found that many different scientific problems can be mapped into something we call scorable tasks.

You can often phrase a scientific problem as, “I really would like to have a piece of code that maximizes some score.” It’s surprising how many different scientific problems you can make a lot of progress on by mapping them into that framework.

A lot of scientists spend time building models. They might be statistical models, or they might be physically based models. If it’s a statistical model, like in machine learning, your scoring function is, “I have some data set, and I’d like to have the fit between the model and the data set go up.” We can talk about overfitting in a minute.

Speaker 2

That was one of our questions.

John Platt

Machine learning is a subset of this sort of scorable task, right? But you can do other kinds of things as well. Michael Brenner, who’s the lead author on the ERA paper, is very skilled because he likes to knock out a scientific paper in an evening with this tool.

There’s something in applied math called asymptotic expansions. You’re asking how an ordinary differential or partial differential equation behaves. Say it’s an ordinary differential equation, and there’s some parameter with an epsilon in it. You’re trying to determine how it behaves as epsilon goes to 0.

It turns out you can turn that into an empirical task by essentially asking it to propose some solutions that are asymptotically correct. You check to see if the asymptotic solution is correct for, say, epsilon equal to 1e-4, and then you check the fit. Then you ask Gemini, which is the core AI underneath it, to do the mathematical reasoning—to try to solve the problem while also maximizing the fit to the data.

There are a lot of tricks you can use because the underlying thing that’s altering the code, or making the decisions, is not a random process. It’s an AI itself that is smart, knows about things, and knows a lot about the world. You can solve a lot of interesting problems because that core inner loop is an AI that has huge amounts of prior knowledge.

So that’s the trick. We’ve been running around trying to map lots of scientific problems into scorable tasks and trying to solve them. It’s really been kind of fun, and I’m happy to talk about the ones that I’ve been involved in, at least.

Speaker 1

Yeah, I would love to hear about some of the more interesting ones. The statistical one is what everyone listening will probably know about. What you just mentioned makes sense. What are some of the other interesting ones?

John Platt

Let’s see. We have interesting ones that aren’t statistical. One that we just put a paper up on arXiv—actually, I think it might be on GitHub—is something you often run into in remote sensing, because there’s always a trade-off.

There are satellites flying above the Earth, and there’s a trade-off between how frequently they can revisit a spot on the Earth, what their spatial resolution is—how big the pixels are—and their spectral resolution, meaning how many bands they have. Ideally, you’d like to have monitoring of the Earth that’s constant: a frame every 5 minutes, at hyperspectral resolution, at 10 centimeters. You can’t get that.

For example, to monitor the atmospheric concentration of CO₂, you can take data from one satellite, such as OCO-2 or OCO-3. OCO-3 is attached to the International Space Station, so it gets you a little strip of CO₂ measurements that are highly accurate and have pretty high resolution.

Because a lot of it is in the infrared, you can use weather satellites. GOES has some infrared bands, and it takes a picture essentially every 5 minutes, but the pixels are very large, and it doesn’t have great spectral resolution because it wasn’t designed to find CO₂.

You can ask one to estimate the other and add in other data, such as the current weather and the long-term albedo. They came up with this very nice model that can do almost like super-resolution—an informed super-resolution—from one satellite to another. So that’s one example.

Speaker 1

Yeah. Okay. So any scientific problem that you can map into this framework—the input to ERA is sort of this mapping, and the output is code?

John Platt

Well, sort of. The way we’ve got it set up in the product is that you just start talking, right? A lot of times, it’s nonobvious how to do this mapping. Although experts like Michael Brenner know how to do it, we wrote an agent that helps you. It talks to you to try to help you define what your scorable task should be.

So already there’s an instance of Gemini sitting there trying to help you write code. That’s almost like an intermediate result. You start talking to it about your problem, and it tries to produce essentially a Python notebook underneath that has a score—a function that’s scorable, which essentially produces a score. Then it starts to mutate that notebook in a clever way because, again, it’s Gemini, and it will try to keep proposing code that maximizes the score.

Speaker 1

So what is different about this than just a general agentic system that can optimize notebooks?

John Platt

Right now, it's essentially its own thing. In modern 2026 parlance—we actually worked on this in 2024 and 2025—it's a specialized harness that runs an algorithm, which in our case was Monte Carlo research.

It's keeping hundreds or thousands of possible instances of notebooks, and then it selects one. I can explain how it selects one. Gemini asks itself, “What can I do to make that notebook better?” Then it makes a new one, tests it, and puts it back into the candidate pool.

You can imagine the candidate pool is tree-structured because every candidate possibly has some children. What you do is pick based on something called the upper confidence bound, or UCB. It's actually a fairly standard algorithm from reinforcement learning.

You essentially pick—it’s an optimistic algorithm. It tries to estimate, say, what the 95th-percentile outcome of a mutation would be, and it picks the one with the highest bound, the highest optimistic bound. In other words, it doesn't always pick the best-performing notebook. It tries to predict the current performance plus 2 sigma of its guess, so it's always trying to hunt around.

Speaker 1

So, high recall, basically?

John Platt

High recall. It's trying to make its bets so that it most efficiently makes progress, which isn't always greedily doing the best candidate. Sometimes it's the fifth-best.

We've also played around with recombining candidates. It takes ideas from 2 candidates, smashes them together, and tries to make a third candidate out of that.

Speaker 1

How does it seed the initial candidate pool?

John Platt

That's the amazing thing: underneath, Gemini is actually good at writing code. You just ask it, “Write me a thing,” because you have a textual description of the problem.

It isn't just, “Here's your scoring function. Start.” You give it a textual description of the function, and you might also give it, in fact—we have some things like this—5 papers that people tried to solve the problem with. It's smart enough to go and read the papers and take a first stab at the code.

It might not be great, and sometimes it has bugs and returns essentially minus infinity, but it will then try to mutate the code and make it better.

Speaker 1

It's pretty cool. You don't actually have to give it starter code. You can, if you want, but you don't have to.

How many agents are you spinning up? I guess maybe not agents—or how many different tree branches are you spinning up at each iteration?

John Platt

At every iteration, there's a trade-off. You'd like to do a lot of parallel work, but if you do too much parallel work, you can't learn from previous things. Right now, the default is about 10 parallel searches, so you try to grow 10 leaves at a time.

Speaker 1

Okay.

John Platt

That seems to be about the right trade-off.

Speaker 1

When you say you can't learn from previous iterations, does that mean that the orchestrator—or is there some sort of... You said that there is some step that is able to recombine or make decisions beyond just the score.

I guess maybe one of the questions is that, as a human, when you're doing some sort of ML project, you don't just look at the one metric that you're trying to optimize. Oftentimes, there are orthogonal metrics. Sometimes there are even insights, such as watching training curves, that can give you intuition about what's going on, or looking into specific examples. Does it do any sort of introspection like this?

John Platt

It has the history. What I meant was, why you can't do too many things in parallel is that if you have 10 parallel searches at once, number 1 can't actually see what numbers 2 through 10 are doing. If you do 1,000 at once, you're using a huge amount of computation without a lot of cross-learning.

Whereas once you finish a little batch, you get the history. Obviously, you have to prune it so it doesn't blow up the context, but you get the history of what it was thinking about as it was writing the code and the results of the code. So it can learn from its previous attempts.

Speaker 1

Okay. And does it learn across?

John Platt

Oh, yes. Essentially, it's one shared context. It is thinking as it goes along. It's not like there are 1,000 completely independent branches.

Speaker 1

You're really pushing Gemini's long-context abilities.

John Platt

That's right. You have to do the right management and everything.

Speaker 1

5. Reward Hacking, Scoring Functions, and the Gemini 2.5 Breakthrough

Yeah. Yeah. Yeah. Okay. Oh, that's cool. Going back to R.J.'s question, the key point here is that the first goal is to identify the specific score that you're trying to optimize, right?

Sometimes that is the hardest part of the problem. I find it interesting because I'm not sure I'd always trust my agent to do that part. That seems like the more human task in the loop.

John Platt

It is. Often, you have to be careful. A lot of what you do is very meta. Everything we do is at a very high level.

One common thing is that you come up with a scoring function—the agent does, or you do it together—and then the iteration finds a way to cheat or exploit a hole in it. You think, “Oh, no, I didn't mean that.” You have to go through and often play with it, have a loop around it, and iterate: “No, no, I didn't mean that.” Or you have to tell it in its instructions, “Don't do this.”

So there are often iterations. Even with agent help, you don't necessarily get the right scoring function from day 1.

It's really neat because, in the old days—in 2024, for example—a lot of grad students would spend a lot of time doing scientific software. It took so much effort just to write code that you might try a few things, or a few things that were very related, and then stop because you had to write your paper or move on to your next experiment.

This thing is relentless. It keeps trying and keeps trying, so the people who use it are now spending almost all their time at the right level—almost at the scientific creativity level. What does it mean to have a cost function? That's almost the essence of the scientific problem.

You're not so much in the details of, “I have to import this CSV file,” or, “I have to get this database to work.” You're thinking deeply and almost philosophically about your actual scientific problem, rather than being down in the grungy goop of worrying about databases.

One cool thing the agent can do is suggest datasets to you: “Have you thought about pulling in this dataset and doing a join?” It makes suggestions about datasets you can join with, which is kind of cool.

Speaker 1

Going back to what you said a second ago, agents love to hack things and reward-hack. Are there any fun or interesting stories about where things comedically ran off the rails?

John Platt

Boy, I'm blanking. I know other folks have run into it. I don't know if I have enough details to express the comedy of it, but it does—you kind of get—

Speaker 1

Surprised. Yeah. I don't know if I have any really concrete examples. Sorry, I'm blanking.

John Platt

No, it's fine. I always like to think of machine learning as being like the old genie stories, before “The Monkey's Paw.” You have to be careful what you wish for because you're going to get it.

Speaker 1

That's right. And you have that, and that happens very much with this. You have to be careful.

John Platt

On the other hand, it has some knowledge. The nice thing is that Gemini knows a lot about many things—more than any one person can know. It at least knows a lot, especially if you point it to papers and say, “Here are 5 papers that tried to do this in some way.”

To some extent, it does have that genie feel, but to some extent it also does sane things. Remember the whole idea of evolutionary coding? It's been around since the 1970s. Everyone has loved doing that: “Let's mutate Lisp code or whatever to do things.”

The reason it hasn't taken off is that random mutation in code space is pretty much worthless. Just like DNA, most mutations are harmful. Here, we can actually find it. It knows interesting gradients to try, which is why the thing works: the underlying loop itself is an AI.

Yes, it can overfit and have funny genie problems, as you allude to, but it also has some amount of sanity because it knows about the world and has world knowledge in it.

Speaker 1

The paper, though, was doing Gemini 2.5, and I think Gemini has advanced quite a bit.

Do you have metrics? Is this a tool that you're continuously using? It sounds like you're improving, and I'm wondering whether you've seen almost a phase transition internally in how effective this tooling has been. How dramatic has the improvement been over the last year or two?

John Platt

Oh, well, year or two? Yeah, amazing. In other words, every half version of it—essentially, I think it would have been impossible under Gemini 2.0. [laughter]

Speaker 1

Yeah, I think so. In other words, it wouldn't have worked. So, you started at 2.5, and that was like just the—

John Platt

Oh, no. We've been trying to experiment with these things for a while. Things just weren't working, then they started to work, and now they're just amazing. The progress on Gemini's major versions has been stunningly amazing.

Speaker 1

Yeah. I think this is an experience a lot of people have been having, where things that seemed impossible are suddenly becoming magically useful, really quickly.

John Platt

And so, if people are—I even say this to scientists, because some people say, "Oh, I tried 2.0 and didn't like it." That was a long time ago. That was a year ago. That was an eternity ago, right?

In fact, we even have one of the preprints where we've sort of combined ERA with Antigravity. The whole Antigravity harness is pretty amazing, too. That's the one where you can pull in lots of papers, and it can write lots of code for you.

Speaker 1

So, yeah, is that publicly available, or is that—

John Platt

The Antigravity? Yeah, yeah.

Speaker 1

Well, Antigravity is certainly public.

John Platt

Oh, sorry. Sorry. The ERA plus Antigravity.

Speaker 1

Not yet.

6. Predictive vs. Descriptive Models: The Limits of AI Discovery

John Platt

Okay, okay. Not yet. Okay.

Speaker 1

I find this area really fascinating because, like you said, there's been some form of code mutation out there since the dawn of computer science, basically. The canonical problem is the overfitting or multiple-hypothesis-testing problem, I think, which is maybe a better match to the problem. My hypothesis is that this algorithm works, and so you run the risk that it has exponentially exploded, because suddenly I have these hyperparameters that I'm optimizing. You have an explosion of state space that you're exploring, and it seems much easier to overfit to a problem.

What are your thoughts about that? On the other hand, empirically, my experience is—I even tried the open-source version of ERA. I strapped it into Claude, and it's running right now, so I can't tell you how well it's working.

John Platt

Okay, I'm curious.

Speaker 1

7. FireSat: Detecting Wildfires from Space and Building Climate Resilience

Yeah, I'll let you know. But I'm curious to know—this is a question that's been on my mind about general AI for science. What are your experiences with this on the ground?

John Platt

I guess there are 2 questions embedded in your question, right? When you talk about multiple-hypothesis testing, there are predictive models and descriptive models. Can you tell I've been doing machine learning for a long time in statistics?

Speaker 1

That's actually a really good point, though. Do you mind explaining and expanding on that? I'm not sure that's something that everyone in the audience would be familiar with.

John Platt

Especially in modern days, I think people are trying to obscure the 2.

Speaker 1

If you started with LLMs, I'm not sure that distinction would be meaningful.

John Platt

That's right. So, a predictive model is like—let's say you have some inputs and some outputs, and you just want to build a piece of code that has the lowest error rate on some data set. That's just a statistical model, right?

A descriptive model is what science is actually trying to get to: it should be able to extrapolate because it has the physics, or an actual description of reality, captured within it. Then you can use it to extrapolate.

It's sort of like—Newton thought about apples and gravity, but gravity isn't actually about apples. If he had just fit—if he had taken a 17th-century machine-learning model and said, "Apples will fall, but how about planets? I don't have any data about planets, so who knows what they do?" [laughter]

Speaker 1

So, when you say "extrapolative," I'm curious. I could think about this in different ways. One is, you said a model of physics or a model of the world. Are you talking about introducing explicit priors based on some human intuition—or, in this case, LLM intuition—or are you talking about the physics actually being learned by the model, or—

John Platt

It's more—

Speaker 1

The underlying process of the world is under the model.

John Platt

The distinction between those is a little bit blurry, right? When a physicist or scientist comes along, they use their intuition—or maybe more than intuition. There may be a solid pile of facts that they know about the world, and then they make sure that whatever model they build is consistent with what's known.

Currently, in ERA, it is LLM intuition, essentially. That's what I was trying to say about having a good gradient underneath: especially if you point it at existing papers, it will try to build models that are kind of sane underneath. If you give it guidance, like, "Be sure to incorporate this and this," or, "Look at these papers," you get these things.

So, you can introduce a bias toward certain model choices, and it will have a bias because its own little bit of world knowledge is accumulated inside itself during pretraining.

Speaker 1

Can you give us examples of what that might look like? Is it modeling something in a way where it's actually—there are different, for example, if you're doing something with partial differential equations, there are formalisms people have, like U-NO neural operators, or where you can encode a differential equation in some sense? I think that's called physics-informed neural networks, or something. That's one for fluid and kind of PDE-type modeling systems.

For, let's say, molecular systems, there's often this idea about equivariance. Is the model picking up on tricks that have been developed in the literature, or are they adding some weights for some physical prior? What does that look like when it introduces some sort of knowledge like that?

John Platt

I don't know if I have enough data to say, "73% of the time it does this," but especially when you point it at existing papers, it will try to adapt the methods described in the papers to the problem. In fact, it will do an amazing job. You can often just recreate or reverse-engineer a paper.

Michael Brenner actually likes to do this. He'll say, "Oh, that sounds like an interesting paper." We did this, actually, for something—it was kind of a hack. I suggested this to Michael: there's this MIT professor who came up with some code to do—essentially, if you have a rooftop with a fixed area and you want to maximize the amount of solar power you capture over a day, or solar energy, you can build it up so it captures more sunlight.

You can have it design a widget involving mirrors or struts or solar panels, at whatever angles or sizes you like, usually with a maximum height, and then let it explore that design space.

I believe Michael—we can ask him—I believe it actually just—I don't think he installed the simulator. I think the code just reproduced the code from the paper and figured it all out. It has a coding agent inside it, so it just reproduced the code from the paper and figured it all out.

So, yes, it's very good, especially if it's given a pointer to what other people have done. It's very good at, "Oh, I haven't seen it try equivariant modeling." That can get very hairy if you know about Clebsch–Gordan coefficients; it's pretty fun. I don't know if it'll do the true equivariant stuff, but it actually knows a lot.

I remember when Gemini 2.5 came out—I know this isn't exactly about Evo—but I remember thinking, "This is a new world." The day 2.5 came out, I said, "Hey, Gemini 2.5, can you write me some boosted decision tree code?" And it did.

Speaker 1

And it worked.

John Platt

It just did. Yes. I said, "This is a new world."

So, to loop back to your question, yes: if you give it guidance about what's important to put in there, it will. It won't necessarily—at least not that we've seen—discover completely new physical models. If you didn't know about Clebsch–Gordan coefficients, go fish; I don't know what you mean. If you didn't know about something, it won't completely discover a new kind of physical model from scratch.

But if you tell it about interesting constraints about the world that are known, it will certainly evolve.

I don't know if I answered your question.

Speaker 1

So, there's still room for humans for the next year or two?

John Platt

Oh, in fact, going back, I think there's totally room for humans. We have an AI co-scientist that tries to help you with hypothesis generation, but really, I still haven't seen the creativity, the philosophy, and the careful rigor. You totally need the humans.

I don't see humans going away. They can make strange suggestions, and I've used the co-scientist for an interesting problem in geochemistry. I learned about a new kind of ion that I didn't realize occurred in magma, so it'll tell you interesting things and you'll learn stuff, but I don't think it substitutes for human creativity.

Speaker 1

Going back, we were just talking about two years ago. You had Gemini 2.0 to 2.5, and you already saw a leap. Now it's been another year or two, and we're at 3.5, and you're saying this is working much better.

Whenever you look at a graph, if something looks like an exponential, it can either be a sigmoid or you can be at the beginning of a takeoff. I guess every exponential turns into a sigmoid eventually.

John Platt

Every exponential turns into—

Speaker 1

But the question is, where are we on that?

John Platt

I guess I'm a big believer in the whole jagged frontier.

Speaker 1

Yeah, the jagged frontier.

John Platt

And so, certainly, at least from what I see—I don't know what's going to happen in a couple of years—there are some big spikes in that jaggedness in terms of coding ability, gathering knowledge, and finding related things. That's huge and wonderful, and I think it's great for scientists.

So far, it's less in terms of rigor. We can talk about things like the International Math Olympiad and math in general, but in terms of philosophy and creativity, I think it's still not there. Maybe it'll get there. Some people are saying everything's going to inflate and pass, but I'm still seeing a lot of very strong jaggedness.

8. Overfitting, Auto-Kaggle, and Keeping Scientists in the Loop

I can see that maybe it'll get extremely good at coding, fitting models, making suggestions, and things like that, but I don't know. So far, you need the humans.

Speaker 1

Yeah, I want to get back to the question about multiple hypothesis testing.

John Platt

Sorry, I totally love the tangent. Multiple hypothesis testing is when you have a descriptive model and you're saying, “This is the way the world works,” and you have a bunch of data, and you take a billion darts and throw them. So you have to be careful. There's something called the false discovery rate.

The question is: Is this finding descriptive models, or is this finding predictive models? Fundamentally, the scientist is there to make sure that whatever it's saying is descriptive. We haven't been able to make a system out of these pieces that can really discover completely new physics or completely new science, but this is a power tool to help you discover completely new science.

Maybe I'm trying to unask your question, but I think this gets at the heart of it. Yes.

Speaker 1

Yes. But then you're saying, what about just pure—okay, let's set aside the fact that it's not trying to figure out a descriptive model of the world. That's still up to the scientist. But what about just plain old overfitting?

John Platt

Yeah.

Speaker 1

So the question of how I don't slice my fingers off is: You need to use the same techniques, but be very careful with them.

John Platt

Yes.

Speaker 1

That's a very clear answer.

John Platt

Yeah. Okay. Yeah. I actually don't think I've heard any guests say that. I think it's a very important skill, maybe one of the most important skills in the new world. People talk a lot about taste.

Speaker 1

Yeah, but maybe this is a variation of taste.

John Platt

Taste is like the other side right now.

Speaker 1

It's like rigor. It's like the—

John Platt

Yes. Yes. In fact, if anything—

Speaker 1

Yeah. Taste or rigor—which one's more important?

John Platt

Well, I don't know. At least the way I'm viewing it, researchers are software developers, and there's a lot of overlap in a lot of fields. I'm seeing that software engineers are—it's almost like, obviously, there's a lot of concern: “Oh no, what am I going to do?” Coding seems to be getting automated.

I think there are sort of both. A lot of people get pulled into, “Well, I'll be the creative source. I'll try to figure out new science. I'll try to figure out new products. I'll try to really be very, very creative.” Again, I'm a strong believer that I don't think that's going to go away.

There are also people who are pulled toward rigor: “I want to make sure this doesn't crash. I want to make sure this scales. I want to make sure this isn't wrong.” I think you need both, and you need people who are really good at both. But they don't necessarily have to be the same people.

I think this is even broader than science. As software engineering evolves, there will be people who bring the creativity and people who bring the rigor, and I think those will be sort of anchors.

Speaker 1

There's some other related work I've seen out there. There's a really cool Claude leaderboard—an agent leaderboard for scientific problems from Stanford. I don't know if you're familiar with it. It seems like a really interesting idea to have different agents competing on the leaderboard.

If you squint a little bit, what ERA is doing is kind of a leaderboard, but it's internal and it's recombining ideas. What are your thoughts about this? Is that something you guys are working on, and are there problems with that or advantages to that?

John Platt

Ironically, the whole ERA project actually started because people may not realize Kaggle is part of Google.

Speaker 1

Oh, yeah.

John Platt

It was called the AutoKaggle problem. That was actually what it was: Let’s try to have a system that can win at Kaggle competitions. So that's sort of why it has the shape. That's sort of how the project started.

It goes back to overfitting, right? If you've ever actually competed in a Kaggle competition—

Speaker 1

I have done Kaggle competitions—or I've done one. It is a really interesting phenomenon because overfitting is rampant. It's really impressive how people can overfit to certain data sets.

John Platt

We even have a fun project that I can talk about more if you like. It tries to mitigate contrail generation. We had a contrail Kaggle competition, and people actually beat us, but they found that we had a half-pixel error in our labels.

It had to do with the center versus the lower left: Where is (0, 0)? Is it in the lower left of the pixel, or is it in the center? They found that and exploited it, squeezing out whatever little bit extra they could, because when you make artificial data and rotate it, you have to make sure that you take into account that half-pixel offset.

So, yes, people themselves will act like these LLMs and try to reward-hack these things. It goes back to Goodhart's law. Any metric that becomes a target is no longer good as a metric.

It's good, but you have to be very, very careful. You have to have layers of rigor: “Okay, we'll do this and optimize for this,” but you have to realize that Goodhart's law now applies, and you have to be careful. That's one of the reasons the whole AI field has been constantly exhausting these things. Goodhart's law applies individually to every leaderboard you make, so you have to step back and be very, very careful.

Speaker 1

That's really useful to my thinking. As we've had guests on, it's been a recurrent theme: How do you manage all this complexity that's introduced by LLMs and AI science?

I think my follow-up question was about overfitting in Kaggle. If you had an AutoKaggle problem, how often was it successful? I assume you probably just ran this on all of your Kaggle competitions or something.

John Platt

We tried it on various competitions, what they call playground competitions, and it did very, very well in those. We've entered it into different competitions.

Some of them, it turns out, in the last few years, the number of leaderboards and competitions and whatnot has exploded far beyond Kaggle. So we've done very well in some of them. One thing we're super proud of is that the CDC set up this competition where you try to predict the number of COVID and flu cases that will happen in every state and territory in the U.S. the following week. You try to predict a week in advance, and ERA did super well on that.

Speaker 1

It's funny because, in some sense, Google invented the concept of using data to track disease progression with Google Flu Trends. So it's kind of funny that you were going full circle 20 years later or something like that.

John Platt

That did very well. In other competitions we've entered, we weren't quite as good, often because how well you do in these competitions is a measure of how much TLC you put into them and how much you're willing to squeeze the last 0.001 or 0.01. ERA did well. It got you pretty close, but we didn't close the gap in the last 30 places or whatever because no one was there to shave the last 0.01 off the thing.

Speaker 1

Is it very iterative? You get to it, and it sort of—I saw the charts in the paper, and you get these step changes as it discovers something, and then it goes flat. So is it very much human-in-the-loop, like, “Okay, you've stalled on the problem. Try this kind of thing”?

John Platt

Yes, at the outer loop, which is, I think, almost like the more fun, creative part. So, yes: “Oh, here, have a look at this paper. Oh, you're doing something bad.” It's almost like having a hyper-eager grad student or something who doesn't sleep. You sort of tell it things and guide it around.

Speaker 1

How often does someone intervene versus what does the outer loop look like?

John Platt

It might run for a few hours and come back and give you some examples, and then you can, as far as you like, keep trying and keep poking at it.

Speaker 1

So that's very much designed to be human-in-the-loop, then.

John Platt

Yes. Yes.

Speaker 1

Interesting, because a lot of the other tools I've tried tend to be very one-shot.

John Platt

Well, I guess it depends on your definition, right? You talk to it and start it, and it'll go for some number of hours and then come back. But that's where the human creativity kicks in, and you're doing the outer loop, where every few hours—you know, it depends if you want to sleep—you go and give it another try.

Speaker 1

What kind of budget are you giving this thing? Like, you blew through a million dollars accidentally kind of thing.

John Platt

I don't actually know, because we're using internal calls to Gemini. So I don't actually know.

Speaker 1

But look, there's token budget, but then there's also solving a problem that's computationally expensive.

John Platt

Oh, yes, that also. Essentially, underneath it, the scoring function itself might have Monte Carlo estimation or whatever. You can end up using a lot of compute just to even do a simulation. If you have a simulator inside, it has to run a simulation. So, yes, you can spend a fair amount of CPU or GPU.

Speaker 1

My little experiment with ERA and Claude is to build a neural network for some classification problems. Obviously, if you have enough data, larger networks work better, but they're more expensive to train, and you start to run into a question: How do I manage my budget if I have a fixed budget so that I'm spending my dollars on the most effective solutions?

9. Contrails: Using AI to Reduce Aviation’s Climate Impact

John Platt

That's right. I think that's still something we need to figure out. But of course, it's no different from if you have a grad student and they're trying to train a very, very large neural network or a very, very large data set. They themselves have to ask things like, “Is there a scaling law? Can I extrapolate?” So it's the same problem, but maybe more urgent because it just runs into this problem because it's so relentless. It runs into the problem much quicker than a grad student could.

Speaker 1

One of the things that you optimized was contrails. Can you talk a little bit about that?

John Platt

Well, let me spend a minute or two talking about the contrails problem.

Speaker 1

Yes, please. For context, contrails, not chemtrails, which is a conspiracy theory.

John Platt

Yes. Although you should also dislike contrails, but maybe not for the same reason. So contrails are—if you've ever seen those white clouds formed behind jets, those are called condensation trails, or contrails. It turns out they add, at least according to the estimates that people have, about 1% of all anthropogenic global warming is caused by contrails. Why is that? I can talk about the physics of that.

It turns out there are actually 2 countervailing effects. Sometimes, if you've ever seen contrails, they streak and then go away. Those don't really do anything, but sometimes they last for a long time. You'll see in the sky almost like a waffle of persistent contrails. There are 2 effects that they have.

Those are thin white clouds, so they reflect sunlight, but that only happens during the day. It turns out all objects emit something called blackbody radiation, and the Earth does too, at whatever its temperature is—about 300 Kelvin. It's in the far infrared, around 10 microns. At those wavelengths, contrails have very low albedo. They're almost essentially black. They'll absorb a little bit of the outgoing infrared radiation and then re-emit it in both directions. So essentially, they'll reflect some of the outgoing heat, and it'll trap heat like a blanket.

Because that happens 24 hours a day, they tend to be warming. It turns out there's some uncertainty about it, but contrail cirrus—cirrus that comes from contrails—might cover a few percent of the actual sky in places like Europe, which has a lot of airline traffic. It adds about 1 watt per square meter of forcing locally, at least. Just to give you a sense, all anthropogenic warming, averaged across the whole globe, is about 3 watts per square meter. So in places with high airplane traffic, it can be a lot of warming locally.

What can you do? It turns out contrails are caused by areas in the atmosphere that are ice-supersaturated. They're a little bit like rock candy. When you have rock candy, you get a water solution that has too much sugar in it, and any little bit of sugar in it will just crystallize all the sugar out. These regions tend to be pancake-shaped and only a few hundred meters tall. If you fly through one, the jet exhaust has a little bit of moisture in it, which will turn into droplets and then freeze. For every gram of water, ice, or soot you put out in this bad region, about 10 kilograms of water gets sucked out. So there's an enormous 10,000-to-1 ratio.

It's a big problem. What you can do is figure out where these regions are. They're invisible, of course—these regions of ice supersaturation—and then tell the plane to go underneath. You only have to drop essentially what they call 2 flight levels, so it costs a little bit of fuel but not very much to avoid these bad regions.

We built a system that looks at satellite images and tries to detect where contrails are. We have essentially a continuous monitoring system, and then we try to build a model because it turns out the weather models aren't quite accurate enough to find these places of ice supersaturation. We build a custom model, again, like a convolutional net or a U-Net, to predict where they're going to happen. Then we give maps to flight-planning software so planes can dodge them and inexpensively reduce the climate impact of aviation by a lot.

Speaker 1

What's the physics behind why you can predict that? Is it just, “I see it in the satellite, and then tomorrow I think it'll be there because planes go to the same place,” or what?

John Platt

Oh, no. It's because you're trying to detect these regions of ice supersaturation, because they're very, very persistent.

Speaker 1

Oh, they're persistent.

John Platt

Oh, yeah. No one knows exactly, but they could last for days. Essentially, they're caused, they think, by warm, moist air being injected at the boundary of the tropopause, just at the bottom of the stratosphere. When humidity gets up there, it sort of sticks there for a long time and then gradually dissipates.

Speaker 1

Got it. So it's just a sort of—

John Platt

They're like bad spots in the atmosphere you don't want to fly through.

Speaker 1

Right. Okay. And so once you've established that, it's probably good for a couple of days at least.

John Platt

Well, you have to keep predicting where that is.

Speaker 1

10. How ERA Cracked a Climate Problem That Had Stalled for Two Years

Yeah.

And the models you're using—you mentioned CNNs or something like that?

11. Discovering Asteroids and Finding a Moon Around One

John Platt

That's right. And we haven't replaced those with ERA-level models yet. But a very interesting problem came up: You want to know how much warming a contrail caused and how much it added to global warming. For example, you might want to find the biggest ones because there's some fuel cost, and it may cost the airplanes a bit of money to avoid them. So you'd like to know how much warming a contrail caused.

But that's what they call a counterfactual problem. You made a contrail, and a certain amount of infrared radiation happened. We can measure that if we're careful, but what would have happened if there hadn't been a contrail there? That's very difficult to estimate because you can't access the universe where it didn't happen. So you have to make these things called counterfactual models.

I don't know if your listeners know this, but counterfactual models are actually pretty tricky to fit and make. Remember, there were 2 effects: reflecting the sunlight, and then there's the infrared. It turns out that measuring the effect of reflecting sunlight is actually more difficult, and we were stuck on it for 2 years.

We had a counterfactual model that worked okay for the outgoing longwave radiation, but not for the reflected sunlight. ERA actually helped us find a model that searched all the confounders and figured out how we could estimate it. We even had test code on artificial data, because you can inject artificial contrails into datasets and figure out how much the effect was because you injected it. Our own attempts didn't even pass our own tests, but ERA did, and it unstuck this problem.

We're in the middle of writing up a paper. We have a paper about the outgoing longwave radiation, but we have another paper that isn't submitted yet. We've talked about it at EGU, I think, and in that paper we actually solve this problem.

Speaker 1

And the models that ERA comes up with—are they just a big monstrosity of code, or are they pretty basic? Did you just need the intuition to develop them?

John Platt

Yes. In this particular case, it was actually more of the latter. It essentially helped identify what the model was. It was a very simple model with a certain number of confounders that we just hadn't tried in that combination before, and it worked very well. It came up with the model, and in retrospect it seemed obvious. So that was, I think, a big win.

Speaker 1

12. CO2, Weather Forecasting, and Why Climate Is Harder to Model

Yeah, that's interesting. I know you've done a lot of work in climate. What other stuff have you done?

John Platt

I think I talked about this, right? The CO2 thing was pretty fun. The reason why estimating CO2 in the atmosphere is an interesting problem is that we actually don't know what the carbon flux is into and out of the biosphere.

We emit a bunch of CO2 into the atmosphere, and some of it gets absorbed into the ocean, mostly through inorganic chemistry. Some gets absorbed by phytoplankton, and a lot of it gets absorbed on land. But the error bars around what happens are moderately large, and the error bars 50 years from now are very large.

In the models for 2100, we don't know how the biosphere will react to ever-increasing temperatures and CO2. So we don't actually know how much CO2 will be absorbed, and the uncertainty is 300 parts per million of CO2 just from the uncertainty around what gets absorbed.

Just to point out, right now there's about 440 to 450 parts per million. So it's huge. It could be seriously, amazingly awful or not great, but 300 parts per million is an enormous uncertainty. It would be really nice to figure out whether we can reduce that. The CO2 concentration is one step toward solving that.

Speaker 1

And I know that Google has made some really big improvements in climate modeling and weather prediction as well, right? I was at NeurIPS this year—this last NeurIPS—and I stopped by the climate track. I may have only had a chance to listen to 2 talks or something, but it really blew my mind, the sort of step change that has happened in the past 10 years or so in climate modeling. I know a lot of that happened at Google. Can you talk a little bit about what has happened at Google and other places that has allowed that really big transition in climate and weather modeling?

John Platt

Okay. People often collapse climate and weather together because they're fundamentally the same physics.

Speaker 1

Yeah.

John Platt

Although, at least for atmospheric physics, they're the same. Obviously, when you start having ice and land, climate is long-term weather. The complexity of a full Earth system model, which is a climate model, is much, much greater than that of an atmospheric model. You have to measure the water flux and the CO2 flux into and out of the land, as well as what will happen with ice.

There has been a step change with weather models. Weather is up to approximately 15 days, because the atmosphere appears to be chaotic.

Speaker 1

Yeah.

John Platt

I'm hoping—I don't know if I should explain chaos. Essentially, it's the butterfly effect: small perturbations, like a butterfly flapping its wings, can make the weather completely different in 2 or 3 weeks. Weather is trying to predict the actual trajectory of the atmosphere over, say, 2 weeks, and that has seen a huge step change.

That was based on a lot of the machine-learning work from around 2018, not even the new large language model work, along with a large amount of data and a large amount of compute. There has been a lot of clever work, much of it from Google, making new weather models. It's been great.

In fact, we had a really neat breakthrough, because now we can apparently predict the tracks of tropical cyclones much more accurately many days in advance. Places like Jamaica got hammered by a terrible hurricane, and a lot of the classic models didn't actually predict it, partially because of the intensification. That is often driven by the surface temperature, because hurricanes, as people might not realize, are essentially heat engines. They convert heat in the ocean into large atmospheric motions.

Weather has been great. Climate is much more difficult because you don't actually care about predicting whether it's going to rain in Seattle in 2070. You're trying to get averages, and what makes it difficult is that it's what they call nonstationary. The underlying physics is changing: plants are behaving differently, ice behaves differently, and so on.

It's very difficult to use classical machine learning on true climate models. That's why the whole discussion you had about descriptive models and multiple-hypothesis testing is incredibly important in climate. We have no data from 30 years in the future, and we don't want to wait 30 or 50 years to find out whether we were right or whether we overfit.

Whatever we do, we have to be careful and try to peel off subproblems. The problem of how to build a big model that can predict into the future while still being constrained by what we know is fascinating. I think it's still unsolved, but it's a great problem to have because of these uncertainties. We really would like to know what will happen to the climate in 60 years.

So far, AI has not revolutionized it because it's very resistant, again because of this data problem. It's a low-data problem.

Speaker 1

Does the butterfly effect—the chaotic nature of weather—also impact climate? Or is the time scale so large that you have a closed system, where maybe it's oscillating between poles or whatever, but when you look at it on that time scale, it's more stationary?

John Platt

It's unfortunately nonstationary in a different way. The original idea of chaos—well, maybe people came up with it elsewhere, but in meteorology it goes back to a person named Lorenz, who had a very simple model of ordinary differential equations.

The difference between climate and weather is that weather is: Where are you on the attractor?

Speaker 1

And climate is about the statistics of the attractor itself. The problem with climate is that we're altering it. So the attractor itself is changing and moving, and there could be—and everyone talks about tipping points—that means the attractor suddenly changes. The trouble is that that's very, very, very difficult to predict. So even the attractor's shape changes quickly.

John Platt

Or could. And the trouble is, when you run a simulator—

Speaker 1

—you don't know whether it went unstable because your model isn't great or because it's an actual physical instability.

John Platt

And it's extremely difficult to tell the difference.

Speaker 1

What do you do? Especially when, to me, it strikes me that not only do you not have future data, you really don't have much past data. You can do some measurements with ice cores and lots of other things to try to reconstruct it, but there was nobody with an instrument 100 years ago.

So if you have annual data or whatever, maybe, if you're lucky, you have 50 data points in any one location. Whatever we do has to be very constrained by what we know. But it's very difficult. I'm just telling you the horns of the dilemma people are in.

John Platt

People make these—in fact, people in applied science in general, and I would say climate is the most extreme—things called process models. What you do is say, "I'm going to be reductionist, and I'm going to take this horribly complicated climate system and boil it down to 1,000 pieces." Then I'm going to find the expert who wrote a paper about piece number 763, where they fit a cubic to some data.

One thing that's very mysterious, which is related to contrails, is how ice behaves in clouds. It turns out that everything is complicated once you dig into it. [laughter] When you make a contrail, how long does it last? It depends, because the way contrails can evaporate is that ice starts to accumulate, as I said, and then the ice crystals get big and fall.

Of course, how quickly they fall depends on their shape, which is not known. How much does the contrail mix the moisture from inside the contrail outward? Again, people have approximations, but they don't know. The uncertainty is very much confounded.

It's not just, "Oh, John cares about contrails." It turns out that the actual microphysics of ice has very strong implications for what climate models do. We just don't know. I'm trying to say that it's very gnarly and it's not a solved problem.

My hope is that with tools—maybe not tools from today's era, but perhaps tomorrow's era—because, remember, they can not only fit data, they can read papers. The question is whether they can read a lot more papers than we can. Maybe we can integrate all the data, or all the knowledge, that people have carefully evaluated—much more than any one person writing a piece of code—and fit data. That would be utterly glorious.

We don't have that today, but that's one of my hopes for an even more amazing tool in the future: something that really can write code in a sane way, or in a much more sane way, because it will be constrained by all the scientific knowledge that we've accumulated so far. That would be amazing. We don't have that today.

Speaker 1

I mean, that strikes me as being very similar to biology.

John Platt

Oh, yes. Oh, boy. [laughter] If you've ever played with biology or even looked at biology, there are so many exceptions and so many hacks in biological systems. Yes, oh boy. It would be amazing if we could have something that could really integrate all known scientific knowledge with data and try to synthesize new models and new things.

Speaker 1

I think what you're saying is that AI can be an unlock here to some extent, because the models are so piecemeal. They necessarily are piecemeal. Being able to assemble the jigsaw puzzle—not to mix metaphors, but to assemble a really intricate jigsaw puzzle—where having scale and capacity actually helps a lot.

John Platt

That's right. One thing these AIs have is that humans—even me, and I think I'm pretty good—just find it difficult to integrate across the N-squared different things. My N is the number of papers I've read in my life, and it's pretty big, but it's difficult for me to do that. Somehow, there's just so much data in those billions and billions of parameters, and you can also give the model access to read PDFs, that it can start to pull things together that people wouldn't.

I'm starting to see little indications of that inside of ERA. I'm not claiming that's what ERA does today, but that's my hope for where this is going to go.

Speaker 1

I think I've heard a lot of people suggest something like this: the route to intelligence is to combine LLMs with some form of search. ERA is doing something like that, right? Something that is maybe a very strong database lookup with a good search algorithm is one way. Of course, people still do the whole RAG thing.

Google itself—the 10 blue links thing—

John Platt

It was, or is, a form of AI even before we had LLMs, right? You could cast yourself back to 2010 or 2015. You could ask Google about literally anything, and it would tell you stuff.

Speaker 1

Surprisingly well, actually.

John Platt

Yeah, surprisingly well, because somebody on the internet has probably written about it. So if you can match that—

In fact, that was one of the reasons why I wanted to come to Google. It was such an amazing thing.

Speaker 1

I wonder how much of our audience did a search before Google and knows how bad that experience was. [laughter]

John Platt

Yeah, I remember 1998, I think. When Google was released, I was using AltaVista. I don't know—Google had just been released, and I used it. I just dropped AltaVista like a hot potato and started immediately using Google.

No, that is sort of a form of AI. It could be that just having access to all of that and keeping it in mind at the same time—

Speaker 1

That model of scientific discovery, insofar as it pans out, is kind of comforting, too, because it is reductionist. You can look at the individual parts and understand them. It found the exact things to assemble, but they're all fundamentally things that people have invented or that it's done iterations on. All those little pieces are individually understandable, and then you can also put them together into a coherent picture.

I think for a lot of problems, like biology or climate science, we wouldn't trust the answer unless it were in that shape.

John Platt

Yeah.

Speaker 1

Because if there were some giant black-box model that said, "This is how a cell works," would I believe it? I don't know if I would believe it, because I can't examine it. But—

John Platt

I mean, to argue against it, though, if it works really well—

Speaker 1

But you'd have to gather—you'd have to test it, obviously. A statistical model has to extrapolate.

John Platt

Yeah.

Speaker 1

Yeah, and it has to extrapolate to the extreme, to the black-swan events.

John Platt

That's right. It's very hard. This is why things like self-driving cars are such a difficult problem. It's all corner cases.

Speaker 1

Yeah.

John Platt

It's kind of amazing how well they've done.

Speaker 1

13. AlphaFold, Simple Baselines, and When to Trust AI Models

That's an interesting point, thinking about coming from the world of physics, where a model was usually a single equation or a small number of equations that uniquely define a system and everything about it. You just find a solution to the system, and you now know everything you need to know.

I think something like AlphaFold was a shift for a lot of people. Before, they thought protein folding was a problem where, if we found the right force field and had the right computational engine, we would solve protein folding.

John Platt

The thought of really solving it in a data-driven way only appeared a few years before AlphaFold 1 came out. It's interesting that it has forced people to reevaluate almost what science is. AlphaFold and similar models are incredibly powerful. They've opened up a lot of things as tools, but at their core, they often don't give intuition in nearly the same way that most physicists historically would have wanted.

There's an old saying: all models are wrong; some are useful. Yes, Box said that.

Speaker 1

Yeah. When do you find the data-driven models to be sufficient, and when do you want something interpretable that humans can actually understand?

John Platt

I think it boils down almost to the difference between weather and climate. If you're in a data-rich regime, like weather—or even proteins, because of the PDB—you can feel, "Oh, yes. I've got enough data to cover this." A statistical model like AlphaFold can do that.

A lot of people happily use AlphaFold. I think it has really revolutionized my understanding. I'm not a biochemist, but people seem to love it. One amazing thing they did is that they exhaustively ran it on all PDB and published it, which is really, really cool.

Speaker 1

It's like six billion protein producers. The vast majority are actually quite accurate.

John Platt

Yeah, that's just amazing. But it feels closed, if you know what I mean.

When it's climate, it's open and nonstationary, or you have to make these big extrapolations, you have to be much more cautious. Biology may be similar. There may be parts of biology where—

There was this Virtual Cell Challenge from the Arc Institute—

Speaker 1

That had a funny result.

John Platt

I know that there may have been some overfitting, or at least that's what people were saying. Sorry.

Speaker 1

I guess, at a high level, the simple baselines—

John Platt

Work very, very well.

Speaker 1

Things, yeah, just like—

John Platt

One of the classic things whenever you do biology is to always start with a simple baseline. Maybe this is just good machine learning in general: analyze and understand your simplest case. In biology, there are many problems that are extremely resistant to anything beyond the simple baseline, even if you have a lot of data.

Speaker 1

That's right. In fact, I tell people the same thing: always fit linear regression. Just fit—

John Platt

Just do it.

Speaker 1

Just do it. Just do linear—

John Platt

Or SVMs. SVMs are just a different form of—

Speaker 1

Linear regression, yes. So when do you need the more process-model-y thing?

John Platt

I think it's when you have—I mean, climate is on one end and, I don't know, weather maybe on the other end. That may be too extreme, but I think it depends on where you are on the data-richness spectrum. When can you feel, "I really have a closed problem, and I think I can actually cover it"?

Speaker 1

14. Quantum Computing: Willow, Quantum Echoes, and Scaling Qubits

A closed problem that your data fully covers—yes, I think that makes a lot of sense in what I've seen as well. One thing I'm curious about is, when you're working on climate modeling, what are the broad things you're trying to accomplish? You've talked about contrails and CO2 predictions. What are the broad goals? Is one of them making interventions, and the other making predictions for things like insurance? How do you help adjust for some sort of climate change?

What are the principal goals, I guess, for you specifically or the community at large?

John Platt

I think, just like any community, there are probably many different goals. For me and my team, we're very, very interested in interventions: which ones are possible at a relative cost? Contrails were kind of amazing because it turns out that the intervention is quite low-cost. One amazing thing about contrails is that they are local, unlike things like CO2, because if a country decides to fix contrails over itself, it actually improves its—I mean, it has global effects, but it mostly improves the climate a little bit over that country. So they like that.

Speaker 1

I guess, though, if you are in a cold climate and you want to warm it up, this is now your own—you could own—

John Platt

Yeah. It turns out that it's a little bit asymmetric. The warming is essentially constant and global. The cooling only happens when the sun is at a good angle over you. So it's very rare that there are contrails where the uncertainty bounds are clear in terms of the warming.

There are many, many contrails, mostly at night, of course, where it's largely warming, and we're very sure—at, like, 2 sigma. There are not very many contrails where you say, “Oh, I know for sure that it's cooling, and I want more of them.” Only over the poles, in polar summer, do you know that the contrails are cooling, and therefore, if you got rid of them, they would warm up. But there are essentially no flights over Antarctica and not that many over the poles in the summer.

Speaker 1

So no one who's in a cold climate is going to use this maliciously.

John Platt

Well, yes. In that case, they wouldn't know for sure whether it was warming or cooling, so they would do things without knowing. We mostly just ignore it; we don't recommend that people fly those.

Speaker 1

You also brought up an interesting point about the economics. I think there was a lot of resistance historically to certain climate-change interventions, which, in some sense, the market has just taken over. At this point, renewables and batteries are almost universally and unambiguously just better than the alternatives.

John Platt

For nonmobile—I mean, you mean mobile? That's a really good point. Yeah, like planes: we do not have a solution to—

Speaker 1

Correct. I mean, there are some battery-powered planes, but they're very small and have very limited range.

John Platt

They probably will never actually be—

Speaker 1

It's hard to imagine the physics would be very, very different unless we came up with something like nuclear batteries. That would be kind of amazing, but we don't know how to do that.

John Platt

Or even if we did, I think people would be too afraid of a nuclear battery going wrong or something.

Speaker 1

Oh, yeah.

John Platt

Yeah. Since we don't know what they are, we don't know what the risk is. I guess we don't know the risk.

Speaker 1

Yeah. So we don't know. That's the problem. I've given talks about climate change, and I talk about the pie chart of badness—the pie chart of sadness.

John Platt

Which is, there's no one silver bullet for climate change, right? There are so many different things that contribute greenhouse gases across our economy. They all sort of have to be fixed, or many, many of them have to be fixed. So there's no one single thing.

15. Fusion Energy: Plasma Control, the Lawson Criterion, and Economics

I mean, I've worked on fusion. Fusion is cool, and it might actually knock a lot of them out if it's cheap enough, which we don't know, because we don't know if it'll work yet.

Speaker 1

Fusion is one of those interesting things where the joke was always that fusion is 30 years away, but I think it's actually now less than 30 years away—maybe.

John Platt

Yeah. No, I think there's a definite probability that someone will make commercially relevant fusion even by the end of this decade. So I think it's 3 years away, not 30 years away.

Speaker 1

This is very real. Interestingly enough, I think a lot of that actually comes down to materials science.

John Platt

Well, yes. Sorry, we could talk about fusion if you—

Speaker 1

Yeah. There were actually 2 things. One is fusion. The other is better control systems, which I think is actually—

John Platt

Yes. In fact, Google DeepMind has been working on control systems for tokamaks to make sure they don't essentially go unstable and go into disruption.

Speaker 1

Disruptions are quite interesting in themselves.

John Platt

Yes.

Speaker 1

Yeah. It's basically all the energy in the tokamak collimates into one little beam, and then—

John Platt

And it hits your vacuum chamber, and you're very, very sad.

Speaker 1

Very sad.

John Platt

Yeah. I think people believe that ITER could be turned on, have a $30 billion disruption, and then basically be a $30 billion brick or something.

Speaker 1

Oh, yeah. I guess you could try to patch it. I remember—

John Platt

Working—we worked with a fusion company called TAE, and I was in their control room. It was kind of sad. You have to be very careful. We were making systems to recommend new experiments, and they were very, very skeptical—and rightly so—about which way they should go, because even under human control, it's—

I'd been there while they were doing an experiment, and then you hear this big bang, and it's like, “Oh, no.” Then the apparatus is down for 2 weeks as they patch some—

Speaker 1

You were there during a disruption?

John Platt

Oh, no. This is—sorry, they had a field-reversed configuration.

Speaker 1

Oh, okay. So what is that? Sorry, I'm not familiar.

John Platt

It turns out tokamaks are perhaps the most studied form of plasma, but there are many different kinds of architectures—essentially, ways to try to stabilize and compress plasmas. There is a shape, essentially a self-contained football plasma, called a field-reversed configuration, where the magnetic fields inside and outside are opposite. So they're separated by something called a separatrix. That is, in theory, unstable but in practice stable. For example, when you run a magnetohydrodynamic, or MHD, code, it's unstable under that assumption, but that's an assumption; that's not the way the real world works.

It was kind of disfavored for many years, but TAE and other people—I think Helion—have FRCs because they are actually relatively robust. You can knock them against walls, and they stay stable. But you can still get discharges and things that punch holes in your vacuum chamber, which is kind of unfortunate.

Speaker 1

For clarification, you have these fusion reactors—or trying-to-be reactors, maybe—and apparatuses, and you create a plasma. The plasma is magnetically charged—

John Platt

Or confined. Yes.

Speaker 1

Or confined. So it's confined by a magnetic field. You have some sort of magnetic system that is tunable by a computer, and then the computer tries to maintain the confinement.

John Platt

FRCs, once you make them, are sort of sustained. Okay, so all of fusion boils down to something called the Lawson criterion. It explains why fusion is hard. You can very easily, on the back of an envelope, show that the density, the temperature, and essentially the energy loss—the confinement time—is 1 over the amount of time it takes for the energy to decay away by 1/e in a plasma.

The product of those 3 numbers has to be bigger than some constant, and then you can get fusion. If you don't, then you don't. The fact that it's a product of 3 numbers explains why fusion is so hard, because every approach has an Achilles' heel where one of those numbers is not very big, and then they desperately try to make that number higher.

Speaker 1

Every approach is different. You have to be a bit skeptical when there are all these breathless news stories about fusion, because they'll say, “Now the confinement time is stable for 10 minutes,” or whatever, and it's talking about 1 of the 3 numbers. But you have to have all 3 numbers before you can get fusion.

I think the whole field is making a lot of progress, and it's very exciting, but you do have to be a little bit cautious about the breathless news articles that only talk about 1 number. So what is the computational part of that?

John Platt

Unfortunately, for better or worse, it depends on the approach. For tokamaks, as Brendan said, it's mostly stable, except that occasionally there's this instability that takes all the energy and smacks it into one place, so you have to keep everything under control. It's a control system.

FRCs themselves have very simple instabilities. For example, they have what they call a Z instability. It's fine; it's stable. It'll just wobble. It'll literally wobble back and forth, but you just make what they call a PID controller that keeps the football in the center of the reactor, and things are fine.

Speaker 1

And it does that by adjusting the magnetic field.

John Platt

Yeah. It actually adjusts, I think, the electric field. It sort of knocks it back and forth. The issue that people have really depends on which sort of plasma architecture they're deciding to use.

Climate is political because of economics, basically—probably mostly, maybe other stuff—but the economics of it, you have to persuade people to somehow spend more, or you have to have a solution that has this happy coincidence where it's both economically better and better for the climate. That's hard.

Speaker 1

Yeah. In many cases, it's not. I mean, in many cases, it hasn't been earned, but—

Well, you think about predicting extreme weather, right? You can prepare, and you could see how that could be economically beneficial. So what kind of work are you doing with interventions, and how does that interact with economics? It sounds like the contrails one—I did an analysis and said, actually, this is great because it's very low economic impact but high value.

John Platt

That's right. So if you try, I think there's sort of an energy intervention. You have to compete with existing forms of energy. That's not trivial unless there's a co-benefit, or there's some sort of clever co-benefit.

This is highly speculative; it wasn't our work. There was a startup—I don't know if you saw the news—last year, I think, where someone figured out that if you inject mercury into a fusion reactor, the neutron flux can actually transmute the mercury into gold, and then you can sell the gold. I thought that was very clever. It might not work, but—

Speaker 1

You know, as a physicist, the one thing I want out of a fusion reactor is helium, but that's a different story. Sorry.

John Platt

Oh, you—oh, helium-3. Well, I mean, helium-4 is kind of boring, although it's getting more valuable because the strategic reserve has been shut down. There's less of it.

Speaker 1

And, of course, I want helium-3. Not even just for a fusion reactor, but to make dilution refrigerators for—

John Platt

Quantum computers, or so. Yeah. So much technology we think about actually just goes out the window if we run out of helium.

Speaker 1

That's true.

John Platt

Sorry, that's like a complete aside.

Speaker 1

The fact that the US had a strategic helium reserve was for our very important blimp fleet.

John Platt

Yeah. [Laughter]

Speaker 1

But they kept it for decades anyway, so that was nice. Then we stopped. We got rid of it all. It all went up in the air.

John Platt

Yes, in balloons and stuff.

Speaker 1

Or out of natural-gas wells.

John Platt

Yeah.

Speaker 1

Sorry, now we're talking about helium.

John Platt

Yeah. So, interventions. What are some of the most exciting, interesting ones?

Speaker 1

Well, I'm very excited by fusion. I don't know if that's an intervention; that's sort of a source of energy. If we can make it work and make it have a low enough capital cost, that will actually help a lot.

At least the current models say renewables are great. Ideally, you'd like to electrify everything, right? That has problems because you can't electrify flights, but you can try to electrify a lot of things. There are EVs. You'd have to figure out how to electrify things like cement or steel. Those are hard, especially things like steel, where you want reduction power anyway. Essentially, you're adding carbon and reducing iron ore. So there are a lot of things that are difficult about electrifying everything.

But if you could electrify everything, the amount of electricity required would grow by a factor of 5. You could try to grow renewables—renewables plus batteries—and trying to squeeze all of it out starts getting ever more expensive because you need ever more batteries. You need a huge number of batteries to cover the last few percent, or even 10% or 20%.

We do need some sort of power that can cover the last 20%—something that's baseload. Fusion might be a thing for that. That's super exciting. Again, there's no one silver bullet that can cover all the cases. I'm happy to talk about any specific case, but—

John Platt

The world is a very complicated place, and the global economy is a very complicated place. It's super hard to talk about interventions in general.

Speaker 1

Yeah.

John Platt

Maybe instead of interventions, one thing I'm curious about is how does this affect decisions? For example, what do we build? How do we build? I think you're from LA, right? Or at least you—

Speaker 1

Well, I spent 11 years. Okay, so, yeah, you spent a lot of your life in LA. I mean, LA—large parts of it—just burned down, maybe probably close to where you used to live.

This is something that I think a lot of people saw coming, maybe partially due to regulatory issues, but partially due to other issues. We were completely unprepared, and it seems like there's a lack of preparation about what to do next or how to adjust for this.

Have you worked on predicting new risk assessments, or suggestions for what we actually change to harden society for what's coming, regardless of whether or not we actually do something to solve the underlying problem?

John Platt

That's right. In fact, there's a big effort at Google into something called climate crisis resilience. We had a very fun project called FireSat. I don't know if you know about this.

It turns out that for wildfires, if you catch them early enough, it's very easy to put out a wildfire the size of this room. But even if it's an acre, it gets much, much harder. Under certain circumstances, they can grow exponentially from the size of this room up to an acre. That might be hard to catch, but they often start small and spend a while that way.

We figured out that if you had a global constellation of low-Earth-orbit satellites that could detect in the midwave infrared—which goes back to the blackbody, essentially, and the temperature of fire—the fires would stand out in the midwave infrared.

We designed a sensor so that, depending on exactly what their orbits are, with roughly 50 to 80 satellites, you could actually find fires about the size of this room—about 5 meters on a side—anywhere on the planet. Depending on how many satellites you had, you could detect them within 15 to 20 minutes. You'd have to put a fair number up—around 80—to get them within 15 minutes.

Then you could actually intervene. You could decide not to, if you wanted to have the fire burn fuel and thought it was safe, but if it was going to grow to something unsafe, you could act.

We worked with a nonprofit called the Earth Fire Alliance that we're part of, and they're starting to—well, we've launched 1 satellite, which is a prototype. We've worked with a company named Muon Space to make the satellite. That's cool.

We have wildfire boundary detection, and we propagate that information out through Google. We can figure out from existing satellites and existing data feeds where the boundaries of fires are, and then we tell people through their Android phones or through Search about fires.

We've worked with the US Forest Service on making new models for how fires propagate, because again, that goes back to these process-based models from the 1970s by a person named Rothermel. We've made a little neural-network proxy model based on it, essentially, to be able to run it very, very quickly. We worked with the Forest Service on that.

We're very interested in trying to minimize the damage because people might not realize that the World Health Organization estimates there are 300,000 excess deaths a year across the world from wildfire smoke.

Speaker 1

Yeah. I remember—it's been a few years since we had a really bad fire season. Maybe 4 or 5 years ago, there was this cloud of smoke that crossed all of the northern US and Canada and caused a lot of respiratory issues, I think.

John Platt

Yeah.

Speaker 1

Yeah. And it's very hard to track. You have to estimate these excess deaths from statistical means.

John Platt

But it's a very serious public-health problem. It's also just very scary. It burns people's houses down, and it's terrible.

Speaker 1

My view is that climate change is sort of like a serious disease. Do you treat the symptoms—that is, do you adapt—or do you try to attack the underlying thing? The answer is, well, if it's serious enough, both, right?

John Platt

Right. So, yes, we take adaptation, especially around climate resilience, very seriously at Google, and we try to give people informational tools to help. That's part of the reason why we're working on weather, and then cyclone prediction and things like that. It all actually hangs together. It's more than just interventions. It's climate resilience, too.

Speaker 1

Yeah. Having lived through 4 or 5 fire seasons on the West Coast, they can be quite nasty. It didn't used to be that way. I have a cabin up in the Sierra Nevada mountains, and it used to be, “Oh, summertime, it's nice.” Now it's not every year, but there's winter, spring, summer, and smoke.

I want to stay to the west of the fire line.

John Platt

Yeah. [laughter]

Speaker 1

Yes. So that is another thing that I'm interested in, and that Google is also very interested in: climate resilience. Are these infrared sensors small enough that they could hitch a ride in a microsatellite grid? Would it make sense—would it be almost cheaper just to hire, I mean, to pay someone who's launching a constellation?

John Platt

Oh, they're not that small. The thing is, you need refrigeration because it's midwave IR, so you have to keep it cool.

Speaker 1

Oh, okay. Okay. So these would have to be their own special—

John Platt

Satellites. They're not super large.

Speaker 1

Okay. They're not like the satellites in geosynchronous orbit. Those are giant monsters because of all the optics and who knows what. But, yeah—

John Platt

And they're basically just IR sensors with a resolution of 5 × 5. No, no, that's the other cute thing: the resolution is about 50 × 50 meters, but you can use super-resolution because it's essentially multispectral, and you sort of know where fires are. So there's a fair sprinkling of AI on top of them to reach that 5 × 5-meter resolution.

Speaker 1

Yeah. Yeah. When you also have a convolution over what you're reading out, right, as well—

John Platt

I forget the frame rate. The satellite is moving. I don't remember what the point-spread function is. I'm sorry. I don't know how fast they move. This also has something called a pushbroom sensor. There's this funny thing where you spread out the spectrum one way, and there's also—it's a somewhat complicated thing. It isn't just like a Polaroid; it's a complicated sensor.

Speaker 1

16. Scientific Taste, Domain Expertise, and Learning the Hard Way

Yeah. Sort of switching gears a little bit: you've been at the intersection of AI and science for quite some time. I think you've wound your way into and out of it, back and forth. How do you see the field evolving? I feel like it's evolving very quickly now. What are the lessons that you've learned, that you think the community has learned, and how do you think this should change? If you are a young scientist or young practitioner, how should this change how you approach the future?

John Platt

Well, I think there has been a phase change in the last 12 to 18 months. A lot of what we used to do, as I said, was build these specialized models to solve individual problems. If you think that's your job, it's kind of fun: you find a problem, you solve it. You find another problem, you solve it.

But now we have these much more general AI systems. I think the whole AI-for-science community is still feeling around. The fact that they're working is so new that collectively we're not sure what's the best thing to do—or maybe there's no one best thing; maybe there's a toolchain. I think we're all trying to figure out what we should do.

There's a question of what young scientists should do. I have a son who just turned 21, and he's really into AI and coding, as well as chemistry. I look at him and think he's doing the right thing because he's both learning a lot and trying to be a domain expert about RNA, but he's also using vibe coding and all the tools. I think that's the right answer, because everyone's still figuring it out. Be deep in a domain.

I don't think domain expertise is going away, because it goes back to what a lot of people have said: it goes back to taste and trying to figure out how people get taste without doing all the grunt work. That's an interesting open question, but develop domain expertise and also try to play with all the different tools that are available.

It's not like, “Oh, yes, we know what's going to happen and the smart old people know what's going to happen.” No, we're experimenting too. So I would say definitely develop domain expertise, try to use these tools, and try to solve big, hard scientific problems as best you can.

There's still the huge open issue about what you actually do about physical lab work. That is not going away, because experiments are the ground truth—

Speaker 1

And a bottleneck.

John Platt

And a bottleneck. I mean, people are talking about a lab in the loop, but that's still very, very, very open, because no one has, I think, as far as I know, a general lab that does everything. There are a lot of very specific labs that are controllable.

I think we've gone through this phase change. It seems super exciting. Again, I would advise people to play with whatever tools are available and develop deep domain expertise and taste to the extent you can. I would also advise people not to be scared and to try stuff. We have student researchers at Google, and they come and do wild and crazy things, and that's always delightful. People should be trying wild and crazy things and see what happens.

Speaker 1

Yeah. This may be a question without an answer, but when I think about how I developed expertise and how a lot of people developed expertise, it was by starting with a simple, defined problem and then hammering it. In that process of exploration, you learn more; in some ways you go broader, in some ways you go deeper. But the process of banging your head against the problem—which now would be instantly solvable—teaches you the skills you need to solve harder problems. What advice would you give to your son for that?

John Platt

You know, I don't know. Maybe it's a bit like hiking. You obviously can't drive everywhere. Or you could drive up the mountain. Or you could hike up the mountain. Maybe it's okay, even fun, to occasionally hike up the mountain, even if you can drive up the mountain.

Speaker 1

Yeah.

John Platt

I mean, in the old days—I'm going to sound like a real old man—in the old days—

Speaker 1

In the old days of 6 months ago.

John Platt

Well, no, no. I was even thinking of the old days of the 1980s and 1990s. A lot of people take it for granted: there are open-source packages, there are libraries, there are all sorts of things we didn't have. I had to write my own numerical library. I had to write my own machine learning. I've written boosting—probably rewritten it 4 times in 4 different languages—and now I know boosting.

Maybe not taking the totally easy route—I mean, obviously there's this trade-off: I want to be as efficient and productive as possible. Yes, but you also have to develop the muscles. It's a little bit like being an athlete. There are times when you're trying to run as fast as you can, and then there's also training time. Maybe people just have to train.

Speaker 1

It's entirely plausible that if you spend time actually hammering away and doing the hard work, even if it goes slower there, that pays dividends into your larger productivity long term. Even if, locally, in that one moment, you are not being maximally productive by not exploiting an LLM, that feeds into something.

John Platt

I hope so. I hope the thing—I don't know—is that I hope people, in their careers, do that. It's hard, because the whole world seems to want to optimize everything.

Speaker 1

And your fault.

John Platt

No, I don't. [laughter] It's my fault, but it's just sort of the cultural impetus. Sometimes you have to set aside time. At Google, especially in my group, we have this concept of 20% time, which I still very, very strongly try to protect in my own group. You can do whatever you want. If you want to learn stuff or try stuff, you don't even have to tell me. In fact, I probably shouldn't tell you. Just do stuff for learning, and also because that's where the creative juices are.

I don't want to occupy people's time so completely that they can't feel like they can play or learn or try new, crazy things. I know 20% time is unusual, and there just seems to be this strong impetus in the world to optimize and squeeze everything out. But you do lose something when you hyperoptimize; you sort of overfit.

Speaker 1

You overfit to productivity, as you're so—

John Platt

That's right. I know my advice might be swimming upstream against perhaps cultural norms.

Speaker 1

The thing that always comes up for me here is that the problem is not stationary. There's a new skill set that will be the right skill set for the future, right? The question in my mind is always just disentangling: is this a skill that's an enduring skill? Maybe it wasn't enduring yesterday, but today it will be enduring because—I've seen that. For just as an obvious example, when you become sort of a manager, your skills transfer. Well—

John Platt

Some of them don't, but a lot of them do. As a manager, you lose track of the details of what's going on, and you trust your people or agents or whatever to manage that, so that they can report up to you and answer the high-level questions and get the judgment about the little things correctly. Is that what we've come to? Are we just middle managers now?

No, I mean, I don't know. Again, I have a little management work, but you don't want to be an empty suit.

Speaker 1

In other words, the things might get it wrong, especially LLMs that are really weird. They don't make the same kinds of mistakes that humans make.

John Platt

And so you can't fully trust them. You have to be rigorous and poke at it and make sure. Although, you should be poking at software that you write yourself, too. You shouldn't trust yourself. That's one thing I've learned. What did Feynman say? You absolutely can't fool yourself, and you're the easiest person to fool.

Somehow, fundamentalness—I mean, really learning a domain that is about the world, for example. This is why I like biology or the physical sciences. I think those are enduring, fundamental things. Math is very enduring, but even things like rigor and checking go back to management: you want to make sure that the LLMs are producing the right things or haven't cheated in some way. But again, you should be doing that to yourself, too.

Speaker 1

Yeah.

John Platt

I think there's some enduring value, and also just the enduring value of creativity and thinking out of the box. I think those are enduring. I don't know. There's something very fundamental about all of those.

Speaker 1

So, you mentioned Feynman. If you don't mind me changing gears—

John Platt

I took a class from Feynman. Yeah, not just any class.

Speaker 1

Yes, his physics of computation class. He did it with, in fact, John Hopfield and Carver Mead.

John Platt

That was fun. At the time, I don't think maybe even Feynman knew. I felt like none of us knew what the problem even was. I guess Feynman was trying to say, “Let's do quantum simulation,” which I guess turned out to be the right answer. But at the time, it was cool. First of all, it was DARPA-funded, so you were supposed to go once a week to get a prime rib dinner, but I just went every week anyway.

Speaker 1

So, just for a little more context, this class was basically the class right after Feynman—and I forget who else—proposed the concept of a quantum computer without really knowing what it was, but knowing that there was some sort of—

John Platt

Oh, that was when he taught it. He did say there was “Plenty of Room at the Bottom,” which I think was in the very early days. I took it in ’82. The way the class was structured, it was like a guest lecture on Tuesday, and then Feynman would stand up on Thursday and explain why that was all wrong, which was pretty fun.

Then we encountered something that other people have also encountered: the Feynman effect. Maybe he was so charismatic or something. He would explain things, and you would say, “Yes, yes, I understand. Yeah.” Then you would walk out and think, “No, no, I did not.”

Speaker 1

Yeah. So it was kind of fun. But a lot of the guest people—I mean, maybe it sort of showed the chaos. There were a lot of interesting guests. I think Danny Hillis came, and there were all these interesting people.

John Platt

The union of all of them showed the mass confusion of what was going on. There was a lot of, “Should we make computers reversible?” We had to make sure—could they even be reversible? Could the bottom limit of the heat per operation be zero, or is there some sort of thermodynamic limit? That was a big deal. I don't think that's a big deal now.

Speaker 1

Wait, so you're talking about the Landauer limit, right? Or not?

John Platt

At the time, there was all this question about whether you could have reversible, billiard-ball computers and whether they could be reversible, and things like that. There was also the question of what kind of computing to pursue. That's why Danny Hillis came. He was in the era of the Connection Machine and the original Thinking Machines—not “Mira’s,” but the original one in the ’80s.

Speaker 1

What was the state of general computation in 1982? I mean, at this point—

John Platt

To rounding error, we had zero. I remember when I got there, I was Carver Mead's system administrator, and we had a VAX 11/750 that maybe did 1 MIPS—1 million operations per second. The whole research group shared an 80-megabyte disk drive that was the size of a dishwasher. It was very exciting.

Speaker 1

So what about complexity theory? Because I know there's a lot of interest in quantum complexity and how it relates to gravity right now. I wondered when complexity—

John Platt

We could try to look it up. I don't think there was a whole hierarchy of quantum complexity classes yet.

I think a lot of those came out in the ’90s or something. I remember reading Nielsen and Chuang, the classic quantum information textbook.

Speaker 1

Yeah, Quantum Computation and Quantum Information, which brings up a lot of those points. That textbook was written in the late ’90s or early 2000s, that's right.

John Platt

And I think that was the first textbook that put down the general knowledge of the field, but I could be wrong. I think, when I was young, I didn't take it for granted that this was the book that everyone used.

Speaker 1

17. Hopfield Networks, NeurIPS, and How GPUs Changed AI

Yeah, but you remember, in the early ’80s, there was a whole bunch of excitement around neural networks, and it's really interesting because, again, we really did not know what we were doing. No one knew what they were doing. There was an interesting cycle: neural networks, then SVMs, then neural networks. Oh, no, no, no, this is before that. In the ’80s, everyone said this had the capability of revolutionizing computing.

John Platt

But what does—well, I mean, there was tremendous excitement around Hopfield networks. In fact, NeurIPS came out of a workshop at Snowbird that was nominally private, but everyone tried to crash it, so they spun up NeurIPS.

Speaker 1

Snowbird is a skiing trip with a computation conference attached.

John Platt

That's right. I learned to ski because I didn't know how to ski, and then I kept going.

Speaker 1

That workshop came out of the Santa Barbara workshop in 1985, which came out of some local things at Caltech called Hopfests.

John Platt

People thought, “Oh, wow.” Ironically, there was something about associative memory. If you actually dig down into what transformers are, they are associative memory. In fact, there was even a paper called “Hopfield Networks Is All You Need.”

Speaker 1

Yeah, I forgot that title was a reference to that paper.

John Platt

It was a reference to that era. The whole thing has come full circle. In fact, I think we have collectively, as a field, revolutionized computer science, but the hopes and dreams completely outstripped the capabilities because, again, we effectively had zero compute.

Speaker 1

Yeah, I think the history of machine learning is about different paradigms—compute versus memory, scaling versus data becoming available—at different levels.

John Platt

That's right. I think people don't realize that neural networks won because they're the one compute-limited thing. Although, now that we have transformers, things are getting memory-limited again. They ride on top of BLAS, so the fact that BLAS was being optimized by things like GPUs mattered. I don't actually know if brains work by matrix multiplication, but the algorithms co-evolved with the hardware.

Speaker 1

And that's why we're here. Who knows? If we'd gone down some other path, where people really cared about some other kind of computing, who knows what architecture we would have ended up with? I don't know. Didn't a lot of this start with people hacking PlayStations to train neural networks, or maybe even before that, for supercomputing?

John Platt

There's a funny story about a friend of mine from grad school, Brian Catanzaro, at NVIDIA. He came to NVIDIA and was having a lot of trouble getting traction. They had doubled and tripled down on gaming, and he basically had a meeting with Jensen and convinced him: “We have all these people using CUDA for BLAS, basically, and for deep learning in particular.” Jensen was convinced, and Brian said it was a 15-minute conversation. They pivoted the whole company the next day, or whatever.

I have friends there—David Kirk, who was the first chief scientist, and Bill Dally were both friends of mine from grad school.

So yes, but I don’t know the details of the way. Remember that people in Hinton’s group were using GPUs for deep learning even around 2010. In 2007–08, they weren’t that much faster than CPUs.

Again, they only started really exceeding CPUs, in fact, I think—not coincidentally—in the era of the original ImageNet and some of the speech-recognition work. So I don’t think it’s a coincidence that they really took off when GPUs passed CPUs.

Speaker 1

18. How a Student Internship Led to an Academy Award

Yeah. I really want to know: how does one get the opportunity to name an asteroid?

John Platt

Well, again, at Caltech there were some wonderful people, Gene and Carolyn Shoemaker, and they were teaching a class in planetary science. I liked planetary science, and as part of that class they took us through the asteroid-discovery process.

This was in the 1980s, so now there are all sorts of amazing systems. For a while there was something called LINEAR, which automated the process, but this was before anything was automated. It was just part of a class.

What you would do is go through the steps. Even though we did them out of order in the class, the steps we used to do this—40 years ago—were that you would go to Palomar. There would be a fast telescope, which is literally now a museum piece in their visitor center, but at the time it was a real thing. You would put a piece of film in it, take a picture, wait a few more minutes, and take the same picture of the same point in the sky.

Then you would put the film in a stereoscope, back at Caltech, and look around to see if anything floated, because it would have moved by a tiny bit. It would pop out, and you could see it because it was different in the two eyes. You would be able to see it as something projected in a different place.

That’s right. It would pop out at you, literally. Then you would go to a measuring microscope and take measurements of known star references and where this floater was. If you got an accurate enough measurement, you would send it off to a person at the Minor Planet Center—his name was Brian Marsden, although I don’t think he’s there anymore.

He had a big software system to piece together what are called apparitions. If you happened to have made an observation that was the final apparition, which allowed his software to connect it into one big orbit, then you would get discovery rights and could name the asteroid.

But now it’s amazing. There’s an observatory in Chile called the Vera C. Rubin Observatory, and there’s this amazing telescope called the Simonyi Survey Telescope. It essentially automates this process. It can just take many frames of the sky. It’s an utterly stunning instrument. They discovered 11,000 asteroids in 6 weeks.

Speaker 1

Oh, wow. How many of them are there, at least?

John Platt

It depends, I guess, on the cutoff, but there are probably millions of asteroids.

Speaker 1

So there’s still a lot of opportunity.

John Platt

It’s true, but I don’t even know if they bother naming them. [laughter]

I don’t know. Maybe you don’t need to. But people are doing occultations. Sorry, I can talk about this forever. There’s some stuff called occultation. One of my asteroids—I was emailing with an amateur astronomer—happened to pass in front of a star, so the star’s brightness dipped, like when they discover exoplanets.

But if you’re super lucky, it will dip and then dip again because there’s a moon. An amateur astronomer found a little moon around one of my asteroids. That was, I think, last year.

Speaker 1

Around the asteroid?

John Platt

Yeah.

Speaker 1

Interesting.

John Platt

So that’s pretty cool. There’s still room for discovery.

Speaker 1

Your asteroid has a moon.

John Platt

Yes. In fact, there’s actually a little bit of a story there, because it’s not really my asteroid. I found 2 of them. I named 1 of them after my dad. I waffled, so Carolyn named it after a professor at the University of Washington, and I said, “Oh, but I wanted to name it.” I whined, and she was nice, so she gave me 1 of hers.

Then they named that 1 after my mom, and that’s the one that has the moon.

Speaker 1

Oh, okay.

John Platt

So the asteroid named after my mom has a moon.

Speaker 1

How big is it?

John Platt

The main body, they think, is about 4 kilometers across. I’m trying to get the units right. The moon is about 1 kilometer, so it’s actually a pretty big binary. They’re not quite twins, but they’re close.

Speaker 1

Did other people in that class discover asteroids?

John Platt

I think there was 1 other person. It’s a little bit of a crapshoot, and the fact that I found 2 was quite unusual.

Speaker 1

Was there a reason why? Was it just pure luck, or was it staying there all day and trying?

John Platt

I tried to be careful, but yeah.

Speaker 1

How does one get an Oscar?

Yeah, an Academy Award.

John Platt

Well, again, maybe I was at the right place. My thesis advisor is named Al Barr, and I was a student intern in 1986 at a place called Schlumberger, which is an oil-services company. They had an AI lab, and we were all doing computer-graphics research.

At the time, this was revolutionary. I realize it’s now considered incredibly boring, but you could actually use physics simulators to make computer-graphics movies. At the time I thought, “Wow, that’s really cool.” I said you could use the theory of elasticity to make floppy things, so I wrote an elastic simulator and made fabric, stretchy things, and so on.

They said, “Oh, wow, that’s cool.” The descendants of that became a lot of the physics simulators that people used at Pixar and in their various movies. I tell interns, sort of half-jokingly, “If you do a really good job as an intern, you can get awarded an Oscar.” [laughter]

Speaker 1

How long was that between the time you did that work and when you received the award?

John Platt

Twenty years.

Speaker 1

It was 20 years?

John Platt

It was 20 years.

Speaker 1

Which is actually not atypical, right? When they give you an Academy Award, they want to make sure that it’s been well used, that everyone uses it, and so on.

John Platt

Yeah, but of course, in that 20 years, everyone started doing it. It became obvious.

Speaker 1

This was specific work done for a specific movie or something?

John Platt

No, it was a paper in SIGGRAPH. Then Pixar became, I think, somewhat important—core to a lot of their work.

Speaker 1

Yes, and in fact, some of my friends—in fact, a lot of my friends—did a lot of these simulators.

John Platt

Yeah.

Speaker 1

I’m curious: since you have been quantum-computing-adjacent and worked on quantum computing directly at Google Applied Sciences for a while—I think you’re not currently on that, but—

John Platt

I’m still dabbling in it.

Speaker 1

You’re still dabbling. Where do you see the trajectory of quantum computing going over the years? This is one of those things that, like fusion, seemed very exciting at first and then seemed like it wasn’t going anywhere for a long time. Maybe now we’re seeing hints again that it’s exciting. Or maybe that’s just my quantum-adjacent perspective.

John Platt

I think a lot of people have gone through that. I tend to average things out over the decades. I’m old enough now that I sort of average things out over the decades, and it’s just progress.

Speaker 1

But going from Feynman trying to figure out what this even means as a concept, all the way up to now, there’s this recent Willow result on quantum error correction, which actually seems genuinely achievable with the right scaling laws and so on. Do you still think that? Would you say that quantum computing is 20 years away from some simple but practical algorithm that actually succeeds? Do you see this accelerating, or do you think this is still going to take a long time?

There are a lot of areas where we saw something advance very quickly and unexpectedly, and I’m wondering whether this is something that you think will be quick or not.

John Platt

I think it’s sort of in between. It’s not a purely software thing, because we need to build systems that are large enough and stable enough to be able to be a quantum computer instead of a quantum apparatus.

We’re in what they call the NISQ era—noisy intermediate-scale quantum. That was coined by John Preskill. It’s not a very good term, sorry, but I guess that’s the term we have.

The quantum team made this wonderful paper, which I think I’m a co-author on—one of many—for something called the Quantum Echoes algorithm, which could be applicable now. Fundamentally, it’s in the style of Feynman’s proposal. The technical term is that you could try to fit a Hamiltonian to observed data, like NMR.

What that means is that you have a physical model that’s parameterized, and you use the quantum computer to adjust the parameters and try to figure out, inside of a loop, what the right parameters are. For example, you could deduce NMR parameters.

So that is a practical thing that's used. Now, the question is whether it will be big enough to make breakthroughs. That's still TBD. The quantum team at Google has been executing amazingly well against a roadmap that Hartmut Neven laid out a few years ago, and they're continuing to march down it as they scale up. When they hit their last milestone, they should be able to have a quantum computer that does amazing things, and they're continuing to advance it.

So, yeah, I think it's on the scale of a few to several years. I haven't kept up on exactly what date they're saying, so you should ask Hartmut exactly when that's going to happen. But, yeah, they're marching along. The main interesting question is whether superconducting qubits will be the winner, or whether one of the other alternative technologies will be. That's still TBD.

I still think superconducting is very promising because it is very scalable. I don't think it's 30 years away, and I don't think it's tomorrow. I don't think there will be—well, even if there's a sudden hardware breakthrough, these are very finicky things. Someone might have a brilliant idea, but it will still be a while, because fundamentally, at least in the NISQ era or for a while, these are analog computers, and they tend to be very, very, very finicky. I wouldn't expect all of a sudden that some phase change happens.

No one has figured out how to scale up qubits in a way that lets them interact at an arbitrary distance. We're still on the order of 100 or 200 qubits, I think.

Speaker 1

You mean in terms of—okay, so—

John Platt

In terms of—

Speaker 1

Well, it's a little like—yeah, yeah, exactly. So, for superconducting qubits, people are working on it, but the ratio of physical qubits to logical qubits is still relatively large for the error rates that you need. Maybe there'll be a breakthrough there; I don't know. There are other things that might mean you don't need that ratio to be so high, because a lot of it has to do with the 2D connectivity of the chips that you lay your qubits out on. You lay your qubits out on a 2D chip, whereas for things like neutral atoms, in theory—in theory—they can connect anything to anything else. But, of course, in practice, we don't really know, and we don't actually know what the limitations of neutral atoms are. Maybe the super experts know.

Superconducting qubits in particular have this sort of tension where you want to have nice, clean resonators. You get a clean resonator by decoupling it from the environment, and then you get good interactions by coupling resonators, which involves coupling to the environment.

John Platt

True, but at least there's the whole environment, and then there are little tiny holes that you want to go through to your neighbors.

Speaker 1

This is not like a fundamental uncertainty relationship or something. It's a technological limitation, a difficulty, or something.

John Platt

Yeah, the hardware team in Google Quantum is very, very skilled. They're very, very skilled. They're really good at making these designs and making these things actually work. I find them impressive.

Speaker 1

Yeah.

John Platt

It'd be exciting to see that advance.

Speaker 1

Yeah. Yeah. Yeah.

John Platt

Yeah.

Speaker 1

I guess we'll just have to hold our breath and wait.

John Platt

Just wait. Yeah. I guess I'm just very patient, so I just start working. So, 20 years ahead of time.

Speaker 1

That's what Dave Bacon, who runs the software team in Google Quantum, always teases me about. He says, “John, like, oh no, I can't work in quantum computing. It was like 10 years ago, because you're always 20 years ahead of time.” So I have to wait for 20 years. Well, 10 years have gone by.

John Platt

Halfway there.

Speaker 1

Halfway there. I don't take that as an absolute. Yeah, but I also started working on fusion 10 years ago.

John Platt

Maybe you're actually causal. You start working on something, and reality just catches up.

Speaker 1

I guess so. Maybe. Who knows? I don't know.

John Platt

I started working on convolutional nets in the early '90s. I loved convolutional nets.

Speaker 1

Yann LeCun claimed it checks out.

John Platt

I coined the term “convolutional net.” As far as I can tell, that may be true, because everyone called it LeNet because it was a very specific thing. They all worked for Yann. I said, “Well, I don't work for Yann. I don't want to call it LeNet.” So I called it a convolutional net. I don't know. That was a more generic term. It's a good term.

Speaker 1

19. The Future of Science and the Dream of an “Everything Lab”

Yeah. Before you go, is there anything you want the audience to take away? Any messages you want to deliver?

John Platt

I think AI is an example of it, but I'm really amazingly excited about the potential of AI for science. I think it's going to be an amazing power tool for scientists. I don't think scientists will be replaced. In fact, I'm hoping that they'll spend all their time on, again, the creative stuff, the rigorous stuff, and the philosophy stuff. I think it's going to be way cool.

Speaker 1

I hope so, too.

John Platt

Yeah. My intuition is that a lot of—I've heard a lot of people say that, and I hope they're right. I hope it's not just a sort of coping with a reality that's uncomfortable.

Speaker 1

The other question we almost forgot to ask: if you could remove a bottleneck in your industry, however you want to define that, by fiat—then, like magic—what would that be?

John Platt

If I could get a magic wish, I would say, “Someone, please make the everything lab,” where you could send a JSON blob and it would do any experiment at all.

Speaker 1

Okay. Automated—

John Platt

But it would have to be anything. So, essentially, I guess we have to solve the sort of AI-complete robotics problem. But if we did, then that would be stunning, because right now things like ERA are all computational. Someone has to gather the data.

Speaker 1

So, yeah, if we could just break that—oh, that would be so amazing. That would be so utterly amazing. Great. I really appreciate you taking the time to see us, and I think you kind of flew in and adjusted your schedule a little bit to—

John Platt

Yeah, I was sort of flying over San Francisco to get home, and I said, “I landed in San Francisco.”

Speaker 1

So you made a big effort to be here. We really appreciate that. It was really fun to talk to you.

John Platt

Mhm.

Speaker 1

Yeah. Listen, it's been a blast.

John Platt

Okay, cool. Thank you for having me.

Speaker 1

Yeah, you're welcome.