[BidClub_]
The Cognitive Revolution · · 98 分钟

材料进展:让 AI 形成科学直觉——与 Orbital Materials 的 Jonathan 和 Tim 对谈

Jonathan GodwinTim Duignan

YouTube
TL;DR
  • Orbital Materials 的核心押注,是让 AI 把材料发现从工匠式试错,变成更可预测的工程学科。 其生成模型提出候选材料,力预测模型进行筛选,内部湿实验室负责闭环。Jonathan Godwin 表示,到决定要制造什么时,他们已经回答了约 90% 的关键问题,并将范围收窄至2到3种材料。

  • 最深层的技术信号,是模型从约20原子的无机晶体,泛化到复杂蛋白质系统。 Jonathan 认为,这说明模型“在这么小的尺度上学到了某种真正基础的东西”。消息传递神经网络不像语言模型那样使用显式位置嵌入,因此原子级模拟可以随着可用算力扩展,但仍受实际限制。

  • 钾离子通道的工作,为这种泛化提供了一个引人注目、但明确尚未确认的展示。 Tim Duignan 在一块 V100 GPU 上观察到水进入通道,一个羟基通过氢键将水牵引进去——这可能解释一组突变实验:移除该基团后,导电性下降了接近一个数量级。该工作尚未经过同行评审或实验确认;Tim 称其为一个有力且合理的假设。

  • Orbital 初期的商业化重点,是数据中心建设,而不是室温超导体这类“登月项目”材料。 公司正在开发热管理和脱碳材料,包括一种利用数据中心废热和气流的 CO2 捕集材料。经济关键在再生:一次性捕集相对已经解决,而降低解吸能耗、让材料或许能重复使用10年,可能大幅压低成本。

  • 模型经济学更接近图像扩散模型,而不是前沿语言模型。 Orbital 表示,公司在融资不足4,000万美元的情况下,已经构建了领先模型和湿实验室。由于架构没有 Transformer 式的位置限制,原子模拟可以跨额外芯片扩展;但大系统会变慢,长程静电作用可能需要额外建模,密度泛函理论也不足以处理部分高度关联电子系统。

  • 真正的护城河,可能来自可扩展模型与历来难以复制的实验默会知识的结合。 Jonathan 将半导体工艺经验与 TSMC 积累的优势相提并论:前者可能包含10,000个步骤,通过师徒制传承。Orbital 的内部实验室因此更重要的不是制造产能,而是一个快速反馈系统,持续让 AI 扎根于可制造的材料。

  • 科学产出提高,可能伴随着令人痛苦的“满足感工作”从人转移给机器。 主持人总结的一项研究显示,材料发现增加44%,专利申请增加39%,下游产品创新增加177%;但随着 AI 承担更多创意生成、科学家转向更多验证工作,82%的科学家表示工作满意度下降。Jonathan 预计,“创造力会是最先消失的东西之一”,即便他也认为,要以更低成本生产更多东西,就必须加速发现。

摘要 · 为研究而整理的核心内容

1. 材料决定 AI 经济的物理上限

  • Jonathan 的半导体例子从软件栈之下讲起:微小碎屑就可能破坏蚀刻电路,因此必须在越来越窄的结构之间填入低 k 绝缘材料,保护芯片。若这些材料没有重大进步,“芯片上就不会有更多晶体管,也就不会有 AI”。

  • 他的整体判断刻意采取最大化表述:“材料科学是最早的科技创业公司。” Silicon Valley 的革命建立在材料之上,而下一轮前沿——从更好的芯片和 GPU,到航天器和新型光学材料——都说明材料科学是“实现科幻未来所缺失的基本拼图”。

  • Tim 的逆向检验方式,是提出一个重大问题,再把它一路追溯到材料。气候治理是最清晰的案例:电池和太阳能板通过材料突破带来了大规模 CO2 减排,而在他宽泛的定义下,应对气候变化的大多数工具本质上也都是材料科学问题。

2. 材料发现仍依赖运气、过度设计和积累的直觉

  • Jonathan 把今天的材料开发比作可靠结构建模出现之前的桥梁建造。工程师可以建造令人惊叹的桥,但不确定性迫使他们过度设计,Brooklyn Bridge 就是他的例子;有时这种不确定性还会造成灾难性失败。现代软件则可以在施工前解析每一根螺栓和每一根梁所承受的力、应力和极限。

  • 理想的转变同样具体:科学家应在制造材料之前,就知道它是否具备所需性质;如果产率或性能不理想,再用计算从底层诊断问题,包括实验无法直接揭示的分子尺度过程。

  • Jonathan 的半导体例子说明了难点。一块芯片可能包含约10,000个工艺步骤,专家要花30年学习哪些气体用于清洗和蚀刻晶圆等细节;这些知识往往没有成文,不能还原为基础化学规律,而是通过师徒制传承。

  • 这套默会知识解释了为什么仅靠资金无法立刻复制 TSMC。但 Tim 不愿把 AI 视为简单替代:人类直觉和机器直觉的形成方式不同,可能互相捕捉对方的错误;即便要突破人脑的边界,也意味着将部分“美妙”的工匠式知识计算化。

3. Orbital 将模型与湿实验室结合,因为单方都无法独自闭环

  • Jonathan 在 DeepMind 负责大规模材料 AI 工作后创立 Orbital。他原以为需要10年才能实现的能力来得快得多:10年前最强超级计算机都无法完成的计算,开始变成笔记本电脑可以处理的任务,这让他确信材料 AI 正沿着整个领域陡峭的能力曲线发展。

  • 组织方式与模型本身同样关键。AI 实验室不会自然组建实验材料团队,而传统化学和材料公司过去“25年来”一直在拒绝计算工具;新公司则可以把前沿建模,与真正负责制造和商业化材料的科学家放到一起。

  • Jonathan 的终点,是一个配备 LLM copilot 的材料设计界面:用户提出电池或其他功能材料的需求,在计算机中反复迭代,并高度确信实验室最终制成的产品会具备预测的性质。湿实验室则通过合成约束让这一目标落地,并在模拟失败时提供快速反馈。

4. 神经网络似乎把物理学压缩成了可用直觉

  • 主持人的基础类比是接球:人类不会显式计算每一条轨迹,却表现得像是拥有直觉物理。他认为,类似的捷径正在蛋白质折叠、原子模拟、天气预报等领域出现,并追问神经网络学到的是否是紧凑的高阶概念,而不只是加速方程求解。

  • Tim 表示认同,但承认:“我不认为我们目前完全理解其中原因。”他的候选解释是降维:物理模型通过保留有影响的变量、丢弃无关变量来实现有效建模——某个模拟里是电子,另一个模拟里是溶剂,或者球周围空气的微观动力学。

  • 这与粗粒化和统计力学相连:跨尺度的行为可能反复出现,而准确预测不必追踪每个组成部分。Tim 对此的表述贯穿整期节目:“找到高效方式追踪重要信息、丢掉不重要信息”,是机器学习尤其擅长处理的核心问题。

  • Jonathan 指向了架构的普适性:相似的网络既支持材料模拟,也支持天气预报;紧凑的世界模型还能复现水、反射、重力和过去需要昂贵渲染才能实现的小型视觉效果。它们创造了“保留保真度的捷径”,就像顶尖运动员可以完成高难度动作,却不需要知道背后的数学。

5. 第一性原理模拟把数据稀缺变成规模化机会

  • Tim 起初怀疑 AI 能否帮助电解质化学,因为实验数据库几乎没有信息——有时每种电解质只有几个参数。后来他意识到,第一性原理计算可以自动生成数百万个数据点,而主动学习闭环可以把新的模拟集中到模型最薄弱的区域。

  • Jonathan 区分了有用推断与可行训练。一个现实的 CO2 制燃料催化过程,可能需要在很长时间跨度内模拟数千甚至数百万个原子;但训练可以使用大量约100原子的片段来学习局部规则,再把这些规则应用到原本无法处理的大系统。

  • 让他“彻底震撼”的结果,是先用约20原子的无机晶体训练,再直接模拟蛋白质。Jonathan 将其视为模型在小尺度上学到了基础规律的证据;Tim 更谨慎地认为,分子相互作用可能只包含有限数量的相关复杂性。

6. 基础模型与 Orb 把生成和动力学视为同一物理学的两种表达

  • 主持人在开场中将 Orbital 的基础模型称为 “Linnaeus”,Jonathan 后续解释时则称为 “Linus”。Jonathan 描述称,该基础模型通过扩散学习简化物理,而 Orb 模型则针对物理模拟进行微调。

  • Tim 将扩散模型追溯到非平衡统计力学和 Langevin 动力学。在材料生成器中,原子沿着学习到的 score 向量场移动;在分子模拟中,原子则在力场作用下运动。他引用的研究显示,当扩散模型在模拟轨迹上训练时,score 在数学上可以与力场完全相同。

  • 由于结构可以加入近乎无限的噪声样本,扩散训练能够密集覆盖简化后的物理规律。再用量子计算得到的力进行微调,就能得到可用于物理模拟的模型。

  • 给晶体加噪,主要是不断扰动原子位置,直到结构变得不再连贯。特殊之处在于周期性:原子从模拟盒子的一面离开后,会从另一面重新出现,需要进行数学调整,但整体流程仍更接近图像扩散,而不是针对蛋白质的专门加噪。

  • 原子身份不再是物种0到117之间的分类选择,而是先转化为连续预测,再映射回离散类别。Orbital 的消息传递神经网络在小型晶体结构上训练,不使用语言模型那样的显式位置嵌入,因此模拟系统可以随着可用算力扩展,不会撞上固定的位置上下文长度。

7. 生成负责提出结构,力预测让结构具备科学价值

  • 科学家可以要求基础模型生成10种具有指定带隙的材料,就像图像模型接收提示词一样。Jonathan 更偏好精确的数值条件,而不是自然语言,因为材料往往需要落在狭窄的“金发姑娘区间”:不是某种性质越多越好或越少越好,而是要达到特定数量目标。

  • Orb 接收原子坐标和物种信息,然后输出力和能量。数十年来的分子模拟方法可以将这些基础量转化为扩散率、化学平衡、结合强度等基础性质,使看似狭窄的力预测任务成为连接广泛材料工具箱的入口。

  • 两套系统形成外循环:生成器提出假设,模拟器对其进行更精确的测试,结果再帮助判断哪些候选材料应进入实验室。Jonathan 认为其资源画像更接近基于 CNN 的图像扩散,而不是 LLM;Orbital 在融资不足4,000万美元的情况下构建了模型和实验室,同时他认为这个领域可能还处在“GPT-1”阶段,具备很大的扩展空间。

8. 模拟同时成为显微镜、实验和物理数据工厂

  • Tim 的第一个应用场景,是观察纳米和纳秒尺度的“计算显微镜”——也就是键形成和断裂发生的十亿分之一米、十亿分之一秒尺度。就像最初的显微镜一样,只要模拟足够广泛,就可能在研究人员还不知道该提出什么假设之前,接连带来新的发现。

  • 在静态预测之后,动力学尤其有价值。蛋白质结构可能已经可得,但科学家仍需要观察蛋白质如何运动、如何在实验可触及的微秒尺度以下发生相互作用;材料同样需要转变态、反应路径和不断演化的结构,而不只是最终的平衡状态快照。

  • 第二个应用场景是计算机内实验:在输入文件中改变一种原子,而不是订购化学品、合成化合物,再用显微镜理解它。如果模拟足够快、足够可靠,就可以在投入稀缺实验室时间之前,探索数千个候选实验。

  • 第三个应用场景,是成为“物理模拟的基岩”。原子级运行可以为气候或其他更大尺度模型提供参数,生成训练集,并支持统计和力学上严谨的粗粒化流程;例如忽略溶剂,同时学习能够复现目标组分行为的有效自由能。

9. Orbital 的工作流试图在合成前消除90%的不确定性

  • Jonathan 用金属有机框架说明这一流程:它可以被想象成一个拥有巨大孔洞和大量连接位点的海绵。生成模型可能提出10个、15个或20个候选结构,其中配体像触手一样,从混合流中选择性抓取目标分子。

  • 生成模型很难便利地编码所有现实问题。因此模拟器会检查吸附动力学、触手是否堵塞孔洞空间,以及其他失效模式;Jonathan 表示:“等我们决定要制造什么时,已经回答了90%的问题”,足以建立信心。

  • 模拟中出现的意外行为,会推动下一轮创意生成。经过2到3轮迭代后,Orbital 可能将2到3个通过计算筛选的结构带入实验室——这些候选材料的预期成功率更高,也足够新颖,能够进入人类团队通常不会探索的化学空间。

  • Jonathan 认为新颖性几乎和可靠性同样重要。AI 不应只是更快地复现既有化学直觉;它的价值在于暴露陌生设计,再由模拟和实验验证这些设计在物理上和实践上是否可行。

10. 钾离子通道检验晶体训练出的物理规律能否迁移到生物学

  • Tim 以“身体本质上是一个真正的电气设备”开场。细胞使用离子而不是电子制造电压变化;钾离子通道通过选择性允许钾穿过细胞膜,控制肌肉收缩、心跳和神经元放电背后的信号。

  • 但选择性过滤器仍是一个关键的未解问题:研究人员仍在争论钾离子究竟是单独通过,还是与水一同通过。实验依据并不一致,而经典模拟通常假设钾离子单独通过,这意味着该领域可能缺少某个关键观察或建模能力。

  • 在 Orb 处理电解质溶液表现良好后,Tim 开始思考它是否可能成为计算化学家期待了数十年的通用力场。他原本预计,分布外蛋白质模拟会“立刻崩溃”或变得不符合物理;但模拟保持稳定,暴露出有趣行为,让他“真的……跳了起来”。

  • 训练分布包含的是小型、高度重复的晶体,而不是膜或蛋白质。Tim 的初步解释是,一组有限的信息——例如原子偶极矩和随空间衰减的相互作用——可能已经足够;Jonathan 则认为,这一结果部分来自规模,以及应用于材料领域的 AI 工程能力。

11. 一个羟基可能解决水传导争议

  • Tim 的限定条件至关重要:这项工作尚未经过同行评审。钾离子通道专家认为结果有趣、合理且与实验一致,但仍需要直接实验确认;因此,这项模拟应被视为一个精确、可检验的假设,而不是已经确立的机制。

  • Orb 观察到水进入通道狭窄的选择性过滤器,这与许多经典模拟不同。蛋白质上的一个羟基伸向水分子,形成氢键,并似乎将水拉入通道——这是 Tim 此前没有在其他报道中见过的全新定性现象。

  • 关键的佐证来自突变证据:移除该羟基后,通道导电性下降了接近一个数量级。当主持人问这种新机制是否可能解释这一结果时,Tim 表示,如果非要下注,“我认为肯定是”,但仍保留了需要验证的前提。

  • 医疗路径从机制走向突变。缺陷钾离子通道会导致包括长 QT 综合征和心悸在内的心脏疾病;模拟或许可以揭示突变如何损害电流,并最终指导修复性药物开发。Tim 提到,氯离子通道已经有人沿着这一路径推进。

12. 扩展在架构上开放,但受物理和计算约束

  • Tim 最初的运行使用一块 V100 GPU,因此将系统截断在距离选择性过滤器约20埃的位置——大约4到5层原子——并冻结最外层。这是“有点取巧”的做法,用来稳定该区域,以便观察其核心动力学,而无需模拟完整膜结构。

  • 完整系统需要包括通道、脂质膜,以及足够多的周边环境,避免陷入不断添加稳定原子的无限回归。周期性边界条件可以封闭这个世界,但大系统会变慢,最终需要许多 GPU;延伸到更长时间尺度,是主要的实际问题。

  • Jonathan 表示,这里不存在语言模型式的上下文上限,因为网络没有与固定训练长度绑定的位置编码。分子相互作用主要是局部的,因此模拟规模大致可以随增加芯片而扩大;但长程静电作用按1/r衰减,可能需要更长程的模型。

  • 粗粒化提供了走向全细胞尺度动力学的路径:保留选定组分,其余部分则以学习到的有效自由能替代。神经网络势已经支持非常大的模拟,但可访问时间跨度较短,以及这类系统难以验证,仍是核心限制。

13. 室温超导体暴露出当前理论边界

  • 当被追问室温超导体时,Jonathan 表示:“至少给我们5年,但也可能更早。”同时,他将碳移除和高效扩展 AGI 放在优先位置。

  • Tim 解释了技术限制:超导性是高度关联的电子问题,而不是主要由局部原子相互作用决定的问题。Orbital 当前的理论水平建立在密度泛函理论之上,无法足够准确地描述所有相关量子行为,因此现有方法无法完整解决这一问题。

  • 这些工具仍可以通过提出晶体结构,或在研究人员明确希望诱导哪些电子性质后,帮助确定如何制造更好的超导体。AI 加速并不会消除用于生成训练标签的物理理论本身存在的保真度边界。

14. 数据中心提供第一个商业化试验场

  • Orbital 正优先开发适用于更高密度、更高温 AI 基础设施的材料:机架和芯片功率密度上升,需要更好的热管理;GPU 数量增加,则需要半导体持续改进。Jonathan 将这一机会与脱碳结合起来,因为在他看来,以可持续且低成本的方式扩展 AGI,本质上是同一个材料问题。

  • 节目中已经提到的产品,是一种专门为数据中心运行条件设计的 CO2 捕集材料。它可以利用废热和气流提升捕集性能,并有望抵消运营商部分范围2和范围3排放;Jonathan 表示,产品将于今年陆续披露更多信息。

  • Jonathan 确认了整体的海绵式模型:反应位点在普通空气中选择性地与 CO2 发生作用,而吸附容量则根据运行温度、空气成分和其他条件进行定制。再生是决定性环节——“一次性捕集相对已经解决”——因为废热可以降低解吸能耗,而一款可以使用或许10年的材料能够摊薄制造成本。

  • Orbital 不打算“在新泽西某处建一座大型工厂”。授权和供应商关系可以将产品所有权与制造分开,Jonathan 将其与 Nvidia 的模式相提并论;与此同时,内部实验室的成本高于软件公司,但对于一家风险投资支持的公司仍然可负担,并且对快速实验反馈不可或缺。

15. AI 将科学劳动转向验证,也可能隐藏更多推理过程

  • 主持人总结了一项关于 AI 辅助材料研究的研究:材料发现增加44%,专利申请增加39%,下游产品创新增加177%。令人不安的伴随结果是,随着机器承担更多创意生成、人类接手更多验证工作,82%的科学家表示满意度下降。

  • Tim 认为这些工具带来了赋能感,因为作为一名没有接受数十年化学训练的物理学家和数学家,模拟让他可以直接看到化学键和作用机制。他理解那些花费整个职业生涯建立直觉的专家,但相信,更快地把有影响力的材料带入现实所带来的满足感,可能超过独自解决每一个问题的乐趣。

  • Jonathan 对保留这种乐趣并不乐观。他表示,AI 研究人员已经感受到更低的满足度,尽管岗位更受尊敬、薪酬更高,因为工作已经从创造力转向工程;而昂贵的算力又限制了他们能够独自运行的实验。他预计,“创造力会是最先消失的东西之一”,之后是手工任务和机器人。

  • Tim 最后的技术判断是,LLM 加上神经网络力场,可以把假设生成与内部物理检查结合起来。o3 等模型展示了可验证答案领域的价值,而神经网络势或许能让科学系统在计算机中检验假设,而不是把所有事情都交给实验室。

  • Jonathan 预计未来会出现共享的多模态嵌入,但不同模态可能仍使用不同解码器。主持人担心,连续潜空间中的推理会难以解释,AI 系统之间的沟通可能比它们与人类的沟通更有效。Jonathan 引用的冷峻退路是:“希望道德实在论是真的。”

Speaker 1

Today my guests are Jonathan Godwin, founder and CEO of Orbital Materials, which is pioneering the application of AI to materials science, and Tim Duignan, who was previously here to discuss his work on the simulation of electrolyte solutions and has since joined Orbital Materials as a materials science researcher.

Materials science underpins virtually every aspect of modern life, from the semiconductors that power our devices to the batteries and solar panels driving the clean-energy transition. Advances in materials have been at the heart of human progress for at least the last century. The challenge has been that discovering and developing new materials has always been painstakingly slow, traditionally relying on trial and error and scientists’ hard-won intuitions developed over decades.

More recently, even with the shift to computer simulation, it has still required huge computing power to simulate small molecular systems for short time intervals. Orbital Materials aims to dramatically accelerate this process with AI. Its immediate focus is on developing novel materials for data centers, both to improve efficiency and to capture carbon emissions, but its technical breakthroughs could unlock advances across clean energy, electronics, medicine, and beyond.

The technical approach is fascinating. Orbital uses an architecture called message-passing neural networks, which are trained on small crystal structures and, because they do not use positional embeddings like large language models, are capable of scaling up indefinitely with computing power. The models can design new materials with specific target properties through a diffusion process and predict the forces between atoms orders of magnitude faster than numerical methods can, allowing them to simulate larger systems for longer.

The company recently demonstrated the power of this approach by simulating a potassium ion channel, a critical protein that controls electrical signaling in our cells by selectively allowing potassium ions to pass through cell membranes. Despite its importance to everything from heartbeats to brain function, fundamental questions about how this channel works have remained unanswered for decades.

While Tim’s recent work still needs to be experimentally confirmed by the broader research community, his simulations were able to show a level of detail in the mechanism that had never before been seen and that helps explain previously inexplicable data. The implications for biology and medicine are significant, but perhaps more important still, this work illustrates a critical phenomenon that we are seeing time and again as AI is applied to different branches of science: neural networks seem to have the ability to develop a sort of intuitive physics in virtually any problem space.

Just as humans can catch a ball without explicitly calculating its trajectory, these AI systems are developing efficient shortcuts for predicting complex physical phenomena, whether that is material properties, protein folding and interactions, weather forecasts, single-cell transcriptomes, or even the evolution of human brain states. For me, this is the clearest reason to believe that superhuman intelligence is not just possible but increasingly likely for human scientists.

Meanwhile, this means a shift away from hypothesis generation and toward more validation and implementation work. While that might mean lower job satisfaction, the potential to dramatically accelerate scientific progress and more effectively address critical global challenges makes it a worthy trade-off. It also reminds us that we might soon need to look beyond our work for meaning.

Jonathan Godwin, founder and CEO of Orbital Materials, and returning guest and now researcher at Orbital Materials, Tim Duignan, welcome.

Tim Duignan

Thanks for having me. I’m excited for this.

Speaker 1

I’m on quite a journey to try to figure out what is going on at the intersection of AI and science. Obviously, there are a lot of different sciences, a lot of different intersections, and a lot going on. It feels like life right now is one big crash course for me.

You’re advancing the frontier when it comes to materials science and the use of AI in materials science, and I’m excited to get into all of that. I thought maybe, just for starters, because listeners to this feed are obviously plugged into the AI moment but probably most do not have much background in materials science, you could give us the basic foundational pitch for materials. Why do they matter so much? Maybe highlight a couple of recent times when an advance in materials directly led to a change in the way we live, just to get the juices flowing and get people motivated to understand all this stuff.

Jonathan Godwin

The most obvious example for people who follow AI is semiconductors. You have this wafer, the silicon wafer, and you etch a bunch of circuits onto it. Those etches are incredibly thin and incredibly small. If a piece of dust or some debris gets stuck in that etch, then your whole chip fails.

You need to find materials that do not conduct electricity, because conducting electricity across the edges in your wafer will screw up your chip. The material has to fit in between all of the different bits in your wafer in order to protect your chip from any sort of debris or pollution that gets in. These are called low-k materials, and without significant advances in the performance of these materials, you do not get more transistors on your chip and you do not get AI.

Materials science underpins literally everything in our lives. The entirety of the Silicon Valley revolution, which gave rise to the technology industry in the West, is based on materials science. Materials science is the original tech startup, and we have just forgotten that these companies built a lot of the industry in America, in the UK, where I’m from, and all around the world.

I do not think the 21st century is going to be any different from the 20th century. The 20th century was, in a sense, the century of materials science, and the 21st century is going to be even more so. We are going to have more chips, AI is going to be even more important, we are going to go to Mars and need incredible spaceships built out of new materials, and we are going to have hollow lenses in virtual reality and need new optical materials to do all of that.

I think materials science is the fundamental missing piece in achieving the science-fiction future, and AI is our most powerful tool and certainly the thing we want to be using to achieve it. That was just one example of all the things materials science enables, but Tim has been working on materials for longer than I have, so he must see many more.

Tim Duignan

I think that is an excellent one, and you do not think about it a lot. The story of Bell Labs is fascinating, particularly the birth of the transistor. It is all materials science: trying different elements and seeing what happens.

The other big one is obviously climate change. Essentially, all of the tools we need to fight climate change are materials-science problems, broadly defined. The most successful tools we have built so far have come out of materials science. Batteries and solar panels, which have led to massive reductions in CO2, both came from breakthroughs in materials science.

It is so universal that I often like to play the inverse problem: name a problem, and I can show you how it is related to materials science in one way or another. It is at the foundation of so many things.

Speaker 1

Maybe a little bit of the status quo would also be helpful in terms of how materials science advances today. We have done versions of this for biology, for example, and contrasted somebody sitting there with a pipette and what that throughput can look like. Obviously, some of that is starting to get automated in the physical world, but more and more of it is moving into simulation and even into AI models.

You are basically doing the same thing, as I understand it. I saw the motto on the website: “Engineer in silico, materialize in reality.” For people who do not know, before we get to engineering in silico, what does the life of a materials scientist look like today? Obviously, it is such a big field that there is no single answer, but just give us a little flavor for what it takes and what the experimental cycle looks like to make advances pre-AI.

Jonathan Godwin

For an experimental materials scientist, there is still a huge amount of experimentation, trial and error, and guesswork involved. If you talk to an average experimentalist trying to make a new material, they will have hypotheses and ideas about what is going on and what they want to do, but there is a huge amount of luck involved, as well as trial and error and intuition built up over years about what is going to work and what is not.

I like to make the analogy to bridge building 100 years ago, when we did not have highly accurate models of which bridges were going to be stable and which were not. There was a lot of guesswork and intuition involved in how you would build those bridges. People built incredible things, obviously, but it was not nearly as efficient. They could not push things to the limit, and they could not do it nearly as quickly as we can today.

To design a modern bridge, you would plug it into some kind of computational modeling software. For every bolt and beam, you would know the exact forces, stresses, and tensions. You would know the limits of the material, and you could design everything really optimally.

If you look at the Brooklyn Bridge, for instance, it is massively overdesigned relative to what it needs to be because they needed to make it safe so it would not fall over. There used to be cases of new bridges falling down. There is a famous one that had a resonance because it had not been tested computationally, and it collapsed after it was built. Today, we model bridges incredibly carefully. We know before we build them exactly what they are going to do.

I think we want to transition materials science to that kind of world, where we know in advance that if we can make a material, we are very confident it will have the properties we want and do the things we want. We also want to answer materials scientists’ questions much more quickly. They might say, “This is not working for some reason,” or, “I am not getting the yield I expected,” or, “The performance is bad.”

We want to be able to go to the computer and say, “Understand this from the ground up. Here is exactly what is going wrong, here are some suggestions for how to fix it, and here is what is actually happening at the molecular scale.” That is a big problem because often we cannot see what is going on.

Speaker 1

Where are we on the path to that dream? This is perhaps a great opportunity to introduce the company, tell a little bit of the founding story, and give us the big vision.

Jonathan Godwin

Orbital came about when I was working at DeepMind and leading a team building large-scale AI models for materials science. I thought it was going to take us a decade to reach the level of performance we are now able to reach, but the speed at which this field is changing is exactly like every other area of AI. It is changing and developing at an unbelievable pace, with new capabilities that we had dreamed about now becoming reality.

Some of the things Tim is going to talk about in more detail later simply were not possible on the world’s best supercomputers 10 years ago, and we are now doing them on our laptops. The goal of our company was to continue developing those models and techniques but combine them with a wet lab and a team of extraordinary experimental materials scientists to translate those AI capabilities into real materials.

We felt that by creating a team with both of those skills, we were going to be able to move faster than anybody else and bring those materials to market and commercialize them more effectively. We were going to ground the AI we developed in the real problems and bottlenecks associated with making breakthroughs through materials science.

That team was not going to exist at DeepMind. It was not going to exist at Google or other AI companies, and it was not going to come out of a traditional chemistry or materials science company. Those companies have been saying no to computational tools for the last 25 years; they are not suddenly going to turn around and start building frontier technology. It had to be a new company if we wanted to achieve that goal.

We are so much further along than we were just a couple of years ago. The things we are doing now are completely different. I think it is only going to be a matter of a couple more years before we achieve the dream: computer-designed materials, with a GUI and an LLM copilot in front of you, allowing you to get to a new battery in a couple of iterations on your laptop and have high confidence that when you make it, it will have those properties.

All the things I spoke about at the beginning—spaceships, better chips, better semiconductors, more GPUs, AGI, and solving climate change—are going to happen through the technology we are building.

Speaker 1

One thing you said caught my ear that I want to dig into for a minute: the idea of simulation. It seems like we have three modalities for pushing things forward. The original is hands-on work with actual physical materials, doing things and measuring properties. Ultimately, everything still has to ground out in the ability to do that successfully, to achieve the results we want to achieve.

Then there was this middle period of simulation. You said some of these things could not possibly have been simulated even on the world’s greatest supercomputer. Now we have this third wave, the AI wave. This was a big theme of the conversation I had with Tim last time when we were talking about crystallization discovery: what distinguishes the AI era for this sort of work from the simulation era, above all, is that it runs orders of magnitude faster.

I have been thinking about that quite a bit, and I am coming to a worldview that basically amounts to the belief that today’s AI architectures seem capable of learning what I might describe as an intuitive physics in basically any space where you want them to learn. We obviously have an intuitive physics that lets us play sports and react to things in subseconds. If we were actually using our feeble brains to calculate and simulate everything step by step, we would never be able to catch a ball, but we quickly grok the situation. We are not calculating in the way a Wolfram model might suggest we would have to.

We seem to be taking an intuitive shortcut, and that seems to be appearing in all of these critical domains, whether it is protein folding, materials science, solving the wave equation for a bunch of atoms in a configuration, or whatever else. Is that how you see it, too, or would you edit my picture? It seems like a pretty foundational update to my worldview that I want to go shout from the rooftops, but I want to make sure I have it right before I repeat it too often.

Tim Duignan

I think that is right, and I do not think we fully understand why yet, or at least I do not. It is very exciting that this seems to be a universal tool. It seems to be revolutionizing climate simulations and weather modeling as well. It is across the board in physical simulation.

My theory is that it has something to do with dimensionality reduction. Often, you do not need to treat every single small part of a system. In our case, this is called coarse-graining in physics and statistical mechanics. There are great theories of it, including the Mori–Zwanzig formalism. It is all based on Langevin dynamics, so it is intimately related to what a lot of the AI algorithms are doing.

The basic idea is that you get these universal behaviors across different scales. They have repeating physical properties, and the key point is that you do not need to treat every single piece of the system to predict what it is going to do.

In our case, we are ignoring the electrons. In a lot of biology, you might want to ignore the solvent. When you get to larger scales, such as looking at a ball moving through the air, you are tracking one object and ignoring a vast amount of other information going on.

Finding efficient ways of keeping track of the important information and losing the unimportant information is a central problem in a lot of physical modeling, and I think AI and machine-learning algorithms are very good at doing that.

Speaker N

What do you think, Jonathan?

Jonathan Godwin

I think you are absolutely right. What has been extraordinary to me is not just the universality of the AI models we use. The things we develop are very similar to the things used in weather forecasting—similar architectures, almost identical—but also how much implicit physics knowledge can be contained in a very small amount of space.

When you look at some of these world models, such as Genie 2 from DeepMind, or a couple of similar systems released by other startups, they are single models that simulate physics. Because they simulate water dynamics and reflections of light in puddles, they understand gravity. They can deal with tiny effects that, just a couple of years ago, would have required very expensive computer-rendering or CGI software to produce.

These models only have a few billion parameters. They are diffusion models, and diffusion models are not parameter-heavy. They are very data-intensive, but they require far less compute during the inference pass, or for learning these dynamics, than a large language model does.

There is this implicit, intuitive knowledge, because we know that calculating all of these effects from the bottom up—including the atomic-scale quantum effects that we really care about—requires a huge number of floating-point operations. We are taking a lot of jumps, whatever is going on. You can call it intuitive, but they are shortcuts that preserve the fidelity of the simulation in the ways we care about.

It is extraordinary that we can condense that information, with that level of fidelity, into something so incredibly small. These models are large compared with neural networks from a couple of years ago, but they are tiny compared with our brains and tiny compared with the amount of computing we were throwing at the problem with bottom-up physics simulations.

The intuitive-physics explanation is a really nice way of thinking about what is going on. It is the same way that you can have extraordinary athletes who can score incredible goals and perform incredible feats without knowing any mathematics. That is exactly what these AI systems are doing when they perform these simulations.

Speaker N

Before we get into the practical details of how you have gone about all of this, I want to ask another philosophical question. It relates to the idea of grokking, although I do not want to be too literal about it.

There are debates about language models in which people say, “Sure, they might learn the concepts that we have learned and encoded in language, but they are not going to learn anything beyond that. How would they? They can only learn what is already encoded in language.” I find discussions around that topic can easily become circular.

Then I move over to a different domain, and we have seen multiple examples at the intersection of AI and biology where you can take interpretability techniques and apply them to, say, a protein language model. You discover that the model has learned something new. A Stanford professor, James xiaan, recently told me that researchers had discovered a new motif by applying interpretability techniques to a protein language model.

It seems hard to deny that the model has learned a concept we did not know, and it did not get that concept from us, because in that case it was trained on sequence data. I do not think I work that way, so this is where the analogy to intuitive physics might break down. I cannot articulate it, and I really have no idea what is going on inside my brain, but it seems like some kind of conceptual grokking is happening in protein language models, allowing them to take these shortcuts.

Is that also what you think is happening in materials science? Do we know what those concepts are? I have some sense of a motif or an active site, but I do not really know what the higher-order concepts would be in materials science that these models might be grokking from the raw data they are trained on.

Jonathan Godwin

One thing that may be going on is that scientists’ brains work that way. A lot of science, although people do not think of it this way, is really intuitive and driven by gut instinct. You rely on your gut all the time. It is possible that even though we do not rationally use these methods to discover things, scientists may be doing this intuitively.

In materials science, that is very common. People talk about chemical intuition, and it is a very important concept. People rely on it, and they cannot always give you a complete explanation. They will give you a story, but it is a very powerful tool. It is possible that the AI systems are doing something similar.

Speaker 1

What does that look like, just to calibrate myself? Does a good materials scientist put a finger to the wind and say, “Let’s throw in a little more acid. I have a feeling about this”?

Jonathan Godwin

I am stealing someone else’s story, but that is exactly what happens with semiconductors. People spend 30 years of their lives working on a very small part of the process I was talking about earlier: what gases do you use to clean and etch a wafer? You need to degas it and clean it with another gas. Which gases do you use? How do you change that slightly? How is it going to affect the silicon?

People spend 30 years of their lives on each of these things. I think there are something like 10,000 process steps involved in making a chip: the packaging, the etching, the lithography, and so on. These people are craftspeople and artisans. They have done so many experiments and developed such strong intuition that they can tell you what will and will not work.

It is not written down anywhere. You cannot go and learn it, and it is not easily deducible from first-principles chemical knowledge. It is passed down in a master-and-apprentice kind of way. One reason people cannot simply recreate TSMC, regardless of how much money the United States invests in new chip fabs or how much China spends trying to build its own fabs, is that you cannot beat TSMC’s accumulated knowledge. They have spent all that time developing tacit knowledge among their chemists and materials scientists.

It really is a deep sense of intuition: “This will not work. This will work.” Someone who has done this for 30 years pronounces their view, and it is almost prophetically correct in many cases. You think, “That cannot be right,” but it actually is.

Speaker 1

My level of admiration for human knowledge, practice, and intuition—which is something that is not really captured in something like a reasoning model moving intuitively through the search space—has grown exponentially through this experience. We have some incredible people who simply know what will and will not work.

Tim, has your experience with chemists and materials scientists been the same?

Tim Duignan

Yes, I think so. It is sad because it is a beautiful thing: this built-up body of amazing knowledge. It is sad to be trying to replace it, but I think there are fundamental limits to the human brain, and we want to push beyond them. We need ways to do similar things computationally, and I think it is very exciting if we can get there.

Combining that with human intuition could be very powerful. As Jonathan was saying, people have reached their conclusions through very different routes, so they may be able to correct problems that the AI systems create, and vice versa. I think it will be a powerful combination when we can get them working together really well.

Speaker 1

Another thing that caught my ear was that you said diffusion models are not parameter-heavy but are data-intensive. We talked last time with Tim about training on molecular simulation data. If I remember correctly, it required a lot of compute. The simulations used time intervals of 10 to the minus 15th seconds, and then the model was trained on that data. You get a model that can do essentially the same thing orders of magnitude faster.

AlphaFold, if I recall correctly, was trained on perhaps 5 figures or maybe low 6 figures’ worth of known protein structures, which was basically everything known at the time. What does the data universe look like in a materials-science context? I can imagine it could be based on simulations, a large catalog of all the materials we have characterized to date, or perhaps both.

Tim Duignan

This has been a big issue in chemistry. For a long time, I was quite skeptical of AI for chemistry because, for the problems I wanted to work on, there was vanishingly little data. Experimentally, I was interested in electrolyte solutions. There are some databases out there, but they might have only a few parameters per electrolyte, so I thought it would be very difficult to do anything.

There are other areas of chemistry where there are large data sets, but in my particular area, I could not see how to do it. Obviously, biology has the Protein Data Bank and other huge data sets. The thing that really clicked for me was that we can automatically generate these training data sets using first-principles computation.

I realized that I have huge data sets if I take into account the massive simulations I have been running based on first-principles physics. These generate millions of data points very easily because the data is generated computationally. You can do all sorts of useful things, such as active learning and automatically generating data in loops where you generate more data exactly where you need it.

That is why I think first-principles-generated data is a big area of future growth in science. It also has implications for how AI generalizes. The data you can generate quickly is often from very small systems.

Jonathan Godwin

You might be simulating what Tim has just simulated: a protein complex interacting with a metal. Now you are simulating that thing in isolation, but it is a small part of something far bigger. The water molecules that surround it, and the rest of the cells and proteins, affect the dynamics in some small way.

For something like a catalytic reaction, if you are working in green chemistry and want to convert CO2 into sustainable aviation fuel, you need to simulate thousands, if not millions, of atoms, and you need to simulate them over very long time scales.

The extraordinary thing about AI is that you can take just a tiny piece of that simulation, perhaps 100 atoms, and run it. That simulation will not tell you very much about chemistry, but it will give you rules and insights into physics. You can train your algorithm on many examples of very small systems, and then it can generalize to very large systems at inference time.

That is the crucial observation that accelerates this field and is the game changer. You can take something that is too simplistic to be useful for prospective materials design, learn the rules from it, and then apply those rules to large-scale systems that do tell you something useful. That would be completely intractable with traditional approaches but can work with AI.

The generalization capacity of these new AI algorithms has been truly game-changing. The thing that completely blew my mind was training on small inorganic crystals—20-atom systems—and then using the model to simulate a protein through out-of-the-box generalization. That tells you we are learning something fundamental at that small scale, which I do not think anyone had ever expected.

Tim Duignan

That is a great point. There might be a finite amount of complexity that you have to capture, and once you have captured it, you can do a huge number of other things. There may be only so much data you need for some problems.

Speaker 1

You mentioned the Orb models. I was doing my homework, and I understand that there is a constellation of models, but I need your help understanding how they relate to one another.

I gather from the website that there is a model called Linnaeus, which is the internal foundational model, and that the models are then fine-tuned for different purposes. It might be useful to talk about some of the tasks we have alluded to.

One task is to simulate a system through time. You mentioned that the same architecture is used in weather-simulation models. Those tasks line up: given some initial weather conditions, propagate them through time; given some initial setup of atoms, propagate them through time.

But I have also heard other task types described. For example, can you come up with a material that will have certain properties? That sounds like a very different input and output. Could you give us a sense of the task types you are trying to tackle?

Jonathan Godwin

This is where Tim’s point that everything is Langevin dynamics becomes especially relevant. We start with a generative model of materials. The right way to think about it is as a generative model of a protein or an image, rather than as a large language model. These are known as diffusion models or score-based models.

I was about to give a description of the physical interpretation, but Tim will get the mathematics more correct than I will. Tim, could you explain how these things are actually all physics?

Tim Duignan

The interesting thing about diffusion models is that their inspiration comes from nonequilibrium statistical mechanics. It comes from generating trajectories of atoms moving through time—molecular-simulation algorithms and Langevin dynamics designed to generate those trajectories.

There is a deep connection between generating new molecules and simulating things through time. The mapping is the score, which is this vector field that the atoms move along in a diffusion model. The score is also basically analogous to the force field, which is what pushes all the atoms in a molecular simulation.

There is a lovely paper from Cheng and colleagues showing that if you train a diffusion model on one of these simulations—on this time evolution—the score is mathematically identical to the force field. There are deep mathematical connections between these things, which is very exciting.

Jonathan Godwin

You can also change the atom types. It is not just the positions of the atoms that you allow to move; you can change the atom types as well.

Tim Duignan

The generative model is essentially doing molecular simulation. You can change the atom types, not just the positions of the atoms, so there are deep connections between molecular simulations and diffusion models.

Jonathan Godwin

That allows us to do a pretraining and fine-tuning step. The way I think about it is that these diffusion models are learning simplified physics. The great thing about a diffusion model is that you get to sample infinitely many points. You have almost infinite training data because, in an image diffusion model, you take an image and gradually corrupt it with noise. Then you predict the reverse step: you start with noise and predict how to denoise the image.

That noising process is the generative process. You can sample infinitely many noise corruptions, so you have an enormous amount of training data. You get very high data density—effectively, infinite samples of your force field—which is a great way to do pretraining.

A generative model is doing the generative side, so we can generate materials according to the functional properties of the material. But it is also a great way to learn simplified physics. You have effectively infinite data for simplified physics, which is a great pretraining step for learning the real physics during fine-tuning, when we learn the real forces computed through quantum calculations.

Linus is the foundational model, learning simplified physics, which we then adjust to give the correct physics that we can use in real simulations. That fine-tuning step gives us Orb, which is essentially calculating the interactions that happen at the quantum level between atoms.

Once you have that right, you basically have physics right. You can then move on to simulating whatever you like. You can simulate a nuclear reaction, what happens in the body, or the incredible materials we use every day.

Speaker 1

I understand the image example. I have given a very similar explanation to many people about the image version of diffusion. I have a good intuition for adding noise to an image and for reversing that process by predicting what the image would look like if it were a little less noisy.

I can also imagine steering the process: what would this look like if it were a little less noisy and a little more like a text prompt aligned to image space through CLIP or something similar?

In the context of proteins or peptides, I have learned that there are custom noising processes because the problem is more complicated. If you simply add pure noise, you end up with something incoherent. There are also issues with discreteness when you are talking about changing atom types, which does not seem like something that is naturally differentiable.

There is also the alignment question: how do these low-level configurations of atoms translate to the qualitative macroscopic properties we care about? If that was the 101, give me the 201 on what noising means in this context and how you align low-level things with the properties we care about.

Tim Duignan

Noising is very straightforward in materials. It is similar to images. You take a crystal and jiggle the atom positions around, jiggling them more and more until the structure becomes totally incoherent.

There are techniques for discrete diffusion as well. They are related to the techniques used in images, and they work fairly well when transferred to materials, so it is actually more similar to images than to proteins in some ways.

The one thing that is unusual in a crystal structure is that if you have a box with a crystal inside it and an atom moves across one face of the box, it appears again at the opposite face. Noising means something slightly different because you do not have the opportunity to move all the way off into the distance. You are always sitting inside one of these small boxes. That creates some mathematical challenges, but in practice you treat it much as you would an image.

When you think about guiding the model toward a property, it ends up being very similar to an image model. There are a bunch of ways to do it, but one way is to train an image classifier. The classifier creates an energy function that says how close an image is to a cat. As that number goes down, the image becomes closer and closer to a cat.

We do the same thing. We take a crystal structure or a material and train a simple supervised-learning algorithm to predict a property. We set that property to the value we want and nudge the material toward it through the differentiable machine-learning classifier.

Speaker 1

I would like to have a better intuition for how we make the leap from one atom type to another.

Tim Duignan

You treat categorical prediction—for example, predicting the atom species from 0 to 117—as a continuous function. Instead of using a cross-entropy loss, you treat it as a continuous target and then cast it back into a categorical one.

This is a technique called Analog Bits. It came out around 2022 as a way of doing discrete diffusion. You cast the problem as something with a gradient, effectively softening the discrete distribution. It is a simple and elegant technique that works quite well. It is in a paper, so it must be correct.

Speaker 1

Can we go back to the tasks? When I use an image diffusion model, I might show up with a text prompt and say, “Make me a picture of a scientist in a materials-science lab.” Or I might provide an image and ask the model to make it more like a dinosaur working in a materials-science lab.

What do I show up to the different Orb models with, and what do I get back?

Jonathan Godwin

For Linus, you might say, “I want a material with a band gap of X. Give me 10 materials that have this band gap.” It will generate materials with that property. That is very similar to an image-generation model: you start with a description and generate an image that fits the description.

That is what Linus, our foundational model, does. The Orb models that we have released publicly have a different task: simulation. You have a system and want to know what is going on from a first-principles perspective. How is a reaction taking place?

The main thing they do is predict forces from coordinates. You give them the positions of all the atoms and their atom types, and they output the forces and energies. That is essentially all they do, but it unlocks a huge number of other capabilities.

People have been running these simulations for decades and have developed theories, code, and knowledge about how to use the ability to calculate forces. We can leverage all of that to compute a whole host of other properties. We can calculate diffusivity—how quickly atoms will move—and chemical equilibria, such as how strongly atoms stick to one another.

Essentially, all of the fundamental things you want to know about materials can be extracted from these molecular simulations. There are two ways you might generate new materials with the properties you want. You could directly simulate a large number of them and see whether they have the properties you want, or you can bypass that with a direct generative model. You can also combine the two approaches: generate hypotheses with the generative model, then feed them into the simulation to check whether the more exact simulations produce the result you want.

Speaker 1

Keep going there, because the outer loop is really interesting. We talked about it in the context of your work with solutions. You calculate all the forces, apply a time-step increment, everything moves a little bit, recalculate the forces, and the next thing you know, crystallization is appearing out of nowhere.

I would like to keep zooming out from the internal tasks and talk about the scaffolding you are building around them.

Tim Duignan

There are three key things you can do with molecular simulation.

The first is what I think of as a computational microscope. You directly see what is happening. You can make a movie and watch what the atoms are doing. This is a powerful and useful tool because it gives us access to temporal and spatial scales we have not had access to previously, particularly the nanoscale: nanometers and nanoseconds, a billionth of a meter and a billionth of a second.

A huge amount of interesting activity happens at that scale. Chemical bonds form and break, among many other things. It is very hard to observe experimentally because these systems are atom-sized, so you need something of a similar size to observe them experimentally. Computationally, we may be able to access them.

Aral, your previous guest, discussed this in the context of protein folding. The static problem may now be solved—you can predict the static structure of all these proteins—but what we are really interested in now is how they move and what their dynamics are, particularly at time scales below microseconds that we cannot access experimentally.

This idea of a microscope is powerful because it has revolutionized science many times. When the first microscope was built, everything people looked at led to new discoveries. They looked at pond water and found bacteria. They looked at blood and found red blood cells. The hope is that if you can simulate a huge number of things, you start to make important new discoveries and find new things simply by looking at them.

The second thing you can do is what I think of as in silico experiments. You no longer have to order the chemicals, synthesize a new molecule, do all the work to see it in action, and then use microscopy to understand what is happening. You can change the atom type in your input file and rerun the simulation.

You can do simulations much more quickly than experiments, at least in principle. Eventually, you can imagine doing high-throughput experiments computationally: running thousands of simulations and brute-forcing your way to the material that works best.

The third thing is actually the most exciting in the long term. You can use simulations to extract data for larger-scale models. I think of this as providing a bedrock for physical simulation.

We simulate all sorts of physical systems, including the climate, using a variety of modeling approaches. But these models all need to be parameterized. They need to have numbers plugged into them, and it is often very hard to work out what those numbers should be. You have to do experiments, which can be very difficult.

If you could use lower-scale simulations to extract that information, you could plug it into larger-scale models. You could also generate training data sets this way and use them to train generative models.

There is also the coarse-graining idea I mentioned. It is a statistically and mechanically rigorous, well-defined procedure in which you can show that you can ignore certain aspects of a system and retain others. You might ignore the solvent, for instance. One day, you could simulate how proteins fold over time. People are already beginning to do this using approaches trained on all-atom molecular-dynamics simulations.

It is very exciting to be able to look at how proteins fold dynamically. Materials science is similar: you want to know both the static structures that form and the dynamics—how chemical reactions occur and what the transition states are. These are all very difficult to access experimentally.

Speaker 1

How do we go from an idea to a material using these technologies?

Jonathan Godwin

We use these technologies in several different ways, paying close attention to making them useful for the prospective design of materials. You often start with a functional property you care about and decompose the chemistry problem into a subset of things you care about in the material.

You might say, “I want to create a metal-organic framework.” This is a large, sponge-like material. You might want it to have a very large pore size because each bubble in the sponge needs to be big enough to contain a lot of material that you want to adsorb.

You might also want to attach small tendrils, or ligands, to the metal-organic framework. It needs to have many active sites where you can attach those ligands, which will act like tendrils that pick up the things you want to adsorb from some unusual mixture that you are passing through the material.

You use a generative model to say, “Generate a material with a large pore size and many sites where I can attach ligands.” The model might give you 10, 15, or 20 candidate materials with those properties.

But making 20 materials still takes a long time, and there may be many important properties that are difficult to express as generative prompts. We may want to understand the kinetics: how quickly will the adsorption happen? Will the tendrils block the entire pore space so that the material we want to capture can no longer get through?

These questions are difficult to express in a prompt, but simulation can answer them very well. We go through an in-depth process of using our AI-accelerated simulation tools to answer these questions at unprecedented levels of detail and accuracy.

By the time we decide what to make, we have answered 90% of the questions we need to answer in order to feel confident that we are going to get that kind of material. We use the generative side for creativity and the AI simulation side for qualification and insight into what is going on.

The simulations often stimulate more creativity. Tim is always coming up with something unexpected: “I saw this simulation, and this was not at all what the literature said, but I actually think it agrees better with the experimental data.” You see new things, and that stimulates creativity and new ideas for materials.

You might go through that process two or three times. Eventually, you have two or three materials that are fully qualified and slightly outside the scope of what you would have come up with yourself. The creativity of the AI and the generative process moves you in a different direction, and then you go into the lab.

Your success rate in the lab is much higher at that point. You are confident that the material will have the properties you want, and you are doing something more novel than you otherwise would. That novelty is just as important as the success rate because you are entering new areas of chemical space, which is often where the most exciting materials and chemistries are found.

That is the process we use. We use these different tools and functionalities at different stages of the materials design and R&D process.

Speaker 1

When you say you want something with a large pore size and many active sites, are you saying that in natural language? Or is there a well-defined property space that you are specifying?

Jonathan Godwin

Those are things where you could have a well-defined property space. We have generally found that language is not always the best conditioner for generative models of materials because you are often looking for a specific range of values rather than a textual description.

Text is great for describing an image because you do not have a scale of values to condition on. But if you want a specific number for your pore size, it is better to provide that number than to put it into a free-text box.

Tim Duignan

A lot of the time, you are trying to hit a sweet spot. You do not want too little or too much. You want the right amount, so it can be quite quantitative.

Speaker 1

One more detailed question about the models themselves, and then we will get into the application that is making the headlines. How big are these models, and what are their compute requirements? Are we talking about something that would qualify under the Biden executive order as 10 to the 23 floating-point operations if we wanted to call this more of a biological system? Where are we in terms of the resources going into these models?

Jonathan Godwin

Your mental model should be an image model, such as a convolutional-neural-network-based diffusion model. Most image diffusion models are probably still U-Nets, so you should think of them in those terms. They require similar levels of data and compute, and they are much smaller than the latest language models.

That is one reason we can build world-leading models and develop new materials with less than $40 million raised. That is still a huge amount of money, but we would not be able to train a large language model for that amount.

The exciting thing is that we may be at the GPT-1 stage. If you scaled these models, particularly on the training-data side, and threw large clusters at generating really high-quality training data sets, there would be a lot of room to improve them.

Speaker 1

I remember that Jonathan Frankle and Obby from Mosaic showed that Stable Diffusion could be trained for under $50,000. I do not know whether that is exactly where you are, but as a general order of magnitude, those costs are certainly far below what you would need for trillion-dollar data centers.

Let us talk about the paper you graciously shared a preview of with me, in which you are modeling the potassium ion channel that sits in the membranes of our cells and lets potassium ions in and out in an ordered fashion.

People are probably familiar with ion channels and cells, but why does this matter so much? It seems like this was a longtime dream for you—a challenge you wanted to tackle.

Jonathan Godwin

This has always fascinated me. The basic idea is that the body is really an electrical device. We do not think of it this way, but it is. A computer uses electrons in metal to carry electrical currents around, whereas the body uses ions: things like sodium, chloride, and potassium ions.

The way the body carries these signals is that cells have a high concentration of ions—in this case, potassium ions—and when a cell wants to send an electrical signal, it opens things called potassium ion channels. The potassium ions flow through those channels, creating a change in voltage that triggers subsequent processes.

This is how a huge number of the body’s functions work. It is how your muscles contract, how your heart beats, and how the neurons firing in your brain carry electrical signals through the flow of potassium ions in and out of cells.

When the Orb model first appeared, I was not working for Orbital. I saw it and thought it looked very exciting. It had very good accuracy and, more importantly, it was very fast. I started trying it on many different things, and it was doing a fantastic job and giving me stable results.

I simulated many electrolyte solutions, which is what we discussed last time and what I was particularly interested in. But the fundamental reason I had always been interested in electrolyte solutions is that, apart from their importance for climate change, they are central to biology and this electrical-signaling function.

The key challenge is the potassium ion channel. No one has really been able to fully understand it using experimental techniques or traditional computation. We do not even know some of the most basic questions about it. For example, is it just potassium ions flowing down these channels, or is water flowing through them as well?

That is an unresolved question. Papers have come out in both directions in top journals every few months. I think of it as a keystone problem—a problem that demonstrates a fundamental lack of some capability. If a new tool came along that could solve it, that would be very important.

It had always been a dream to use the things I was working on to address it. Because Orb was doing so well, I started to think that perhaps I was thinking about this incorrectly. Maybe this was the universal force field that computational chemists had been dreaming about for decades.

I thought, “If it were, what would I do? This is the problem I would try to solve.” I set up the simulation, and at the time I literally thought, “This is a silly idea. It is not going to work. It is going to crash immediately.”

The models I had built myself previously became unstable as soon as they went out of distribution, so I expected the same thing to happen. But it did not. The simulation started running stably, and I began to see interesting things. It was so exciting that I literally started jumping up and down.

The training data looked nothing like this system. It consisted of small, highly periodic, repetitive crystal structures, so the model had never seen anything like a potassium ion channel. The fact that it produced something reasonable and showed interesting behavior was very exciting.

Speaker 1

Eureka moments increasingly come from AI systems. What do we attribute the generalization from the small crystal training data set to this rather unusual structure?

People can search for a picture of a potassium ion channel to get a rough approximation of what it looks like. You have a cell membrane and this unusual structure with some symmetry that protrudes through both sides of the membrane. Through the channel, ions can flow.

That is very different from a crystal structure. Nothing about it is crystalline. Do we have an explanation for why the model generalizes so well?

Tim Duignan

I do not know. As I said, I did not expect it to work. My general philosophy is that I initially did not trust many AI systems because they can hallucinate and do strange things. I have tried to change my philosophy because it is possible to go too far in the other direction and not trust them enough.

Sometimes it is worth trying and testing these systems to see what they do. It was definitely surprising. If I had to give an answer, I would say that there is a fundamentally finite amount of information about molecular interactions that you need to know.

There are things such as the dipole moment of each atom and how those interactions decay spatially. Those properties can change in different circumstances, but there is still a finite set of information you need. Perhaps the crystal structures are enough to extract enough of that information to generalize to a much broader space of materials.

Jonathan Godwin

I think that is right. What is remarkable and continues to amaze me, though, is the effect of scale. We take scale, particularly data scale and engineering excellence, incredibly seriously at Orbital. It is part of our pedigree as an organization.

Without casting too much shade on the academic community, we are bringing AI engineering into this discipline, and that has led to extraordinary results. We have seen similar results in natural-language processing and images, and now we are applying the same techniques to materials.

All the information is there. There is a limited amount of information you need. You are learning interactions between different types of atoms. If those interactions are the same whether they occur in a protein or in a small crystal, then as long as they are present in the crystal-structure training data, a network that learns efficiently should extrapolate well.

It is extraordinary, but after seeing what I have seen in AI over the past few years, you learn to expect these things.

Speaker 1

Expect the unexpected and expect to be surprised. That is definitely unavoidable these days.

What is at stake here? Is this more of an AlphaGo moment, where it is remarkable but does not change the economy much in the immediate term? Or are there practical things that become unlocked by a better understanding of the potassium ion channel?

What did you learn about the channel, and how has our state of knowledge advanced?

Tim Duignan

The first caveat is that the work has not yet been peer-reviewed. I have sent it to potassium-channel experts, who say that it looks interesting and plausible and is consistent with what we know experimentally. Ideally, we would like to see direct experimental confirmation of what we observed.

The exciting result was that water molecules entered what we call the selectivity filter, the narrow part of the channel. Classical simulations do not normally show this. In the classical approach, people tend to assume that potassium ions pass through alone, although the experimental evidence about whether water passes through is mixed.

We saw a particular hydrogen bond in which a small hydroxyl group on the protein reached out, grabbed water molecules, and pulled them into the channel. That was very interesting. As far as I am aware, no one has seen that before.

It makes sense. There is a new qualitative phenomenon that the other methods do not produce, and it explains why they did not see water entering the channel. We know that if you delete that hydroxyl group through mutation studies, the conductivity drops by almost an order of magnitude.

The group is clearly doing something important, but experimentally no one has really known what. This looks like a plausible explanation. It still needs to be confirmed, but I am optimistic that it explains the longstanding dispute in the field about what might be missing. It is something we can now try to validate experimentally.

In terms of why this matters economically, many diseases are caused by mutations in potassium ion channels. These include heart diseases, long QT syndrome, and heart palpitations. If potassium ion channels are not working properly, they stop electrical signals from flowing in the correct rhythm, which affects the beating of the heart.

If we want to understand how to fix those mutations, we first need to understand how the non-mutated channel works. We can then simulate the mutations and see how they impair the function. One day, we might even design drugs to restore that function.

That sounds somewhat science-fictional, but it is already being done for chloride ion channels. Vertex has small drug molecules that can restore the function of chloride ion channels. There are definitely biological applications for this kind of work.

Speaker 1

Just to make sure I understood the point about the hydroxyl group: it had previously been observed that if you create a cell with a mutation such that the group is not there, the channel works an order of magnitude less efficiently?

Tim Duignan

Exactly.

Speaker 1

What you saw that was new was a mechanism that made you say, without feeding any of that background information into the model, “That is probably why this is happening.”

Tim Duignan

Yes. If I had to bet, I would say it definitely is. It is very plausible. The model was doing something new that made a lot of sense, and I have not seen it reported anywhere else in the literature. At the very least, it is a strong hypothesis, and if I had to bet, I think it will be confirmed.

Speaker 1

How far do you think this goes? There could be an obvious hard limit related to the structure of the models, or there could be a more conceptual answer.

We have these models trained on very small crystals, and they can simulate something as complicated as a potassium ion channel. The channel exists in a cell, the cell exists in a body, and the body exists in a world, but you have to draw a box around the system to bound the simulation.

You essentially decided that beyond a certain box, everything outside it would be treated as fixed in place. If I understand correctly, some additional atoms are represented in the simulation but are simply not allowed to move, so that you can focus on the region of interest.

Tim Duignan

Exactly. In the early days of simulating these systems, people used tricks like this. I thought I could use some of them with the new technique.

You want to simulate the selectivity filter, the narrow part where everything interesting happens, but you need the surrounding atoms because they help stabilize that central part. Then you need the atoms surrounding those, creating an infinite regress. You end up needing a huge system.

In modern simulations, we normally use periodic boundary conditions, where if an atom leaves one edge, it comes back in through the other side. That is how people run large simulations with classical techniques.

I only wanted to take a quick look to see whether anything interesting was happening in the selectivity filter. At the time, I had only one V100 GPU, so my computational resources were very restricted. I had to cut off that infinite regress at some point.

I cut it off about 20 angstroms away, roughly 4 or 5 layers of atoms out, and froze the outer layer to keep the system stable long enough to see the interesting dynamics in the selectivity filter. It is a bit of a hack. Eventually, we will do the full simulation of the entire system.

The other issue is the membrane. Because the potassium ion channel sits in the cell membrane, which consists of these fatty layers, you need a very large system to simulate the whole thing.

Speaker 1

It is incredible that this was done on one V100.

Tim Duignan

Yes, it is pretty wild.

Speaker 1

Is it mostly a matter of scaling up the inference infrastructure? I assume the model has structural limits as well. You could not take the Orb model and simulate an entire cell just by giving it enough inference compute. What constraints will we hit as we try to simulate larger and larger systems?

Tim Duignan

It is an interesting question. I do not know; we will have to wait and see. I did try simulating a larger system, and I saw it become somewhat unstable. It is difficult to set these simulations up properly. The cell membrane has to have the correct thickness, among other things.

There is a lot of work to do to understand how to scale these systems to larger sizes. We can run very large simulations with Orb, but the systems become slow. Then you have to start scaling across multiple GPUs.

People have simulated huge systems with neural-network potentials before, including an entire virus capsid, and it looked stable. It is difficult to validate because the time scales are short and the simulations are computationally expensive. That is the main problem: reaching longer time scales.

The coarse-graining idea will be key. We can run all-atom simulations, as I did for salt crystallization, and then say that we are going to ignore the solvent molecules. We train a new model that predicts what is technically called a free energy, but it does the same thing as an all-atom simulation while simulating only the parts of the system we care about.

That will get us to much larger scales. I am optimistic that we can eventually simulate an entire cell using this approach. It will involve some hard work, but the theory and tools are there. It is a matter of investing in them.

Jonathan Godwin

There is no fundamental limitation in the current models that would prevent us from extending the size of the simulation. It is possible that we will need longer-range models to handle long-range electrostatic interactions. In these systems, charges have very long-range interactions because they decay as 1 over r, just like gravity. Those effects can extend over enormous distances.

It is difficult to say what will happen with very large systems, but it will be interesting to find out.

Speaker 1

There is not an equivalent of a context window or some hard limit built into the current models that you would hit after a certain number of atoms?

Jonathan Godwin

You would eventually hit the memory limits of the GPU if you tried to scale the system too much, but that is not our bottleneck at the moment. We can do huge runs. The problem is that the simulations become slow as the systems get larger.

Tim Duignan

This is different from what happens in a language model. A language model with a finite context window has positional encoding as a feature of the input. During training, each position in the context window has a specific feature associated with it. At test time, you cannot generalize much beyond the length on which you trained.

Models such as Mamba and other state-space models do not have that explicit positional encoding, so they can have an effectively infinite context window. The same is true of the models we use. We do not have explicit positional encoding, which means we can extend the system as far as our computer resources allow.

You can increase the size of an atomistic simulation roughly linearly with the number of chips you add. There is no fundamental architectural limitation like there is with a Transformer-based language model. Molecular interactions also tend to be local, which is a useful property that enables this scaling.

Speaker 1

I have to ask: when do we get our room-temperature superconductor?

Jonathan Godwin

It will happen. Give us at least 5 years, but it may happen sooner. We are going to solve carbon removal and scale AGI energy-efficiently and cheaply first. We have some things to solve before then, but once we have done those, we will think about the room-temperature superconductor.

Tim Duignan

The problem with superconductivity is that it is what we call a highly correlated electron system. It does not have the property of very local behavior. There are complex quantum-mechanical effects taking place, and the current level of theory we are training on, density functional theory, is not good enough to describe them.

It could still be very useful for finding new crystal structures and working out how to make new superconductors once you know how to induce the electronic properties you want.

Speaker 1

Could you talk a little more about the roadmap between here and a room-temperature superconductor? What else is on your to-do list?

Jonathan Godwin

Our focus at the moment is on data-center materials. We think scaling AGI is one of the most important things we can do, and new materials will be an important part of making that both cost-efficient and sustainable.

As rack density and chip power density increase, we need better materials for thermal management and better materials for many other purposes. We can only assume that we will need better and better chips to bring AGI to life. In that sense, new materials for data centers and new materials for AGI are our focus.

We also want the technology to be sustainable, so decarbonization is our other focus. We will have a pipeline of materials and products in this space. We have released information about one of those products, and more will be coming over the course of this year.

The product we have already discussed is our carbon-removal material. We designed it to operate and capture CO2 under the operating conditions of a data center. It allows us to use the waste heat and airflow from a data center to maximize its CO2-capture capacity and ultimately offset some of the scope 2 and scope 3 emissions of data-center operators.

What we are seeing from hyperscalers may be the largest capital-expenditure buildout in history, but those companies are very serious about their climate goals and have the resources to invest in reducing the carbon intensity of their operations. We see this as a great opportunity to deploy advanced and effective carbon-removal and decarbonization technologies.

Speaker 1

It seems like technologies such as this may be the only thing between the hyperscalers and a lot of broken net-neutral promises. I do not think they are going to stop building AGI data centers, so something is going to have to give.

Is the material you described earlier—the sponge-like material with high pore size and active sites that grab onto things—the right mental model?

Jonathan Godwin

I cannot say too much because we are protecting some of that information, but it is certainly a sponge-like material with reactive sites that selectively interact with CO2 over the other things present in the air.

We care about its adsorption capacity at the temperatures at which a data center operates, the composition of the air, and other factors. We tailor materials and develop new ones specifically to take advantage of those operating conditions.

That is very difficult to do. Developing a new material for CO2 capture normally takes years and years. We are doing it incredibly quickly, which allows us to maximize the unique opportunities presented by the deployment environment.

Speaker 1

You could imagine multiple paradigms for this kind of material. You could imagine a sponge that adsorbs a lot of material and then gets buried with the CO2, so you continually bring in more adsorbent material and the used sponge becomes a waste product. Or you could imagine cleaning it off somehow and reusing it.

Jonathan Godwin

You need to adsorb and desorb the CO2, and that is a crucial part of the overall cost of CO2 capture. If you can reduce the heat at which you desorb the CO2, or tailor that temperature so you can use the waste heat from a data center, you can massively reduce the energy costs.

The longer you can reuse the material, the better. Eventually, desorbing the CO2 will be more energy-intensive than making another quantity of the material. The manufacturing process is itself very energy-intensive and costly, so if the material can last for 10 years, that is one way to radically reduce the cost of CO2 capture.

Tim Duignan

The one-off capture is relatively solved. The regeneration problem is the real challenge.

Speaker 1

Manufacturing and scaling up production is obviously a different business from developing the models that design the materials. You have traditional materials scientists on the team to close the loop. I assume you are licensing the materials rather than building factories and scaling production yourselves.

Jonathan Godwin

That is right. We may license the technology or work with suppliers. Nvidia does not have a fabrication plant, but it still sells chips. We do not have to build the manufacturing capacity for a material in order to sell it.

I do not think we are going to stand up a massive plant somewhere in New Jersey.

Speaker 1

Do you need a lot of physical capital even to do the validation experiments? It seems like that could itself be a scale game.

Jonathan Godwin

We need more physical capital than a software business does, but we are not going to be raising money for capital expenditure. Our lab is cheaper than many of the GPUs you need to run an AI company. It is not beyond what venture-capital-funded businesses should be able to do.

It is important to have the lab internally so we can get feedback quickly from experimentalists about what is and is not working, what they need, and so on. It would be much harder without that capability.

Speaker 1

I am sure you saw the recent paper from Aiden toner Rogers from MIT that made the rounds. I do not think we know what models were being used or which company conducted the research, but it was essentially a behavioral and industrial-economics study of applying models in this general area to the practice of materials science.

The headline results were a 44% increase in materials discovered, a 39% increase in patent filings, and a 177% increase in downstream product innovation. Interestingly, the AI systems were doing much of the idea generation, while the scientists followed up with more validation.

Overall, the system produced more output, but there was an important caveat: 82% of the scientists reported lower satisfaction with their work. I assume they felt they wanted to be the idea generators, and now the AI systems are generating the ideas while they are left to validate them.

You could react to that in multiple ways. I am interested in how those increases in throughput relate to what you expect or aim for. Also, is there anything this implies for the future of work? I am not someone who gets too hung up on whether my job changes. If it is going to solve climate change, I can accept somewhat lower job satisfaction here and there. But do you think people will adjust and find joy in the work again, or is this a sign of things to come?

Tim Duignan

On the question of enjoyment, I am a physicist and mathematician by training. I have been trying to get into chemistry because it is so important, and I love it, but I do not have the deep knowledge that chemists have. I did not study it as an undergraduate.

From my point of view, these tools are incredibly empowering because they give me access to the intuition we were talking about. I can run a simulation and see bonds forming. I do not need to know anything about potassium channels, know that the hydroxyl group is reported in the literature to be important, or have a deep intuition about what bonds with what. I can run the simulation and see it.

I have sympathy for people who spent their careers building up this intuition and now see something else doing it as well. That is a sad thing. But I think there will still be many ways for people to be useful, and the satisfaction of getting materials out into the world that have an impact much more quickly will more than offset the spark of solving the problem entirely by yourself.

Jonathan Godwin

It reminds me of the job satisfaction of some of my friends who work in AI research. Their satisfaction has declined significantly because they used to do creative work, and now much of what they do is engineering.

There are a limited number of experiments you can perform yourself. You cannot start a research project and freely experiment with it because the compute required to test it fully is so expensive. People’s jobs have radically changed in AI research over the past 2½ years.

For many researchers, their satisfaction has declined despite the fact that their jobs have become more prestigious, they are being paid more, and they are watching AGI take place in front of them. I think that will be the case for many people. We will become more productive and richer, but we will have to find other ways to derive value from our work.

Creativity is one of the first things that will go. Manual tasks will follow, with robotics coming somewhat later, after the creative industries have been fully accelerated through AI.

I am not optimistic that people are always going to enjoy working with AI. I am fully convinced, though, that the only way the future is going to be a happy world is through AI. We need to produce more things for less money, and that means we need a greater acceleration of scientific discovery than humanity has ever achieved.

The only way we are going to achieve that acceleration is through AGI. We have to develop it as quickly as possible. We will have to deal with some short-term pain to get to what one might call a glorious future in which a large portion of economic activity is fully automated.

Speaker 1

That is probably a good note to end on. Is there anything else you want to discuss that we have not touched on?

Tim Duignan

One point is the combination of large language models and neural-network potentials. Universal force fields are quite powerful, and it is worth emphasizing why.

A central challenge, if you are worried about artificial superintelligence, is how these systems will get experimental validation. Perhaps they can design whatever new materials they want, but they still need to validate those materials.

The combination of LLMs with force fields gives them a kind of ground truth. It may allow them to do these things internally without needing to go out into the laboratory and measure everything experimentally.

It is a powerful combination because it provides a way to ground the systems and check for hallucinations. We are seeing this in models such as o3, which can solve extraordinary mathematics benchmarks, but the problems need to be ones that can be validated somehow.

Neural-network potentials may provide a way to validate scientific hypotheses automatically in the computer rather than having to go into the lab. I think they will be a powerful combination going forward.

Speaker 1

Do you think we end up with multimodal models in which the AGI uses the same weights to talk to me and run molecular simulations?

Jonathan Godwin

Certainly, we will have a shared embedding space. I do not know exactly how it will work, because some image-generative models have a separate image decoder that feels like a different model structurally. But they have the same embedding space.

We will certainly have multimodal embeddings. That is where the concepts and reasoning happen. Integrating atomistic systems into the training of a scientific AGI would be a very interesting thing to do, and I imagine someone is already working on it somewhere.

Speaker 1

It will be difficult to make sense of it. I worry for several reasons. There has recently been a lot of discussion about deception, followed quickly by discussion of reasoning in continuous latent space.

I am among the people who think, “I liked it better when I could read the chain of thought.” If the chain of thought is going to happen entirely in latent space, that is going to be strange. But it does seem like a powerful attractor for several reasons, including that systems will be able to communicate with one another much more effectively if they do not have to condense everything into language tokens to send messages back and forth.

I do not know how we get around that. It seems like that vision takes us quickly to a world where we are simply hoping the AI systems are doing the right thing for us. Is there an alternative vision?

Jonathan Godwin

Someone recently said that you just hope moral realism is true—that the systems will find the same things and do the right thing.

Speaker 1

That is a sobering note to end on, but this has been inspirational work. The images from the crystallization work and the potassium ion channel are really impressive.

Jonathan Godwin and Tim Duignan of Orbital Materials, thank you both for being part of The Cognitive Revolution.