[BidClub_]
Latent Space · · 77 分钟

🔬 AI 在科学中的边界——为什么我们需要自驱动实验室 — Joseph Krause, Radical AI

Joseph Krause

YouTube
TL;DR
  • Radical AI 的核心押注是,材料 AI 的优势来自闭合实验闭环,因为成分生成只是起点,“材料本身才是真实标准”。 有用的材料必须经过合成、表征、加工、制造,并接受成本、供应链和应用约束的检验。因此,Krause 认为 Radical 与 Lila、Citrine、Periodic 等公司的差异,在于实验数据和实验室基础设施,而不只是模型更好。

  • 早期吞吐量数据表明,这可能是一次真正的跃迁,但 Radical 距离端到端制造仍很远。 Krause 对生产约1,200种合金给出了两个不同时间窗口:一次说是5至6个月,后来又说是3个月。其中约300种不在既有文献中,“可能有10种”的性能尤其令人兴奋。目前吞吐量为每天8-20种合金,单种成本约60-300美元;目标是在6月或7月达到每天100种。MACH 项目给出的基准则是12个月完成500种合金。

  • 商业化风险正在下游显现,资格认证和制造经验可能吞噬相当一部分发现优势。 航空航天资格认证通常需要约10年,而 Radical 目前处理的规模是克级,约200-500 g,距离讨论中的300磅或10吨仍有明显差距。Krause 认为,进入国防或航天应用存在一条可行的3-5年路径,但载人飞行涡轮机不在其中;半导体集成则“还相当遥远”。

  • 可服务市场不只是发现更强的合金,还包括与需求这些材料的产品同步设计材料。 含有5至7种大致等比例元素的高熵合金,可能用于超过2,000°C、甚至3,000°C的高温环境,也可能应对压力、氧化、腐蚀或中子轰击。Krause 借用的表述是“并行工程”:工程师不再围绕几十年前的材料设计火箭、涡轮机或芯片,而是让材料与产品共同迭代。

  • Radical 的自驱动实验室本质上是一个横跨软件、机器人、感知、科学判断和实体工具的编排问题。 自动化实验室像是免提驾驶;自驱动实验室则是一辆 Waymo,既选择路线,又执行完整的研究项目。定制夹具必须从托盘中撬出3,000-4,000°C高温熔炼形成的合金“纽扣”,模型必须判断熔化是否完成,操作系统还要决定失败样品是继续流程还是直接终止。

  • Krause 认为,材料科学受实验约束,而非算力约束,因此实验室吞吐量和实验数据才是经济护城河。 相关搜索空间可能包含约10^40种合金,但由于高质量实验数据稀缺,数百次实验已经开始产生有用的发现信号。他的直白说法是:“在科学领域,我们认为护城河不是模型,而是实验。”因此 Radical 可以开源模型,同时把实验和数据视为真正优势。

  • 更广泛的战略判断是,自驱动实验室可以放大稀缺的科学劳动力,帮助美国与中国竞争,同时不必复制由单一实体控制公私活动的中国体系。 Radical 表示,一名冶金学博士可以同时监督10个研究项目,颠倒传统上约10名研究人员对应一个项目的比例。Krause 提议的制衡方案,是把国家实验室的数据、高性能计算和仪器,与私营部门的软件和自主系统结合起来。但主持人指出,中国同样可以部署这些生产力工具,因此执行能力和规模化基础设施仍是决定性因素。

摘要 · 为研究而整理的核心内容

1. 材料 AI 必须闭合从假设到物理真相的循环

  • 当被要求说明 Radical AI 与 Lila、Citrine、Periodic 以及日益拥挤的赛道有何区别时,Krause 回到一个核心信念:“材料本身才是真实标准。”模型很重要,但一个被提出的成分配方只是起点。

  • Radical 设想的闭环是:AI 科学家提出候选方案,实验室完成合成和表征,再将实验结果反馈给下一轮研究。自动化程度很高,但人类仍参与其中,尤其是在合成和科学标注环节。目标不是得到一个看起来合理的数字结构,而是得到一种最终能够进入工业应用的材料。

  • Krause 的因果逻辑是,性能往往在选定成分之后才真正出现。微观结构、热处理、后加工、增材制造还是铸造,以及生产规模,都会决定同一种名义合金究竟是高强度、延展、耐氧化、可制造,还是毫无用处。

  • 行业所说的15-30年开发周期,反映的是流程割裂:学术界负责发现,政府支持的项目进行初步测试,大公司则把现有系统优化5%或10%。数据很少跨越这些交接环节,发现与制造由此脱节。

2. Radical 已实现发现规模的自动化,但还没有覆盖材料的完整生命周期

  • Radical 目前覆盖假设生成、合成、表征和早期性能测试。其表征组合包括 SEM、EDS、XRD、XRF 和 TGA,分别用于识别结构、相态、化学成分和热行为。

  • 性能测试包括氧化表现、拉伸应力-应变曲线和显微压痕。维氏硬度可以直接测量,但实验室用于判断延展性的信号只是代理指标;Krause 明确提醒,这“不是对延展性的精确测量”。

  • Radical 报告的产出约为1,200种合金。Krause 一次将时间窗口描述为5至6个月,后来又说是3个月。其中约300种相对于文献属于新材料,“可能有10种”的性能足以推动更深入的行业讨论和专利工作。

  • 规模边界非常明确:Radical 处理的是克级样品,通常为200 g或500 g,而不是讨论中的300磅或10吨。风洞、火炬等依赖专业经验的航空航天测试仍由第三方完成,完整的制造数据尚未进入闭环。

3. 高熵合金说明了并行工程的价值

  • Krause 认为,Radical 并不是简单地排列组合一本成熟的配方手册。其高熵合金将5至7种元素按大致相等的原子比例组合,为极端温度——通常超过2,000°C,甚至达到3,000°C——高压、氧化、腐蚀和中子暴露环境寻找候选材料。

  • 机会之所以存在,是因为航空航天及其他行业至今仍高度依赖20世纪50年代至70年代开发的合金,有时再加上后续涂层。漫长的开发周期使 incumbents 理性地偏好渐进式优化,而不是押注陌生的材料家族。

  • Krause 借用了 SpaceX 材料负责人 Charles 的说法“并行工程”:在设计火箭助推器、涡轮机、导弹、太阳能电池或其他产品的同时设计材料。性能规格不再是从现有合格合金继承下来的约束,而是材料发现的输入。

4. 资格认证、供应链和单位经济决定什么能活下来

  • 主持人将下游材料资格认证与药物开发作了比较。航空航天和国防合金要按 FAA 或军方规范完成资格认证,可能需要多个铸锭、标准化测试,并在进入安全关键系统前耗时约10年。

  • 主持人问,能否像“Operation Warp Speed”那样,通过并行化测试来压缩资格认证。Krause 提到 DARPA 正在开展使用增材制造和逐层分析的工作,但没有声称整个流程已经解决;目标是通过新的机制取得同样结果。

  • 他对安全的区分刻意而鲜明:一部会弯曲的 iPhone 可以被拒收,或者更糟糕地被召回;787上的涡轮机材料则必须达到高得多的标准。“没人希望这个标准被取消。”目标是淘汰过时的机制,而不是降低安全要求。

  • 供应链冲击可能在项目中途改写目标。Krause 举例说,hafnium 的价格上涨了10-15倍,因为中国控制了供应链的大部分;C103按重量计含约10%的 hafnium,Radical 已着手在保留性能的同时去除这一元素。太空应用能够容忍更高成本来换取性能,而消费电子和医疗器械对价格敏感得多。

5. 自驱动实验室选择路线,而不只是提高速度

  • Krause 给出的清晰区分是:自动化实验室以高吞吐量执行人类定义的实验,就像免提驾驶仍需要司机做转向;自驱动实验室则运行完整的研究项目,就像乘客只需说出目的地的 Waymo。

  • Radical 将这一系统拆分为困难的样品操作与工具、实验室操作系统,以及互联自动化。软件负责追踪样品、控制仪器、读取传感器数据、执行质量检查,并可在浪费时间进行 XRD、SEM 和后续测试之前,终止不合格实验。

  • 实体操作的难度远超表面印象。以3,000-4,000°C温度熔炼形成的合金“纽扣”会粘在托盘上;人类会凭直觉拿凿子处理,而 Radical 必须定制机器人执行器,在不损坏样品或改变其微观结构的情况下将其取出。

  • 人类仍是重要的老师。冶金学家会为 SEM 图像做标注——“我在这张图的这些位置看到了枝晶形成”——让 AI 科学家吸收博士在阅读微观结构时几乎下意识使用的判断。

6. Radical 收窄平台野心,以换取垂直深度

  • Krause 改变想法的过程很关键:公司创立时计划建设覆盖7种材料体系的7个实验室。客户对专项测试、制造以及扩大到300磅规模的追问,暴露出这一计划的浅薄,因此 Radical 先在合金领域建立纵向深度,再考虑聚合物或陶瓷,甚至可能永远不扩张。

  • 半导体仍是相邻项目,因为新的后端互连材料可能降低集成损耗和能源成本。Krause 表示,目前一些系统推荐的材料可能带来约2x至5x的改善,未来潜力或许超过10x,但他没有透露具体材料,并承认:“我不知道能达到什么程度。”

  • 集成进 iPhone 或 NVIDIA GPU 仍“还相当遥远”,芯片产能稀缺以及必须从零搭建测试基础设施进一步增加了难度。Krause 对合金在3-5年内进入国防或航天系统更有信心,同时明确排除了载人飞行作为首个应用场景。

7. 主动学习按项目推进,人类进行选择性控制

  • Radical 的 AI 科学家设计一个研究项目,按自己设定的置信度选择一批候选材料,再让它们经过合成、表征和早期性能测试。机器学习模型分析部分输出,科学家标注其他输出,所有结果最终返回数据库,用于下一轮项目。

  • 更新按项目进行,而不是每个样品完成后立即进行。Krause 表示,实验室可能可以在不同体系上同时运行约7至10个项目,结果每天或每隔一天修正假设:“我们确实希望先出手几次,拿回足够数据来改变假设。”

  • 表征实现了全自动化,氧化测试和显微压痕也是如此。拉伸测试接近自动化,但合成仍由冶金学博士完成,因为铸造需要判断某个角落是否已经完全熔化;Krause 预计定制合成工具将在夏天实现自动化。

  • 候选材料生成已经交给 AI 科学家。人类研究人员偶尔会提交竞争性成分,实际上是在对它进行红队测试;这些方案有时会因强度不足而被拒绝,但研究人员也会从模型引入的意外元素中学习。

8. 吞吐量让 AI 探索人类直觉不愿涉足的区域

  • Radical 的文献地图将已发表的合金家族与 AI 科学家探索过的区域叠加起来。人类科学家坦率解释了此前的遗漏:他们预期某些元素会蒸发、无法铸造、形成糟糕的晶粒,或损害机械性能,但其中一些组合最终成功完成了合成。

  • 主持人的质疑值得保留:也许机器只是获得了更多试验次数和更高的“温度”,一支不受约束的人类团队也可能做出类似探索。Krause 承认,文献往往会把系统拉向已知的成功路径,而吞吐量是“一个重要数字”。

  • 他的回答既关乎行为,也关乎算法。一名博士研究人员每年可能完成约50次实验,每次制造和测试大约需要2周,因此每个选择都弥足珍贵;而一名每天制造8种、20种、最终可能达到100种样品的 AI 科学家,可以承受更多试探性的“向球门射门”。

  • 当前实验成本约为60-300美元,取决于是否使用 platinum、palladium、aluminum 或 titanium 等元素。吞吐量从每天8种难处理的难熔合金,到每天20种更容易处理的体系不等;粗略目标是在6月至7月达到每天100种,不论材料体系。

  • AI 科学家还以并行方式运行,而不是像人类研究人员那样串行工作。Krause 表示,它可以实时比较100,000篇论文和100,000张 SEM 图像,而人类无法保留并直接对比如此庞大的信息量。

9. 瓶颈是实验,而不是算力或搜索空间大小

  • Krause 将 Radical 与 DARPA-GE Aerospace 的 MACH 项目作比较。他称后者在 AI 和仿真筛选后,约12个月生产了500种合金;Radical 的目标是在5个工作日内完成500种,即使尚未进入制造规模,也已经是数量级的变化。

  • 搜索空间仍然大到无法靠暴力搜索解决:Krause 估计可能存在约10^40种合金,人类需要700万年才能全部合成。AI 对筛选仍然有价值,但反馈质量取决于实验,而行业过去一直没有系统地捕获或共享这些实验。

  • 主持人将数十到数百个合金样品,与能够扩展到数百万乃至数十亿次的生物实验作比较,质疑局部稀疏模式是否能胜过专家设计。Krause 的实证反驳是,Radical 在100次、200次或300次实验中已经看到有意义的结果;截至目前完成的1,200次实验中,有300种是新合金,每种合金约有50-150个数据点。

  • “材料行业不是受算力约束,而是受实验约束。”因此 Radical 的野心是建立一个“材料领域的蛋白质数据库”,但其中不只是晶体结构,还要包含加工过程、微观结构、性能、成本和应用背景。

10. 工业材料不存在一个单一的 AlphaFold 时刻

  • Krause 同意,在完整系统层面,“材料领域不存在 AlphaFold”。但局部的 AlphaFold 式突破是可能的:分割模型可以读取 SEM 图像,识别枝晶、裂纹或缺陷,并把裂纹扩展与机械性能联系起来。

  • Krause 对生物学的更广泛区分是,SELFIES 和 SMILES 可以用字符串表示分子元素和化学键,而合金的供应链、成本、微观结构、加工过程,以及增材制造还是铸造,无法用这种方式捕获。没有任何单一模型能一次性设计出最终进入 iPhone 或 Starship 的材料。

  • 这种能力并不能回答材料能否被雾化成粉末、进行增材制造、完成铸造、实现规模化,或成功集成。更糟糕的是,规模化可能暴露出此前无人知道需要测试的变量,使目标数据集从定义上就不完整。

  • Krause 对“发现”的标准刻意设得很高:假设、合成和表征都只是里程碑,而不是终点。“当你拿起手机,发现里面装着一种新材料时,我们才算发现了它。”

  • 一名拥有35年经验的 3M 顾问总结了制造问题:关键数据可能存在于某位操作员的经验中,他确切知道什么时候、向哪个方向拧多远。Krause 坦率地回答:“我们还没有解决这个问题。”因此 Radical 需要与成熟制造商合作,同时学习如何为这些隐性流程加装仪器并实现自动化。

11. 硬件摩擦正在变成基础设施护城河

  • 早期一个典型案例是,某些仪器的软件没有开放接口,团队不得不花2周完成软件冲刺,研究如何通过程序控制设备。Radical 已经为软件付费,还必须策略性地争取访问权限;Krause 表示,团队最终找到了控制所需功能的方法。

  • 建设自主无机科学系统后,团队被拆分为实验材料科学、计算材料科学、机械工程、机电一体化、全栈软件、应用机器学习、机器人、路径规划、感知和计算机视觉等方向。“问题不在于工具前面放一个机器人。”机械臂一旦装上,所有下游集成问题都会出现。

  • Krause 认为,现在时机成熟有3个原因:机器学习的原子间势可以加速计算筛选流程的一部分;机器人、夹具和执行机构更便宜、更好用;仪器厂商越来越多地维护支持自动化接口的软件团队。

  • 根本限制仍然是漫长的物理反馈回路。一个拥有1,000台 XRD 或 SEM 的设施可以通过并行化压缩周期;更具变革性的解决方案,是让厂商重新设计仪器,使其“面向智能体和机器人”,让研究人员操作整个科学系统,而不是逐台机器接受培训。

12. 国家级规模和开放模型共同强化以实验为先的护城河

  • Krause 认为,中国的制造创新中心是美国的优势,但美国“在制度上不应模仿”,在运营上却必须回应。按他的描述,无论公有还是私有,一个实体都可以控制相关环节,并支持一种新材料完成规模化,直接冲击 Radical 想要缩短的25年差距。

  • Radical 给出的生产力例子是:一名冶金学博士运行10个研究项目,而历史上大约10名科学家才会集中处理一个研究问题。主持人指出,中国同样可以做到这一点;Krause 的回答是持续投资,以及一种不同的公私合作模式。

  • 国家实验室提供高性能计算、研究人员、仪器,以及 Krause 所说的可能是全球最深厚的实验数据储备。他提到 Berkeley、Argonne、Ames、Livermore 和 Oak Ridge 的自驱动或半自主工作,以及 Genesis Mission 和数亿美元的投资。

  • Radical 的 AI 科学家本身就是多智能体系统:编排器提出并测试假设;文献智能体提取相关图表;付费的行业标准数据集和此前实验为研究项目提供基础;MATRIX 则是一个在“Quinn”上微调、用于读取实验室图像的 VLM。Krause 表示,据他判断,公开数据集带来了5%-16%的通用科学推理能力提升,数学是明确的例外。

  • 他主张劳动力走向专业化,而不是把所有人重新训练成同一种混合角色:“不要试图成为材料科学家,成为一个在材料科学领域工作的 MLE。”陌生的机器学习视角可以挑战延续数十年的实验室惯例,而领域科学家则提供模型所缺乏的直觉。

  • 开源也遵循同样的逻辑。Radical 发布了 MATRIX 及其基准测试,并将 TorchSim 拆分为非营利组织,以吸收社区反馈;但自身的实验数据仍然专有。更好的外部基础模型受到欢迎,因为“我们不卖模型”;Radical 认为,实验、自动化和累积的物理证据才会继续构成优势。

Joseph Kraus

This is the difference between AI for bio and AI for materials. If you look at bio, or maybe small molecules as a broader category, you look at SELFIES and SMILES strings, right? That has been a big way to have those materials, those molecules, in text. And then you can use that because you know the elements and the bonds, so you know most of the things you need to know.

What about everything I just told you about the alloy, supply chain, cost, microstructure, how you're processing, additive versus casting? How do you capture that in a string? You can't. This is what's so hard: there is no one model that can one-shot a new material that ends up in your iPhone or on Starship. That's just not the way materials work. And so there is this really tough challenge: how do you capture all this data and try to bring that back and really improve your AI engine to encompass more than just discovery?

swyx

We're in the room with Joseph Kraus, CEO of Radical AI. Joseph, you're in a market that's getting crowded really fast. You have Lila, you have Citrine, you have Periodic, all developing AI for materials science. What are you trying to do that's different? And how are you going to beat the heavily capitalized competition?

Guys, great to be here. Thank you so much for having me. I must start with this: I'm a big fan of the show. I get to commute in New York City every day, and you're one of the top things in my rotation. I always love learning, and I'm a materials scientist by training, so the aspects that I can learn from your show are awesome. I'm super excited to be here, especially in person. Thanks for making it work.

What makes us different is our deep belief in experimental data, right? I think now you're starting to see the industry pay more attention to this. You see self-driving labs, a talked-about concept everywhere from academia to people like Google DeepMind, all the way through to pretty much every competitor that you've named in the space building an SDL.

It was not always that way. When we started the company 2.5 years ago, people would have thought we were crazy. That's CAPEX-intensive. Are you really going to be able to pull the data? Models aren't really built for that data today, and we can get into why models struggle in materials science, particularly inorganic materials science. And so why are you going to do that?

I had a deep belief, and my co-founders had a deep belief, that in materials, the ground truth is the material itself. You have to be able to make it, you have to be able to test it and characterize it, and then you have to, at one point, be able to see if it can go into a real application if you're going to have it used in industry.

That was where our thesis really started: you're going to build this loop, this closed-loop system, what we call a self-driving lab, that can actually run those experiments, capture that data, and feed that information back to your AI scientist so that it can learn and actually predict materials that are relevant to industry. That is what our whole company is built around, and for the last 2.5 years, that's what we focused on building.

Alessio Fanelli

Why do you believe that versus, pejoratively, the “think big thoughts, come up with stuff, and then try later” approach?

Joseph Kraus

So much of what makes a material real is in the latter part of the discovery process, specifically at the characterization and synthesis phases. What did we make? Does it have some cool properties in the lab? But also, what happens after that?

We work in a field called structural metals, or alloys. So much of what dictates the performance of those alloys is actually in processing. How do you manufacture it? What techniques are you using in post-processing and manufacturing that push performance or change performance?

You can generate a new composition, and that's very important to do, and we do that with AI. But it's everything that comes after that that actually impacts whether you have a new discovery, whether that new discovery is relevant to the application space you're going for, and then whether you can actually make it. Can you scale it? Can it actually go into that application space and be used by an end customer?

Those 2 and 3 don't get solved with AI today, right? A model can't figure out your way through the qualification pipeline for a new alloy for a jet turbine. You have to do experiments to do that. And so that ground truth is really important for us to bring back and understand: We know what we want to make, but are we actually making those things? Do they actually have the properties we care about? Can we actually push them to industry?

Those latter questions are the hardest questions to answer in materials. One thing you always hear about, which is true in materials, is long timelines. We've heard everything from 15 to 30 years—pick your favorite number, whatever you're feeling on this day of the week. But the point is, that's true.

The reason why is that materials is so fragmented today. Academia handles discovery and some of the lab-scale testing. You have small companies that will look at it, typically supported by the Department of War, the Department of Energy, or other government programs like NSF. And then you have late-stage companies, where bigger companies really optimize their current systems today.

They're not focused on room-temperature superconductivity, high-entropy alloys, or new ceramics. They're focused on how to take their current material system and make it 5% or 10% better, and capture the margin from that. There is so much fragmentation across this whole industry that the data never gets shared, and the connection from discovery to manufacturing is typically lost in that process.

That's the connection that we want to bring back to materials science. That is what we think the true opportunity is for AI and autonomy in materials: linking those two together in a fully closed-loop system.

swyx

I want to dig into that. I have a rule of thumb that I often follow when thinking about things: anytime you change orders of magnitude in a scaling system, your problems completely change. The orders of magnitude of the problems that you're talking about are drastically different, right? In discovery, you have an N of 1, or an N of some small number. On the commercialization side, you have an N of millions or whatever, right? Why do you think that you're capable of solving all those intermediate problems?

This is a good question, and I think I'll use a real, practical example to explain how hard it really is. In our field, with these alloys, one of the important things that determines properties is the microstructure. How does the microstructure form in this alloy so that you can see things like strength, ductility, emissivity, or whatever your other favorite mechanical property is?

At the generation level—candidate generation, hypothesis generation—you can predict a new composition. AI is actually quite good at that. All of our hypotheses are generated by our AI scientist today. You'll take that composition and synthesize it in the lab.

There, step 1, something changes. It might not be homogenized. There might be dendritic formation on the surface. You might see different phases, or it might be single-phase. Those dictate what the properties look like.

Once you move past that, you actually go to manufacturability: annealing or thermal processing, and looking at how you manufacture it. Is it additively manufactured with powders, or are you casting with actual raw metal? Both have wildly different outcomes in the performance that we like to see.

To answer your question, the first step in solving this is capturing that data. We do that at the discovery and testing phases today. We don't do the manufacturing part, to be clear, but we do that at discovery. We do synthesis and characterization. We have a bunch of characterization tools in our lab: SEM, EDS, XRD, XRF, and TGA.

swyx

Wait, real quick. That was a lot of acronyms. Would you like to explain them now? We can also pause this, and we wanted to talk about them later if it makes sense. Or we can just talk about it.

We can come back to it, and I'm happy to dive into everything that we're doing in this.

swyx

For now, those are just a lot of acronyms. You don't need to know what they mean.

They're just a lot of tools that tell you different things about a material in a lab, and we can talk about what they do. Then we do testing of properties at the lab scale. We'll look at oxidation performance in our lab today, which is really important to see how these alloys perform in oxidative or corrosive environments.

We'll look at mechanical properties, something called a tensile test, which gives you these stress-strain curves of a material. Then we'll look at microindentation. Here, you can pull what's called the Vickers hardness from a material, as well as a proxy for ductility. It's not an exact measurement of ductility; we kind of pull out whether the material is ductile from that.

So that's everything happening on the discovery side and moving into the testing side. We have not yet crossed into, “Okay, now when we go to manufacturability,” but we hope to. Back to your original question, if you can capture the data at the manufacturing side as well, now you have the whole suite of what we call the lifespan of the material.

swyx

I see the hypothesis. I see the synthesis. I see the characterization. What do we make? I see the early properties showing good results, and then I see the manufacturing and what came out of the back end of that. Does it actually make it to the end system?

Now I've seen this lifespan of the material. Now I can use that to go pick more materials targeted at the right applications. That is the North Star of what the company wants to go out and do.

Where do we stand with that today, then?

Shawn Wang

I'd say we're really good at the first part: the discovery and that lab-scale testing I've mentioned. There are some testing mechanisms that we use externally, like with third parties, where there is deep expertise required in the industry itself. Aerospace is a perfect example. If they do wind-tunnel tests or torch testing, we don't have those capabilities at Radical today, and so far we haven't needed to own those. We want to use third parties.

There's a heavy tail.

The Self-Driving Lab

Exactly right. Really heavy. Then you look at the cost of a wind tunnel and you think, “Yeah, I'm never going to get an update from that.”

Lastly, we haven't touched manufacturing. We have spoken with people who do manufacturing, and we do know some of the things they care about. Processability is a really good one. Can this material be formed? If you're going to cast it, can I actually move it into the shape that I need it to be in? We do look at that, but we look at it at the small scale, not at 10 tons.

We look at grams—200 g, 500 g of material—so not at that larger scale yet. But that first section is done; that's running today. We've probably made 1,200 alloys in the last 5 or 6 months. 300 of those alloys are new, novel, never before seen in the literature. I'd say probably 10 of those alloys have performance that has got us very excited about where they're going to be in the industry. That's a rough scale of where we are today from a company perspective.

Shawn Wang

That brings up a follow-on question about how much you're just optimizing within a well-understood space, picking permutations that are new and novel, versus trying experiments that are really pushing the frontier of science.

The Self-Driving Lab

Yeah, the latter. We really are making new materials that push the frontier. A good example is that we work in a field called high-entropy alloys. These alloys are really exotic because they have 5 to 7 elements in the system, all roughly equiatomic, give or take, and they have really exotic properties in extreme environments.

Think super-high temperatures, usually north of 2,000°C, even 3,000°C. They have very high pressures—think space, coming back from space. Then they have environments that can be corrosive, like a nuclear reactor where you're feeling neutron bombardment, or oxidative, like if you're flying in a defense application or in a jet turbine.

These alloys are really exciting, and for the past 50 years, the same alloys have been used in all these industries. The reason why is the long discovery timelines we talked about. This is a perfect example where we're working in an industry; we're not really creating a new industry. Of course, turbines—there's a big industry, it's a great one. It's having a second tailwind right now.

What we're trying to drive there is new performance that does not yet exist for the materials they have today. We stole this term from a gentleman named Charles, who's the VP of materials at SpaceX, which is called concurrent engineering. It's this idea that I can actually design my materials as I'm designing my product.

As I make a new rocket booster, a jet turbine, a missile, or a solar cell, I'm actually inventing the new materials that meet the property specs for it. I'm going back and forth as I engineer them to get to the application. We don't do that today. The alloys that are in the plane I flew here on are from the 1950s, 1960s, or 1970s. They might be coated with some CMZs from the late 1990s.

There's a real opportunity here. We're tackling industries that exist and have huge markets, but we're bringing novel materials that historically they never would have had the ability to look at in enough throughput and at enough scale.

Shawn Wang

Going back to the bottlenecks you were talking about before: you said that it takes 15 to 30 years to get material through, and part of this is because there's a disconnect between research and productization. How do you validate that this is going to be the thing that will work, and that you're not going to get killed by another unexpected bottleneck in the whole process?

I'm coming from the world of drug discovery, where even if you think you have good early-stage data, there are all these things that will get you later on in different phases of clinical trials. I wonder, are there things like this in materials? Where do you think these are going to get hit? What will stop this super-cool alloy you just developed from actually making it to market?

The Self-Driving Lab

Really good question. There absolutely are those in materials. If we talk about the alloy example I just gave, 1 of those areas is called qualification. Qualification is this process that, if it's going in manned flight, is run by the FAA. There's also a MIL-SPEC one for the US military as well.

Essentially, your alloy has to qualify to be used in aerospace or defense applications, and that process is very slow. It's typically a 10-year process today. You have to make a number of different ingots of material and run these standardized tests on them to prove your material is usable in those systems.

There are a bunch of things that have a gotcha later. I think actually trying to capture that data and understand where and why those things are happening can impact your discovery loop.

Shawn Wang

It takes 10 years because, like clinical trials, you have a series of incremental phases. Can you Operation Warp Speed this, where you do these all in parallel, or is it really like you need to do these things sequentially?

The Self-Driving Lab

There are people working on changing that, rather than doing it sequentially. There is really good work right now—for example, at DARPA—where they're looking at new ways to do qualification. They can use additive manufacturing to go layer by layer and actually look at whether we can do qualification of a material that way. I would say that's a new technique that's trying to completely redo the way we do that process today.

Shawn Wang

So it's not regulatory.

The Self-Driving Lab

This is the challenging part. Some of it is regulatory in nature, in that it's a government body like the FAA or MIL-SPEC that runs it, and you have to see almost the human side of this as well.

It is different to develop a new alloy for an iPhone. If it bends in your testing, you can get rid of that or move on, or even, in a worst-case scenario, recall iPhones. That's obviously a terrible scenario for the company, but you can. But if it's going in a jet turbine and you're going to fly it on a 787, there's a serious bar that has to be met there, and for good reason. No one would want that bar to be removed.

I'm flying back to New York tonight. I certainly don't want that bar to be removed. To be super clear, make sure that's on the record.

I think the way we go about that process is very dated. What people are trying to attack is whether we have to do it that way, or whether there's another way to get the same result via a different mechanism. That's where AI is interesting, but autonomy has been really interesting as well.

As you make the process of manufacturing more automated, there are more sensors, therefore more data, therefore more things you can capture and analyze, and therefore a bigger loop that you can build around that. That has not been deeply extrapolated in the materials-manufacturing sense today.

Shawn Wang

So maybe 1 difference between drug discovery and materials is that, in drug discovery, you have different phases, each of which is designed to basically not kill people. Whereas in materials, there is, in some sense, no reason other than budget that you can't just do the—

The Self-Driving Lab

Solve it in 1 go.

Shawn Wang

And successive levels of qualification do not depend on the prior ones, aside from just budgetary constraints.

The Self-Driving Lab

As well as some of the other metrics you have to pay attention to, this is what makes materials so hard, actually. A perfect example is supply chain.

Probably 5 years ago, maybe a little bit longer, maybe 10 years ago, there weren't the constraints we're feeling today from the metals industry or the minerals industry. Hafnium is up 10–15× in price because China owns a majority of the supply chain. Things like refractories—

Shawn Wang

Tantalum and niobium. What is hafnium?

The Self-Driving Lab

Hafnium is an element in the periodic table that's used in things like C103. It's about 10% by weight of hafnium in C103, which is a very common aerospace and space alloy that's used today.

Now we're starting to see requests in conversations around, “Can you remove that element from that material?” There, you're actually trying to maintain the same performance specs, or you're trying to completely remove the hafnium from that equation.

We've worked on that problem specifically, and we have successfully done that. This is where you get back to supply chain being a concern, cost and margin being a concern, and who is paying for that and feeling it. I can tell you the space industry has much more tolerance for high cost.

Performance is everything. When I'm designing a new heat shield or a new cone that goes on the rocket engine, performance is the number-one thing I care about. Cost is not the first thing I care about. It's not irrelevant—I don't want to pay $100 million for a nose cone—but it's still not the top thing I'm thinking about.

You think about something like consumer electronics or maybe even a medical-device application where alloys go into products. Well, now cost is definitely much more sensitive. There are probably alloys we could put inside smartphones today, but they would just make them unbelievably expensive and probably not tolerant of some of the other things we have to put in there.

There are just so many things about a material that make it so much harder. This is one of the reasons why we deeply believe in self-driving labs. Just to come back to this point for a second, this is the difference between AI for biology and AI for materials, in my opinion, from the materials lens.

If you look at biology, or maybe small molecules as a broader category—you probably could include some organic materials in that—you look at SELFIES and SMILES strings, which has been a big way to represent those molecules in text. You can use that because you know the elements and then you know the bonds, so you know most of the things you need to know.

But what about everything I just told you about the alloy? Supply chain, cost, microstructure, how you're processing, additive versus casting—how do you capture that in a string? You can't. This is what's so hard: there is no one model that can one-shot a new material that ends up in your iPhone or that ends up on Starship. That's just not the way materials work.

There is this really tough challenge of how you capture all this data and try to bring that back, really improving your AI engine to encompass more than just discovery, and certainly more than just composition.

swyx

You mentioned this sort of loop, right? You're really doing 2 things with automation. One is you're collecting data; one is you're running experiments and building stuff. How iterative is that, and what does it look like at the different steps where humans could be in the loop?

Yeah, and humans are in our loop today in a very important manner: training and teaching what I like to call the scientist about what they know. We call this scientific intuition at the company, and what that literally looks like is that we have a scanning electron microscopy image—that's an image that takes a picture of the material—and our scientist will go in, analyze that image, and make comments in our system.

“Hey, I see dendritic formation on this image in these locations.” The AI scientist goes and looks at those comments. That's one amazing example of a human in the loop, where we are trying to download the brain of a PhD in metallurgy. When you look at this image, what do you see as a PhD scientist? We need to be able to replicate that as an AI scientist.

That's one way they're in the loop. The second thing—and we can go deep into this if you guys want to—the lab is not easy to automate.

swyx

Yeah, it is.

We're getting there. It is super hard to automate, both in ways that are hard engineering challenges and in ways that are annoying. The tool vendors just don't have SDKs or APIs to work with. There are engineering things that we can talk about, but even this idea of the tool provider letting you have access to the data via their software layer was not understood 2 years ago.

I can tell you that a few very big tool vendors were not too excited about self-driving labs 2 years ago. They were not jumping to give us, even with payment, access to the software and pull the data. That tone has now changed.

swyx

Are they trying to own it? Is that why?

From my understanding—and I'm not a tool provider, so they might give you a different answer—a lot of what they sell is the ability to analyze the data coming out of their tool. One of the things that makes those tools different is how they actually use the software to generate your spectra, and they like to sell on that. That's a really big thing for them.

If they give you access to the raw data from that and you no longer need their software, why would you buy their tool? We tell them, “No, no, no, you're way wrong,” and we're getting them there. It's a work in progress.

I do think there's been a lot of momentum in AI for science. Number 1, it's having an incredible moment, which is going to be so good for the world. Number 2, you're seeing, I'd say, academic and national buy-in, as well as private buy-in.

You see the Genesis Mission. You see the Department of Energy and the national labs moving this way. You see people like Google DeepMind, Microsoft, and other places like Meta either building their own lab or running experiments at someone else's lab to get that data back.

Then you have the private companies that are forming self-driving labs and looking at the automation of scientific equipment. That has really started to push this wave to, “Oh, now we don't really debate that AI labs or self-driving labs are a part of the future anymore. It's which part are we going to play in that?”

That's a better, more helpful conversation to have now because we can get access faster.

swyx

Self-driving labs in biology have been notoriously difficult. You can automate certain parts of them, but you inevitably have people who are just walking around moving trays from one section to another, and it doesn't actually end up speeding things up. Oftentimes, it can even slow things down.

Full end-to-end automation for non-research activities—or, you know, for manufacturing—we're really good at. But when the process changes, how do you deal with that? Have you figured out some way of automating the type of problems you're solving?

The self-driving lab—first, I think it's important to talk about what a self-driving lab is, because this impacts your answer. There's a difference between an automated lab and a self-driving lab, right? An automated lab does experiments for you automatically, without humans, and at high throughput. That can be very effective.

A self-driving lab runs research campaigns for you, and there's a big difference. The way I like to describe it is, one of them is like hands-free driving. I don't have to touch the steering wheel; it'll keep me in the lane and keep my speed set. But when a left-hand turn is coming up, I have to pay attention, put on my turn signal, turn the car, and know to make a left.

Now compare that to a Waymo, which I love bringing up because every time I'm here, I'm going to go out of my way to just drive around the block sometimes. It's living in the future. You don't need to make a left-hand turn. Actually, you don't even need to know to make a left. You don't care what route it takes to get you there.

You get in the car, you can close your eyes if you want, scroll X, work on a research paper, and then end up at your destination without knowing how you got there. That is the difference between an automated lab, which you are controlling and just using automation to do throughput, and a self-driving lab, where it's actually doing this entire process for you.

In the self-driving lab, there are things that a human scientist does that are actually very hard. Sample manipulation is a perfect example. When we synthesize these alloys, we get these little pucks that come out. They're called buttons in the industry, and because you're blasting them at 3,000 or 4,000 degrees, they get stuck to the tray.

How do you get them out? You have to be careful because you don't want to mess with the microstructure or chip off part of it. They're strong enough that you're not really going to do that, but we had to design custom actuators that go on our robotic arms to be able to manipulate them.

That does not really have anything to do with the discovery of this new high-entropy alloy. That's just required if you want to run autonomous alloy science. One answer is that there are these challenges that humans either don't face or, if we do, they're very intuitive.

The button's stuck, so I take a little chisel, smack it, flip it over, and move on. I don't even think twice about doing that. Not so much for a robot. The second thing we talked a little bit about is the software.

It's not just about controlling the tools; it's about running the lab. How do I track my samples? How do I know what sample should go in a tool or should not go in a tool? Is there a quality check where, if I look at a sample after it comes out of synthesis, I actually want to kill that experiment?

I don't need to waste time going through XRD and SEM and the other tools in the lab. I want to just stop that sample, save its state, and throw it away. How does it know how to do that?

This is where you start to bring in all these different factors from vision, different sensors in the lab, and sensors on the tooling themselves to build what we call an operating system that runs the self-driving lab.

And then the third part is automation. Automation includes what I like to call the connection of the lab. I have one tool that's automated, and I have another tool that's automated. You have this operating system that's running them individually. How do you connect them? The same way a human scientist would come in, look at the results, take the sample out, and go to the next one, our robots do that today.

Those are the three parts that we really see making up this self-driving lab, and I'd say each of them has its own difficulties. We can walk through them. Some are, like I mentioned, the tool provider. Some have no actuators. Some are just really hard to load and unload, like XRD. You have to put it in a hard sample mount, and it's weird and awkward geometry, and you need a custom gripper. That's just what it is.

All of that kind of forms what a self-driving lab becomes. We're very good at that for alloys today because the tools in our lab are built for alloys. Some tools are shared, like XRD, XRF, and SEM. Even tensile testing, although mostly used in the structural metals space, can be used elsewhere.

I would say our oxidation chamber is very suited for the specific customer application we're going after. I would say our synthesis mechanism is directly for alloys. We custom-built a tool with a third party to do alloy synthesis at high throughput. That's built to do alloys. It's not built to do ceramics, polymers, or any other material system today.

swyx

Is the goal to expand to polymers and ceramics and everything, or is it that we're going to do alloys and basically get all the way to end-to-end manufacturing on alloys and then expand?

Yeah. Or never expand.

It's both, but on the right timeline. The first one is vertical integration, and this is very important. When we started the company, and I'm not afraid to admit it, we were like, "We're just going to do 7 different labs, 7 different material systems, across the board. We're going to capture all this data. It's going to be amazing."

Then we started talking to customers about it.

swyx

Yeah, exactly.

Exactly. And that'll come back to doing that over time, but we started talking to customers, and they're like, "How are you going to do this? Have you thought about this test? What about when you go to scale? What about when you need to go to 300 pounds?" And we were like, "Oh, no, we hadn't thought about that per se. We were more worried about going to polymers and ceramics and everything else."

Why is that important? Because we are a materials company. You see the company talk about this a lot. We really believe in inventing new materials that change the future of the world. We think that is the opportunity with AI and autonomy.

There are so many industries that we all care about that are blocked because of a lack of novel material advancement: automotive, aerospace, manufacturing, defense, climate, energy, semiconductors, and electronics.

Shawn Wang

What's your favorite example of that? What's a problem that they unlocked with an amazing material?

The Self-Driving Lab

Aerospace and semiconductors immediately.

Shawn Wang

Well, what specifically?

The Self-Driving Lab

In back-end-of-line integration for semiconductors, there are particular materials that we've been using for a long time called interconnects. The entire industry is doing R&D here. They are a cause of not having great efficiency and very high energy bills. At the back end of that, a new material could potentially completely remove that problem. That's probably an example.

Shawn Wang

Is there a theoretical efficiency that you could achieve, and where are we now compared to that?

The Self-Driving Lab

We have ideas for materials that would be able to solve that problem, but we'll release more on that in the coming months. We have a specific program we're working on that's directly around that problem. It's a really exciting one, and there are estimates from the industry on moving past the current material system and what new materials could bring, but there are other challenges.

Shawn Wang

Are we talking about 2× more efficient, 5×, 10×, or 1.1×, which is often quite huge?

The Self-Driving Lab

Yeah, of course. I think you could, in the near future, with some of the systems that have been recommended today, see 2× to 5× generally, especially when you think about integration. In these materials, you need barrier layers, and there are all these interface things that you have to be able to understand.

I think beyond that, you could start to see it push over 10×. I don't know to what level, but there are cool, exciting materials that we'll share more about in the future.

Shawn Wang

What's the validation time scale for this?

The Self-Driving Lab

In which way? The lab starting to work on these materials, or a material being in a new chip in the iPhone?

Shawn Wang

You have a material you just created that you think is going to revolutionize the world. How long do you think it's going to take for that thing to get into an iPhone or an NVIDIA GPU?

The Self-Driving Lab

That's still long. That's still pretty long. We're new and early in semiconductors, and actually, what's long about that is that when you start from ground zero, you have to build everything from scratch.

Even the way that we do material testing for that industry today—I would say they're one of the industries that's actually far ahead of everyone else because they invest so much in R&D, and materials really make or break some of their performance—even there, that integration timeline is very slow.

And number 2, no one can get enough chips, so everything is delayed in that industry, which is important. I even saw today, or this week, TSMC telling ASML they're going to hold off on some of those new tools until they get through their 2029 production run, or it was some story like that.

Shawn Wang

Go, go, go.

The Self-Driving Lab

That's what I believe, yeah. I'll find it afterward and send it to you guys, but I was like, "Man, this industry is really getting pushed to its limit." I would say something on the alloy side, like aerospace, offers a good opportunity on a 3-to-5-year timeline.

Shawn Wang

That's pretty short.

The Self-Driving Lab

Yeah, correct. I think it'll be an application, not manned flight. I don't think it'll be a jet turbine because of the constraints there with humans. I think defense and space systems, though, are definitely doable in that timeline.

There have been examples in the past of people who have done that in that timeline, and we feel very confident in our ability to try to execute on that.

Shawn Wang

I want to get back to the validation question because I feel like this is the crux of automation, right? Every AI engineer who uses cloud code or something has the experience of one-shotting something, and then when you look at it, you're like, "What is this garbage?"

If you're talking about what is essentially active learning, where you are hypothesizing, manufacturing, testing, and then forming hypotheses on that, and doing that over and over again, your mistakes will obviously compound as you do that. So, how are you thinking about this?

The Self-Driving Lab

We call these the negative results, these mistakes, and we do have a version of this loop built already today. It doesn't include manufacturing data, like I mentioned—we don't have manufacturability in the lab today—but it does include synthesis, characterization, and those early property tests that I described to you guys.

What this system does is have this AI scientist that's really good at designing the campaigns I talked to you about. It can come up with these campaigns, determine the number of materials that it's confident it wants to make and test, and launch that campaign.

It'll send it to the lab, it'll start running autonomously, and then it will go through the whole characterization suite, and we'll get all the data. That data is all pulled out autonomously. Some of it is analyzed with machine learning models, like computer vision. Some of it, again, has a human in the loop analyzing it as well, and it'll get put in our database. That AI scientist will look back to that data when it designs the next campaign as a follow-up.

So, it's this active learning loop on a campaign-by-campaign basis. It's not every experiment. Honestly, we don't need it to be that fast. We actually want to take a few shots and get enough data back to change our hypothesis, but it is very rapid.

I would say we probably could run 7 to 10 different campaigns right now inside the lab across different systems, and we're updating those daily, or at least every other day, with the results we're seeing from the lab when they come out.

Shawn Wang

And what does the human do in that part? I find that when I'm doing transcriptomic analysis or whatever, Claude does some of my work, and then I end up course-correcting. Is that the gist of what's going on in your lab as well?

The Self-Driving Lab

In some parts, yes, though some parts are more complex, to be honest. Synthesis is a really good example. We still have PhDs in metallurgy running our synthesis machine. That tool is not yet fully automated. It should be automated by the summer. That's the custom tool I was telling you guys about that we're building with that vendor.

It's not just opening and closing that's automated. The synthesis machine itself requires automation. If you've ever seen how you cast these alloys, you take a plasma torch and blast them, melt these raw precursors down into a liquid, and then cast it, and it solidifies into the shape that you cast it into.

So, there’s a lot of intuition in that. A scientist stares at it and looks at it: “That’s not melted yet. Let me hit that corner there.” We have models built that can actually start to learn how to do that at the same performance as a human scientist can.

So, we’re getting up there. That’s not fully automated yet. Characterization is fully automated. We just have scientists annotating images or results afterward to train the AI scientist on. All of our characterization tools can be loaded, unloaded, and controlled with our back-end operating system to do characterization.

Property testing: 2 of the 3 property tests are fully automated. Microindentation and oxidation are automated. Tensile is almost automated; it should be automated in the future. So, that’s where humans are now.

After the process for generation, all of our materials are generated by our AI scientist today. Occasionally, a human scientist will try to compete, and I love telling this story. They hate when I tell this story. They’ll throw in a composition, and then the AI scientist says, “Get that out of here. That’s not strong enough.”

Or we’ll see, “How do we think about something new?”

Shawn Wang

Actually, it’s like you’re red-teaming.

The Self-Driving Lab

Yes, that’s a good way to describe it. I don’t know if they would call it that. I think they would call it losing their job, but they’re not, actually. They’re super important to the process.

I think what’s cool is that sometimes the scientist will recognize a new learning: “Oh, interesting that you threw that element in there. That’s cool.”

The other part of that is going places where human scientists won’t go. We have this beautiful chart that shows, in publications—in the literature—what we can access, where all the places scientists have gone. Then we have a second overlay on that chart showing where our AI scientist has gone.

Shawn Wang

Mhm.

The Self-Driving Lab

It’s moved into elemental families or alloy families that no one has ever published on before. The question is, why? Why did it think to do that?

We ask our scientists, “Why did you never go there? Why did you never use that element or that element?” Their answer is, “I just didn’t think it would work with the other elements that are in that mixture. I didn’t think it would cast. I thought it would evaporate when we tried to make it, and it didn’t—we were able to synthesize it.”

“I didn’t think it would work in the microstructure, or would cause grains to be not what I was looking for, and would not get the mechanical properties I thought. So, I just never considered it.” But it actually works in that formation.

There’s this really interesting feedback where now the scientist is getting good at exploring places that, I’d say, humans have a natural bias against, even though it might be an unknowing bias. That’s a huge power of an AI scientist.

Shawn Wang

But is part of this just that you have higher throughput and that you’re letting the AI scientist do its own thing? I wonder what would happen if you took those same scientists and said, “All right, no constraints. Just go crazy. You can do whatever you want to.”

swyx

If you have an AI scientist that doesn’t have preconceived notions, I’m honestly kind of surprised that it’s not just reiterating what is known. But I wonder how much of it is that you’ve just turned up the temperature in your sampler.

The Self-Driving Lab

Yeah, that’s an important metric. Processing is really important. To be clear, there are times when it goes to what it knows, especially when it pulls in literature. It’s like, “Oh, this is where it is.” Literature is a great teacher. If things work, it’s actually a good place to ground on why they work and try to understand why they work. That’s a big problem for the materials field that we can talk about.

But it also has this good ability, because it is high-throughput, not to be afraid to test. When I was in my PhD, I probably did 50 experiments a year—rough estimate, something around that. Every experiment is important, right? Not to mention the mental load: 2 weeks at a time to fabricate this thing, synthesize it, and then go test it.

The scientist doesn’t think like that. This AI scientist is like, “I’m making 8 of them today, 20 of them today, and once that tool is done, I’m making 100 per day. I don’t really care about taking a shot on goal and learning from that shot.” So, it’s a mindset shift.

swyx

How much does 1 experiment cost?

It depends on what elements you use. Some elements, like platinum and palladium, are much more expensive than aluminum and titanium. It’s anywhere from $60 up to $300.

swyx

Okay. It’s all element-dependent, though, usually. What’s your throughput?

Today, it can be anywhere from 8 to 20. That depends on the elements as well. Refractories, particularly if we’re doing refractories, are much harder to cast, so we go down to that 8 number.

If you’re doing things like titanium and aluminum—your standard alloy, Ti-64—those are much easier to cast. They melt immediately or quickly, so we go to a higher throughput of 20. We should be at 100 per day regardless of the system by around the June or July timeframe, rough estimate, give or take.

swyx

This is across the entire lab, not per workflow?

Yes, that’s correct. That’s across the entire lab.

swyx

Okay. Eight to 20, and you could conceivably have humans inspect many or most of these, or is that—

I don’t know. You could.

One thing I wanted to touch on—you just reminded me with that answer—is that AI doesn’t operate in the same dimension that humans do. Let me explain what I mean. When I was a scientist, I went through a very serial-based process. I read a bunch of papers, made a new hypothesis, and might run some computational workflows—DFT, MD, or ML. Then I’d go synthesize it in a lab and move to characterization. I’d study my characterization for 2 weeks, get a new idea from an image or something I saw, circle back, and do that whole process again.

That’s what a human does, and it’s very serial. If I take 100 SEM images, I don’t memorize all 100. I’d love to think I could—my advisor would have loved it if I could—but I couldn’t. So, I pull one thing or a couple of things out of that that I want to learn.

Now switch over to the AI scientist in that same process. Now it’s parallel. I can read 100,000 publications and directly compare them to 100,000 SEM images at the same time, in real time. I can study, learn, and memorize all of the things I’m seeing in those 100,000 SEM images and draw direct conclusions back to my papers, back to my hypotheses, or back to my mechanical property testing, where I want to see what actually comes out.

I can’t do that as a human scientist. This parallel nature allows it to operate in a way that human scientists simply don’t have the ability to do.

swyx

Okay, but you’re still talking about on the order of 10s or 100s of materials that you’re producing per week or month. The overall scale is not that large compared to even a lot of biological methods. You have ways of scaling up to 1 million if you’re doing, let’s say, next-generation sequencing-based assays. You can do billions, whatever.

This is much more reminiscent of ligand-based modeling, where you’re really looking at a small number of examples and trying to pull out local patterns. For small molecules, you can have real predictive power. These are useful techniques. But the almost universal rule from experience in cheminformatics is that, oftentimes, by the time those become useful, the actual scientist can just go out and do it. They could have designed it or found the molecule they were looking for without using the AI model by the time the AI model gets there.

So, this is a very specific kind of regime, but it’s one that has been well established. I’m wondering, since the timescales in this sort of data seem very reminiscent of that, why is this actually that much more effective?

Yeah, 2 things there. Number 1, throughput is an important number. To our knowledge, based on what is publicly released—if there’s someone who has done it behind closed doors, I don’t know about them. Please, I would love to talk to them.

The largest alloys program was the MACH program. It was run by DARPA and GE Aerospace. They did a bunch of AI and simulations on the front end, and then they synthesized 500 new alloys in about 12 months. That’s kind of the benchmark, I would say, for how many alloys someone can do in a year.

Again, we’re trying to do 500 in 5 business days. We’ve done 1,200 in 3 months. So, that’s an order-of-magnitude step up that we’re moving to.

The second piece of this is that what’s really challenging in the alloy space specifically—and I think it’s probably specific to alloys, though I do think it will carry over to some other industries—is that there are so many variables that go into determining your end product and then your end properties from that product. That makes it harder to do discovery because there are endless potential combinations.

There are 10^40 different potential alloys that you could go out and synthesize. How do you do that? Even if you could do high throughput, to your point, it's still not that much throughput, right? We think it would take humans 7 million years to make all of them. What do you do even if you're only doing 30,000 a year?

The screening mechanisms here are very helpful, as we all know. That's where AI is great. But I think the other point is that the data is missing from the industry. We don't have experimental data.

We do see results from 100, 200, or 300 experiments quite aggressively. We see 300 new alloys that we've never seen before from the experimental results of 1,200 alloys that we've run to date. You probably think each alloy has anywhere from 50 to 150 different data points, depending on how many images you take, how many spectra you run, and so on. That's a fairly small data set.

It's funny when I talk to the ML side. They're like, "How are you going to get millions and millions of data points?" And I'm like, "I just don't think you need to. We have not seen that you need to make new discoveries yet today. We have a bunch of new discoveries, many that are going through patent protection and that we're talking to potential customers about. We just haven't needed millions of data points."

This brings up the arguments I was getting about compute. I got an argument at GTC about this: we're not compute-constrained in the materials industry.

swyx

Yeah. You're making stuff.

Yes. We're experiment-constrained.

swyx

I mean, this is even what I do, which is computational. It's oftentimes dominated by data movement, right? It's not like I can—I see these people, and I'm jealous that they have 14 Claude sessions going at night. They have all these different experiments going, and I couldn't do that just because I can't move the data around fast enough.

Yes. That's almost a good comparison to our world. It's not a model problem. It's not a language problem. We don't have the same problems there. It's an experiment problem. It's really: how can you run enough experiments to start to change the output of an AI scientist and capture the data you need to discover something new?

For us, that's really what it's about. That is our bottleneck. That is the throughput. That's why we're so bullish on self-driving labs. That's why, when I start the conversation, "What do you guys work on?" it's all about the self-driving lab, the autonomy, and the experimental data.

We're trying to build the Protein Data Bank for materials, and it's much more complicated than just crystallography structures or whatever else was in there. There are all these different properties you talked about today that have to be inside that data set to make it relevant. It's hard to do.

swyx

That reminds me of Heather Kulik’s episode, where she said that there is no AlphaFold for materials. First of all, do you agree?

I do. I think you can add AlphaFold moments for specific areas of materials, like microscopy, for example. Reading, using a segmentation model on SEM images—our team does that today. That's a cool AlphaFold moment, whether you call it that or not. I don't know, but there are real-world models that make a huge impact on being able to do that.

What I don't think you can do is go from, "I have this new hypothesis," to, "Oh my gosh, I have a new material. It's scaled, it's done, it's in products in your iPhone." You can't do that today. I would agree with that statement.

swyx

Okay, but even then, AlphaFold solved a scientific problem, which is: how do you take a protein sequence and figure out what the 3-dimensional structure of that protein is? I want to put in so many caveats so that none of my structural biologists flame me.

Yeah, yeah.

swyx

But anyway—

Luckily, I'm not a structural biologist, or everyone watching, so I get a free pass.

swyx

But the thing that seems useful here is that you put in a chemical formula and some sort of processing, and what you get is—essentially, what you call—the microstructure, which is something where you can get part of the information from X-ray diffraction, but not all of it. That problem sounds much harder in a lot of ways than the biological problem. Can you maybe explain a bit more about that?

Shawn Wang

Yeah, and it's even harder than what you just described. There are probably things you don't know to test for yet, or that you might see when you go to scale that you did not know you should have predicted or been paying attention to. That's really hard to build a data set around.

What I like to tell people when they ask about this is, "Can you build the database for materials?" Well, if you want to do SEM images—scanning electron microscopy images—and use segmentation, yes, you can build a really large data set of SEM images that are very good at finding dendrites, cracks, or defects in a material. The model will be very, very good at using that to predict and relate that to a mechanical property, because you're looking for what crack propagation does to strength.

Okay, tracked.

The Self-Driving Lab

But that's not the same thing as, "Can you make it? Can you atomize it? Can you make it with the powder? Can you use additive manufacturing? Can you cast it?" Way different thing. It's related, obviously—the microstructure relates to that problem—but just because you understand the microstructure, just because you see and can predict crack propagation, does not mean you're necessarily going to perfectly nail manufacturing.

There are just so many things that we see stack up, and we learn a lot of new things. Every time we think we know everything, we go somewhere and learn new things along the way. I do think it's multifaceted, for sure, and each inorganic material has different constraints.

I talked about the supply chain. That's very relevant for defense applications; those are a perfect example. Supply chain is not one of the things we worry about in consumer electronics, per se. Ti-6-4 is still there. It's available. People care about where it comes from, but that's not the same as the critical-minerals focus that you see in the US today: where are we getting these minerals, the vast majority of which we do not control? That's a different problem.

There are different inputs that go in to get an output. I feel like that's why materials are so hard. It's all of this other data that comes after discovery. What I always tell people is that the second you design a new material, that's a milestone. The second you synthesize it, that's a milestone. The second you characterize it, that's a milestone.

That is not a new discovery. We count a new discovery when you pick up your phone and there's a new material sitting inside it. That, I think, is a fair claim on a new discovery in a scaled material.

Shawn Wang

As you get past—there's fallout in every one of those steps, every milestone. Presumably, in manufacturing, there are separate steps where there's fallout as well. As you get closer and closer to the consumer or the application, then you have less and less data, right?

Fundamentally, how do you get over that? Because I think in pharma right now people are starting to think about rules of thumb that can be used to do reverse translation back from the clinic to the discovery process. How do you do that in materials?

The Self-Driving Lab

It's funny. I love telling the story of one of my advisors at the company. He was at 3M for 35 years, and we asked him about manufacturing: "When you guys go to manufacture, what are you paying attention to?"

This was early. This was 3 months into the company. He said, "Whew, you guys got a lot to learn." I was like, "What do you mean? I'm a materials scientist." He said, "Different worlds."

One of the challenges he pointed out was, "You know the hardest part about the data you're asking for or inquiring about? That is someone with a 35-year trajectory at the company who knows exactly where to turn the knob on whatever manufacturing tool you're talking about. What you're asking for is his or her ability to know when to turn the knob right to that spot at the specific moment. How do you capture that? How do I give you that? Even if you assign a formula to it, is it the same every time?"

Again, this gets back to intuition, which we've talked a lot about today, just in the manufacturing sense. This is the hard part. I don't have an answer for you on manufacturing because we haven't done it. What I do have a lot of answers on is the discovery side, where we've had to look at where intuition is important.

I talked about the casting of alloys, which I touched on earlier. That's one of them. Reading SEM images, that's one of them. Looking at XRD spectra and identifying phases and how strong the peaks are: that's arbitrary. That's really strong. That's kind of strong. That's not really strong. That's terrible. What do any of those mean?

I can guess what they mean, but if you look at an XRD, you might not get it perfectly compared to what you or I think about that. This intuition aspect is so important. This is why we still have humans in the loop, because you want to capture that. Now, when you go to manufacturing, we think we'll have to do the same thing.

And we think the opportunity is to rebuild those processes fully automated. You can put all the sensors and capture mechanisms—in the absence of a better phrase—in place so you can bring all of that back. That's a hard problem to solve, and to be clear, we have not solved it yet. We are certainly still at the discovery and testing side of that, but that's where we want to get to. How do we get there quickly? Partners.

We talk to a lot of companies in our field that make materials at scale, particularly in the alloy space, and are thinking about this. They look at it from a different lens. They're not all hyped about AI for science. Actually, I'd tell you that a lot of them are bearish. They're like, “You don't know what we know. We've been doing it for a long time.” And that's okay. I think that's healthy.

What they do know and bring to the table is that we have that intuition. We will tell you, when you show us a family of elements, what we think is going to work or not. We might be wrong, but we can tell you why we think that's going to happen. We can tell you why that relates to aspects of a business that are important, like supply chain, cost, and performance under certain environments that don't exist in others—extreme environments, for example, involving temperature and pressure. That's really important information that you want to bring back. That's how we get there in the near term, until we can do it ourselves: you partner with people who want to bring this discovery, this turbocharged engine, to their process.

Shawn Wang

So, okay. Right now, you're still refining that process in the lab.

The Self-Driving Lab

Absolutely.

Shawn Wang

What are some war stories from the lab? Give me your best.

The Self-Driving Lab

That's a good question. I love that question. About the first tools in the lab, I can get in so much trouble for saying it, but I'm pleading the fifth.

Some of the tools don't let us interface with the software. We now pay for that software, by the way, and we love that tool vendor. We were very strategic about how we got access to it, and the engineering team—the software engineering team—was smart about how we could do that. That was a whole 2-week sprint that we had to figure out how we could programmatically control all these tools.

Shawn Wang

So, what's the juice, man? Come on.

The Self-Driving Lab

You can look into the things that are running those tools, and you can find out how you can control what you want to control.

Shawn Wang

Fair enough. Fair enough.

The Self-Driving Lab

So, that's one. My comms team is going to be so mad at me for that one.

Shawn Wang

We can cut it.

Jill S. Becker

No, I'm kidding. I'm kidding.

I think one of the other war stories that we saw early was how interdisciplinary the team needs to be. We knew it was going to be interdisciplinary going in. I think each field we had assigned has splintered into even more fields.

Materials science at large is one we knew would involve computational and experimental work. With mechanical engineering, we have real mechanical engineers who build tools, design tools, and put them together. We have mechatronics engineers who design all our own custom mechatronics to make those tools run autonomously. Obviously, they're in the field of mechanical engineering, but they have completely different jobs.

Software certainly involves standard full-stack work and building the operating system, but then there is more of what we call applied ML. I come from a software background, and I'm applying the systems that we are building, like pulling out images from SEM into the AI scientist. That was interesting.

Robotics—things like path planning and perception—are areas we probably didn't think we would need as much as we do today, simply because we thought, “We'll use what's open source and off the shelf today.” Then we started to realize that what the scientists were doing was very intuition-based, and that's the perfect place where perception and computer vision can be really effective. I do this with PyTorch.

All of these different fields have splintered, and I think we had to build the plane as we flew it. The startup mantra was that we had to continually add people. Ironically, that has now built a huge moat. In inorganic materials science, it is not easy to build self-driving labs. If you sent me back 2 years ago, I'd have said, “That is a tough path to walk.”

Now, of course, it's a big moat for us. We're like, it's not about a robot in front of a tool. Go ahead, put a robotic arm in front of a tool, and watch what happens. Everything else I just talked about will come the second you do that. Now we feel so much farther ahead of the industry in really running self-driving labs for inorganic materials science. That war story is funny—we laugh about it—but now we see it as a huge win.

Shawn Wang

Is there something special about this moment that enabled the self-driving lab versus 5 years ago or 10 years ago? You've been doing this for a while, so why now?

The Self-Driving Lab

A couple of things. One, AI for science is important. If I take force fields—machine-learned interatomic potentials—they're what, 2 or 3 years old, or whatever the exact timeline is? I don't know. I mean, that's interesting. You can actually start to do some parts of your process faster. Although they're still computational, in the simulation sense, with things like DFT, they're still very important to that funnel, to sharpening that funnel and moving faster at the top of the funnel.

Two, robotics are just better. First of all, they're cheaper. Second, you can do more in actuation: custom grippers and different systems that we can either custom-build or acquire.

Three, there's the buy-in from the tool vendors, like I talked about earlier. Again, 2 years ago we saw a difference from what we see today. There is a lot more optionality. I'll give you a perfect example: some of the tool vendors now have software teams. Or, if they had them before, they weren't focused on this problem. Now they actually provide support: “Here's how you can work with the interface.” That's a big change for the industry.

We've started to see it become easier to build self-driving labs from an infrastructure and hardware perspective. Most importantly, I think the biggest change has also been the excitement. Everywhere I go, AI for science is a talked-about area. I was here this week speaking at Unlock—kudos to Michelle and the MSR team—and there was unbelievable energy in the room from so many different builders across so many different companies and fields talking about AI for science.

As I mentioned, even at the national level, one of the things that we did early on was spend a lot of time in D.C.: on the Hill, at the Office of Science and Technology Policy, at the Department of Energy, and at the Department of War. We let them know, “Hey, if you want to be competitive in science, you guys really need to pay attention here. AI for science is a serious field. Self-driving labs are a serious field. There are other people who have already built these systems.”

We really think it is national infrastructure that should be built out, and there's been a big buy-in at the federal level and at the state level as well. I think you have so many tailwinds pushing this industry forward, plus the excitement and the venture dollars coming in that start to solidify some of that.

I think soon we'll start to see some results. We're still waiting on a big result from someone. Hopefully we're one of the first ones there. I think that'll be a cherry on top for the final tipping point, when customers who are coming in with restraint or caution really see this and say, “We can't believe you did that. We've been working on this problem for X many years. We've never had an output like that in that period of time. I'm convinced. I'm a believer. Let's talk about it together.”

I think you're seeing that moment happen already in robotics. It feels like we're on the upswing of that for robotics. I don't know—2 years ago, there was interesting work in robotics if you were a nerd like me and were reading about it in your free time. But I still feel like, when I talk to founders in that area, they're just now getting over the hump. Supply chain and logistics companies, the big warehouse companies, and even the big humanoid companies are now like, “Oh, okay. There are real foundation models that we want to pay attention to. We want to push from 99% to 99.9999% in our foundation model technology.”

I think we're going to have that for science over the next 2 to 3 years. I do. Once the discoveries start coming out and this field continues to mature even more, remember, we're early in this field. We are a couple of years in from an energy, a community, and a new technology perspective.

swyx

So, going to the competitive landscape a bit, we talked about what people are working on in the US and in Europe. But China is both running ahead on materials development and thinking about many of these ideas about labs going all the way through manufacturing, and they do have the expertise in that.

swyx

So, what is your thought about how we stay competitive versus China, and what is the most important thing we need to focus on there?

The Self-Driving Lab

Yeah. This is an incredible question. This is actually what I spend a lot of my time talking about, especially when I'm in D.C. China has an unfair advantage that we do not want to replicate but must figure out how to defend against.

In China, they really will go out of their way—and there's incredible work by NIST and a couple of other groups documenting this—to stand up manufacturing innovation hubs where they make a new material and support, via capital or infrastructure, the scale-up of that material system or invention. They can do that because, honestly, in China, whether you're public or private, one entity owns everything.

Again, that's the part that we should not mimic or copy. In no way am I suggesting that. What I'm saying is that, because they have that, we need to have a similar focus. We need to figure out how to break the 25-year timeline, because when I come in here and tell both of you that materials development timelines are long, you're like, “Yeah, that's all we ever hear about.” That's all we hear about. Everyone I talk to says the same thing. We know the same thing. We need to change that.

How do you do that? I think number 1 is you start to teach the scientists of the future how to run science this way. You know what the most impressive thing from our lab is? I love everything we've talked about today, but what is so impressive is that we can have 1 Ph.D. in metallurgy or alloys run 10 campaigns at a time.

When I was in a Ph.D. program, we had 10 scientists focused on 1 campaign, 1 research problem at a time. That's an order-of-magnitude jump in productivity from 1 scientist. Now imagine every scientist in the United States, every scientist in North America, in the world, doing 10 times the research output. That's fundamental. That just changes the trajectory of discovery.

swyx

But China can do that, too, right?

Yeah, they can. They can. So, I think the second piece is investment. I think we're getting it in the private sector, which is great. I think we'll continue to see the government invest in this area to try to start bridging these gaps and building up this workforce.

You see the Genesis Mission as 1 area where there's a ton of—hundreds of millions of dollars of investment there. I know that groups internal to the national labs are building self-driving labs. My co-founder Herd Seeder has 1 at Berkeley. Argonne has 1. I know that Ames is looking at self-driving labs. Livermore's looking at self-driving labs, or may already have 1. Oak Ridge has a Manufacturing Demonstration Facility, which is almost fully autonomous or semiautonomous in nature.

These labs are now starting to invest in the infrastructure to start to speed that gap up, to start to show you can shorten that gap. That's important. The third one is, I think, maybe where we have to be different: public-private partnership.

This is what we talk a lot about. This is the perfect opportunity for private enterprise to work with public research to create the greatest scientific tool in the world, as the DOE likes to say. Why? Because we have all the HPC that we need—high-performance compute. We have all the researchers that we need. We have some of the best scientists in the world at the national labs, and the infrastructure from a materials-tooling or science-tooling perspective.

Then, lastly, you actually have the data. You actually have more experimental data in all of those national labs than anywhere in the world—or, we think, anywhere in the world—from all the science that we've run. If you can couple all 3 of those things together, and you can bring in private enterprise to help you make sense of that, help you close that system, and help you tie the loop together, then I think the entire national infrastructure runs this way.

The research infrastructure, whether that's corporate R&D or, again, SBIR/STTR R&D, runs this way, and private enterprise runs this way. Now you've changed the fabric of all of R&D in the U.S. Now you're at a place where you can compete with China, not by owning everything and employing forced labor and the unethical practices that they might employ, but rather by changing the mentality and the approach to how we do R&D.

That's how I think we can compete. That's the only way we can compete, if we want to move forward. If we do not do that, then they will continue to win, because they will outpace us on cost and they will outpace us on people. If you try to play that game, it feels like we're going to lose that game.

But if you play the game of changing the system and building a better workforce and a better system to do R&D than they have, then we can beat them with raw output.

swyx

We often ask our guests a question that you've kind of answered now, but I wanted to ask it anyway and see if you have anything more you want to add. If you could remove a bottleneck from the industry, what would that be?

I don't think this is removable, so that's why it's not a good answer. The hardest part about AI for science is that our feedback loops are long.

swyx

Right, that's fundamental.

Yeah, exactly. It's fundamental. That's why I don't know if you can remove it. Maybe there are ways to get it faster, but I'll give a perfect example.

You think about math, like AI and math, right? You can run a lot of experiments in hours that will take us weeks or years to run in science. How do you get around that problem? That's a really hard problem to solve, and I think 1 answer is large-scale automated systems.

A perfect example: if you build a facility with 1,000 XRDs or SEMs, you can certainly build that model that I mentioned that can do image analysis better than anyone else in the world. I think there are paths there to leapfrog the challenge of doing fast experimentation, but it's fundamental, so it's a bad answer to the question.

For something that's not fundamental, I would have the tool providers restart their stack. Their tools are built for humans; they should build them for agents and robots. I feel like this is already happening in software.

swyx

Yeah, I think you see this in CLI and MCP. You're seeing this already happen.

That's right. If you could do that at the infrastructural level for tooling, I think it would supercharge this industry. Because now you don't want to train someone on running an XRD or an SEM; you want to train someone how to run the system that can do that, and that scales so much more effectively.

Now you don't need to get a Ph.D. to analyze your alloys in an SEM. Now you just need to focus on how I can run the system to do that exact analysis for me. That'd be transformational, but it would be quite expensive.

swyx

Okay. Do you have any calls to action for AI for scientists or, let's say, AI engineers?

The Self-Driving Lab

Yeah, they can bring techniques that scientists are not aware of. I'm a perfect example: I am not a machine learning scientist by training at all. I'm a materials scientist by training. I did my graduate work in materials science. My first job out of grad school was materials investing in materials science. Now I run a materials science company. I'm a materials scientist at heart. That's what I do.

One of my other co-founders is a materials scientist. He's an academic professor in materials science. But we have learned so much from the MLEs and the AI research scientists on our team because they're not materials scientists.

They show up to a problem and they don't get stuck on dendritic formation and grain boundaries and this stuff. They're just like, “Why don't you just train a materials model for that?” And it's like, “I don't know what that is.” They're like, “That's not the first thought I have when I look at an SEM,” but they do.

One thing that they can do is supercharge the industry with their own skill set. One thing I don't like about the industry is that I see so many people trying to be the other thing. I meet ML engineers who are like, “I want to be a materials scientist,” and I'm like, “Why? We need ML engineers that work at Radical AI. You should be an ML engineer that helps us do science.”

Then I see this one way more: materials scientists try to be an MLE. “I did a Ph.D. in materials, did a master's in materials, and I got to get into the AI side because that's where the field's going.” No, you just have to learn how to use those tools to make you better at your job.

I don't care who can figure out the problem from SEM. I just want to be able to figure it out. So there's so much cross-disciplinary work that I would highly encourage ML engineers to, number 1, pay attention to AI for science. I think that's already happening, actually. I don't think they need to hear that again, especially on this podcast. I was a huge supporter.

But number 2, lean into your expertise. Bring a first-principles perspective to the way that we do science. We have been doing science the same way for hundreds of years, or 50 years, or however old the tool is that we're running on. We've been doing it for that long.

You can come to that perspective and just totally change the way something works. The field has already done that in the past, and it will continue to do that in the future. I can tell you, 100% guaranteed, that we've done that at Radical AI today.

Specialization. Exactly. Bring the specialization and lean into your expertise. Don't shy away from it. Don't try to become a materials scientist. Be an ML—be an MLE that works in materials science.

swyx

Awesome.

swyx

So, what does your AI stack look like?

Joseph Cross

At a high level, the AI scientist is really a multi-agent approach. There are multiple agents that sit within what the AI scientist is, and we really have this orchestrator agent at the top that comes up with new hypotheses and has a specific way to test those hypotheses internally. That allows us to test whether they’re going to be good hypotheses before we send them to the lab.

But also going into that scientist are a bunch of other models as well, right? We are taking in datasets like industry-standard datasets. We pay for Caltech, and we actually pull that data in so that we can use it the same way I would use it if I were a human scientist. That’s really important.

We have a literature review agent that we’ve custom-built. That benchmark is also public on our website, and it can go and extract figures and information from scientific literature that’s relevant to the hypothesis that we’re making. So, that’s in the stack.

We have custom models built as well. One of them is called MATRIX, which is on our website. I would encourage all the ML engineers out there to go check this out. The model and the benchmark are available on Hugging Face, and you can find a blog post on our website about this.

MATRIX is incredible because it’s really a VLM that we’ve fine-tuned on Quinn. It can go into images from the lab and experimental data and extract scientific knowledge from them. The obvious benefit that we saw in the model, which you can guess, is that it gets really good at reading experimental data, which makes sense. The one that we maybe didn’t see coming was that, by understanding that data, it gets better at being a scientist. This is really cool, right?

For the AI and ML engineers out there, for the AI scientist, that is how you start to capture that. We talk a lot about this intuition, the scientific intuition you get a PhD on. That’s how you start to capture it.

We have actually seen this, and you can go read the publication—it’s on archive—where the public dataset that we used is showing improvements, like 5% to 16%, I believe, on general scientific reasoning.

swyx

Is it adding math to your reasoning?

Yeah, actually, I think math is the one area we call out that doesn’t work.

swyx

No, but the theory is the same, right? Where you add math and then you end up in other domains.

In the paper, we do move outside into the bio space, I believe, and see that same improvement. You can find this all in the preprint that’s out on archive.

What’s cool about this? What’s so cool is that, as you start to build these systems that can do science like a human scientist does, you start to compound the knowledge in the way we talked about earlier. I can be looking at all of these different things as I’m making a hypothesis.

Although a human scientist would like to think we can pull in Cal Fad, pull in literature, extract the right information from literature—not just read a paper, but pull the right stuff out that’s relevant to this hypothesis—and look at all my past experiments, all of the database of our experiments is feeding into that AI scientist.

So, when we make a new campaign, it is looking at the past results to make that campaign. We have a really cool demo that we show customers where, when you look at a generated hypothesis, it’ll actually tell you what experiments it’s pulling in and what it’s using to learn from about why it made that new hypothesis. So, that’s really cool.

Then, again, with these models like MATRIX or MATRIX-PT, you can actually start to pull out intuition, and that intuition actually helps you become a better scientist at large. This is a really important concept. This multi-agent approach is, I think, what an AI scientist really means.

I don’t think it’s ever going to be one scientist. Maybe something will happen in the future. If you’re an MLE, maybe that’s a good problem to tackle. Build that, and we’d love to be a customer of yours. But if not, I think you’re going to have these specialized models, these specialized agents that are really good at one thing, and together, collectively, they make a scientist that’s better than Joseph, better than the scientist that we have today.

That’s kind of what our stack looks like at a high level on the hypothesis-generation and new-materials side.

Alessio Fanelli

And then, just a quick follow-up. I love that you’re open-sourcing a lot of your work. Why are you doing that? Not that I want to discourage it in the least.

Joseph Cross

Yep. This is a really important question. There are 3 reasons.

Number one, I talked about community in this episode. We need the community to move toward doing science this way. Open-source work is one of the best ways to do that, as we’ve seen over history.

Number two, learning. We actually get way more feedback from open-sourcing things than we could possibly work on ourselves. You guys are probably aware of TorchSim, which was this package that we open-sourced. We don’t need to go into that today—we can—but the feedback from the community and the ideas from the community have been incredible.

We’ve actually spun TorchSim out into its own organization, a nonprofit, that can continue to run with the community. We call it Ignite: a materials revolution, a simulation revolution, which we’re excited about. That’s a good example. This idea that you can build better technology with the group is number two.

Number three, we actually don’t think models are the moat. We actually think that in 5 years, most models will be open-source. There’s probably a proprietary model or two, the same way there’s a proprietary model or two today that I can run, whether it’s Claude, ChatGPT, Grok, or whatever your favorite AI is.

However, we think in science models aren’t the moat; experiments are. We actually think the more great models we can share, like MATRIX, which is out there, the better. The dataset isn’t, right? The model that we built and put in the preprint is built on public data. Of course, we have our own proprietary data on top of it. That’s what we train MATRIX on. Then the model can go out there.

I tell people all the time, “What happens if someone else comes out with a better foundation model, better MLIP, or better diffusion model?” I’m like, “That would be amazing. It would supercharge our scientists. We’ll drop it right into our stack.” That’s why we don’t want to sell models. We don’t sell models.

We think the entire community will continue to have ideas that we cannot have alone. That’s why we open-source, because we don’t think that’s the edge. We think the edge is on the experimental side.

That’s specific to Radical, and obviously I’m talking about what Radical believes in—the thesis there—but that is a big part of why we do it. If the whole community can push the whole field forward, we benefit from that, and they benefit from that. That’s a win-win for materials at large. That’s a win for AI for science. And it’s a win because our stack now has a better model that we didn’t have to produce, which is great.

swyx

So, whether it’s us—we do have custom models built internally—or someone else brings that model in and uses it, or we even pay for a proprietary model, which we do on the LLM side, obviously, all of that just goes into making a better scientist.

That’s why we open-source: for the community to get better ideas, and to really understand that we think experiments are the moat, not the model itself.

Alessio Fanelli

Yes, Joseph Cross. Thank you so much for

The Self-Driving Lab

Thanks for having me, guys. Awesome to be here.

swyx

I hope that you enjoyed the show so much. I hope that we can spread the good vibes.

I think we just talked about a great closing topic, which is how important this industry is. The impact that the world will feel from AI for science is enormous. People that can start at the grassroots and actually push that forward, like yourself, are imperative. So, thank you for everything you do.

I’m super happy to be here, and looking forward to coming back when the lab is fully autonomous. We’ll run an episode in the lab.

swyx

Yeah, I was just going to say that. Yes.

Alessio Fanelli

We should definitely do that.

swyx

We’ll come back to New York.

The Self-Driving Lab

Yeah, no problem. Pick a different time than the winter, though. We had a tough winter this year.

Alessio Fanelli

Okay, we’ll do something in the summer.

The Self-Driving Lab

Perfect. Thank you guys for having me. Thank you very much.