材料丰裕:Radical AI 的闭环实验室实现科学发现自动化
Radical AI 的核心论点是,材料创新被困在代价高昂的商业化鸿沟中:一个新系统可能需要“超过1亿美元”和“10年以上”,有时甚至20至25年才能上市。 学术界追求基础理解,企业研发则瞄准1%、2%或5%的改进,因而留下了一个与商业需求直接挂钩、完全敞开的突破空间。
其拟构建的护城河,是将 AI 推荐与一座每天进行100项物理实验的机器人实验室连接起来的闭环数据飞轮。 Krause 在 Army Research Lab 工作期间,全年大约完成50项实验;即便政府项目每年能完成400至500项,也只相当于 Radical AI 运行约1周。合成结果、失败记录、显微图像和性能测试会通过主动学习反馈给下一项实验。
商业模式将延伸至规模化和制造,因为发现与商业化不可能长期割裂。 一块四分之一大小的实验室“按钮”样品,可能需要变成重达400磅的部件,同时保持原有性能。因此,Radical AI 计划规模化销售材料:在有用时保护材料配方,但将低成本加工诀窍视为更深层的商业秘密。
语言模型位于 Radical AI 复现科学直觉的押注中心,背后由基于 GNN 的 MLIP、生成模型、计算机视觉、仿真和实验室分析模型提供支持。 关键任务不只是预测原子尺度的力,而是综合论文、专利、实验、图像、成功与失败记录,判断下一步该做什么实验。Colindres 称,科学实验室是可能最好的“真实基准”评测。
高熵合金是最初的商业切入口,因为它们结合了适用于高超音速、聚变及其他极端环境的多种性能。 高超音速材料必须同时承受超过 Mach 5 的速度、高温、高压、氧化、成本约束和供应链限制。Radical AI 最近获得的 Air Force Direct-to-Phase-II 项目,正是将高通量实验应用于高超音速系统合金。
实验数据,而不是又一个边际改进的模型架构,被视为稀缺的战略资产。 Krause 估计,他在 Army Research Lab 的工作中有“90%”没有成功,也从未被记录下来;即使现有实验笔记,也往往没有标签、信息不完整,若不知道科学家的意图,几乎无法解读。Radical AI 认为,掌控实验室就能把这些原本会丢失的痕迹转化为结构化、带上下文的数据。
管理层为科学上行空间设定了激进时间表,同时保留了必要的条件限制。 Colindres 预计,如果“我们按计划推进并实现目标和路线图”,未来12至24个月内可能出现材料领域的 AlphaGo Move 37;他预计6个月内形成原生多模态内部模型,并可能在6至12个月内发表或发布。更长期的目标,是在不为每个材料类别准备数百万样本的情况下实现跨类别迁移。
主要执行风险仍集中在物理自动化、模型是否真正理解,以及工厂级可复现性。 一些老旧仪器仍需科学家操作,有些使用数十年的设备无法让软件控制其泵、腔室或热源;Colindres 也表示,他们尚不能证明模型拥有基于原理的理解,而非记忆式启发。Krause 坚持:“除非能在实验室里做出它,否则它就不是新材料”——而要成为商业材料,还必须实现规模化。
1. 材料发现被困在学术科学与渐进式研发之间
Krause 给出的基准异常严峻:开发一个材料系统可能耗资“超过1亿美元”,耗时“10年以上”,有时要到“20、25年”之后才能商业化。材料支撑着汽车、航空航天、制造、国防、气候、能源、半导体和电子产业,因此这种延迟会传导至整个工业经济。
关键在于两端之间的割裂机制。学术研究者优化基础理解,通常不以商业应用为目标;企业实验室则优化已经上市的产品,追求“1%、2%、5%的性能提升”,改善利润率并能向华尔街交代。
两者之间存在 Krause 所说的“完全敞开的空白地带”:从一开始就围绕商业价值筛选基础发现。Radical AI 正将自己定位于此,而不是做科研软件供应商,或另一个企业内部优化实验室。
Labenz 追问,成本上升是否只是因为低垂果实已经被摘尽,就像药物发现一样。创始人的回答是:两者都面临无法穷举的搜索空间和割裂的开发流程,但材料领域以规模化和加工取代临床试验,难点在于把一个小型反应转化为吨级、可重复的产品。
2. 规模化是材料行业的临床试验瓶颈
Colindres 认为,最难的问题“从字面上说就在规模化”。研究样品可能只是四分之一大小的“按钮”,商业目标却可能是一块重达400磅的部件。工艺一旦发生变化——包括环境和化学条件的变化——就可能抹掉最初发现的价值所在。
Radical AI 希望让数据从按钮样品一路贯穿到逐步放大的形态。如果性能下降,系统应识别工艺中引入了哪些“环境、化学因素,或者诸如此类的因素”,再调整下一轮工艺,而不是把制造当作彼此割裂的下游交接。
Krause 说,这可能会占据公司大部分工作量:候选材料会在计算、首次合成、复现和工业制造之间逐步夭折。即便是硅用于晶体管,也经历了漫长的成熟过程,并随着半导体制程节点持续演进;更具投机性的材料,甚至可能撑不过第一次复现。
Labenz 问,如果启动一个足够紧迫的“超高速”项目,能否把流程压缩到1年。Krause 开玩笑说,Colindres 已经让团队把低于1年的开发周期当成硬性要求;但实质答案是,自主实验必须同时加速发现,以及加工过程中所需的学习。
3. LK-99 说明合成细节决定发现是否真正存在
Labenz 挑战了 Krause 对 LK-99 的最初表述:按他的理解,所谓超导性是一个错误,并不是一种等待复现的真实材料。Krause 同意,自己读到的最新解释涉及一种意外生成的杂质,可能是硫化铜,但仍保留限定语:“LK-99 到底是不是真的,我不知道。”
更广泛的教训依然成立。报道中的温度范围、容器细节和其他合成条件说明不足,而每一项都可能改变结果。材料并不由配方成分单独定义;它经历的时间、热量和设备轨迹,也是最终制成之物的一部分。
Krause 用厨房作比喻:给两个人完全相同的面包原料,但允许其中一人给面团划口、让酵头发酵,另一个人却把它在1,000度下匆忙烤30分钟。名义配方相近,前者做出面包,后者做出的却是“一块石头”。
他的标准是绝对的:“除非能在实验室里做出一种新材料,否则它就不是新材料;除非能把在实验室里做出的东西规模化,否则它就不是新的商业材料。”无法复现的主张,还不能算发现。
4. 客户需求可以把沉睡数十年的科学变成平台材料
Colindres 最喜欢的例子是 Corning 的 Chemcor:它在1950年代或1960年代初开发出来,在货架上沉睡了大约半个世纪。2007年前后,Steve Jobs 想要一种触感和外观都优于当时手机塑料屏幕的玻璃,于是 Chemcor 被改造成 Gorilla Glass。
按他的说法,Gorilla Glass 随后进入 Apple、Samsung 和 LG 的产品,并可能成为 Corning 最大的收入来源。重点不在于底层科学突然出现,而在于客户终于明确提出了材料属性和产品场景,让既有材料获得了价值。
因此,Radical AI 不想“在真空中制造材料”。客户要求会在模型优化之前就进入系统:目标元素、抗氧化性、强度、延展性、工作环境、成本和供应链,都被视为科学目标的一部分,而不是事后补上的商业化细节。
5. 赋能型材料比略胜一筹的涂料更重要
公司的北极星是能够创造或解锁一个行业的“赋能型材料”,而不是仅仅改善既有产品的材料。Krause 将其与企业研发追求的小幅百分比提升作对比;Colindres 描绘的愿景包括悬浮列车、星际旅行,以及无需充电即可从纽约开到波士顿的电动车。
室温超导体是他们的典型例子,因为其二阶、三阶影响无法事先穷举。Colindres 说,就像触控玻璃无法直接推导出 Uber 或 DoorDash 一样,“我们其实不会知道”室温超导体出现后世界会变成什么样,但它将允许前所未有的技术出现。
这种雄心也要求聚焦。Krause 承认,早期公司不可能在每种材料和每个行业上“煮沸海洋”;它需要一个滩头阵地、执行能力,然后再扩张。室温超导体仍在长期路线图中,并非公司当前的运营项目。
6. 高熵合金将材料设计变成多目标优化
Radical AI 当前的滩头阵地是高熵合金。Krause 说,这类材料能够同时结合机械性能和热性能,而不必进行简单取舍。其吸引力在于可在高温下保持高强度,并有望用于多种原本彼此不同的极端环境场景。
对高超音速系统而言,材料必须在 Mach 5 以上运行,同时承受高温、高压、大气氧化和腐蚀环境。在聚变反应堆中,相关合金可能还要承受持续辐射和侵蚀,而当前材料如钨无法长时间承受这些条件。
客户几乎从不要求把某一个性能单独做到最大。真正的规格,是在特定温度和压力下的机械性能,同时具备所需的抗氧化性、可接受的成本,以及构成元素的可行供应链。Krause 强调:“你优化的绝不只是热膨胀。”
应用范围可以从 Lululemon 运动衫中的银纳米颗粒,延伸到为火星电子设备阻挡辐射的屏蔽材料。Krause 说,他们瞄准的不是“5%的改进游戏”,而是可能带来“100倍更好”或“50倍更好”性能、从而催生新应用的材料。
7. 飞轮将性能需求连接到每天100项实验
在最高层面,这套系统由 AI 引擎和机器人自动驾驶实验室组成。客户提供目标性能和约束条件;引擎进行“性能驱动优化”,提出材料成分和合成流程,并通过 Radical AI 的操作系统下发实验设计。
实验室随后合成材料,并对产物进行表征。技术包括 X 射线衍射和扫描电子显微镜,之后进行硬度、强度、应力-应变行为、蠕变和裂纹扩展等性能测试——这些测量把成分和微观结构与应用性能连接起来。
机器学习子模型实时分析这些数据流。形成的洞见回到 AI 引擎,利用第1项实验的结果修改第2项实验。Colindres 强调,这“并不是对科学方法做某种疯狂的新改变”,而是更快地提取和复用科学方法中的信息。
吞吐量改变了可行的学习曲线。Radical AI 的目标是每天完成100项实验;Krause 估计,自己在 Army Research Lab 全年大约完成50项。即便有针对性的政府项目每年达到400至500项,也只相当于 Radical AI 计划运行约1周。
8. 语言模型协调异构的科学模型栈
Colindres 列出的组件包括用于原子尺度建模的 GNN、材料科学机器学习原子间势或 MLIP、生成模型、GPT 系列语言模型,以及实验室中的计算机视觉。不同任务需要不同的输入、输出、架构和物理结构层级。
他最有把握的是计算层:MLIP 已较为成熟,并且已经具备实用价值。Radical AI 真正差异化的押注在语言层——推理科学知识、整合多模态和工具、生成假设、设计实验,并从实验的成功与失败中学习。
因此,这套架构更像一名负责协调的科学家,而不是一个万能预测器。LLM 可能调用 GNN 解决原子尺度的不确定性,查阅文献或专利,再通过计算机视觉检查 SEM 图像,最后推荐合成流程。其他模型在语言系统形成科学直觉的过程中,为其提供“增强”。
9. 人类科学直觉本质上是无人记录的失败累积
Colindres 指出 Radical AI 的第三位联合创始人 Gerbrand Ceder 是杰出的材料科学家,“知道一些我们其他人不知道的东西”。但 Ceder 自己的解释很平实:大约40年的实验、学生研究和同事论文,积累成了一套关于什么有效、什么无效的庞大内部目录。
Krause 从自己转向 Army Research Lab 神经形态计算的经历中给出一个更具体的版本。在研究离子栅半导体通道1年后,他说,自己在思考如何为该应用设计新的半导体材料方面,可能已经超过了99%的其他科学家。
真正缺失的资产是失败工作。Krause 估计,“我做的东西有90%都没成功,姑且这么说”,而这些电子束、PVD 或 ALD 实验没有一个变成可复用的数据集。只有最终成功的材料离开了科学家的头脑,进入正式记录。
Colindres 将这套潜在语料称为“一套百亿亿级数据”,但历史上获取它的方法,是投入40或50年的学术生涯,专注于大约一两个领域。Radical AI 希望把这种直觉的积累压缩到不到10年,并将其迁移到更多问题上。
10. 发现系统必须因意外而非仅因拟合获得奖励
Krause 关于单层薄膜的故事说明了视觉证据和偶然证据为何重要。他的团队注意到,更透明的二维半导体薄膜导电性更好,却不知道原因;直到一位外部专家认出它们是单层 TMD 薄膜——此前他们从未想过要在那个实验场景中寻找这种解释。
Colindres 认为,这体现了机器学习的根本张力:在分布上训练的模型会自然排斥异常值,但重大科学发现可能藏在“化学宇宙中某个所有人都觉得有点丑陋的奇怪角落”。系统必须学会正常结构,同时保持“随时准备迎接新颖性和意外”。
Radical AI 使用贝叶斯方法、蒙特卡洛方法、多样化、噪声采样、主动学习、强化学习,并在早期保留人在回路中,以支持 Colindres 所说的“追踪直觉”。目标是在不把过程变成无方向的随机搜索的前提下,以科学家10倍、20倍或100倍的规模追逐异常结果。
11. 物理实验室既是真实基准,也是方向盘
Labenz 将新颖性与基础模型、RLHF 和仅奖励强化学习之间的更广泛张力联系起来:模仿可能磨平创造力边缘,不受约束的奖励优化则可能产生不安全或病态策略。Colindres 的回答重点不在意识形态,而在评测质量。
“老实说,没有比实验室更好的评测了。”他认为,物理结果可以确定模型是否实现了目标性能,以及是否牺牲了其他要求。在一个安全同样重要的领域,这种真实基准支持主动学习和面向客户的调节。
更新机制必须比定期重新训练权重更动态。Colindres 描述了一层代理或智能体系统,利用当前模型,对新的实验数据立即作出反应,并在底层模型持续改进的同时选择工具。机制可解释性则用于揭示推荐为何朝向或偏离既定目标。
12. 材料数据稀缺、无标签,且无法与意图分离
Krause 将材料与语言和软件作对比:“我们没有 Newton 的笔记本”,也不存在类似 GitHub 的开源社区,发布数百万条结构化实验轨迹。论文呈现的是成功结论,而不是那些教会科学家如何走到结论的失败尝试。
Labenz 建议付钱给研究生,让他们提供实验笔记;但 Krause 的回答表明,即便拿到完整档案,也需要采访作者,因为观察没有标签、引用缺失,而且不同专业对相同测量的含义并不相同。
他举的 XRD 例子很具体:研究晶体材料的科学家会寻找衍射峰,而研究非晶材料的人可能恰恰把没有峰作为确认结果。同一个输出,对一名研究者来说可能是空白,对另一名研究者来说却是成功;除非附带目标和解释,否则无法判断。
即便时间戳也可能误导。Krause 有时只记录自己关掉热源、让样品不受控冷却时的中点时间;如果没有过程上下文,“1分6秒”毫无意义。仿真数据更容易整理,但 Colindres 认为,它只解决了“约10%的问题”。
13. 搜索空间大到天文数字,但客户性能要求提供了索引
Krause 以宇宙中约 (10^{80}) 个可观测原子作为组合空间的直觉参照,把目标材料形容为一片地球大小海滩上的“一粒沙”。他还将其与约 (10^{130}) 的蛋白质搜索空间作对比,后者更大,但同样无法穷举。
仅合金开发就可能涉及约60种元素。Krause 说,在等比例框架下,仅从5种元素出发,就有“550万多种”潜在组合;如果允许每种元素比例在0%至100%之间变化,再加入小数级成分变化或改变加工路径,空间还会大得多。
Radical AI 明确拒绝把暴力组合搜索当作充分的科学方法。性能要求先排除不相容区域;专门模型再缩小候选范围,随后主动学习才进入昂贵的实验室闭环。目标是“给搜索空间建立索引”,而不是测试每一种名义成分。
14. 今天限制自主性的,是老旧科学仪器,而非通用机器人
Krause 说,难点通常不是漫游车或样品搬运轨道,而是科学仪器本身。供应商花费25年、50年、75年甚至100年围绕人工操作设计设备,配备实体开关、旋钮和接口;这些接口能暴露输出,却无法提供完整的软件控制。
名义上的接口可能返回测量结果,却不允许控制涡轮泵、热源、真空腔室、压力、装载或污染管理。Radical AI 会改造这些功能,与供应商定制设备,或预计在未来10年内淘汰特别僵化的工具。
公司不会重新设计 X 射线衍射或其他表征方法背后的物理原理。它自动化的是装载、腔室状态、激光时序和样品移动,因为“Radical AI 不是一家材料工具公司”;其目标产品是被发现并制造出来的材料。
今天仍有一些科学家在协调实验、衔接仪器。Krause 的限定是时间维度上的:每件工具都从科学家演示其运行方式开始,并应逐步走向自主使用;如果改造的价值不够,Radical AI 就会与其他人合作定制替代设备。
15. 主动学习在每个尺度上重复探索与利用的决策
Colindres 解释,不存在一条统一的推理时扩展规则,因为每一层的问题都不同。原子尺度的探索要问哪些结构和力预测值得进一步计算;实验室中的利用则要问哪种成分真正能够合成,以及下一步应调整什么压力、温度或工艺。
在原子尺度,预测力的不确定性可能触发 DFT 计算,从而更新模型。在实验室中,一项实验可能显示当日压力或温度与假定条件略有不同,系统便立即调整下一轮运行。
仿真应在消耗稀缺物理实验能力之前筛选假设,就像科学家进入实验室前会先思考、写作和建模。但 Colindres 反复保留限定:科学并不完全逻辑化,也不总是平滑连续的,因此每一层都必须“遵守一般规则”,同时保留打破规则的能力。
16. Radical AI 为 Move 37 设定了时间表,但尚未证明模型真正理解
当被问及模型是否已经从记忆式启发式方法进化到对材料的原理性表征时,Colindres 给出了诚实的否定回答:“我不会说我们已经处在能够明确、确定地指出模型拥有理解的阶段。”幻觉和记忆现象显而易见;真正的顿悟尚未得到证明。
Radical AI 与 Goodfire 及 CEO Eric Ho 的合作,旨在让模型习得的表征变得可检查、可操控。Krause 对可解释性可能揭示出科学家从未知道该如何询问的有用启发式方法感到兴奋,其中包括能在不同领域迁移、又不携带科学家既有偏见的关系。
目前还没有出现材料领域足以震动世界的 AlphaGo Move 37。Colindres 预测,如果实现目标并按路线图推进,Radical AI 可能在12至24个月内做出一个;更长期的“圣杯”是把合金实验中的能力迁移到超导体等领域,而不需要数百万个类别专属样本。
Colindres 还预计,具备相关科学数据类型编码器的原生多模态系统,将在6个月内用于内部工作,并可能在6至12个月内发表或发布。他没有声称一个模型会永久取代所有专用工具。
17. 垂直整合让加工诀窍成为持久业务
创始人从材料问题本身的要求出发,而不是从最容易获得风险投资支持的产品出发。他们的结论很直接:如果公司只卖软件,“你永远赚不到真正的钱”;而为每项发现都授权,则要求不断产生新发现,并且始终暴露在专利到期和逆向工程风险之下。
这种模式更接近 3M、Applied Materials、BASF 和 Materion 等大型材料公司:规模化销售材料。Radical AI 的差异,在于它追求的是能够赋能聚变、新型交通、高超音速或星际系统的材料,而不是另一种涂料变体。
材料成分可以申请专利,但 Krause 认为更深层的知识产权在加工流程:以可靠性最高、成本最低的方式制造材料所需的商业秘密。知道准确成分,并不等于知道如何在工业规模化过程中保持其微观结构和性能。
模型政策也遵循这一商业结构。一些机器学习系统可能公开发表,另一些则会保留在内部;持久的专有资产是实验数据集、与客户绑定的发现闭环、加工配方和制造能力。
18. Air Force 项目是平台的首次执行测试
Radical AI 最近宣布的 Air Force Direct-to-Phase-II 项目,目标是面向高超音速应用开发高熵合金。项目将高通量实验与候选材料测试、优化方法结合起来,之后再由 Air Force 评估其是否适用于高超音速系统。
Krause 直接描述了战略环境:中国和俄罗斯已经部署高超音速系统,而中国将围绕一种新材料建设制造中心,专门学习如何实现规模化。他认为,美国在重要项目的材料基础上,长期未能保持同等力度的投入。
公司正与 Office of Science and Technology Policy、National Science Foundation、Department of Defense、Department of Energy 以及国会推动公私协调。无论应用是高超音速、聚变还是模型训练数据,Krause 都认为政府需求与私人执行必须协同。
公司的运营文化围绕反复实验失败而设计。Krause 说:“我们每天都会失败,而且未来10年还会继续有意地每天失败。”他的招聘标准同样明确:只是想找一份工作的人不该申请;准备把工作奉献给这项使命的人,才应该加入。
Today my guests are Joseph Krause and Jorge Colindres, co-founders of Radical AI, a company that's building a materials flywheel: an integrated system combining frontier AI models with almost fully autonomous laboratories to dramatically accelerate the discovery and production of new materials.
As someone who spent a year as an undergraduate research assistant weighing out small amounts of fine powders and running very tedious experiments, I have long dreamed of the day when science would start to be automated. That day, I'm pleased to say, is now arriving. Radical AI's lab can run 100 experiments per day, roughly doubling both what I did in a full year as an undergrad and what Joseph did in a full year at the Army Research Lab. They're using that unprecedented throughput to tackle some of humanity's most important materials challenges.
The upside here is truly enormous. As Joseph and Jorge explain, advanced materials are the foundation of nearly every major modern industry, from semiconductors and aerospace to energy and defense. Yet developing new materials remains painfully slow and expensive, typically requiring over $100 million and more than a decade to go from discovery to commercialization. This creates a so-called valley of death between academic research, which focuses on fundamental understanding, and corporate R&D, which, for the most part, focuses on incremental improvements.
In that gap lie potentially transformative materials that could enable futuristic technology such as floating trains, interplanetary travel, and hypersonic flight, but which nobody has successfully brought to market. Joseph and Jorge walk us through their entire closed-loop system: how they use property-driven optimization to navigate the vast space of possible materials, how their AI engine incorporates multimodal data from papers to microscopic images, how their robotic lab conducts synthesis and characterization, why they believe that capturing experimental data at scale is the key unlock, and, crucially, how they're vertically integrating all the way through to manufacturing.
Along the way, we also discuss why materials science has become more expensive over time, the LK-99 saga and what it reveals about the challenges of synthesis and reproducibility, how the Radical AI system learns from both successes and failures while ultimately building toward a sort of scientific intuition, and their recent Air Force contract to develop high-entropy alloys for hypersonic applications.
What struck me most about this conversation is the scale of Radical AI's ambition. This is not a company trying to make paint a bit more durable. On the contrary, they're trying to create the materials that will define entirely new industries.
Joseph Krause and Jorge Colindres, co-founders of Radical AI, welcome to The Cognitive Revolution.
Thanks for having us. Excited to be here.
Yeah, likewise. You guys are doing some really interesting frontier work in the application of AI to materials science, and there's a lot to unpack.
We were just joking beforehand that automating science is something that, as an undergraduate research assistant, I dreamed of for many hours as I was sitting there doing the tedious work of weighing out small amounts of fine powder. I was told at the time that it was very difficult to automate and probably wouldn't happen in my lifetime. The fact that it is starting to happen, as part of a bigger vision, gets me excited right off the bat.
For starters, our audience is very interested in AI by definition, but probably mostly doesn't know a ton about materials science broadly. Maybe you could just set a baseline for us. Why is materials science hard? What's going on? One of the observations from doing my homework on the company is that the cost of materials science breakthroughs seems to be going up. What's the lay of the land today that you guys are entering into with a new paradigm?
Luckily for you, Nathan, your dreams are about to become, if not already, a reality in today's world. It's going to be an incredibly exciting 5 to 10 years. We think, and fundamentally believe as a company, that science will never look the same over the next decade. We're very excited to be helping lead that charge and build that.
If you think about why materials are so exciting—and I'm sure we'll get into this in the founding story of the company—they really impact so many different industries in the world: automotive and aerospace, manufacturing and defense, climate, energy, semiconductors, and electronics. The most important industries in the world are all a direct result of materials R&D.
But we always see 3 problems. First, cost: as you called out, north of $100 million to develop a new system, from that computational framework through to really scaling that process. Second is time. It takes a long time to do a novel discovery. On average, we typically see 10-plus years to really push a novel discovery into commercialization, and sometimes it can be much longer, spanning 20 to 25 years.
Then 3, fragmentation. That's probably one of the biggest problems contributing to both the cost and time that I just alluded to.
In materials science, we always describe it as a company as 2 different paradigms. On the one end, you have this very fundamental approach to science, what we call the academic pursuit of science. It drives your understanding of science, is very fundamental in nature, and is usually not focused on commercial application. There can be some tie-in, of course, but that's not the objective of academic research. The objective is to drive your understanding of what you're studying.
On the other end is really commercialized R&D, or corporate R&D. They're entirely focused on commercialization, which is great, and entirely focused on optimization as well. They're driving toward 1%, 2%, or 5% performance gains so that they can change their margin and report that to Wall Street. They're not focused on discovering something like a room-temperature superconductor, which I'm sure we'll get into.
This is the problem. In the middle is this wide-open white space to do a novel fundamental discovery while being entirely focused on commercial applications, and that's exactly what we at Radical AI are trying to do. So that's where materials science sits today, how we view the landscape, and the opportunity that's sitting before us.
Was it always that way? The obvious analogy that comes to mind for this sort of discovery is drug discovery. Broadly, the story there seems to be 2 parts. One is that it used to be easier because there was lower-hanging fruit, and maybe that low-hanging fruit has been picked. The other side is that clinical trials are really expensive, so even once you have something that's really great, you still have a long way to go before you can actually get it to market. Rightly or wrongly, we want to be somewhat safe with that stuff, at least.
Is it similar in materials science? Would you tell a low-hanging-fruit depletion story? Is there something analogous to the clinical trial process that makes things harder than it would naively seem like it should be?
Yeah, absolutely. There are very strong similarities, and then there are very core differences as well.
One important thing about materials is that they don't go into the human body. Well, most don't go into the human body. There is a big delineation there. Some actually do, and there are classes of materials that, of course, have to be bio-friendly. But there are a lot of similarities in the size of the search space. Things like proteins or other things in biology are bigger, but materials are quite literally impossible to move through as well. They're both large, and then there is this challenge of going to full scale, what we call vertically integrating across.
It's about actually going from the novel discovery and pushing that all the way through to producing it. We have different steps in that process for biology or drug discovery versus materials, but they both equally take around the same amount of time. Whereas biology might be going through clinical trials, materials are going through scale-up and processing. How do you actually take this small reaction that you've run in the lab and figure out how to make tons of that material so it can be used in a commercial application or industry?
Again, different use cases and different processes to approach that, but you still get this high cost, this long timeline, and, candidly, this fragmentation, where different people are doing different parts of that process.
Just like in drug discovery, you might have a company focused on the discovery of novel drugs, and then big pharma thinking, “I’ll pay for the clinical trial. I’ll go through that process.” You put all this process together, whether you’re going through clinical trials or worrying about optimizing the microstructure for a novel alloy.
Both of these are time-consuming, incredibly expensive to operate and go through, and they lead to this really low value that doesn’t allow a lot of discoveries to proliferate into real commercial applications. There’s a strong corollary between drug discovery and materials, and the differences we think are actually less explored inside materials today. It’s the exact opportunity for a company like Radical.
It’s interesting because I would think we’d be getting better at this sort of scale-up and process over time, right? Whereas on the drug-discovery side, with clinical trials, you can sort of tell yourself a story like, “Well, we’re becoming ever more safety-conscious, ever more bureaucratic. Everybody’s got even more boxes to check on all the steps of that process.” Again, maybe to some extent for good reason.
But on the materials side, it seems like we should be getting better. Is there a Warp Speed? If there were a sufficiently great breakthrough or sufficiently great market need, do we have the ability to say, “We’re going to do this in a year”? Or is there just some reason that that kind of focus and energy can’t be brought to bear on material scale-up in the same way it was for COVID vaccines?
Whether there’s an answer to that question or not, I can assure you that Jorge makes the entire team aware that it needs to be done in under a year, and our technology needs to be going much faster than that. No, to your point, we think this is where autonomous experimentation is incredibly important, and we’ll get into this in deep detail. Jorge will walk through why experimental results are so important to the AI framework and why our flywheel is so differentiated.
But it’s not just about changing the way we do discovery from an AI lens—meaning, how do we recommend more things that are actually more attuned to the application we’re trying to solve for? It is those things, but we actually can just do more experimentation faster with self-driving labs today.
As we’re doing that high-throughput experimentation, we are not just running something in high throughput repeatedly over and over. We are using an active-learning approach so that every run we do incorporates the learnings from the last one, just like a human scientist would do today.
In the lab that we are building at Radical AI, this full self-driving lab, we are targeting 100 experiments per day. Right now, I probably did 50 experiments per year when I was at the Army Research Laboratory as a materials scientist. In some extreme cases, directed programs from agencies in the government, you might push up to 400 or 500 experiments a year, like the MACH program, for example. But that’s a week of research for us at Radical.
So it’s this compounding effect of speeding up experimentation and then this huge unlock of the experimental data. Jorge, I can kick that to you—how relevant that becomes on the other side of the engine.
Yeah, I mean, experimental data is really the big unlock we think for all of the big materials problems. But really, it goes even beyond experiments. It actually, to your question, Nathan, goes all the way down into scale-up, into processing, and then ultimately into manufacturing.
To answer your question around where this big challenge is, it is quite literally in scale-up. You start at the experimental scale with a material that’s the size of a quarter, and then you’re trying to create a 400-pound piece of material from that thing, while at the same time retaining all of the properties, even though you’re changing it in so many different ways. That’s really, really hard to do and a different kind of challenge from what exists in drugs, for example.
A big part of our belief is that we want to be able to capture experimental data at the research level, at that quarter-sized level. We call them buttons. Then we want to be able to thread that all the way through as we continue to scale up the material, to make sure that we are retaining the properties that we want.
If we’re not, we want to understand why we lost those properties, what we introduced—environmentally, chemically, or so on and so forth—to change those properties, and what we can do about it in order not to change them in the future. That’s a big part of the unlock that we want to go after.
So, do we then have the equivalent of orphan drugs in the materials world? Do we have a bunch of orphan materials that are created at this small scale but just have not ever been scaled up because of all these practical difficulties?
I’ll let Joseph answer that as the scientist.
Yes, I would say that’s probably the bulk of the work that we do, and there are different lenses to that, right? If you look back at some of the biggest materials that have come out, they just might have taken a longer time, like a transistor, right? And the use of silicon in semiconductors—obviously one of the most talked-about, most impactful, most-used material discoveries today. It also took a very long time for that to really push through, and continues to exponentially increase today, with the node sizes that we’re pushing into in semiconductors.
On the complete flip side of that, you have something like a room-temperature superconductor. It’s not a new idea; people have been talking about it for decades. It is one of the more exciting material discoveries you could probably ever make. The applications of that are really incredible.
But we struggle to understand the synthesis process of that. In the most recent one, the room-temperature superconductor LK-99, we failed to be able to make it experimentally—and that was even just going from theory and one experiment to replication.
Now take exactly what Jorge said and not only be able to reproduce it at a quarter size, but then go make, in that case, millions of tons, because the world wants that at large scale. That is unbelievably complex.
There are a lot of these things—from computational discovery to just making it in the lab repeatedly—that go to die. And that’s not even stepping through the long process of manufacturing scale-up and then really putting it into industry. What product does that material end up in? How is it used? It’s a very complex process.
To Jorge’s point, when you can capture all that information and use it to design better materials, better compositions, and better structures from those compositions, we think we can really 1000×—in some cases, 300–400×—the speed of discovery today.
This customer aspect of it is huge in our minds, and it really does get back to that white space that Joseph was talking about a couple minutes ago. Honestly, one of the prime examples that we always point to is Gorilla Glass, which a lot of people don’t know actually originated as a product for Corning in the ’50s.
Corning developed something called Chemcor in the ’50s, maybe the early ’60s, and it sat on a shelf for half a century. It wasn’t until about 2007 when a guy named Steve Jobs walked over to the offices in upstate New York and said, “Hey, I’m developing a phone. It’s going to completely change the world. All of our phones right now have a plastic screen, and I’m Apple. I care about design. I don’t want that. I want something that feels better, that looks better. It just has these properties that I’m looking for.”
Corning said, “You know what? I’ve got a product, and it’s been sitting on the shelf for a while.” Eventually, they took Chemcor and turned it into Gorilla Glass. That sits in every Apple product that they’ve created over the last two decades. It also sits in the products of Samsung and LG, and I think it’s probably the biggest revenue driver for Corning at this point.
The reason we love this story is because it really alludes to this concept: number 1, we don’t really understand the materials we make a lot of the time. We don’t really know what their good properties are. And number 2, we actually need a customer to tell us what they’re looking for in order to map those properties to customer expectations and where value can exist.
So for us, a big part of it is, again, we want to bring in those customer expectations. We want to bring in those value drivers very early into the process so that as we’re designing materials, experimenting and doing R&D, and then ultimately scaling these up and manufacturing them, we’re always keeping the customer in mind and always satisfying their needs. That’s a really important part for us.
I’m very eager to get into the flywheel, but just a couple more foundational understanding questions before we do. One, on the LK-99 saga, it almost sounded like you were saying that was maybe real, but we just don’t know how to reproduce it. I had come to the conclusion that it was a mistake, like some measurement error, and that they never actually had anything. How would you describe the field’s best understanding of where that landed?
No, I think you’re correct. I think the most recent thing I read was that it was an accidental creation of an impurity. Copper sulfide, I think, was one of the things that played out from that. However, there is this notion that we didn’t even know how to reproduce it. Things like the temperature ranges, as well as what we were making it in—the container, because of course that’s going to affect experimental results—were unknown or undisclosed.
I think the best example is to think about a cook in a kitchen. If I give 2 people the same ingredients to make a loaf of bread, and one scores the bread, lets the starter rise, and uses the right temperature and bake time, they have a really good loaf of bread at the end of it. If someone rushes the process in 30 minutes at 1,000° to speed it up, they’re not going to make a good loaf of bread. I think it’s going to be a rock, and it’s going to be pretty bad to eat. Same exact ingredients, same exact process: I’m going to heat it, I’m going to score it, right? But even just changing the time variable there and the temperature gives you an entirely different output from that.
So whether LK-99 is real or not, I don’t know. I don’t think I have enough experimental results that we’ve run to determine that. What I do know is that the proposed method of synthesis and the lacking details from that method don’t qualify as a new discovery, because a new discovery needs to be reproduced. You don’t have a new material until you can make it in a lab, and you don’t have a new commercial material until you can take what you made in a lab and scale that up. And so I think we have not done that yet in our superconductors. Whether that material exists or not, I think time will tell.
Fascinating.
And has the field just moved on from that for the moment? We’re all just doing other stuff now? I mean, it seems like we would want to run that to ground collectively, right?
The field might have moved on, but I can tell you that we think about that problem quite frequently. And when we think about the vision for Radical AI, look, we have a beachhead that we can get into. We have to; otherwise, you’ll boil the ocean with materials science. It impacts every industry, and as an early-stage startup, you cannot do that. You must be laser-focused and really drive toward execution, then scale.
But when we think about what our technology can really do, our materials flywheel is attempting to remove materials as the biggest bottleneck to our most important industries. We want to build a world where you go from a human-driven process to an AI- and autonomy-driven process. And this approach allows us to create what we call enabling materials. We actually create the industries that we sell into; we don’t just optimize the current industries today.
Just like Steve Jobs created the touchscreen for smartphones and other products today, that’s the exact type of application that we want to enable in the long run as a north star for the company. And clearly, our superconductors, with their immense applications, are somewhere situated on that roadmap.
Yeah. One of the things you should understand about the company and why we founded it is that it’s not because we thought there was a big opportunity in materials. That’s definitely true, and it’s a big part of what we’re going after, but we seriously wanted to create a company that was going to move humanity steps forward.
We want to live in a world where we can ride on a floating train. We want to live in a world where we can actually travel to a different planet. We want to live in a world where you can do some of the things that we think should be basic, like being able to drive a Tesla from New York to Boston without having to recharge it. Those are the kinds of things that we feel like should be available to us and within our generation.
And so that’s the whole point of Radical AI. We want to make those things happen, but we can’t because materials are constantly getting in the way. So that’s the real driver for the company, and a superconductor is probably, from a materials perspective, the epitome of an enabler. It will create so many different things, and we like to think about the second- and third-order effects of what will come out of the materials we create, in the sense that we can’t actually predict what will come out of a superconductor.
We won’t know what sorts of things it will allow. We won’t know what sort of products will come out of it. We won’t actually know what the world will look like when we have room-temperature superconductors. But what we do know is that it will change the world, and it will make us able to do things and approach technology in an unprecedented fashion. That’s the whole point of Radical AI.
Could you maybe give a little bit more of a range of what those things are? I’m conscious that you just said we don’t know exactly how the world will look given certain breakthroughs. Obviously, nobody would have immediately known—the classic example from smartphone glass, right?—that people were going to have Uber and DoorDash as soon as we could touch the screen. So I certainly get that it’s going to take time for these things to play out.
I get the sense that, obviously, a room-temperature superconductor is one. How many of these canonical things do we have where we’re like, if we could just get this property that we don’t have today, it would be a massive step-change unlock? I guess I have a couple layers to this question. What is the range of those things? What specific examples are kind of forefront in mind?
A layer down is: How precisely do those get articulated today? Is it like this thing needs to be 4% lighter and then we cross some threshold, or is it like lighter is better in general? Then I always think back to the Elon Musk “mass begets mass.” I’m sure he didn’t coin that, but I heard him go on about that.
So I’m also interested in general mental models. If people just want to think about the material world around them in a more sophisticated way, what other things are there along the lines of “mass begets mass,” where you can just develop your intuition for why materials matter and how materials getting better translates into material progress? There’s a lot there, but take your time.
So there are a lot of ways to think about stepwise innovation, right? There are a couple of classes of materials that I would say everyone is attuned to: RTAP superconductors, high-entropy alloys, ultra-high-temperature ceramics, CMCs, and semiconductor-based materials. People have been thinking about these industries and the novel materials that will drive fundamental impact in all of them.
And really, if you notice from that answer, it comes down to the end application. Jorge alluded to this a little bit earlier: We don’t make materials in a vacuum at Radical AI, because if we do, then we actually don’t know the problems that we’re trying to solve. If you’re an academic research group or a national lab or a nonprofit, that could be a great approach to take.
We want to make materials that leave a fundamental impact on the human race by directly solving a huge problem that exists in industry today or helping enable a new industry in the future. And so, if you’re going to take that perspective, a lot of the things that we think we can do, like the examples I just gave, are grounded in, well, what do you want to come from that? I’ll give you a perfect example of the versatility to your second question: How versatile are these materials?
Something like a high-entropy alloy, which is a core area of focus for Radical AI today, has really incredible performance across mechanical and thermal properties, and in its ability to blend them together. Meaning you get a really high strength-to-weight ratio while still being effective at high temperature. It’s not one or the other, and that’s important.
These high-entropy alloys can be used in something like a hypersonic application, where you are flying at speeds above Mach 5, very high temperature ranges, very high pressure, dealing with things like oxidation inside the atmosphere that are really challenging for a material to solve. We have struggled to solve that problem as a nation today. And then the exact same material can actually be used in a nuclear reactor—a nuclear fusion reactor—that is actually getting bombarded with radiation from this power source.
Our current materials today, like tungsten, can’t withstand that radiation and erosion over long periods of time. And so these high-entropy alloys are really one of the options. Amorphous metal being another option. That’s a subcategory of a high-entropy alloy that can try to solve these problems.
The versatility is very big because the properties are so desirable. These are, again, what we call enabling technologies. These are unlocking not a 5% improvement game, or even a 50% improvement game; these are 100× better, 50× better than what exists today. This is what we get really excited about.
To go to your third question and think about the way that these can extrapolate out and get very defined, every industry in the world can define a materials problem. If you go into the athletics and clothing industries, they're looking for novel materials to wick away sweat and reduce odor. There are silver nanoparticles inside lululemon workout shirts today, so that's a real application, and they're trying to optimize there.
If you go to the exact other side of the spectrum and look at advanced hypersonic aircraft or space travel, they're looking for highly advanced materials that can block extreme heat, block radiation, and withstand immense pressure, all while doing that in an atmosphere that is quite corrosive to the material system. Different problems, different defining mechanisms for them, but it's really important to be able to straddle all those constraints.
In summary, the real core factor that you have to think about with optimization is: What are the most important properties? When we're defining materials today, rarely is it one property. There are probably a few examples you could think about, if you wanted to, that would be a single property.
We always get asked by customers, "Yeah, I need it to have this mechanical performance in this environment, at this temperature and pressure, at this cost, with this supply chain." For example, inside the defense or aerospace industry, you're not just optimizing for thermal expansion. You're not just optimizing for resistance to oxidation. You're actually optimizing for the cost of that solution, and then for the supply chain to be able to employ that solution, given where these raw elements are coming from.
We're always seeing this optimization across a multitude of different properties. Because of that, there become so many different combinations, so many different industries, and so many different applications in those industries that materials like a high-entropy alloy can be useful.
In our opinion, there is an endless amount of problems to solve inside materials science. You can go very specific, like optimizing the silver nanoparticles in a workout shirt, or you can go very large, like building radiation-blocking shields so that, when we land on the surface of Mars, radiation doesn't become a problem for communications and electronics. Everything in between is very, very dense, and that's why we think it's one of the biggest opportunities and markets in the world.
So maybe that's a great jumping-off point to talk about the flywheel. If a customer comes to you with one of these problems—"We want to make a Mach 5 aircraft"—to do that, it's going to have to be no heavier, or maybe even lighter, and stronger, and able to handle heat, and maybe able to handle radiation. Who knows? Maybe this thing's going out into the outer reaches of space.
Take me through the loop that you guys get into to work your way up to actually achieving that. Maybe starting also with just how often there is that Steve Jobs opportunity where something maybe does exist out there: How would you even go and search? And then, if it does exist, where does that allow you to shortcut into the flywheel versus if no such thing exists and you have to come up with your hypothesis of never-before-seen materials and work from that point?
Yeah. So the flywheel consists, at a high level, of 2 different buckets. You have an AI engine on one end and then this fully robotic, self-driving lab on the other, and those work together. I'll let Jorge start on the AI front and what that looks like. It gets into a lot of how we do property-driven optimization, directly to your question, and we can talk about the experiments after that.
When you think of the way scientists do science today, there's really a lot of different things that come into the picture. You're looking at images out of a microscope. You're reading a bunch of papers. You're using your intuition from previous experiments. You're maybe doing some computational modeling. That's really all of the tools in the toolset that a scientist uses to eventually come to a confidence level where they want to go out and run an experiment.
We try to figure out how to create an AI engine that incorporates all of that—all of the different modalities from a data perspective that come into play there, all of the different architectures that we need to use at a modeling level, and this sort of scientific-intuition aspect, which is another part of it that we try to capture.
We bring all of those things together into one piece of technology that we then throw at these problems. The way we do it, as Joseph said, is by doing what we call property-driven optimization. We start with the properties that we're looking for. Those typically come from a customer.
It's the things that you mentioned: "I would like to use these sets of elements. I would like to have it hit this oxidation level. I would like to have it hit this ductility or this strength." Then we give all of that to the AI engine, prompt it, give it some input as well, and let it go to work. We let it design experiments that we think are going to be successful against those properties.
That's when it hands it off to the robotic lab, which I'll let Joseph talk about.
This is where, to Jorge's point, it's really important. I love the way he said, "How do you use the tools at your disposal?" What is an experimental scientist's job? You come up with a bunch of hypotheses. You think about the ways you're going to test them, like Jorge just mentioned. All of that's happening on the AI side.
Then you go and run experiments. You make things in the lab. You test the things that you made in the lab. You look at all the results. That is all happening inside our self-driving lab today.
Our lab will receive, through our operating system, a design-of-experiment approach: Here's the composition, here's how we're going to make it, here's what we're looking for, and it will start executing on said experiment. We will make a material. That material will come out. It'll go through material characterization.
We do a bunch of different characterization techniques to understand what we made and really what it looks like, which is important to microstructure and therefore mechanical properties. Then we'll actually run property-driven tests. How strong is the material? What is the stress-strain curve of the material? How does it handle under creep or crack propagation? How hard is the material? We will do those property tests inside the lab as well.
When we collect all of this data from both the material characterization and the property characterization, we actually use machine learning to analyze all this data in real time. We take this information in and train these submodels to learn how to analyze just like a scientist would.
We start to build this intuition, start to build this understanding: If I'm optimizing for strength, then I need to pay attention to grain growth and microstructure formation. I'm going to look at that evolving in a scanning electron microscope or in a tensile test that I actually do on said alloy.
As we analyze all this information, we pull out the insight from that and then feed that, as an active-learning approach, back to the AI engine. You can think about it as a scientist taking back the results and saying, "Oh, that was the outcome of experiment 1. Let me now run experiment 2, updating or tweaking that process based on the results that I just saw."
If you think about what I just described to you, that's what a human scientist does.
It's not some insane new change to the way we do the scientific method, but we are changing the way we extract information from the scientific method. When I ran that process, I did all of that in a serial approach. I looked at 1 XRD scan, 1 SEM scan, and 1 tensile test, thought about those, and then went back and saw the next thing. I might have remembered there was a paper I read, and I might have looked into a patent database—or I might not.
When our AI scientist looks at all that information, it is analyzing and comparing those results to millions of XRD scans that it might be pulling from publications. It might have past experimental results in our lab, and it might have context on other experimental datasets that exist today. Then it uses all of that to make a new, informed hypothesis that operates on a dimension that, candidly, the human brain just cannot handle.
This is where the power of the flywheel becomes so impactful. It's not just an AI scientist or an AI engine, as I described, and it's not just a self-driving lab, as I just described. It's the combination of these techniques together that builds this iterative approach to not only developing and designing, but then making and testing novel materials. That is how we can actually take a step forward in discovery and, in the future, scale these materials to larger systems.
I want to dig into this on a bunch of different levels. Maybe for starters, when you're developing the AIs, you mentioned submodels. I'd be interested to understand the range of models that you have. We're seeing increasingly more things posted from GPT-5 on Twitter that are like, "This thing helped me think through a frontier science problem." So I wonder to what degree those sorts of things are in this flow.
Obviously, there are going to be more specialized models. I can imagine a lot of different approaches. We've seen many things that are graph neural network structures, with some sort of radius around particular atoms: what are they connected to, and how do you model things in a more literally physical way? I've also seen interesting structures where it's a bunch of different things, and then you mask out different tokens and ask the model to fill them in, right?
You could have the chemical composition, certain properties, and how we synthesize it, and mask those things out to train the model to fill in those details. Can you outline the family of different kinds of models that you're using, and what the inputs and outputs are for those?
We use a bunch of different models. We use GNNs, like you mentioned, to do atomistic modeling. We have generative models, and we use language models, including the GPT family of models, to do everything from understanding and extracting information from scientific literature to even generating recommendations. We use computer vision in the lab, and we use a bunch of different things in a bunch of different ways.
Obviously, the inputs and outputs of all those models are slightly different, depending on the use case and the model itself. But I think the core thing for us is that there are so many different tasks and things that have to be done from a materials science perspective that we actually have to use different architectures and different models in different ways. I'm happy to go into more detail on any of those.
Yeah, please do. I'm not even sure exactly what you would consider to be the most core task or few tasks. Obviously, in anything, at least in my experience, there's always 1 or a very small number of tasks where you're like, "If we could get this working super well, we know we'd be successful from there." Then there's a lot of auxiliary things where you're more confident they're going to work, and it's more about efficiency and ironing out details.
What are the really core ones that are the biggest challenges? What are the inputs and outputs of those?
The one where we're most confident it will work is on the computational side. That's using something like a GNN. In the materials science world, we call them MLIPs. There are a lot of MLIPs. They're all fairly successful in different ways, and they work in different manners. We know that's a core part of the strategy. We use them, and a lot of people have been using them for a long time.
But where we would want to place our bets, where we really think there's going to be a big differentiator, is actually on the language side. We think that what we're trying to do is understand the vast base of scientific knowledge that's out there and then reason about that, incorporating different modalities and different tools in order to come to good conclusions. That allows us to generate hypotheses, design new experiments, send those down into the lab, learn from the successes and failures of those experiments, and iterate and adapt over time in order to improve the recommendations that we send down there.
We do think that, at the core, language models are really what's going to drive a lot of this scientific intuition—the reasoning about what makes a good experiment and what makes a not-so-great experiment over the long run. Where we think the other models will come into play is to bolster those models.
You can imagine a GNN being tapped by an LLM, for example, in order to run a computational workload because it just needs to reaffirm some uncertainty or fill in a gap around some knowledge that it doesn't quite have, even though it's scanned thousands of papers. The same thing could be said for some analysis coming out of the lab, right? We may want to look at a bunch of SEM images and do some computer vision in order to update our intuition and our understanding of what kinds of experiments we're running in the lab and how they're performing at the microstructure level.
That's how we use all the different models. But again, if we were to say there's 1 thing we have to nail, it's this sort of replication of scientific intuition and the designing of what would ultimately become an experiment.
As an aside, what's the best anyone has ever been at human intuition on this? What is the ceiling of human ability? People are very obviously enamored with John von Neumann. Is there a John von Neumann of this?
I also remember this story—I don't know if it's even true—that the structure of benzene originally came to someone in a dream. They envisioned a snake eating its tail or something, and they were like, "Oh, it's a ring structure." We've moved a lot past that, and we've got a lot more concrete knowledge to base our intuitions on. But are there people in the world who are just an obvious cut above everybody else in terms of intuiting what material is going to have what property, or how to make it successfully when nobody else knows how to do it? What does that human landscape look like?
It's a great question. The answer, I think, is yes. One of the best examples I could point to is actually our third co-founder, Gerbrand Ceder. He wouldn't say this, but I think we could say that he's probably one of the world's leading materials scientists—top 3, if not the very best.
When you speak to him, you just get this feeling: this guy knows things that the rest of us don't know. He can understand things and connect dots that are difficult for the average person, even a great scientist. So absolutely, there are people who have this intuition.
The funny thing about Gerbrand is that when you talk to him and really probe where this intuition comes from, he breaks it down pretty simply for you. He says, "Look, I've been doing this for 40 years, and so I have run a bunch of experiments over my career. I've had a bunch of students run those experiments, and I've learned alongside them. I've read papers from all of my colleagues and the experiments that they've run."
This scientific intuition—yes, I'm sure there's a part of it that is innate, and there are some people who are just talented at this stuff and can become better than the rest of us—but a lot of it is actually learned. It is learned from going out and doing these things and updating your understanding of what works and what doesn't work. That really is the scientific intuition that we're trying to capture.
Joseph, you probably have some thoughts on this as well.
Well said. There's so much information in the scientific process that a scientist will use that comes from learned experience. Whether that's at the beginning of your career or even taking something you've worked on for 20, 30, or 40 years and applying that to a new industry, you have to ask: How do they like to test that material? What properties do they care about? Is there a different way that they need to consider manufacturing?
I spent the better part of my undergraduate and early PhD days working on 2D materials for optoelectronic sensors. Then I ended up getting a fellowship at the Army Research Lab and went there to work on neuromorphic computing. To your point about von Neumann, I was working on getting rid of that architecture inside semiconductors.
The testing requirements—the saturation current, the on-off ratio, the conductivity, the resistance, and the memory states that we were building—all of those were learned things. After a year at ARL, I was better than probably 99% of other scientists at thinking about how to design new semiconductor materials for neuromorphic computing that directly used ion-gated channels.
There was a group of people who were experts in that as well, all from intuition and experience. But exactly to Jorge's point, it comes from this learned process through experimental research. And that is very, very hard to capture. Actually, we don't capture it.
Of the year of work I did at ARL, 90% of the stuff I did didn't work, quote-unquote. Those negative results we always talk about—but there is no recollection of that. There is no capture of that. There's no data set that exists of, “Here are all the e-beam, PVD, and ALD experiments Joseph ran over the last 12 months.” There is just 1 resulting material, which we pushed out and shared with people.
And so, if you can't capture that, and the only place it exists is up here, and there's actually no framework in place today to build this intuition except with exactly what Jorge just really walked through, that's what gets us so excited. So, to Jorge's point around here, that's why, 1, he is such an important part of what we're doing and why, when we met him, we said, “We have to build this company together.” But, 2, he has this really unique ability. How do you replicate that? How do you build scientific intelligence?
Intelligence that is not just general, like, “I know the English language, and I know what science is, and I know the periodic table and electron orbitals and the way we actually make new elements or compositions,” but, “How do I go deeper than that and build this intuition, this scientific intelligence around, ‘Yeah, I'm going to use these 4 elements specifically because I am targeting an X property, and that property requires them, and I just know that’?” How do you teach an AI that? That's quite complex to do.
Yeah. And I think there's sort of 2 layers of this. There's the easy aspect of it, which is, “Okay, yeah, science is really hard to do. It's super slow. It's very manual. We don't record it, and so we should automate that, and we should learn how to do that with AI.” That's not a big leap for people to take. That's an obvious thing, and we obviously do that.
I think the bigger leap here, and really the one that we're after, is this understanding that there is an exascale set of data that's out there but is virtually impossible to capture—at least, it has been in the past. The only way to get at it is through a really, really inefficient process, which is being a 40-year academic. That's why we are bottlenecked by all these different things, and that's why we don't have materials that solve some of the biggest problems, right?
Not only are you stuck doing it over a 4- or 5-decade time span, but in that 4- to 5-decade time span, you can really only focus on 1 area, maybe 2 areas. And so you're just limited by the amount of things that you can do, on top of it being an inefficient process. And yet, like Joseph talked about, we know the data is out there.
A big part of our belief is that if we can build a system that can actually tap into that data and then figure out how to use it in a good and efficient way, we can couple that with all of the advancements that we've had in machine learning to really drive at this problem that takes 40 to 50 years and shrink that down so that it's under a decade. And, of course, that will translate to solving materials problems in under a decade as well. That's a big part of our belief.
For the people that have this intuition, how language-grounded is it? How much can they articulate it versus how much does it live in a fundamentally different space than language?
Because I could imagine somebody might say, “I can't really tell you why I think this is going to work. I just think that if we do it this way, it might work.” Or you could have, obviously, a very mechanistic and clearly articulable theory of why it's going to work.
In my brief chemistry experience, it was honestly a little bit more of the former. I had a limited time there and a limited view, but the 1 reaction that I worked on developing for a year—we never really had a clear mechanism. It was a palladium-catalyzed hydrocarbon oxygenation at room temperature or slightly warmer, under mild conditions.
Anyway, the point was, we never really knew what was going on at the reaction center. We had a vague sense of that, and so it was very much, “Maybe we'll try this, maybe we'll try that,” but without any precise articulation. How common is that, and what does that imply for whether GPT is going to be able to chain-of-thought its way into making these things happen versus whether there's just a different modality of data that has to be collected and learned from?
Yeah, you took 1 of the answers there, which is that modalities of data are very important. We'll talk about how Jorge can go into deep detail around where that comes in the engine and why it's very, very important.
Language is an important piece. It's not just language, as you just alluded to. Scientists learn a lot from the different data streams that they're pulling inside their experimental process.
There's 1 other thing that is a form of data but is not maybe what you call structured data from an experimental tool or from an output in the lab, and that is this intuition, for lack of a better word, or experience you have in running experiments.
I'll give you a perfect example. When I was working on these 2D semiconductor materials, we would always look for the transparency of the film. The technical reason why, which we later came to figure out, was that they were monolayer 2D TMD materials, not multilayer 2D TMD materials. And that drove way better conductivity, etc. We didn't know that up front.
You could see this in an optical microscope. We weren't looking for this in an optical microscope. We weren't taking pictures of transparency and asking, “What are the changes in transparency?” until we were racking our brains around, “Why are we getting such better conductivity in these materials? They're the same material, and we're making them the same way.”
It actually came from an expert in the 2D materials field who said, “Oh, those are monolayer films.” And we were like, “What? We know what monolayer films are. We worked on monolayer films for 3 years, but never in that setting had we thought about looking for monolayer-based films.”
This randomness, this accidental procedure that you hear all about in famous science stories, is really important to capture in this negative-result-based process. That is almost like a separate form of data where we deeply believe it's not just making the things that work. It's not just using the intuition from multimodal data sets that actually allow you to predict new things.
It is looking at, exploring, and capturing the 1-off, random, unexpected outcomes that might have a key insight in them that a human scientist will look right past, but an AI engine would never look past. There's a very important piece of what is, I guess, intuition as well. But the multimodal data is also equally as important, and there's a huge approach we take on that side.
Yeah. The multimodal aspect of it is enormous because, again, it gets at this idea that the data is out there. It's just in so many different ways and forms and shapes, and it's so unorganized. If you can bring a little bit of order to it and a little bit of cleanliness to it, then you can start to really enrich your representations and get to better predictions and generations over time.
But I do want to go back to what Joseph talked about because this is, I think, something that's actually quite hard to do with machine learning. When you have models that effectively learn from a distribution of data, it becomes very, very challenging to have them do things that are out of distribution, which is effectively what Joseph is talking about.
How do you go about this sort of randomness when the entire purpose of the model is to learn from a distribution where those random things are generally just ignored by the model? It sort of says, “Now, that's out of distribution, and I'm not going to veer in that direction.”
And so we think about being very purposeful with that. We think about saying, “Look, in science, it's a little bit different. We have to do the normal thing that we always do, and we have to learn from a distribution of data. If we can improve that data over time and better and better enrich the modalities of data that come into it, that will be good.”
That's not entirely true. We also have to design systems that are effectively primed toward novelty and surprise, because that is really where some of the biggest discoveries in science come from. They come from some weird corner of the chemical universe that everyone thought was kind of ugly and didn't want to look at, and lo and behold, here's this amazing scientific discovery. Or pure accident, sometimes.
We have to design systems that not only learn really well from good distributions of data, but also figure out how to basically go hunch hunting: how to develop hunches and then say, “You know what? I'm going to go out and pursue that thing. I'm actually going to go deeper here.”
There's a bunch of different ways we think about doing that. You can do some sort of randomization or diversification, which we do through Bayesian approaches and Monte Carlo and stuff like that. You can do noise sampling. Active learning is a huge part of this for us.
But then there's another side of it, which is, again, the underlying design of the system itself really has to be geared toward capturing scientific intuition at the end of the day, because that is such an important part of it. And so that's what would drive a scientist to pursue something that is a little bit weird but, just like Joseph said, at 10 or 20 times or even 100 times the scale.
That’s really important to us. We have a bunch of techniques that we use for that as well: reinforcement learning and things like that. Human-in-the-loop is a big part of it as well, in the early days for us, in capturing this uncapturable data like Joseph was talking about too. So there are a bunch of different techniques that we think are really important that go beyond just the normal, “Hey, train a model on a good distribution of data.”
Yeah, that’s really fascinating. I mean, that’s also basically the core problem in all of AI, I would say, in some sense, right? For better or worse, are we going to take these things past the point where they model the past and into the realm where they can figure out the future? I’d love to hear more about that in any way you want to go deeper on it, but I want to try to map a little bit of what you have experimented with and learned onto the fully general problem.
How do you see base models, so to speak, versus RLHF versus just pure success-signal RL fitting into that? Because it seems like people are finding that middle ground of RLHF is pretty good for the human assistant. There’s part of me that is concerned with big-picture, let’s-say-AI-safety questions. As much as we can get the things to imitate us, I’d really like that, because I do worry about, if we just go super hard on RL and give nothing but reward signal, what are these things going to learn, and what reward hacks might we find ourselves encountering?
But the flip side, I would guess, is that if you just imitate humans too much or just get the thumbs-up from the human evaluator too much, you maybe sand down a lot of these rough edges. That certainly seems to be what we see from the chatbots, right? People are like, “The base model is more creative.” I like that more in some ways, especially if I want to get outside-the-box ideas.
RLHF kind of gets boring, and then the RL gets interesting again, but sometimes in kind of scary ways, depending on exactly what sort of problem you’ve put in front of it. How would you compare and contrast your learning against that stylized story from the—
From the chatbot space?
You know, I think a lot of that is definitely true. For us, a big part of it is: do you have good evals? If you can develop good evals, a lot of it is undiscovered, and a lot of it is true research on the ML side—trying to understand how to bring these things together and move them in a direction that you want over time, with, like you said, not forgetting about safety and some of the other things that come into play as well. Obviously, in materials science and chemistry, that is an important thing.
Evals are hugely important, and for us there honestly is no better eval than the lab. I mean, that is as ground truth as it gets, right? If you’re looking for a good benchmark, a scientific lab is pretty good as a benchmark. It’s a really good way to understand: Did my model actually successfully do what it was supposed to do, and did it do so in a way that didn’t come at the expense or sacrifice of some of these other key things that I really wanted to hold?
That’s a really important part of it for us, and it brings us back to some of these active-learning loops that we were talking about earlier. How do we smartly incorporate that so that we can update these models, not just at the weight level? Because it needs to go beyond that. It needs to be a much more dynamic process than, “Oh, we’re just going to retrain the model.”
There really needs to be a sort of surrogate or agentic process that sits on top of it and says, “Yeah, I’m going to use my model to get things going, and hopefully we’re going to improve that model over the long run. But in the meantime, I need to be able to be adaptive and responsive to the new learnings that are coming back out of, again, what is a best-in-class benchmark or best-in-class eval, aka the scientific lab.”
That’s a big part of it for us. I think the other thing that I would bring up is that we do truly believe in trying to understand how these models work. We’re making a ton of investments in mechanistic interpretability, for example, because we think that’s going to drive a lot of performance and a lot of gain and enable the correct steering of these models toward the objectives that we have.
Again, if you think about what Radical AI is trying to do and the core purpose of the company, it’s to solve really big materials problems for people. If our models are just not going in the direction we need them to, we need to be able to move them in that direction. If they’re going against the expectations or the wishes of an end application or an end customer, then we need to be able to curtail those things too.
It goes a lot further for us than just, “Well, give it to the model and let it do its thing. It’s okay if it’s a black box, and we’ll just learn how to improve that black box over time.” We really need to be able to use it as a tool over the long run if we want to solve some of these big, hard problems.
Especially going back to one of the things that you brought up earlier, this is the nexus, or this is one of the big problems in AI. Coupling on top of that the fact that we don’t have historical data sets to tap into—like, even in the English language, you can digitize old books and bring them in and learn about how Shakespeare wrote and build intuition or knowledge off of that.
We don’t have Newton’s notebook. If we did, there’d be scratches of all the things that they tried to do. There wouldn’t be a written text that we could directly learn from and transcribe. Coupling on top of that problem is this idea that not only do we have to know where they need to go, but what can we bring in so that they know the different directions that they can go?
We don’t have a different text for the way we write books today and the way that Shakespeare wrote books back in the day. So if you don’t have that intuition, we don’t even have that understanding. It becomes really hard to think about how to index that.
I think you’ve seen this in software, right? A lot of the best software is private. A lot of the coding agents use what’s open source to train on, and that’s gotten really good because there is an open-source community that pushes really good software out. So it can learn what really good software looks like and then utilize that data in an effective way.
There is no open-source science community. There’s open-source science from publications. I mean, there is no multimillion-data-point experimental set of, “Oh, these are all the scientists who have opted into publishing their lab notebooks.” We don’t even have that, let alone structured data like pulling a GitHub repo and being able to learn from, “Okay, here’s how professional engineers write code; here’s how it’s organized.” We can’t do that.
If you layer that on top of exactly what Jorge just talked about, it is really important to know which way our models are going so that we can get in there and bring a scientist in to try to drive them to better discoveries. Exactly like Jorge said, our objective is better discoveries that fundamentally impact humanity.
It’s not just about producing or recommending new things that are marginally improved. That’s not the north star for the company.
So I definitely want to get into the mechanistic side in a second. Shout-out to Eric Ho, CEO of Goodfire, who originally introduced us.
But maybe first, on data: you’ve said a couple of times that data is the limiting factor, and the big 3—the holy trinity—of AI progress is data, compute, and algorithms. I was struck, in looking at some of the specialized models, by just how small they are: 30-some million parameters in some cases, orders of magnitude smaller—4 orders of magnitude smaller than some of the really big ones, maybe even 5. That’s a lot of orders of magnitude.
That would reflect, I guess, that obviously data must be scarce; otherwise, I would assume you’d have a bigger model. Where does that future data come from? I mean, you said we don’t have the history.
I was kind of wondering, and I think of an analogy a little bit also to the computer-use agents. I’ve remarked many times—and still nobody has come forward to pay me to watch me use my computer—so I’m like, why is nobody just offering me $100 a month or whatever to install some observer on my computer? There could be different reasons: it’s too messy or whatever; it’s maybe not worth the trouble.
The sources of data I had in mind for you guys were, one, literally going and paying grad students to let you digitize their notebooks. I don’t know what the going rate for that would be. Grad students would probably be interested in the money, I would think. I don’t know if there are IP concerns there or whatever.
Has that happened? Why not? Then there’s just the straight simulation data, which I understand is at the core of a lot of—or at least filling in gaps for some of—these small, kind of force-prediction models: what’s the next time step in the evolution of a collection of atoms kind of model.
That would be one way you could invest heavily in compute and get data, but I’m not sure what gaps that may leave or why that falls short. Then, obviously, the lab is the final stop. Maybe I’m even missing categories.
What do you think the portfolio is of contributions from those different types of data going forward? Are you ultimately thinking, “Our progress is going to be rate-limited by how many of these autonomous factories we can build and how many experiments we can run per day,” or is there any other way to convert financial resources into progress that doesn’t go through that sort of factory buildout?
Yeah, I mean, data is always going to be the limiting factor.
Certainly, in our field, where we're so data-limited, but I would even argue, generally speaking, across all areas of machine learning, data is going to be your biggest driver of success. If you can figure out a way to get over the hurdle on the data side, then you're going to be in a good spot.
Yeah, improvements in models will help, and certainly access to compute and being able to do more with compute is a big part of it, too. But if you asked anyone to choose between each of those 3, they would almost universally say, “Yeah, give me data. Give me data at the end of the day.” And the same would be said for us. Joseph can probably answer the mix of data and where we think that will come into play for us.
Yep. It's even harder than just a mix of data, too, because it's all unstructured. It's not even labeled, actually. If we could get access to every lab notebook in North America today, that would be great. We'll take it, by the way, if anyone's offering.
Yeah, we will. We'll take it because, to your exact point, we'll always take data, and then we probably have to go hire an entire team to sit down with everyone who wrote those notes and understand what they wrote and what it's in reference to. Even generalized, simple things like running an X-ray diffraction tool—the peaks mean things, and people might interpret them differently for their research than we interpret them. When we're doing amorphous-based materials, we're using XRD as a confirmation that there is no crystal structure. We're actually not even looking for relevant data.
So if you gave that to someone who had never worked in amorphous materials, they'd probably be like, “I don't—there's no peaks here. I don't get it,” right? Like, “This is irrelevant to me. I can't train a model on learning XRD.” It's like, but you can train one on learning amorphous behavior because there isn't, of course, a crystal structure.
So maybe that's a very technical example, but it's very important. Even if we had access to all of this, it's not labeled, it's not organized, it's not structured, and it's interpreted in wildly different ways based on the discipline that you're working within. Most of it is only a piece of the data. For example, I used to always write down the timestamp of halfway through my reaction because it didn't matter when I started, and I would always wind down the temperature. I would turn the temperature knob off at halfway, and then I couldn't control the cooling. So however long it took to cool to the end of the experiment was irrelevant.
So I gave you a timestamp: 1 minute, 1 minute and 6 seconds, 1 minute 4 seconds, 58 seconds. What does that mean for us doing the same experiments in our lab? Those might be relevant metrics; they might not. If I sit there and tell you, “Oh, that's the halfway point, and I cut the temperature exactly at that point,” well, now, yes, of course, that data becomes more relevant to what we're trying to do.
But this challenge in materials is not only that it's multimodal, not only that it's not captured anywhere, not only that it's really important to, as Jorge said, establish ground-truth benchmarks—we don't even have it structured and annotated to be able to use it, even if all 3 of those things I just mentioned existed. And so this data problem, I think, will actually spur an entire industry. We have a lot of ideas on the way to get around that, which we can dive into. From a scientist's side, understanding the data output I produce as a materials scientist is, in and of itself, quite a great challenge.
Yeah. On the computer-use side, I think maybe one of the reasons that it hasn't been done is sort of analogous to what you're describing, which is that at any given time, it's not immediately clear what I'm trying to even do as I'm using my computer. So there would have to be a sort of—or probably would need to be—you could guess, obviously, but you would be greatly benefited if you could also capture the internal monologue that's going on when the person is like—
Exactly.
—going tab to tab or whatever. What is it that you're even trying to do here? So I take the point that that is generally the case, and I can think back to my own not-super-well-documented, super-tidy, or thorough notebook notes myself, too.
And how do you judge quality there, right? If my grandmother records her use of a computer and you record yours, I hope they're going to be different. I think they're going to be different. How do you delineate that? What sets the bar for a not-great computer user and a great computer user?
Now bring that into a subjective field of science. Was I a good scientist or a bad one? My boss and the people I used to work with think I was great. Other scientists in the field might think I was terrible. I would not have tried those experiments. I would have gone in a way different direction. I would have never used molybdenum. I would have used tungsten instead, or whatever the property is. So who sets that bar? Who determines how it moves?
How do you even begin to build it when there are so many different fields inside materials science? Every single bar or benchmark is different. There are these deep problems that come up in even just annotating the data that are really hard to answer.
And, like, in the computer example, I don't know how I would answer that. I would generally start to say, well, someone who's very proficient at moving around and understands where all their applications are and can get their task done quickly. But what if someone goes to email and opens a new email, pops it out into a new tab, goes somewhere, writes it in a Notes app, copies that, and then pastes it in Gmail's email box? It's like, okay, well, I wouldn't do that, right? That's not the most efficient way to write an email.
So how do you delineate that? There's probably someone who would argue, “Oh, I make too many errors when I send emails. I have to write it in Notes with Grammarly installed and check it first.” Who's right? So there's just so much ambiguity inside classifying, or I guess you could say labeling, the data, too. It just becomes really an exponential problem.
Yeah. I think what we're hitting on is that it's pretty abundantly clear that the highest-value data comes out of an experimental setting, not out of a lab.
And yet it's also where the biggest challenge lies, right? Just getting access to the data and doing it in a way that's orchestrated and clean is really, really challenging, not to mention expensive. That's why we don't have the data. We have not seen people really build these datasets out because it is so hard to do.
Actually, if you flip to the other side of the equation, the computational side, that's actually a lot more straightforward. We've seen a major push to develop datasets there. But our big thought is, look, that only solves 10% of the problem. You're not there. It's because of this intuition again: we know that the highest value is in the lab, anyway. It's just really hard to get to.
So yes, computational data is important and it's helpful to us, and we're going to do that, too. But we're kind of missing the real big picture here, which is: let's figure out the lab. Let's figure out how to get experimental data and incorporate it alongside the computational stuff that we're already getting pretty good at. That is really going to push things forward. That's really what we're trying to do here.
And if you do that, it's worth adding on, you actually control the problems I just brought up. It's no longer unstructured. It no longer lacks context. It's no longer unlabeled. All of those things become a reality.
So not only are we building the dataset, but we actually can use the dataset. Those 2 things are incredibly effective, in our opinion, at really developing new materials, to Jorge's point.
Yeah, that's really interesting. You mentioned earlier that the search space is vast. Could you maybe help me develop my intuition for the vastness of the search space? How many different dimensions are there, or are there sort of big branches in it that are just critical to understand—forks in the road as you're going down a thought process or, ultimately, a material development process?
And then, in terms of how much of that space you can effectively explore today as you build out the capability of lab automation, are there parts of it where we can't automate yet? That's still kind of like—we can't automate the plumbing in my 100-year-old house, right?
Again, I remember doing chemistry stuff. One of the things I observed in my grad student mentor was, like, this dude just knows how to do certain things in a very intuitive way that I don't know and that certainly, at that time, no robot had.
As you think about pushing the frontier of what you can automate experimentally, some of it is just horizontal buildout at some point, right, to do more throughput. But presumably, there are also frontiers of certain kinds of experiments or reactions or what have you that are just not yet automatable, that become automatable.
So again, take as long as you want. How would you describe the vastness of this space and what parts of it can we explore or can't we explore? What frontiers do you think you'll be pushing the automation forward to actually be able to get into and properly explore?
Yes, it is an incredibly vast search space. It is impossible. It is improbable to think you're going to test every single potential combination. We always refer, when we're talking about the company, to Eddington number, I believe it's called: 10^80 different observable atoms in the universe.
And that is the number we use for the potential combinations. If you go into something like alloys specifically, which is, of course, our area, there are 60 or so, give or take, elements on the periodic table that you're going to use in novel alloy development. When we talk about high-entropy alloys, we're doing 5- or 6-element systems. They're called multicomponent alloy systems. If you just take 5 elements and use that, it's like 5.5 million and change in potential combinations that you can develop. And that's if they're all equally balanced, meaning 20%, 20%, 20%, 20%, and 20% inside that alloy.
No alloy, to my knowledge, is that equally balanced. Within each individual element, of course, you have 0% to 100%, and then you can go into as many decimal places as you want to optimize. It is massive. The way we always describe it is a single grain of sand on an incredibly large beach, a beach that's the size of the Earth. That's what you're really trying to identify and find when you're looking for a new alloy or a new material.
There are spaces that can be a little bit bigger than that. Proteins are one to call out that everyone tracks and everyone thinks about how big they are. I think it's something like 10^130, that the search space is there. That's probably just a single grain across multiple Earths or multiple universes, even. Maybe that's huge. We think that's a huge advantage for materials. It's actually more likely that we'll hopefully find something in our beach versus the bigger beach, but it is really big.
When we think about how we take that big number and start testing things to get to the point of discovery, the first step is indexing the search space. All of that sits within the AI engine: how do we actually do property-driven optimization to weed down and filter out the best materials that we want to try?
When we think about the challenge on the experimental side, it's really the material tools themselves. It is not robotics or different automation systems, like our RoverTrack, which is the materials-handling system that we build. It's the tooling themselves that we need to conduct scientific experiments. Those companies have been around for a very long time—some 25, some 50, some 75, and some 100 years. They have been human-operated for decades, so they're built entirely around a human scientist, with no automation in consideration.
The ability to then drive those tools with software is, in some cases, nonexistent. In the cases where it does exist, there's some type of interface that's very, very limited in the actuation you can drive. You might be able to get outputs from the tool, but you can't control the turbo pump, the heat source, or the vacuum chamber, or the pressure inside the chamber, therefore rendering it ineffective from an automation standpoint.
A lot of our time on the automation front is spent figuring out how to take existing tooling infrastructure and make it autonomous—make the tool itself able to operate in an autonomous fashion. We do a lot of custom development ourselves. We have multiple tools in the lab that are entirely custom-designed around the materials tool to make it automated. We work with companies, and we have multiple ongoing conversations with companies where we want to build a custom tool together. That's taking that outside automation and bringing it into the real engineering and design of the tool.
There are some tools that we think will probably be pushed out because of their lack of automation, or rebuilt from the ground up by companies, because the way they were built was entirely human-driven. We have examples in our lab today where they're built in a very mechanical way, meaning humans flip switches, move knobs, and turn dials. We think those things will be replaced over the next 10 years as these new tools come out.
From an experimental-setting standpoint, a lot of our effort is spent taking existing tools today and making them automated. Remember, Radical AI is not a materials-tooling company. We don't want to sell XRDs; we want to sell materials. We work with companies and vendors and do a lot of that work ourselves. That's how we do it on the experimental side.
On the AI side, the search question is challenging because we do not do a combinatorial-based approach. We really don't think combinatorial-based science has driven enough value. Again, we move toward this property-driven optimization that has an active-learning component, is using a bunch of different tools, and is generating new systems. Jorge can go into the details if you want.
The big aspect of it is definitely the active-learning side. Even in this sort of universe where we have dozens of elements that we might want to use, how do we go about doing that? Some of that just goes away, right? For example, if I really, really care about weight, I'm not going to use aluminum. You could just remove aluminum.
Some of these things do naturally start to shrink down. You still end up with a very, very, very large number. This is why you need to have multiple machine-learning models that can solve multiple problems all along that trajectory, so that by the time you get to the active-learning loop, you're sending really good stuff down into the lab. Yes, you're learning how to make them in the right way, but hopefully the things that you send down to the lab are already good to begin with.
I guess one more brief follow-up on that, and then there are so many different interesting directions to go. Are there things right now where you have to either patch with a human? I know Emerald Cloud Lab, from what I understand, has certain aspects of the process where a person is told, “Go to this machine, pick up this file, move it to this other machine, and put it in,” and they're almost sort of Chinese-rooming it to a degree, with a person who doesn't necessarily even have to understand what's going on but is just following very point-level instructions.
Yep.
Are there things like that that you're bridging because of missing tools? Are there still certain kinds of things that you just can't get to—a purely automated process could do—that are outside the bounds of what you can bring into the closed loop as it exists today?
The answer to that is yes today, meaning there are some things we don't have fully automated. We do have scientists in the lab who help orchestrate experiments. But the answer is also no, in that while they end at that solution, that's our starting point. We are actually beginning our learning process from there.
Of all the tools we have in the lab, everyone starts at the place I just described: a human scientist comes in, shows how you operate it, runs it, and then it moves to fully autonomous, with no human inside running that tool specifically. For all of our tools, we will reach that place. If a tool cannot be automated, or we think the value is not there to make it automated, we will work with someone else to custom-build a tool specifically to our needs.
I don't mean custom-building the science part of the tool. We usually don't touch the science part of a tool. How the tool operates—we don't get into that. We are touching the loading chamber, the control of the vacuum, the control of when the laser turns on and when it turns off, the shuttling of samples, and the prevention of contamination if there are multiple samples inside a chamber. Those are all mechanically driven, optimization-focused—or, I guess you could say, autonomy-focused—functions of the tool.
We are not redesigning how to do X-ray diffraction. That's not what we do as a company. Again, we're not a materials-tools company. We are a materials-discovery and, therefore, materials-manufacturing company, and that's what we want to move toward.
The way that we can get there can look different ways, and we are going through that process with all of our tools today. But no, we don't have every single tool we've ever bought fully autonomous today. There are some tools that are going through that custom-design and retrofitting process right now that aren't fully automated yet.
Cool. You mentioned the idea of hopefully what you send down to the lab is good. A couple of questions on that one. Is there a sort of inference-time scaling law that you can tap into? If we run 10 times more experiments, or, in silico, use 10 times more compute, does that translate to a certain higher hit rate?
How does that relate to the sort of loss landscape, if you will, of materials? We have this desired property and this sort of constraint around inputs: generate me a bunch of protocols, or generate me the experiment—what am I going to do it with, or whatever, right? It could have a lot of different guesses. Is there a known relationship between using the models 10 times more, or using 10 times more compute, and getting a higher hit rate?
I'm not sure if my intuition should be that if I come up with a material that has pretty good but not quite the desired properties, does that mean I'm close enough that I can explore the very local space around that particular thing and expect that there's going to be something better that I can gradient-descent my way into?
Or I could also imagine that there could just be a ton of weird discontinuities in the material space, such that I might take one step in a particular direction and be in a totally different property regime.
Yeah. I mean, there’s always this explore-exploit conversation coming into play. Regardless of what stage you’re at, you could be at the purely atomistic level, where you’re talking about trillions and trillions of different atomic systems that you might want to look at, and you’re going to want to explore and exploit there, right? You’re going to want to explore the vastness of it, and then when you find something that feels pretty good, you’re going to want to narrow down and go deep there.
As you go down the ladder and eventually get yourself into the lab, it’s going to be the same question over and over again. A big part of it is that at each layer, you need to do that slightly differently, because the types of questions you’re trying to answer in order to understand where you should be exploiting are actually different. What you’re trying to answer at the atomistic level is very, very different from, “What should I be synthesizing in the lab?” Are they related? 100%. They are absolutely related, but they’re a different set of questions, and you need to have different techniques in order to answer those robustly and soundly.
That’s just something that we’re constantly doing over and over again in terms of inference time and all of that stuff. Generally, the way I like to think about it is: if you forget Radical AI, forget machine learning, forget automation, and just say, “Hey, as a scientist, would it make sense for me to think about a bunch of things in my head, come up with a bunch of hypotheses, narrow those down in my own head, maybe on some scratch paper, maybe do a little bit of modeling, and then walk into the lab? Or should I just say, ‘You know what, I’m going straight into the lab’?” Which one is likely to yield the best result? It’s probably the former.
I think that same thing holds and applies. Again, you do have to always keep in mind that science doesn’t always work in a logical, rational way. There is just a natural degree of randomness and counterintuition that has to come into play, and you cannot forget that. We have to design every single system with this idea in mind.
We have to abide by the general rules, but also understand that we need to be able to break them, too. Everything is designed around that core thought process. But, yeah, to answer your question, if we can offload some of these things into simulation, that naturally makes a lot of sense, and we will do that. You just have to do it in different ways at different stages along the process.
Can you talk a little bit more about the active learning? We’ve got these few-shot-prompting-type things. Again, there’s test-time training. There are a bunch of different paradigms where people are trying to push this, but it seems like, again, this is one of the core questions: Can we really get them over the hump to where they can take on this new data that is genuinely out of distribution and do an effective update on that?
What more can you tell us about how you’re getting that to work?
Yeah, we do active learning at every stage. At the atomistic level, we do active learning, and that’s different active learning from what we do down in the lab, but we do it all the time to improve our recommendations and improve the likelihood that, as we move into the next phase, it’s better and better.
At the atomistic level, we might be doing some active learning around uncertainty. We might be trying to generate these recommendations around forces, and if we’re not feeling pretty good, we might run some DFT and then update our understanding of how we’re making our predictions at the atomistic level. In the lab, that’s a totally different thing, but we’re still doing active learning.
We might run an experiment and realize, “Ah, you know what? Even though we scanned a bunch of literature, even though we’ve done all of these calculations, we didn’t actually realize that the pressure and temperature are just a little bit different today. And so we actually need to update it for that next experimental run.” We do active learning all of the time.
That’s mostly because science is not necessarily predictable to a perfect degree. It actually does change in a bunch of different ways. It comes back to this core thing around 2 main aspects. Number 1, we just don’t have the data. It’s out there, but we don’t really have it. The more of it we can get, and the more of it we can get in a good fashion, the more likely we are to move successfully through that process.
But, number 2, never forget that there always has to be some degree of randomness that we have to lean into in order to open ourselves up to greater discoveries. We do active learning all along the way. We think it’s just an important part of the process that naturally happens in science, and if we’re trying to replicate, improve, and upscale the way that we do science, we have to include that in it as well.
Yeah, that’s the core stuff these days. We’ve made it pretty far in terms of learning the past distribution. How to take a surprise on board and really do the right thing in response to that is central in all areas, in terms of interpretability.
A bit earlier, Joseph, you told the story of the grizzled vet who was like, “Oh, that’s a monolayer film you’ve got there. That’s why that’s happening.” To tell a little hacky story there, it seemed like you had gradient-descented your way into a heuristic. If I understand correctly, you had realized without a theory that when they look more transparent, that seems good, right? And it seems like a lot of models today, certainly thinking at the language-model level at least, kind of operate in this heuristic zone where they’ve learned the patterns, but they haven’t really grokked the underlying laws.
Then you’ve got this veteran who comes in and says, “Oh, well, I’ve grokked that situation, and I can actually tell you why your intuition has any value with something that’s rooted in the actual physics of the question.” I basically think that a lot of that is going on in models, as far as I can tell.
I always think back to the original grokking paper from OpenAI, where you get heuristics really fast. You can memorize the training set really fast. It takes orders of magnitude longer in that original experiment to actually do the grokking and really implement an algorithm that does modular division. It took interpretability to even know what was going on. Is it still just guessing, or does it have some actual right way that we can trust to solve this problem? Can we be confident that it will actually work for all unseen cases?
So where are you in the end? Where do you think the models are, and to what degree have you been able to do the interpretability at this point to have a sense for where the models are on these grokking curves? How much of the time are you getting stuff out that’s still heuristic and not really reflecting any real understanding, versus are there things that you can point to and say, “Demonstrably, we see real understanding here that is principled,” in any particular model or any particular area?
So I wouldn’t say that we are in a place where we can demonstrably point to a very clear and definitive sense of understanding. What I would say is that, depending on the models that we’re working with, you see varying levels of this. It’s almost like in the inverse, where you see signs of memorization, and you see signs of hallucination and that kind of stuff as well.
I think the fact that you can see the inverse means that there should be a way to get to the opposite end of it as well. This is why we are so bullish on interpretability really playing a big role here, because we think that there is a realistic world where these models can actually develop a rich representation of a materials problem, and that we can then learn how to exploit that model and steer it in the right direction. But I wouldn’t say that we have any proof or anything like that just yet.
Yeah. And the one thing I would add from a scientist’s perspective—and this is one of the reasons we got excited about working with Eric and the Goodfire team—is this potential ability to try to unearth things like this.
We actually—I don’t know if we know. Actually, I do know that we don’t know what—
I know that we know.
Yes, what we might want to look at to help impact our discovery. The monolayer film on an optical microscope and seeing the transparency of that—if you had asked me before, I would never have said to even look for that.
Out of the machine-learning world, at a base level, at a scientific level, we don’t even think we fully know which areas we can look into or what heuristics we can try to capture that might be impactful, and in some cases might not. There’s probably a world where the transparency of something does not matter to the resulting property that you’re going after.
So, therefore, it’s not effective, but maybe it’s transferable, right? Maybe there is something that you can actually learn from that process that comes into a different heuristic that you use. I think at a scientific level, as a human scientist, this is why Goodfire, interpretability, and this concept are so exciting to me: this ability for AI to maybe teach you what those heuristics are, for lack of a very simplistic phrase, and actually use those to drive different results that you would never be able to anticipate.
They’re numbers, their weights and biases. In a way, that’s really cool to me, because you remove the human bias, the human perspective from that, and it’s just math and the underlying models. That’s a really cool concept. I’m not sure we’ve had too much of that ever in science per se, maybe much more in theory, where you can have this optionality and do different types of simulations, but certainly limited and experimental.
I think that’s a really, really cool perspective to bring into it from a purely scientific perspective—not deep in the machine-learning weeds of what it means and how it impacts models, but from a scientist’s perspective of what the output of those models is and what interpretability gives me. I just think that’s such a cool concept that is very infinite in its exploration, and we believe, as J said, it will be a really important field for science at large over the next 2 to 5 years.
Have there been any Move 37s in the materials science domain? Or is the alternative that there’s a lot of grinding work done a lot faster and at a scale that just couldn’t happen otherwise? Have there been any? Would you be confident if there were one? Would you immediately know it if there was one?
If there haven’t been any yet, is that a sort of 2- to 5-year timeline in your mind, as you just scale out data? Eventually, you think you’ll get there? What’s the outlook from materials science to Move 37?
Great question. I think we’re moving toward that. I’m not sure if there has been an earth-shattering one yet. I mean, like an earth-shattering superconductor, just to use an example we talked about earlier, where everyone in the world is going to realize it and then feel the direct result of that discovery that comes out of a lab. I can’t say that a Move 37 is there yet.
There probably are examples of materials that are being optimized and discovered in the computational realm, in the AI realm. I think it’s still very lacking on the experimental side, hence why we think our flywheel is imperative for materials discovery.
One cool thing we always tell is the AlphaGo story. We think it’s a really good story about indexing information. How do you bring that to science? It had learned Go games—I forget the exact number, millions of Go games or whatever the exact number was—and made this one-in-10,000 move. Then the subsequent versions of that AI didn’t need to learn Go games. It actually made similar moves without having millions of learnings in Go.
I’m excited for us to get to that place as well. I think we are building the technology to do the first version of AlphaGo today. That infamous Move 37—everyone watching AI thought it was a mistake. It was a one-in-10,000 move. We didn’t see it coming.
I think Radical AI will have something like that in the next 12 to 24 months if we follow and hit our objectives and our roadmap. Once those become par for the course, once they become more normal, then we’ll move into, “We didn’t look at experimental data,” or, “We used experimental data from alloys,” and start making discoveries in room-temperature superconductors. That’s what we think is going to happen.
We’re going to amass the largest experimental data set ever built in human history. That’s what we think, and so if we do that, we are very, very excited about the learnings that come from that, and then the other areas that we can make Move 37s in that don’t need millions of trained examples exactly in that specific material space.
To us, that is the holy grail of materials science. That is removing materials as the biggest bottleneck to our most important industries. That is enabling a world that is not limited by materials, but only limited by imagination and the laws of physics.
That’s what we want to build, and we are so excited to watch that proliferate and, candidly, very excited to be leading the charge at the intersection of AI and discovery and fully robotic, self-driving labs to supercharge and inform that discovery. We think that flywheel together will really craft a world that, as we always say to the team, you can’t imagine today—the amount of Ubers and Airbnbs and Facebooks and even things like NVIDIAs that come from this world that we can enable.
We think it’ll be one of the most important companies to ever exist because of that. And that’s the world that we’re excited to live in.
In many podcasts, I would say that’s a great place to end it. But in this one, I have a couple more questions, because that is an inspiring vision. How about the future of the integration of modalities?
This is kind of a hobbyhorse question of mine. But I do, to be very real about it, lower-impact work. I started this company, Waymark. We helped small businesses create video content for local advertising. We had basically a pre-AI version of this, and then we’ve tried to infuse AI into everything that we do.
One of the biggest things that I’ve observed is that the integration of language and visual understanding has been a real step change, when you no longer have to go through a language bottleneck to say, “Okay, I have this image. I want to change it to this other image,” or, “I just want an image.” If I have to say all that in a 100-word prompt, and even if I train a language model, which I’ve done, to prompt effectively, that’s just such a lossy thing, right?
What has made some of the more recent models go viral and amaze people so much more is that they clearly have this joint understanding. Now I can turn my image into anime, whatever, but it still retains the meaning of it while layering on this other coat of paint, whatever.
That depth of understanding, where it knows what you’re asking for in language but it also really clearly gets the image itself at a pixel level, has seemingly created a major step change. What I understand today about the system that you guys have built is that there are a lot of models doing a lot of subparts.
Do you think this eventually goes the Tesla route, or the Gemini 2.5 Flash Image route, where there’s one joint embedding space and language and a sort of intuitive physics—and potentially even the physics of materials—all exist together? Or is that too fanciful to expect at this point?
I don’t think it’s fanciful at all. It’s probably going too far to say that we’ll only ever need that one model to do everything, but internally we are developing something like what you described. It will be natively multimodal in nature. We’ll have encoders for all of the different modalities that we want to work with, and we’ll run it through this model. We’ll be able to have things output from it that are quite useful to us.
So no, I don’t think it’s fanciful at all to say that we’ll have something like that. In fact, I think in the next 6 months, we’ll be using that internally here at Radical AI, and certainly within the next 6 to 12 months, we will likely publish and release something along those lines.
The interpretability on that is going to start to get super interesting.
That is one of Eric’s prized things. We have been talking about it for a while. If we could work on something like that.
Yeah, cool. And that sounds like it’s not super far off. How do you think about the IP associated with this? You’re going to accumulate the world’s biggest experimental data set, as you said, and I don’t know if you meant to open-source the model that you’re creating there when you talk about publishing it. I also just don’t know a lot about the IP that exists in the materials science space.
How do you guys think about what data and what value you want to keep internal to the company versus what you can share? Again, I don’t even know if you can—is it possible to patent an alloy? Is it, especially, something amorphous? Would you be patenting the process that creates it?
You can’t really patent something that’s amorphous, right, and has no repeating structure. I would think that would be hard. Maybe the process to create it. So, yeah, I guess what is the sort of outlook for what is going to be core IP, not shared, and what can be shared?
How do we diffuse the benefits of this while still giving you guys the ROI needed to scale out the autonomous systems, to actually scale up the flywheel, the data set, and eventually get those Move 37s?
It’s an awesome question. When we started the company, we spent a lot of time thinking, “What do we want to build?” And not just, “What do we want to build?” What is needed if we are trying to solve the materials science problem? Not what’s the easy, low-hanging fruit, not what’s a business model that we know VCs will get behind—what is needed to build a company that does the vision you’ve heard us articulate today?
What we netted out was, well, if you just sell software, you’re never going to make real money. And so that can’t be it. We’re not going to do that. There have been a bunch of companies that have tried and failed doing that.
If you try to license every material you make, that can be a business model. But the problem is then you need to be creating new discoveries multiple times a year for decades to keep up with and continue to drive revenue, because, of course, reverse engineering and expiration of patents becomes a serious problem when you’re just licensing.
And so then we looked at what the largest material companies do in the world. What do 3M, Applied Materials, BASF, Materion, et cetera, do?
What do they do? They sell materials at scale. That's what their business models are. Except they're doing it in areas that, frankly, to be completely blunt, don't impact society.
Look, I understand the need to have a new Sherwin-Williams paint color every year. My wife is a massive fan. I know I'm a customer, but that's not why we're building Radical AI.
Honestly, they can continue to handle that. We are trying to make the species interplanetary, multiplanetary. We're trying to drive low-cost nuclear energy. We're trying to rethink transportation with floating bullet trains. We're trying to create a world that we want to live in.
And so, if all of that is true, then not only do you need to sell materials at scale, but you need to push the materials that you're selling at scale into enabling the industries and creating the industries of the future. When we do that, our focus is on the IP of making it.
In materials science, yes, you get patents on composition, but the IP is really around the trade secret of how you make it at scale and for the lowest cost at scale. We call it processing in materials science; it's all in that process. That's where all of the real meat of it is.
If I tell you the exact composition of a high-entropy alloy that we have today, you're not really going to be able to make it at scale, per se. There's this really important piece of this. If that is the business model, then that changes the way we think about the AI models. I'll let Jorge articulate that specifically.
No, I think Joseph hit it on the head. The machine learning models will change. They will. They're going to get better and better over time.
We're going to create some of them. Other people will create some of them as well. We may release some of them. Some of those we probably won't release, to be entirely honest with you.
But the one thing that Radical AI will never not do is go after the hardest problems that we think are going to be needle-movers for the world. At the end of the day, that's the whole point of the company. That will stay core to us through and through.
Love it. You want to talk a little bit about the—I know we're just about at time here, and I really appreciate all the time you've shared with us. Just yesterday, you put out a news item about working with the U.S. Air Force. Do you want to tell us a little bit about that?
Yeah, I'm super excited about this. It is in high-entropy alloys for hypersonic applications. We are conducting a lot of really high-throughput experimentation and pushing into how to test and optimize these materials for hypersonic applications.
Materials, again, are one of the biggest bottlenecks in the most important industries. In the hypersonic industry, China and Russia both field hypersonic systems today. They do that because they never stop investing in materials science.
Actually, it's worth mentioning that when China makes a new discovery in a material, it will set up an entire manufacturing hub around that material just to figure out how to scale it. It is an unbelievable emphasis on putting materials at the bedrock of all the most important innovations.
We've struggled with that in the U.S. and other places, and we are trying to drive that nexus back through. We spend a lot of time in D.C., a lot of time with the Office of Science and Technology Policy, the National Science Foundation, the Department of Defense, and the Department of Energy. We even go to the Hill and tell people in government, "Look, a public-private partnership and a focus on materials is imperative."
Whether you care about hypersonics, nuclear fusion, or the data sets that are going to go inside the ML models to generate a bunch of other new materials, that needs to be driven from a government perspective. You need to allow and rely on private enterprise to do exactly that, which is where we fit in.
The Direct-to-Phase-II is, in our opinion, kind of our beachhead for trying to push this forward and, separately from that whole plan I just mentioned, really an ability to execute on the first material system that we've chosen.
We had a thesis that high-entropy alloys can be very impactful for hypersonic technology. They're very hard to discover. It's a very large search space, and there aren't a lot of experimental results.
Our work with the Air Force is directly around trying to solve that problem and really generate these new materials for them to then explore putting into hypersonic systems. That is what we've been working on.
You can find a lot more information in our press release. We're really excited about that first project, which is again specific to the Air Force and hypersonics. But we're not just stopping there.
We have a lot of other things in the works with the U.S. government that we are trying to push forward so that we can identify—look, materials are not a subsector of industry or innovation. Materials are at the cornerstone of innovation.
Whether it's the Stone Age, the Bronze Age, or the Silicon Age, every modern technological wave has come from novel material advancement. The future will be exactly the same. That's the perspective that we're trying to push inside government and inside policy, and to explain why materials are really important as a nation, as a society, and as the human race.
That's great. That could be the right place to end on. Is there anything we didn't touch on that you would want to make sure people know about before we go?
One thing we call out is that we are really intense about our culture at Radical AI because we're trying to solve some of the hardest problems in the world. You have to have people who are insatiably curious, who question everything, who have this ability to go to first principles and ask why, but who are also unbelievably relentless.
We fail every single day, and we will continue to fail purposefully every single day for the next decade and beyond that. That's actually where the unlocks come from: building on those failures, of course, to reach success, as clichéd as it is.
If there are people out there listening who really have an appetite to challenge everything down to the laws of physics and want to do that in a way that is trying to tackle our most important problems, recognizing that there's going to be a lot of failures along the way, but that in the chance that you succeed, you leave a fundamental stamp on the human race through impact and innovation, then please reach out.
We are actively recruiting right now. We are looking for not only the best people in what they do, but people who are willing to dedicate their life's work to a mission.
If you're looking for a job, this is not the place for you. You should not apply to Radical AI. You will not fit in here. If you're looking to work on a mission, then please reach out. We'd love to talk to you. Jorge, anything you want to add on that topic?
No, I think that was perfectly well said. Nothing to add.
Amazing. Well, this has been an excellent conversation. I've learned a lot from it, and I think people are going to find it fascinating. Obviously, I wish you guys nothing but success in terms of breakthrough after breakthrough, and I look forward to changing material reality around us as we unlock one thing after another.
Joseph Krause and Jorge Colindres, co-founders of Radical AI, thank you for being part of The Cognitive Revolution.
Thank you.