[BidClub_]
The Cognitive Revolution · · 91 分钟

数学超级智能:Harmonic 的 Vlad 与 Tudor 谈 IMO 金牌与万物理论

Nathan LabenzVlad TenevTudor Achim

YouTube
TL;DR
  • Harmonic 的核心押注是:数学就是推理,而经过形式化验证的输出,能把 AI 能力变成用户可以信任的东西。 Aristotle 生成带注释的 Lean 代码,其每一步都由小型内核依据3条基本公理检查,但前提是内核和定理陈述本身设置正确。产品目标是打造一台“惊人的计算器”:拥有前沿模型的表达能力,同时具备计算器般的可靠性。

  • Aristotle 在2025年 IMO 达到金牌水平,支持了 Harmonic 关于强化学习可以围绕可验证奖励实现异常高效扩展的判断。 Harmonic、OpenAI 和 Google DeepMind 都未解出第6题;Achim 估计,即便对人类而言,这道题可能也难5倍,而且异常依赖空间推理,但 Harmonic 在继续运行中看到了“生命迹象”。创始人预计能力将大致平滑地呈指数增长,并称相较于规模更大的实验室,Harmonic 已经“远超自身体量出拳”。

  • Lean 4 和 Mathlib 可能以计算认证和开源分发,替代数学同行评审中相当大的一部分工作。 Mathlib 被描述为“以自洽方式合并成一个整体的全世界每一本数学教材”,而 Lean 允许贡献者通过类似 GitHub 的工作流提交证明,正确性靠测试验证,而不是靠社会声望背书。数学声望可能从期刊把关人转向 stars、forks、依赖关系和复用率,让精英机构之外的贡献者也能参与严肃数学。

  • 形式化验证可能成为 AI 生成软件的控制层,首先从错误代价最高的领域开始。 创始人介绍称,API 用户正在检查密码学实现是否具备碰撞性质,并考虑自动驾驶控制器是否存在导致系统不稳定的输入序列;他们还表示,用户已经在用 Aristotle 检查安全关键软件。更长期看,创始人质疑 AI 为什么还要编写 Python 或 Java——这些语言主要针对人类可读性进行了优化。如果智能体能生成约150万行的浏览器或5000页的证明,人工审查就无法继续扩展,由此可能从手工形式化方法走向“形式化氛围编程”。

  • Harmonic 正把开放访问同时作为分发策略和数学品味的去中心化机制。 Harmonic 不雇用内部团队来决定 Navier–Stokes 是否比 P versus NP 更重要,而是通过 API 和网页界面开放 Aristotle,让社区需求来分配算力。创始人更偏好一个由数百万获得工具赋能的研究者组成的未来,而不是“一个拥有2吉瓦数据中心的巨型 AI 实验室”垄断每项发现及其价值。

  • 其训练理念偏向可扩展的搜索,而不是依赖人类对美感的监督,同时把幻觉视为必要的探索机制。 Harmonic 几乎没有做过数学家小组参与的优雅证明 A/B 测试;相反,研究人员优化 Achim 所说的“未来证明的净现值”,惩罚那些能靠蛮力解决简单任务、却无法建立可复用能力的方法。预训练模型仍然是有用的起点,也可能与更高熵、较少受人类方法束缚的系统结合:“幻觉让模型能够探索此前从未被人类编码过的东西。”

  • 2030年的愿景不是立即获得全知,而是理论上的丰饶,随后迎来新的数据瓶颈。 Achim 想象,也许会出现5套内部自洽、统一量子力学与广义相对论的理论,科学家需要越来越高能的实验来区分它们——这是“基本上能够解释一切的理论”,但不是无需观测就拥有知识。Harmonic 当前仅允许 Lean 文件作为行动空间,因此运营风险有限;创始人预计,一旦这类系统获得 API 和自主性,网络安全风险会升高,并坚持“人类应该掌舵并作出决定”。

摘要 · 为研究而整理的核心内容

1. 数学就是推理,而抽象的效用往往姗姗来迟

  • Achim 对数学的核心定义不止于定理证明:“数学就是推理”,也就是把一个解释拆解成“由一连串微小逻辑步骤组成”的过程,让其他人能够逐步检查。物理学、税务、宇宙学和工程学的研究对象不同,但最终都需要一套从明确事实出发、内部自洽的解释。

  • Achim 是通过物理学进入数学的:先读 Stephen Hawking 和 Brian Greene,之后进入 Stanford。关于大爆炸、引力,以及今天的各种力是否都从“起初的同一个东西”中分化出来的问题,反复把他从物理学引向微分几何和纯数学;他的实际链条是数学通往物理,物理通往工程,工程最终通往飞行器、晶体管、GPS 及其他有价值的系统。

  • “数学非理性的有效性”构成了 Harmonic 对“数学必须立即产生效用”这一要求的回答。关于流形的抽象研究后来成为相对论的关键,数论也从一项深奥的学术追求变成安全数字经济的基础。因此 Achim 的建议带有投资组合思维:“你只管做数学”,因为其中一部分最终会比任何人最初设想的更有用。

  • 创造一个重大新思想的数学家极其罕见,往往5到10年才有一次;另一类则是高产的综合者,他们吸收论文、记住技术,并把方法迁移到不同领域。AI 未来或许能加速前一类人的工作,但 GPT-4 已经显示出对后一类人的杠杆效应:搜索庞大的文献,找回相关技巧,再以任何个人读者都无法匹敌的速度组合知识。

2. Lean 把证明变成可执行的证书

  • Achim 称 Lean 是“有史以来最好的编程语言”,因为它同时覆盖普通程序和逻辑命题。作为依赖类型语言,Lean 能在编译阶段表达复杂性质,而不是仅仅运行代码后再检查结果;在他的光谱中,JavaScript 几乎不检查任何东西,而 Lean 可以在执行前指定并验证异常丰富的不变量。

  • Lean 建立在构造演算和3条公理之上:命题外延性、公理化商集的可靠性,以及选择公理。最后一条公理表示可以从非空集合中选出一个元素;3条公理都短到一条推文足以容纳。创始人认为,这个极简基础足以支撑数学、计算机科学,以及物理、经济学、统计学和生物学中的定量建模。

  • Nathan Labenz 的国际象棋类比依然成立:一个定理说明棋盘可以从起始局面走到目标局面,而证明则列出中间的每一步。Lean 的内核依据规则手册检查每个声称的步骤,并验证最终状态确实已经到达。这里的步骤不是棋步,但证书具有同样的结构:不断重复“这一步是正确的”,直到抵达终点。

  • 学习曲线可以从研究数学以下的层级开始。Achim 推荐 Natural Number Game,用户可以在那里推导加法和乘法的性质;他还提到一款实分析游戏,内容从数列和实数逐步推进到微积分基础。他更大的判断是,数学教育将从黑板迁移到计算机实验室,甚至可能触及中学生。

3. Mathlib 可能把数学从期刊带到 GitHub

  • Tenev 记得自己在2000年代末做数学时,几乎完全依赖黑板、白板和沙发,电脑主要用于把完成的成果排版。Lean 把工作过程迁移到 VS Code、Cursor 和 GitHub,使地理上分散的协作成为可能:几十个人可以共同形式化研究,或者一起攻克费马大定理这样的项目。

  • Mathlib 在 Lean 的小型内核之上提供可复用的高层次操作。它覆盖代数、几何、实分析、统计学等领域,把定义和已经证明的结果整合起来,调用方式类似软件库中的函数。Achim 所设想的终点是“以自洽方式合并成一个整体的全世界每一本数学教材”,并且可以在普通电脑上从基础重新构建。

  • 形式化改变了信任机制,但有两项前提必须保留:Lean 内核必须没有漏洞,定理也必须忠实表达原本想表达的命题。在这两个条件成立时,经过检查的证明不再需要一位有声望的数学家为每一步背书。“公民数学家”可以直接贡献一份证书,而不必说服精英院系或期刊来证明其正确性。

  • Achim 把这一转变类比为开源软件从“大教堂”走向“集市”。任何人都可以提交证明;Lean 负责测试正确性,维护者可以评判风格,而 stars、forks、依赖关系和复用率则可以衡量影响力。时间点同样关键:Lean 3 仍然像测试版软件,Lean 4 和 GPT-4 则在 Harmonic 2023年成立前后变得真正可用。

4. Aristotle 在多个尺度上搜索,而不只是蛮力穷举

  • Labenz 将 IMO 时代的系统拆解为蒙特卡洛树搜索、非正式引理生成器和几何专家,但 Achim 修正了“树搜索只是蛮力”的理解。语言模型会推断具有决定性意义的高层步骤,并解决困难的子问题;只有深入搜索树后,工作才会变成例行的分类收尾。

  • 搜索机制本身由 LLM 构成:一个组件提出步骤,另一个组件为步骤打分,二者围绕引理协作,最终拼装出完整的 Lean 证明。非正式推理器更适合理解为上下文管理器,而不是永远正确的登山向导。它会产生“数量巨大的错误”,包括错误的方案,或根本无法形式化的方案,但有用的中间落点仍可能从中浮现。

  • 几何模块更接近真正的“磨床”。Labenz 把它比作 AlphaGeometry;Achim 说,它先探索高层步骤,再用算法把后果逐一磨出来。几何较早被解决,部分原因是其对象和关系受到约束——点只能产生有限数量的角度;但当图中出现10个或15个点时,搜索空间仍会迅速爆炸。

  • 另外两个不太显眼的组件承载着相当大的产品价值。自动形式化负责把自然语言提示忠实翻译成 Lean;理论构建则创造 Mathlib 中尚不存在的结构,并在证明搜索过程中将其纳入。Achim 提醒,IMO 系统之后,架构已经经历大幅整合和修改;技术报告描述的是一个不断变化的截面,而不是固定不变的产品。

5. 当人们无法就有效推理达成共识,可证性就到此为止

  • Labenz 用刻意别扭的提示测试非正式模式。让 Aristotle 证明“万物皆爱”,它把命题归为哲学问题,认为超出 Lean 4 的范围;让它证明“Epstein 没有自杀”,它则把该陈述当作时事问题,而不是形式化定理。这些拒答暴露出证明之前更困难的一步:判断一个混乱的现实世界陈述究竟应当被忠实表达成什么命题。

  • Achim 认为,最终边界会落在“人们能够就何种推理序列算作有效达成共识”的地方。数学和国际象棋天然符合;定量软件行为也符合,因为输入、分支、循环和输出之间存在可检查的关系。他不太相信历史论文能够变得客观可验证,暂时把目标限定为“任何本质上属于定量和逻辑的东西”。

  • 现有 API 的使用说明了软件领域的延伸。一段密码学实现可以被检查两个输入是否可能发生碰撞,尤其是在算法要求唯一性的情况下;自动驾驶控制器也可以通过符号方法测试是否存在会制造不稳定死区的输入序列。单元测试是在代码存在之后抽样执行,而形式化推理则可以尝试证明:对于满足既定假设的全部输入,系统都会表现出某种行为。

  • Tenev 在创始人的分歧中站得更开放一侧。早期 Aristotle 或许可以围绕“万物皆爱”发明一套理论,而用户已经提交了生物学、医学、经济学和金融数学问题;当前限制主要是把客户引向可靠价值。他可以想象,把互联网事实明确写成公理;而天文学问题——例如 Palo Alto 50英里范围内下一次日全食何时发生——已经显示出,第一性原理计算为何胜过信息检索。

6. 数学品味来自社区需求,而非内部评审小组

  • 在 Aristotle 达到 IMO 金牌水平后,Harmonic 面临一个选择:把系统留在内部,招募顶尖数学家,定期宣布私下完成的成果;还是把它广泛开放。公司选择了 API、终端界面,随后又推出网页界面,让社区的“显 revealed preference”决定哪些问题值得投入算力,而不是由公司宣布 Navier–Stokes 比 P versus NP 更重要。

  • 这一选择暴露出 Harmonic 内部不会优先处理的工作,包括计算学习理论、图论猜想、计算机科学、密码学和数论的若干分支。Tenev 把战略分叉说得很直接:发现可以从一个捕获全部价值的2吉瓦实验室中产生,也可以来自数百万拥有工具、独立工作、彼此协作、获得署名并保留更多上行收益的人。

  • Labenz 追问的是另一种品味:Aristotle 的证明是否优雅,而不只是正确。Harmonic 几乎没有通过数学家小组进行证明 A/B 测试。Achim 则把目标描述为“未来证明的净现值”——偏好那些能降低未来解决日益困难问题的计算成本的方法,而不是只追求今天让人类评审者印象深刻的证明。

  • 这个目标内含一种刻意保留的张力。蛮力可能在简单问题上高效地产生短期答案,却什么可复用的东西也没有教会系统,未来反而变得昂贵;对于真正经过推理的解法,则偏好更短、更高效的证明。Harmonic 遵循“苦涩教训”:有用时从预训练模型出发,施加尽可能少的先验,扩大强化学习规模,并可能混入更高熵、较少受既有人类数学偏见影响的系统。

7. 可验证输出取代对模型内部的理解

  • Labenz 追问,可解释性是否能揭示数学模型内部的新抽象,或者 Aristotle 最终是否会与一个更模糊的世界模型合并——他戏谑式的终点是,由一个仁慈的“安全 HAL 9000”驾驶宇宙飞船。背后的担忧是,正确性证书或许只能告诉我们模型证明了什么,却不能解释其内部表征如何生成这一洞见。

  • Achim 回应说,人们往往把可解释性当作可信度的代理指标,而 Harmonic 在公司成立初期就通过要求形式化输出解决了这个问题。Lean 是“最具可解释性的输出”:机器可以检查它,人类则可以反复使用“转到定义”,像浏览代码库一样遍历整个证明。模型在机制层面或许仍然不透明,但其会对外产生后果的推理是可检查、可认证的。

  • 因此,Harmonic 优先追求廉价、经过验证的证明,而不是打开神经网络的黑箱。一位创始人提出——同时承认“我可能错了”——更容易理解的路径也许是研究 Aristotle 如何把3个数学子领域的技术组合成一份前所未有的解答。这种行为层面的解释,可能比挖掘权重、寻找他们尚未观察到的内部“超级智能火花”更有用。

8. IMO 金牌与从非正式到形式化的阶段跃迁同时发生

  • Harmonic、OpenAI 和 Google DeepMind 都在2025年 IMO 达到金牌水平,也都未解出第6题。Achim 估计,这道题即便对人类而言也可能难5倍,其中包含许多步骤和难以形式化的空间推理。继续运行后出现了“生命迹象”,因此他认为最终解出这道题是可能的,而不是把它视为架构遇到的天花板。

  • 第3题和第5题很可能在仅仅一年前还会超过大多数模型的能力,这强化了创始人对能力大致平滑指数增长的判断。社区工作流也开始呈现组合性:Labenz 看到 GPT-5.2 Pro 在 token 空间生成候选证明,随后由 Aristotle 将其形式化并检查。与此同时,Aristotle 用户已经解决了一些曾经开放了30年或40年的中等难度问题。

  • Achim 认为真正的断点发生在别处:“向形式化跃迁的阶段转变”已经发生。数学家不久前还会觉得可笑的工作——上传一篇完整的数论论文,自动将其翻译成 Lean——如今通过多次运行 Aristotle 已经可以实现。Harmonic 甚至考虑过一个“Ralph 按钮”,只要按下就让系统持续运行直到形式化完成,让人类负责选择问题、评估技术,而不必检查每一个推导。

  • 规模让创始人对终点给出明确判断。没有人会手工审查一份5000页的证明,更不用说那份在 Harmonic 2023年成立时激发其想象的、假设有100000页的黎曼猜想证明。Labenz 指出,DeepMind 在2024年从形式化的 AlphaProof 转向了今年的非正式 Gemini,而 OpenAI 的 IMO 系统也是非正式的;训练可能继续采用混合方式,但 Harmonic 认为形式化与非形式化输出之争已经“尘埃落定”。

9. AI 编写软件带来与 AI 数学相同的验证瓶颈

  • 创始人把长证明的逻辑直接延伸到了代码。Cursor 团队的一项实验生成了一个约150万行、兼容 Chromium 的浏览器;当智能体自主工作数周后,无论人类还是协作模型,都不可能以经济可行的方式检查每一行代码是否存在漏洞。如果这类系统要成为可靠的基础设施,验证成本的增长速度就必须远低于生成复杂度。

  • Tenev 质疑,AI 生成的代码是否还应继续使用 Python 或 Java——这些语言的设计部分考虑了人类可读性。如果机器是主要作者,形式上可验证的语言或许是更好的目标,因为系统可以直接建立性质,而不是从测试和代码审查中推断性质。Lean 是 Harmonic 最喜欢的语言,但创始人也承认,采用情况仍然是开放问题。

  • 扩散应当从任务关键型系统开始,因为这些系统的缺陷代价格外高,团队也已经在聘用 Lean、Rocq 或 Isabelle 专家,进行耗时的验证。Aristotle 可以先加速这些专家,再把实践扩展到“形式化氛围编程”。设想中的终点不只是更快的编程,而是越来越少漏洞的软件,其安全性和可靠性声明都带有机器可检查的证书。

10. 新抽象需要熵,而不是逃离逻辑

  • Labenz 最尖锐的形而上学反驳是:在今天的形式系统中训练,是否会把 AI 困在今天的抽象里?Einstein 推翻了直觉上的三维世界观,因此数学超级智能或许也必须“打破第四面墙”。Achim 的回答是,Einstein 仍然可以通过微分几何严谨表达这一突破;任何通过独立可检查的推导推进的新理论,本身都可以编码进 Lean。

  • Lean 的公理旨在描述极其基本的推理操作,而不是某一种特定的物理范式。因此 Achim 把形式化推理视为非正式推理的一个极其细致、可由计算机检查的版本,而不是概念牢笼。哥德尔式不完备性会把真实但不可证明、不可判定的命题留在前沿,但创始人并不认为这些边界情况会妨碍证明“绝大多数有用的东西”。

  • 在这一框架内,熵仍然不可或缺。Aristotle 会尝试许多最终失败的路径,而这些失败让它能够触及人类从未写下的思想;“幻觉是推理系统的关键组成部分”。理想设计不是一个零幻觉生成器,而是一个与形式验证器耦合的高熵探索者:它能够提出错误命题,发现命题错误,然后只保留经过认证的结论。

11. 终局是受实验约束的理论丰饶

  • Harmonic 的既定路径始于2023年成立,经过2025年 IMO 金牌水平表现、年终基准测试登顶,以及 API 用户解决那些曾经开放了30年或40年的中等难度问题。下一个飞轮很直接:自动形式化扩充 Mathlib,更大的认证库降低困难证明的成本,用户继续追逐更重要的猜想,而几乎每天的产品改进则降低提交工作的摩擦和成本。

  • Achim 对2030年的渐近状态描述为“基本上能够解释一切的理论”。科学界不再需要艰难寻找量子力学与广义相对论之间任何一个自洽的桥梁,而可能拥有5种能够解释全部现有观测的竞争性统一理论。届时瓶颈将重新回到数据:研究人员需要越来越高能的实验,甚至可能需要新的对撞机,才能判断哪一种数学上自洽的解释描述了真实宇宙。

  • 这个判断明确不等于全知。模型只能从有现实依据的假设和观测出发进行推理;宇宙的某些性质“只能运行实验,亲自找出答案”。但 Achim 预计,在达到那个渐近状态之前,系统就会产生巨大的效用:消除能够持续进行高水平逻辑推理的人才短缺,可能带来一场科学复兴。

  • 今天的安全性部分来自受限的行动空间:Aristotle 输出 Lean 文件,而不是自主接触 Gmail、iMessage 或运营系统。Tenev 预计,一旦智能体与更多系统连接,最初的严重风险将类似于由 API 和自主执行引发的网络安全事件。随着 Harmonic 最终把模型接入现实世界,Achim 表示,公司必须更加认真地对待这些风险,同时保留其治理原则:“人类应该掌舵并作出决定。”

Vlad Tenev

Thanks for having us. Greetings and salutations.

Nathan Labenz

Thank you. This is going to be, I think, a fascinating conversation. It’s probably going to be more metaphysical than most of our episodes, but there’s also a lot of practicality because what you guys are doing certainly has aspirations to go beyond the pursuit of mathematical superintelligence.

Maybe just for starters, how do you guys understand what math is? That was something I was really wrestling with in preparing for this. To make that a little bit more practical, what would you say are the core cognitive skills that people who are good at math really develop and excel at? And how do those skills fare when we look at the performance of the frontier large language models that all of our listeners are familiar with today?

Vlad Tenev

Well, look, first, thanks for having us. It’s really great to be here. When you ask, “What is math? What is it useful for? What are the core skills?” it gets to one of the core theses of our company, which is that mathematics is reasoning.

A lot of people think of mathematics as this really esoteric thing. You’re thinking maybe about group theory stuff you’ve seen in movies like Good Will Hunting. But mathematics, at its core, is the process by which humans understand the world by breaking their understanding down into small sequences of logical steps that other people can understand and verify for themselves.

So when you’re solving a physics problem, doing your taxes, or thinking about what happened at the beginning of the universe, ultimately you have to have an explanation that is self-consistent, that follows from other facts, and that your colleagues or other humans can check.

When we talk about what it takes to be good at math, the question is what it takes to be good at reasoning. Again, that’s the ability to break this out into steps. It turns out math is really useful for understanding the universe and building lots of engineering things, but ultimately it’s just about reasoning.

Nathan Labenz

I watched the podcast you did with Sequoia, maybe 16 months ago or so now, and I recall Vlad’s story: basically, “I thought that if I got good at math, then I’d probably be good at other things,” and it sort of worked for me.

That’s one way to, in a very practical sense, unpack the idea that math is reasoning. It certainly seems to help people generalize to at least related domains and be really effective, for example, in entrepreneurship.

But I’m not entirely clear still on whether you’re making a more almost Platonic claim there. It seems like there’s the very simple notion that, okay, I should teach my kid a lot of math because then they’ll be smart generally—and again, that works for humans—but is there something that you see as a more fundamental law of the universe, a sort of correspondence between what we are doing in math and what we are doing in these other domains?

Because it doesn’t seem like we have the same sort of verifiability in almost anything else. We do have it a little bit in computer science, but even in physics, we’ve still got very fundamental questions about whether the paradigm is even right, or what it would mean for it to be proven right. I don’t think that stuff is at all agreed upon.

So maybe you guys throw up your hands at this mystery too, or maybe you feel like you have an intuition for what the answer is.

Tudor Achim

Yeah, I can give you my perspective. I got into math through physics. When I first came to Stanford as an undergrad, I had read Brian Greene’s The Elegant Universe, which was sort of the first popular string theory book.

When I was a kid, one of the earliest memories—one of the first full English books that I read—was A Brief History of Time by Hawking. I’ve always been interested in the big questions: What happened before the Big Bang? How did the laws of physics come about? Is there just one law, one particle, one force that eventually, as the universe cooled and expanded, splintered into all the different forces we have today, like gravity, electromagnetism, and the strong and weak forces?

Because, back in the day, that was not obvious. We thought electricity was separate from magnetism, and it was probably one of the greatest achievements of science to figure out that these two are actually two sides of the same coin.

And then the big question is: What’s going on with gravity? Is it the same? In the middle of all this, we found out that the weak force and the electromagnetic force were also splintered off of one electroweak force.

It kind of feels like there was just one thing at the beginning, and we have to understand what that thing is. What I found when I became a physics major at Stanford and started asking all these questions was that eventually they’d send me over to the math department, and they’d say, “Well, in order to understand string theory, you have to understand all of these other things. And if you want to understand general relativity, you’ve got to get into differential geometry.” So that’s how I became a pure math major and ended up doing a PhD.

The impetus was actually trying to understand the real world through physics. If you think about the usefulness of physics, all of the big inventions humanity has that really push us forward are physics inventions. When you think about flight, rocketry, computers, transistors, and GPS—obviously, one of the main examples of why relativity is useful—they’re physics things. The real reason to do math is that math is interesting and beautiful, and there’s an art aspect to it, but it helps you understand physics.

Physics helps you understand engineering, and then you can create things that have huge value.

Tudor Achim

You were asking how math works in other fields where things are not as precise. I think math shows up just a little more subtly than people think. There was this physicist, Eugene Wigner, who wrote a famous essay called “The Unreasonable Effectiveness of Mathematics in the Natural Sciences,” which commented on a really interesting phenomenon.

Vlad mentioned differential geometry and special relativity. It turns out that when Einstein was creating that theory, he relied on these thought experiments from the 19th century around how to think about certain manifolds and their properties. That was actually the key tool that we used to explain what special relativity is and then develop it into general relativity. I think it’s a perfectly representative case because those thought experiments in the 19th century were almost preposterous. It made no sense to think about them, because how could you possibly apply these concepts to the real three-dimensional world?

Then it turns out that it’s very useful for understanding the four-dimensional world when you include time and curvature. There are myriad examples like this. If you consider number theory, for a long time that was really seen as an incredibly esoteric branch of math with no practical implications, but people really pushed on that theory for a long time. Then it turns out that it’s the key tool you need to create a secure digital economy.

So now essentially all of human civilization has a digital economy that is based on this branch of math. I think it’s almost the wrong question to ask, “Well, there’s a lot of math out there. How is it useful?” The point is, you just do the math, and then eventually some of it—not all of it—will be more useful than you possibly could have imagined. The investment in math is not just to build a really smart system. It’s to create a lot of new math that we can then figure out ways to apply later.

One interesting thing that the conversation reminded me of when you first asked, “Okay, what is math? What does it look like?” is that I think one of the reasons we got excited about applying AI to this domain is that there are lots of different things that mathematicians do. Some of them are very creative, almost like artists, and maybe they’re not super prolific, but they come up with something new once every 5 to 10 years. That can be just an amazing accomplishment in the field—like Grigori Perelman, for example.

Others are just machines. They can read more papers and comprehend more papers per unit of time than other people. What they’re doing is basically synthesizing all the knowledge, figuring out all the tricks, applying those tricks quickly to new domains, and reusing these things. We’re very excited about the prospect of AI accelerating the former. We think that’ll happen.

But the latter is something that AI is already really good at today and was good at to some degree when we got the idea for Harmonic. You look at GPT-4, which had just come out when we started, and it excelled at pulling information and doing these types of needle-in-a-haystack things: Can you quickly go through all the literature and pull out things that might be relevant?

I would say you can be an amazing mathematician if you’re in that category. A lot of the work could be accelerated if you just knew all the math that was being done and could pick out the relevant things for an unsolved problem that you have at hand. I think the problem itself lends itself really well to what AI is already good at.

Okay, that’s quite helpful. I think coming into this, I had focused my own mind on two modes of math. I guess one would be the Einstein-like mode—obviously, that’s a high-level example of a eureka moment, of having some insight that, hey, this highly abstract and seemingly perhaps very esoteric formalism can actually unlock major understanding. That’s kind of amazing. Very amazing.

Then there’s also this sort of grind-it-out mode: “I’ve got this thing that I want to prove, and I’m going to perhaps stumble my way even through the space of possible logical moves until I finally chart a path there.” Then you’re adding a third layer, which is problem selection in the first place. I guess that’s pretty related to the Einstein thing, but certainly distinct in some ways.

Let’s take a minute before we get into the Aristotle system and how it works and how you trained it and all that stuff to just talk about Lean. Lean is basically a programming language that does this very bit-by-bit logical maneuvering, right? You have certain assumptions coming in, you’re going to take these various steps, and the goal is to get to a certain outcome.

Tell us, because I’m just learning about this in the context of preparing for this and a couple of other podcasts, and I think most people don’t know anything about it. Maybe give us a little bit more of an intuitive understanding of what Lean is. I’d be keen to understand it on a little bit of a practical level, too.

How many symbols are there? How many axioms are we starting with? How many rules are there that we can apply? How big is the space that we’re manipulating our way through?

Tudor Achim

Lean, in my view, is the best programming language ever created. In Lean, you can write any program you would write in Python, C, or C++, but you can also express essentially any logical concept.

If we’re okay getting into a bit of the detail, it’s a dependently typed programming language. That means that, at compile time, you can express very complicated properties of the program that you can check before ever running it. On one end of the spectrum, you have something like JavaScript, where you can check basically nothing, and on the other end you have Lean.

The really cool thing is that you asked about axioms. When Aristotle produces any output, it’s produced as annotated Lean code. There’s the programming language Lean; we write theorems, we write programs, and we prove things. There are also a lot of comments explaining to the person reading it what it’s doing.

When we talk about proving things, you end up relying on 3 axioms, in addition to the basic concept of the calculus of constructions, which is what the programming language is based on. 2 of them are extremely technical: one is propositional extensionality, and one is something about quotient soundness. The third one is the axiom of choice.

Just as an example to show what an axiom means, the axiom of choice isn’t saying anything controversial. It says that if you have a nonempty set, it’s possible to choose an element from it. From these 3 extremely basic axioms, it turns out you can build all of mathematics, all of computer science, and all of mathematical modeling in physics, economics, statistics, and biology. It’s all based on this core set of axioms.

The goal of a system that outputs Lean is to find interesting statements and programs, then prove things that just depend on these axioms. That’s really where the difficulty lies. As you alluded to, sometimes you have to make big logical leaps, and sometimes you have to grind through a lot of math. Both of those are essential, so you can’t really skip either of those steps.

Lean itself is incredible. You can express so many ideas in it, prove so many things, and use it as a programming language, too. It’s really up there for me among programming languages.

Vlad Tenev

I started playing with Lean when Tudor and I started making a plan for this business, and we had a pretty early decision about whether we wanted to go formal or informal. One thing that struck me about it is that, as a former mathematician, I barely used the computer when I was doing math.

I was doing my PhD in the late 2000s, and the only time you’d really use a computer when doing math was when you wanted to type up your homework or your research paper or something. All the thinking about it would happen on a chalkboard or a whiteboard. All the collaboration about it would happen in person, at conferences, or on a chalkboard in one’s office.

For a while, it was as if maybe mathematics would always be this pure thing that would remain untouched by technology. But what Lean has done is transform mathematics from chalkboard and couch to now being in VS Code. You can do it in Cursor, and you’re putting your math on GitHub, where you can now run these large collaboration projects.

Even when you subtract out AI, I think Lean by itself, without AI, changed how people do mathematics. Now you’re seeing extremely prolific, famous mathematicians running these large projects where they’re collaborating with dozens of people around the world, trying to do things like formalize research or formalize the proof of Fermat’s Last Theorem. More and more people are adopting Lean as an accelerant.

I think it’s changing how mathematics is being done. It accelerates collaboration, accelerates progress, and sort of removes this notion of peer review. If you’re a mathematician and you want to prove something, a big part of the process is getting someone to read it and spend the time to tell you if it’s correct.

You have the proof of Fermat’s Last Theorem, which took many years to be proved. What happened was that this collection of people got together, and when they all agreed that the proof was complete, it was sort of ordained that the thing was proven. I think another thing formalization does is make that unnecessary.

If the proof checks—and assuming there’s no bug in the Lean kernel or in how you’ve set up the statement—you obviate the need for manual human verification. The implications of that are pretty interesting, too. You have all of these potential citizen mathematicians who now, with AI, can solve unsolved problems, and they don’t need to get anyone at a PhD program or a leading institution interested in their problem in order to establish that it’s correct. They just have to have the Lean certificate, and the proof is correct.

I think that’s a powerful thing. If you think about journals, journals in math exist for this purpose: the prestige of the review board tells you whether you should read something or trust it. I think the notion of trust is fundamentally changing with tools like Lean.

Tudor Achim

Yeah, I think the open-source software community solved this problem a long time ago. If you go on GitHub, someone can simply open a pull request on a repository, and if it passes the tests and the author of the repository agrees with your style, it gets merged. Now you’ve contributed.

That element of trust isn’t as present. You can just run the tests. When you talk about impact and prestige, you can look at the number of stars you have. If a repository is very popular, it gets forked a lot and gets a lot of stars. You’ve essentially disintermediated any gatekeeper here. It’s totally open source, there’s no more trust required, and there’s a measure of impact.

I think math is going to start going the same way. Previously, mathematicians relied on their social networks to figure out who tends to do the right thing and who tends not to make mistakes. With Lean, you can have a big math project, and anybody can come and contribute a proof. If Lean accepts it, then it’s right.

If a lot of other mathematicians start to depend on that result, we’re going to notice a lot of forks, a lot of dependency graphs, and a lot of stars on it. You start to measure prestige that way. It would be very interesting if Lean were the one tool that allowed you to go from the cathedral style of development, where you have very closed networks, to more bazaar-style development, where it’s kind of the Wild West, but Lean is the computational certificate that everything is correct.

Nathan Labenz

I wish I understood a little bit better, or had a more intuitive sense, of what exactly is going on with Lean. This is going to be hard, I think, but in doing my research, one thing that stands out is that the kernel is really small. In terms of what you need to trust, it’s a pretty small amount of core code that has been thoroughly vetted many times by many people.

There’s that level of understanding, but I would still love to have a little bit better sense. When you mention the 3 axioms, for example, it’s a little weird for people outside the field to hear, “There are 2 that are kind of bizarre and technical, and then there’s this one that says if you have a nonempty set, you can choose an element from it.” I’m like, that seems like common sense, but why was that ever controversial?

Is there a way to describe the space of legal moves in math or in Lean? I don’t usually like analogies, to be honest. I often try to set this up as an analogy-free zone, but I think I—and a lot of others—am going to struggle with the very literal understanding. Maybe this is a time for an exception to my no-analogies rule.

Is there some sort of chess analogy where you could say, “Here are the pieces, and here are the legal moves that you can make,” to give people a better sense of what it actually means to move through these spaces?

Tudor Achim

I think the chess example is perfect. A theorem in Lean is something like: given this starting configuration of a chessboard, it’s possible to get to this configuration. A proof of this theorem would be listing the sequence of moves.

What the kernel is doing in Lean is saying, for every single move that you claim is valid, “Does this rule exist in my rulebook?” The theorem says you can get from A to B. The sequence of moves is, “Okay, here’s the sequence,” and the kernel is just saying, “Yes, this step is right. This step is right. This step is right.” Now it has confirmed that you’ve ended up in a target state.

Lean is doing that, but of course the individual steps are different. They’re mathematical steps, and they depend on 1 or more of these 3 axioms. The 3 axioms, although they’re technical, are very short. If you write them down as mathematical statements, each of them is under a tweet in length.

The axiom of choice definition in Lean is maybe 10 characters, and the other ones are maybe 100. They’re not very complicated; they’re just a little bit annoying to write in math. Then people say, “Okay, if we assume these axioms are true—and they’re also common sense, just a bit more complicated—and we’ve checked every single step against those axioms, then we say the whole proof is correct.”

Nathan Labenz

Could you give a few examples of the pieces and the moves? Obviously, we can't come anywhere close to being exhaustive, but what are the primitives?

Tudor Achim

I'll give a mathematical but simpler example of a primitive. Let's consider first-order logic. The deduction rules you have are: if A, then B.

Let's say you have a proof that says, “If I have A, and I know that if A then B and if B then C, the theorem says C is true.” The proof says that A is true, and I have “if A then B,” which means B is true. Then I have the step that B is true, and I know that if B then C, so I can conclude that C is true.

That's first-order logic. It's not quite the same as what we're talking about in Lean. You can do more advanced types of logical statements there, but ultimately, that's what's happening.

I think the next step beyond that is just getting to Lean, the Calculus of Constructions, and these axioms. One thing I learned is that people are also exploring the use of Lean to teach math. I think it's now practical at the high school level, but you could see a world where it extends to middle school and maybe even younger if someone is precocious enough.

I think mathematics education will go from the chalkboard to the computer lab. There's this thing called the Natural Number Game, where you learn Lean by deducing properties of multiplication and addition. For example, the commutative law is that a + b = b + a, or the distributive law, where a × (b + c) = a × b + a × c.

You can discover and prove these fairly basic facts using the core axioms and the Lean language. That's a good way, if anyone just wants to ask, “What is this Lean thing? Why is it useful if I'm not a research mathematician?” Dip your feet into it. I would recommend that.

That's been extended to harder things, too. I think there's now the Real Analysis Game, which is for learning real analysis. It's very proof-based and is essentially the foundation of calculus. You can start with basic facts about what a sequence is, what a real number is, how many of these numbers there are, and how big the sets are. Then you can keep proving more and more complex things.

Nathan Labenz

That's a great tip. I'm definitely going to bookmark the Real Analysis Game and see if I can get my soon-to-be seven-year-old into it.

Tudor Achim

We hadn't really talked about Mathlib, but the Lean kernel is quite small. There's an open-source project called Mathlib, which you can think of as the largest digital repository of mathematical knowledge. A lot of the famous theorems and results can be found in Mathlib, and those give you almost like additional complex moves or algorithms to prove your thing.

You can apply a theorem, and it's almost like applying a function from a library. That can help you get to the goal. You can think of Mathlib as an amalgam of the knowledge written in English, Russian, or German, and all the textbooks out there, consolidated in Lean and available open source.

You can browse it on GitHub by category. There's algebra, real analysis, geometry, and statistics. Part of doing this is building this great repository that people are then building on and proving novel results on top of.

You can understand what it is better by thinking of it as every math textbook in the world merged into one in a self-consistent way. Eventually, all mathematical knowledge will be in this one repository. If you hit Build on your computer, you'll be able to check it all from the foundations.

If you have any question about any math concept, you just search for it. You click on “Go to definition,” and you can jump around. It's really going to be the new foundation for math in the future.

Nathan Labenz

It's pretty exciting. I think mathematics is certainly going to change fundamentally—how it's done and how fast it moves—and, to a large degree, it already has. AI is just going to accelerate it.

Vlad Tenev

The great thing about our timing is that Harmonic really started when both of these things matured to a level of capability where you could start doing interesting stuff. Lean basically went from being essentially beta software—not appropriate for real mission-critical use cases—with Lean 3 to Lean 4, and that was about the same month we launched the company.

Also, GPT-4 started to show glimmers of being really good at synthesizing information and at the starting points of reasoning. That came out around the same time. Both of these matured to the level where you could start putting them together and doing really cool things. I think we were just the first to see that, and that's how we came up with this concept of mathematical superintelligence, which really means the combination of formal verification and formal tools with artificial intelligence.

Tudor Achim

Funny story: I was using Aristotle a little bit to try to wrap my head around all of this. I don't have the sophistication to pose any really interesting problems, so one challenge I gave it was to prove that 2 + 2 = 4.

Then I had to laugh when it came back citing something from Mathlib. It was like, “This is already proved in Mathlib,” where the theorem is literally the 2 + 2 = 4 theorem. So I was like, “It's done.” That wasn't exactly what I was looking for, but I guess I got what I deserved for asking it such a basic question.

Nathan Labenz

Did you use the web interface, or the terminal UI? I started by having Claude Code install the terminal, and then I was using that a little bit. Somehow, it tipped me off to the fact that there was a web interface, so after that I moved over to the web interface.

Tudor Achim

Yeah, that came out last week, and it's probably a little bit more appropriate for those types of questions. We wanted to roll it out on the terminal because I think it makes it a little bit more clear what the tool is great at. Lots of things can answer 2 + 2 = 4, but—

Nathan Labenz

Even I can answer that.

Tudor Achim

Using a calculator. Yeah. For a while, we were talking about how to describe what Aristotle is. It's kind of like an amazing calculator where you can imagine that you could just talk to your calculator.

It has both the reliability—you know that if your calculator gives you an answer, it's correct—but it's not very expressive at the same time. Something like ChatGPT or Claude is very expressive, but sometimes you have to double-check its work because it doesn't always have the verification.

The intent is to put those together, and it turns out that the first things people really want to be sure about and verify are more complicated things.

So I think you probably found this out, but the complicated things are where you really start to have aha moments when you're using it.

Nathan Labenz

Yeah. Let's get into Aristotle. I appreciate the time spent in remedial education. I think it's beneficial not just for me, but hopefully everybody will now be able to grow what we're about to get into much better with the foundation we've laid.

Aristotle has 3 core parts. I'll sketch them, and then you can give me the double-click on them. First, there is this Monte Carlo tree search–type thing. I think of that as an AlphaGo-like structure, where we're systematically exploring the space of moves. I guess that's where I got the chess analogy: I was making this equivalence between Aristotle, at least that part of Aristotle, and AlphaGo.

So maybe I can make this move, and then there's this learned scoring function that's like, “Okay, does that move seem promising? Does this path, this branch of all the possible moves that I could make, seem promising? Do I seem like I'm getting closer to my goal?” With that, you can grind things out and run a deep tree search.

The second part, in some ways, jumped out to me as even more interesting, and I really want to dig into the metaphysics of it a bit, because this is the lemma-based informal reasoning system. I take that to be saying, “Okay, if I have some really big mountain to climb, and it's maybe so big that I can't just grind my way through it—it becomes impractical to grind my way through all these small, localized steps—then it's sort of guessing what the base camps are that I would want to get to along the way.”

Those are the really good waypoints such that, if I can get there, then I know I've made it somewhere. But that's really interesting because it strikes me as behaving a little bit more like a language model, where it's guessing and not so formal. It says in the technical report that it is an informal reasoning system.

Then there's a third part, which we maybe don't have time to go as deep on, specifically dedicated to geometry. In the technical report, you describe that as being like AlphaGeometry, which I think DeepMind developed. Correct any misconceptions that I have there, and give me the double-click on what more I should understand about how this thing works.

Tudor Achim

Sure. I think you covered the components pretty accurately. One thing I have to say is that we revamp our systems pretty often here, so I think Aristotle now looks quite different from Aristotle for the IMO. A lot of things are consolidated and improved.

You made this point about the Monte Carlo tree search being more of a grinder. I wouldn't quite characterize it that way. The Monte Carlo tree search is actually doing a lot of inference on its own about high-level steps. The lemmas we're talking about are much closer to solving a challenging math problem than they are to proving that 2² = 4. There's a lot of reasoning that goes into them.

In some sense, it's grinding once you get low enough in the search tree, because you're just closing out cases or easy subproblems, but it's really solving hard problems on its own. When we combine it with the informal reasoning system, you can almost think of it as a form of context management.

Ultimately, you need to end up with a Lean proof, and that's going to involve big steps and small steps. When you're focusing on the smaller steps, it's helpful not to have to remember the entire context of the bigger steps.

It turns out the informal reasoning system itself makes enormous quantities of mistakes. One should not think of it as, “Oh, it's a really smart human that's laying out the steps to base camp.” It's more like a system that can propose lots of things that are wrong, don't have to be formalizable, or are not even correct, and you try to assemble things from that.

You can think of both of them as doing the same thing at slightly different scales and complementing each other. They're actually all LLMs. As we described in the technical report, the tree search itself is driven by language models. Part of the language model proposes steps, and part of it scores steps, but they work in concert to solve the lemmas and, eventually, the full problems.

As you mentioned, AlphaGeometry is a slightly different system. We're exploring high-level steps and then trying to use an algorithm to grind through the rest of it. If we're talking about systems grinding through a lot of math, I would say AlphaGeometry and the deductive reasoning system are really grinders.

They're really trying to find every possible conclusion of a geometry diagram. I would say there's not too much pattern-recognition intelligence going on there.

Nathan Labenz

Yeah. And that's because geometry, if you think about it, is more constrained. You basically have points. If you have 3 points, there's only so many angles involved. Obviously, if you go to 10 or 15 points, things blow up pretty quickly, but it also becomes hard for humans to solve. I think that's why geometry was among the first classes of competition problems to fall to AI and automation.

I think there are also a couple of other components that might seem simple but are nontrivial, that the Aristotle system handles and that are improving independently. One is autoformalization: taking input that you provide in natural language and faithfully translating it into Lean in the best possible way. Relative to our competitors, at least, I'm not aware of anything that's as good at that as we are.

There's also theory building. Sometimes, in the course of solving something, you have to create new theories and new structures that might not exist in Mathlib. Aristotle has the capability of building those on the fly and incorporating them into the proving process.

Another funny anecdote. You're referring to what I discovered as informal mode, right? I think real users would not do this, but you can provide anything—any natural-language input—and the system will then try to prove it. I asked it to prove “All is love,” and it came back and said, “This is a philosophical statement and outside the scope of the Lean 4 kernel's ability to prove.”

I also asked it to prove, “Epstein did not kill himself,” and it came back and said, “This is a statement about current events,” and again, it's outside Lean 4's ability to prove. This does get back to the metaphysical question that I find so perplexing: the translation from the messy real world of human affairs and intuitions to the formal definitions of, “Okay, this is actually the thing that we would want proof of.”

I found that very interesting—that you had such a thing at all. What do you have a sense for? I do want to get into a little more detail about how you technically created the models and all that stuff, but on my spectrum from 2 + 2 = 4 to “All is love,” how do you think about the intuition for what the boundary is of what is inside the scope of what it can prove?

When I listened to your previous interview with the Sequoia folks, it seemed like you had the sense that eventually, as the system and systems like this become capable enough, more and more things that are of interest to everyday people will start to become the sorts of things they can do. How do you think about that boundary, and how does that boundary expand over time?

Tudor Achim

I think the ultimate boundary of a system like Aristotle is reasoning through any problem where people can also agree on what it means to be a valid sequence of reasoning steps. Right now, you have math, which is one obvious example. When we talk about mathematics being the same as reasoning, that chess example you gave is a perfect one. You can express the logic of a chess game and then check it, right? Then you can reason about it.

One area that's really going to touch a lot of people's lives is that you can use the same reasoning approach to think about software. When people write software, they write things called unit tests and integration tests. That's having the computer run the program and check the output against what they expect. But that's what they do after they've written the code.

It turns out that when engineers are writing code, they're thinking logically: “Okay, if I have this range in my input, as I go through this for loop and these if statements, it implies certain things about the output.” That itself is logical and mathematical reasoning.

We're starting to see API users reason about programs in the same way that they can reason about math. People are writing cryptography implementations and then checking, “Hey, is there any possibility that 2 inputs might give me the same output?” That would violate a certain principle of the cryptographic algorithm.

They might be implementing a controller for an autopilot and asking, “Is there any sequence of inputs for which I'll have an unstable dead zone or something?” I think the same kind of reasoning that was good for math will then go to software and help take us to a bug-free software future.

Now, Vlad and I disagree a little bit. It's not clear to me if we'll be writing history essays or something. Maybe there's a way to value them objectively, but I think the boundary is really in anything that's quantitative and logical in nature.

Vlad Tenev

Yeah. I think in the first version of Aristotle, it would actually formalize and build a theory for your “all is love” example, and it would give you a correct proof that it’s probably true. I think that surprised us. People were asking all sorts of questions. We had people asking biology questions, and we’ve had people ask it medical questions. Of course, economics and financial math, and Tudor mentioned computer science.

I think it actually surprised us how broad a set of things it can successfully create a theory around and formalize. The constraints we put in place were just that, when you’re building a product, you want to make sure that you deliver value. At this point, I don’t think we provide the most value if you want to write a history essay, right? So we’re trying to nudge people to the point where they can discover what Aristotle is really, really good at as easily, quickly, and simply as possible.

Over time, you should expect the surface area to increase, and we’ll start formalizing more things. I don’t think it’s inconceivable that, at some point, it pulls current events and news from the internet, puts out the axioms, and can sort of fact-check and make conclusions based on real-world events. That’s not our focus right now, but I don’t think it’s a crazy thought.

I ask a question sometimes. I’m interested in astronomy, right? I wanted to know when the next total solar eclipse is that I can see from within 50 miles of Palo Alto, California. The models usually struggle with this type of stuff because nobody’s asked that identical question on the internet, so they can’t pull it. You actually have to do some math. You can imagine there’s a spectrum, and there are questions like this that a model that can reason from first principles is going to be way better at.

Nathan Labenz

Okay, let’s talk about just how you created this thing a little bit, and how your experience and lessons learned relate to some of the live questions more broadly in the AI space. I think you can take it on faith that folks listening to this show will be familiar with things like reinforcement learning from verifiable rewards and will certainly understand how the ability to generate synthetic data feeds into a system like that. I’m sure that’s part of what you’re doing.

But what more can you tell us? Would it make sense to start training something like this from an off-the-shelf pretrained model, or does the messiness that those LLMs start with corrupt or pollute the purity of the mathematical reasoning too much? Can you tell us anything about the size of the models, whether that’s parameters, tokens, or whatever?

I’m also interested in whether there’s any role for taste in this process. Obviously, mathematicians are very interested in correct proofs, but they’re also interested in these eureka moments and the sense of elegance in a proof. There’s a sense of beauty that matters as much, I think, to many people as correctness—or maybe not as much, but it’s certainly heavily weighted.

I also noticed that test-time training is part of this, and I think that’s a huge trend that I’m watching in general. You can swing at or take any of those pitches, but what do you think are the most interesting next-level details that people can use to inform their own AI worldview?

Vlad Tenev

Well, first I have to say that if your audience knows about reinforcement learning from verifiable rewards, you’ve got a great audience. That’s synthetic data.

Tudor Achim

I think that is a safe assumption. Nobody was talking about that stuff, right? It was like science fiction almost. It’s cool to see it entering the popular consciousness.

I want to address the taste question because that actually strikes at a key thing that companies can decide on. We got gold at the IMO. We had a very powerful system, and it was obvious that we had to give it to people. There are 2 ways you can do it.

One way is to say, “We’re going to keep this in-house. We’re going to recruit some great mathematicians to come in-house and work in secret on problems.” As they make progress, we say, “Well, Aristotle’s now done X, Y, and Z.” That’s one way of expressing taste in the research.

The other way, which we ultimately decided to do—and we think it’s been great for the community—is to say, “We’re not going to be the ones to decide what’s important in math. We’re going to make Aristotle accessible to everyone.” So we opened up the API and the web interface. There are a lot of great features coming.

In this scenario, taste is expressed by the community through the revealed preference of what they submit to the API. We don’t choose what kind of math they do. We’re not saying, “Hey, Navier–Stokes is more important than P versus NP.” It’s the mathematicians that have the credits on the API who can say, “We care about X,” or some other thing.

That’s why we’ve seen so much interest in computer science, crypto, and certain branches of number theory. For a while, there were people doing a lot of interesting conjectures in graph theory on the platform. I think that’s actually the right way for companies to engage with the community: You open the system and let people decide where they want to allocate those compute resources. I think that’s an important decision. We’ve come down on one side of it, but I think that’s the right long-term approach.

I think there’s a philosophical question in there, too. Are we headed for a future where the AI labs themselves are going to generate all the discoveries? Will the cure for cancer or diabetes look like a giant AI lab with a 2-gigawatt data center just turning on this problem, and then it comes out and they capture all the value?

Or does it look more like millions of people empowered with these tools, working independently and collaborating? In that world, they’ll get the credit, and the value will largely accrue to them. I think we believe that the second world is more interesting, and it’s probably the one that’s more likely. The first one is rather dystopian and less likely.

We noticed that because, when we rolled out Aristotle, we had one view of what people would use it for, but then we started getting all these AIME problem results and things like that. We’re not going to run on all the AIME problems. We’re not going to do computational learning theory formalizations in-house.

The amount of cool things being done with it just explodes if you make it generally available. I think it’s not only right from a business strategy standpoint, but the world that we build assuming this path is a better world that I would like to live in.

Nathan Labenz

So that speaks to taste in terms of problem selection. But I was also just thinking in terms of training the model. You’ve got the correctness signal, but maybe one heuristic for elegance would be just brevity, which is one way of trying to send an elegance signal through a deterministic mechanism.

I’d be very interested to know if there’s a panel of mathematicians that you have reviewing solutions for elegance to make sure that this thing is not just a pure grinder long-term, but really has more of a eureka flavor to it.

Tudor Achim

Well, if brevity is the definition of elegance, then our 2 + 2 = 4 proof probably takes the cake, right?

Nathan Labenz

Yeah, I can’t get any shorter than that. I would feel bad for any mathematician whose job it was to compare AI proofs. That’s certainly not the job I’d want.

Vlad Tenev

It’s a big business these days.

Nathan Labenz

Across all domains, right? Many billions are spent on expert validation of AI outputs.

Tudor Achim

Yeah, we’ve done essentially zero of that in the 2 years we’ve been around. I think the metric we optimize for is the net present value of future proofs, or the computational cost of future proofs. That guards very naturally against certain phenomena.

When you’re solving easy problems early on in reinforcement learning, you absolutely can solve them with grinding. You can say, “Let me just do brute force.” But you know that if you do that, it’s going to cause issues later because you haven’t learned how to do more complicated things.

In contrast, if you’re given 2 proofs that are not grinding, but one is drastically longer and more inefficient than the other, you prefer the more efficient one. There’s a tension there because you can get more efficient by grinding, right? But that messes you up in the future.

It’s a balance that our AI researchers strike based on their intuitions about what will be helpful long-term. But we’ve never had panels of mathematicians do A/B testing on proofs or anything like that.

Really, you want to give your system as few priors as possible and just run reinforcement learning at scale. There’s a famous essay called “The Bitter Lesson,” which I’m sure your viewers are familiar with. We really believe in that at Harmonic.

To get to your question about how we started, sometimes we’ll start from pretrained models. Ultimately, you want to do whatever optimizes that net present value of future cost of proof. Pretrained models are great for that.

At some point, you might ask the question: Is that going to bias you too much toward how humans do math? You might want to mix in reasoning systems that are not trained from human knowledge, right? They have more entropy and more complementary knowledge.

So, that kind of thing we always play with, but it hasn't really been the limiting factor so far. I think that pretrained models are a great starting point.

Nathan Labenz

Cool. Goodfire just announced today that they raised a bunch of money at a unicorn valuation. I was a very small-scale early supporter of theirs, and it has me thinking. This also kind of connects to Vlad's comment, where you said that the system can sort of invent new theory.

One big thing that people have said AI can't do, or AI can never do—which is always a dangerous position to take—is come up with new abstractions. Sure, they can learn from what we have done and what we've encoded into language, but will they ever come up with their own abstractions? I think that's increasingly a hard position to defend.

What's so interesting with Goodfire is that they're now starting to look at model internals and unlock new kinds of understanding based on looking at what the model has learned. The famous one they just put out is new biomarkers for Alzheimer's that people didn't know about, but the model was able to figure them out, and they were able to figure out what the model had learned by looking internally.

I'm wondering: Have you guys done any interpretability work on your models? Do you think there's a different kind of latent space that you're tapping into? Do you see hybrids as part of the future?

One thing I could imagine happening is starting to stitch together a mathematical superintelligence with a more fuzzy, associative, understand-the-world superintelligence—perhaps later in post-training, or later in the training process—to try to get the best of both worlds.

One of the things I'm very excited about is eventually Aristotle powering a spacecraft, right? Much like HAL 9000, but a benevolent one—one that doesn't go crazy. So, yeah, I think eventually you'll see it expanding into more real-world things. I don't know if you're as excited about that. A safe HAL 9000.

Tudor Achim

A safe HAL 9000, I think, would be very valuable. To your question on interpretability, I think interpretability is often used as a proxy for trustworthiness. A lot of the reason people explore interpretability technology is so they can make sure that the system does the right thing or aligns with the user's intent.

When it comes to trustworthiness, we made the explicit decision at the very beginning of the company to focus on Lean by outputting our reasoning in a formally verified way. That is the most interpretable possible output. The computer can check it; if a human wants to understand how the proof works, they just keep hitting “go to definition.” It's almost like navigating through a codebase. There's no more interpretable way to output math than in Lean, really. That's the maximal version.

So now the question is, how interpretable is the model? I think, in the context of the bitter lesson, we just focus on letting the system do whatever it can to optimize for computationally cheap proofs of more and more complex things, with the caveat that it has to output in a way that's verifiable. I think down the road we're very curious: How does it do math? How is it so smart? We'll look into that, but for us, we've solved the trustworthiness question up front by focusing on formally verified output.

Nathan Labenz

Yeah. Okay. That's quite interesting. I do feel like I have this one kind of mental model. Mathematicians are famous for visualizing things. My kind of visualization of what is happening in a large model is sort of like shrink-wrapping reality. You've wrapped in plastic all of the internet data, or all of whatever domain it is that you're trying to learn at scale, and you're just sucking all the air out of it and gradually shrinking down to whatever, hopefully, is the true structure.

It strikes me that, in math in particular, that structure might be amazingly simple. There might be really interesting things to learn by running that process and then cracking it open and seeing what is inside. I would expect it to be maybe a lot more interpretable internally than something that has had to learn all of internet data and can recite Wikipedia and all that sort of stuff.

Tudor Achim

I actually think what these models are doing is interesting because they're smashing together all of the techniques that all mathematicians have used before. While I haven't seen the spark of superintelligence yet—where it's some breakthrough, eureka idea that's incomprehensible—I would say that if you push it toward learning how the models do things, you kind of ask it to solve more and more complex problems and just see, “How did it pull together these 3 subfields of math in a way that no human has done before?”

I think that'll be a lot more interpretable and comprehensible than trying to dig through it. I might be wrong, but that's probably where I start to interpret how it does things.

Nathan Labenz

Yeah. So, does that mean maybe we can look at different levels of problem difficulty? We've got the AIME problems. There's definitely a phenomenon happening right now where people are using either Aristotle by itself, or—and I've also seen increasingly more examples of GPT-5.2 Pro—to generate a proof in token space, then bring it over to Aristotle for formalization.

Then, of course, we've got the IMO. If I understand correctly, everybody who—and I think it was just 3, right?—you guys, OpenAI, and DeepMind got gold-level performance. I think everybody missed the same 1 question, which is really interesting to me. I'd be interested in your thoughts on why that was so consistent.

Then, of course, we've got these extreme problems where you would need this sort of Move 37 moment to solve them. Maybe sketch out where we are on this curve of problem difficulty. Are we just going to ride a smooth exponential—METR task-length style—all the way up to Millennium Prize Problems, or do you think there are going to be breakpoints of some sort, where you might need a new architecture, a new insight, or a new learning method to get from 1 range of problem difficulties to something that's qualitatively different?

Tudor Achim

I think on the IMO, the 3 labs that announced gold-medal performance—you know, us, DeepMind, and OpenAI—all missed question 6. I think it wasn't super surprising to us because question 6 is probably, I don't know, 5× harder even for humans, right? It's just a more complex question with lots of steps, and it requires this type of spatial reasoning that's more difficult to encode formally.

We were running our system on it quite a bit, and we felt like we saw signs of life. So it's definitely not inconceivable that before too long, question 6 is going to fall and be gobbled up just like the other questions. Even 1 year before, questions 3 and 5 would probably have been well beyond the reach of most models. So I think it does appear to be more or less a smooth exponential.

Vlad Tenev

Yeah, I agree with that. I want to highlight, though, that there are 2 aspects of this. I think we're continuing to see a smooth exponential in terms of AI capabilities in math. What I think is a little more interesting, and was less predictable before, was that there's now definitively been a phase transition to formal.

Years ago, if you'd asked someone, “Hey, could you automatically formalize a number theory paper in Lean, Rocq, or Isabelle—these other languages?” you would have been laughed out of any room of mathematicians you were in. Today, we're seeing people upload the full text of a math paper and run Aristotle a few times. We're thinking of adding a Ralph button to just keep going, keep going, keep going, and then you get a formal version of it.

I think that phase transition has essentially come and gone now because of Aristotle. In the next couple of years, as AI keeps improving, the fact that we can now formalize the AI's arguments obviates the need for humans to just be the verifiers—just sitting there and checking whether some output is correct—and turns us into the tastemakers. So we're the ones setting what problems to work on and whether we're happy with the techniques used. I think that's the interesting transition that's happened.

Smooth exponential capabilities, but I think we've gone from 0 to 1 on verification. I think that's such a great point because there was some debate about this at the beginning. In a way, if you look at DeepMind, they started with formal, with AlphaProof, which was the silver-medal-winning model back in 2024. It was a great result at that time, and that was a formal model. Then they went back to informal for Gemini this year, and I'm sure they ran AlphaProof. Maybe it was just that AlphaProof didn't do as well. OpenAI, obviously, was informal.

But if you think about it, let's say we go to a world 5 years from now and autonomous math being done by AIs increases. Instead of 5- to 10-page proofs, you're starting to produce 5,000-page proofs—which you should assume, as these models can autonomously reason more and get more efficient, they'll produce longer and longer output per unit time, and it's going to be a proxy for complexity. Who's going to review that? Nobody's reading a 5,000-page math proof.

So I think it's becoming even clearer that the future is formal, because you have this problem: someone has to validate it and check it.

And we want to make sure that the time to validate a check doesn’t actually grow linearly with the complexity of the proof.

Tudor Achim

Yeah, that was really the founding thought experiment of Harmonic. We asked ourselves in 2023, “Okay, so these models can do high school math poorly, but they could do elementary school math poorly a year ago. What happens in 10 years if we ask it to prove the Riemann hypothesis?” Any model will make an attempt at it and give you 100,000 pages of output, which you might as well throw in the trash for 2 reasons. First of all, there’s probably a mistake somewhere. Second of all, you can’t process it. There’s just nothing to do with it. You can’t wrap your head around what is going on in that proof.

There were 2 hypotheses, both of which have been proven out. First of all, outputting math formally makes it digestible for humans, and there’s a high level of certainty and trust. Secondly, it’ll lead to more efficient ways to do reinforcement learning for math, which is what we saw proved out. If you compare the resourcing we’ve had compared to the big labs, we’re punching well above our weight at the IMO.

I think, in our view, the debate on formal versus informal is settled. Clearly, it’s going to be formal. One can debate what the most efficient way to train a model is. There are some aspects to informal reasoning that are helpful, but I don’t think we’re ever going back to a world where we say, “Oh, it’s just going to be informal from here on out.”

I think the interesting question, though, is to extend this to software, because the same things actually hold for software that hold for math. Let’s say AIs are getting to the point where they can autonomously work and create a software project over a period of a week or multiple weeks. The Cursor team ran this and generated a Chromium-compatible browser, right? It was something like 1.5 million lines of code. Incredible.

So who’s going to read that code and find all the security vulnerabilities and bugs? Is that code in the future that’s generated by AIs going to be in Python and Java anymore? Why would it be in Python and Java? Those are just languages optimized for human readability.

The answer, we think, to humans reading and trusting something—or even an AI that the model is collaborating with checking something—is the same. You want to make the cost of verification as low as possible, and that makes us believe that the future of software is formal as well. More and more software will be written in formally verifiable languages.

Vlad Tenev

Yeah. I think Lean is our favorite language. It would be amazing if everyone could write in Lean. I think that as AI writes more and more code, it will be easier for people to accept that. But we’ll see.

Tudor Achim

Yeah. I’ll start with mission-critical, important stuff, where bugs are much more serious and much more costly. There are a bunch of domains that are already doing formal verification for software, but they’re doing it in a very artisanal way. They’re hiring Lean, Rocq, or Isabelle experts and painstakingly formalizing stuff.

I think you’ll start to see it accelerating the work of those people first, but then it’ll just diffuse. You’ll see formal vibe coding before too long.

Nathan Labenz

Yeah, I love the term “vibe proving,” by the way. I think that vision is an incredibly compelling one, and it’s also one that I’m still kind of wrapping my head around.

For listeners who haven’t already heard it, I did an episode with Kathleen Fisher, who was at RAND, I think, and now has just moved to ARIA in the UK to lead their whole operation, and Byron Cook, who’s a legend of the formal methods field at AWS. They’re right there with you, envisioning this world of basically totally verified, bug-free software, starting with mission-critical stuff but potentially extending to everything over time.

One thing—I don’t know if it’s a worry that I have, or what exactly, but I’ll just frame it as a question—is, if we are training an AI to be superhuman at formal reasoning within the formal reasoning system that we have, how do we get new abstractions from that? How do we get a sort of Einstein kind of moment?

It seems that at some point we all thought the world was just naturally 3D, and that was obviously intuitive. It’s come to light now that this was an adaptive understanding of the world that served us well as monkeys and allowed us to survive, but at the end of the day, we now know that it’s a lossy approximation of true physics.

So I’m wondering: Do we have any room for doubt or worry that the math we have now, as sophisticated as it has become, might also at some point prove to be not quite the right paradigm? If you’re training in this purely formal way, is there any way to punch your way out of the box, as Einstein did? He seems to have broken the fourth wall conceptually.

Tudor Achim

The key thing to remember is that he was able to describe his theory rigorously and formally in the framework of differential geometry. The point I was making earlier about math being reasoning is the point I’ll appeal to now, which is to say that no matter what complicated theory somebody might come up with to explain how the universe works in the future, if it’s going to be based on a series of logical deductions that can be explained to someone else and checked independently, that is itself a logic that can be encoded with Lean or other languages like it and then verified.

The axioms that Lean is based on are so minimal. They’re just expressing the most basic possible common sense about how reasoning should happen: One thing might follow from another, or if 2 things look the same, they are the same. That’s the level of axioms we’re talking about.

I really don’t think there’s any conflict here. One should just think about formal reasoning as an especially detailed version of informal reasoning that a computer can check automatically. There’s no limitation to it. Sometimes it might be a little more verbose than you’d want, so you want to write tactics and things to cut down on that, but there’s really no fundamental tension between the 2.

You might also be thinking about Gödel’s incompleteness—the fact that in any sort of axiomatic system, there are statements that are true and unprovable. There are also statements that are undecidable and independent. There are a bunch of edge cases here, but I don’t think that prevents us from making a lot of progress and proving the lion’s share of useful things.

There could be things that are unprovable but true that are very, very useful to know as well. But there’s no way to know unless you explore the frontiers.

Nathan Labenz

Do you think there’s always going to be a role for entropy of some sort in these systems?

Tudor Achim

I think hallucinations are a key part of a reasoning system. Hallucinations are what allow a model to explore something that has never been encoded by a human before.

When we run Aristotle, whether it was at the IMO or now, it makes a lot of mistakes. It tries a lot of paths that don’t work, but that exploration is the very thing that lets you get the right answer after enough attempts. Entropy is crucial.

I think this whole notion of seeking fundamentally hallucination-free LLMs doesn’t really make much sense. Of course, you want to pair them with a system like Aristotle that can verify things in tandem. But entropy and hallucinations are a key part of the training process for models like this. You’ve got to be able to pose false statements in order to prove that they’re false.

Vlad Tenev

You learn like humans. You know, you try.

Tudor Achim

Some of the most creative humans are the ones that hallucinate the most.

Nathan Labenz

So, what’s kind of the latest progress on the path to superintelligence? You said—and I think this is true of all good frontier AI companies, whether at the application layer or the model layer or any hybrid of those—that you’re updating your systems frequently.

It sounds like there’s kind of a convergence of some sort going on between the tree-search part and the informal lemma guesser that you described in the technical report. What can you tell us about what the trends are right now?

Tudor Achim

I think a lot of the—well, just to review the progress—we started in 2023, and then in 2025, we achieved gold performance at the IMO. We topped out this Vina benchmark [?] at the end of the year. With our public API, users started solving hard-ish problems that were unsolved for 30 or 40 years.

I think there’s a very clear trend in capabilities. The phase transition I mentioned has also happened. I think what’s next for Harmonic and for the field at large is a couple of things.

We can expect Mathlib to grow. Think of Mathlib like the Wikipedia for math that’s computationally certified. As Aristotle makes it possible to auto-formalize a lot of math, you can expect that users will start contributing a lot of pull requests to Mathlib, and that makes it possible to solve more and more problems on top of that base.

When we look at how mathematicians are using our API, people are certainly starting to work on more important unsolved conjectures that a lot of people would care about. You can think about conjectures as, “Okay, there’s a conjecture that’s technically been open, but nobody really cares about it.”

So it's not like people are trying all the time. But now you might have some conjectures that a mathematician might try once or twice a year, just taking a shot at it. Maybe 100 mathematicians would try it, and then eventually—well, the Millennium Prize Problems are where any mathematician would be happy to spend years if they might be able to solve one. So I think what you can expect from Aristotle and other systems is that more and more problems get picked off.

It becomes easier to use, and it extends to software, as I mentioned. We have users using it to check safety-critical software, whether in Lean or other languages. Overall, if I had to pick out just 1 trend, it's really just that formal reasoning goes more and more mainstream. As more stuff is produced with AI, I think you'll see, complementarily, more formal reasoning to verify all of it.

On the product side, we've gotten a lot of feedback from the folks using it. Obviously, whenever you've got customers using a technology like this, they're very passionate. So there are lots of ways in which they're still complaining about things and improving the ergonomics of it, making it so people don't have to hop between so many different tools, and so we can solve their problem as simply as possible and at the lowest possible cost. You should see that continue to improve.

There have been updates to the system pretty much on a daily basis. Maybe you've seen some of them as you've been experimenting yourself, but that's going to continue, and you should expect it to get exponentially more useful over time.

Nathan Labenz

So maybe a good place to close is the vision for what that looks like as you succeed. Obviously, 1 thing is solving the Millennium Prize Problems, but I'd love to get a little bit more of an intuitive understanding than that. One dichotomy that comes to mind is this very formal-reasoning-based paradigm versus what I think of as intuitive physics.

It does seem like models are very good at developing intuitive physics in any number of spaces. Folding a protein with a model is not something that's done in a formal way; it's just something where whatever mess of heuristics they've learned, they can do a protein fold orders of magnitude faster than we would be able to do it if we were going to do it through a sort of physics-based simulation approach.

When we think of there being no limit to math, and what mathematical superintelligence looks like, I also think of Eliezer, who once famously wrote—at least famous to me—that a real superintelligence in his mind could look at 1 still image and deduce all of physics from the information contained in that 1 still image. That also connects, I guess, to test-time training.

What is your vision? You can bounce off any of those concepts, but what is your vision of how this thing evolves? Is it an ever-bigger tower of formal statements? Is there some role for new kinds of intuition, new abstractions that emerge out of that, that aren't so strictly defined but are potentially useful? What is this thing doing in 2030, once all the Millennium Prize Problems are solved?

Tudor Achim

I think that by 2030 we will have theoretical explanations for basically everything. If you look at the history of science, there are leaps of intellect and leaps of data. The microscope comes along, and all of a sudden you can build a lot more theories of biology. Then the electron microscope comes along, and you can build more theories of chemistry.

Right now, there's really been a shortage of people who are able to reason logically at the highest level. So when you think about unifying general relativity and quantum mechanics, it's just a very hard thing to do. I think what you'll see is that anything that can be posed mathematically—which is what underlies all of science—we're just going to get theories for everything that are self-consistent and make sense.

I think we're then going to go back into a regime where we're data-limited. So we're going to have maybe 5 theories that unify QM and GR, and we're going to have to run very high-energy experiments to figure out which one is right. We'll have to wait a while to build those colliders. But at the very least, we won't be bottlenecked anymore on wondering whether we can explain something. We'll have a system that can explain anything perfectly correctly.

So it really will be a renaissance of science. I think you just remove the intellectual bottleneck in everything.

Nathan Labenz

So do I understand that correctly? Basically, you're envisioning multiple grand unified theories that all explain all the data that we have, and then it becomes a problem for the collider experiments to figure out which one of these is, in fact, right?

Tudor Achim

Yeah. AI is not omniscient. Whether it's our model or others, they'll be able to reason about anything that they can ground in their own logical deduction rules. But ultimately, I think there are aspects of the universe where you just have to run the experiment and find out how it really works.

Nathan Labenz

Yeah, wow. I just want to be clear: I think there's a lot of utility before you get there. If I analyze asymptotically where we get to, that's my point.

We've heard about centuries of scientific progress collapsing into 5 years. That sounds more like a few thousand years perhaps of scientific progress collapsing, and then you just have to get more data, but you'll have a superintelligent system that can help you.

Wow. Okay. That's about as grand of a vision as I've heard anywhere. Do you guys worry about the safety of these systems? We haven't talked about that really at all in this context, but I've done many explorations of different safety concerns. Eliezer, when he described the model—whatever AI he was envisioning—when he described it understanding all of physics from a single image, he also thought that was going to be super dangerous because it would be so powerful. How do you guys think about that aspect of this whole thing? I mean, we're talking about a lot of stuff in the next 5 years.

Tudor Achim

I think right now we're not so worried about it because the outputs of our system are constrained. I think the first dangers will probably look a lot like cybersecurity incidents, right? You have models that are making API calls, running autonomously, and interacting with other systems. So that creates API-level cybersecurity holes and the mechanisms to exploit those. I think you're likely to see a lot of those.

For our model, since the interface to the outside world is tightly constrained and it's not just going to fire off a request to your Gmail account or the iMessage APIs, we're a little bit further away from that. But you can imagine we're going to have to start taking that much more seriously when we do get to a point where we're connecting the model to the outside world and it's speaking through interfaces that are not just Lean files being output.

Nathan Labenz

Yeah, I do think a constrained action space is certainly 1 of my favorite paradigms for keeping things under control. But the whole Moltbook–Moltbot thing has been fascinating to watch, and I think we're entering a strange new world for sure. I think the benefit is that we're probably not at the danger frontier, so we'll have the opportunity to learn from others' mistakes, and hopefully they don't screw up too badly for us to learn.

Yeah. Okay. Fascinating stuff. This has been fascinating stuff, guys. I really think the approach is really interesting. The vision for how far we can expect—or even somewhat entertain the possibility of being—in 2030 is arresting, and both inspiring and, for me, a little bit scary. Anything else you want to leave people with before we break?

Vlad Tenev

I think for me—and you can kind of see this in the values that we put on our website, in terms of what we care about—we believe in a future where humans are going to be at the center of all this progress. I think that we'll definitely accelerate it, but humans should be in charge and calling the shots.

I think that's also why we care so much about putting this into people's hands and making them use it, and not just being a lab that runs things secretly and makes big proclamations, because I think humans need to be at the center of everything and still calling the shots. That's what we believe in, and in the future that we're helping bring to life.

Tudor Achim

Yeah, yeah. And I think just to add to that, for me, when I started using Aristotle, it was very different to have an experience where the output is always correct. And so I think if people haven't experienced that before, they should just try it out. It's a free to sign up for.

Nathan Labenz

Cool. Well, I'm sure there'll be plenty of ways to monetize mathematical superintelligence when the time comes.

Speaker 1

We might do ads, you know.

Nathan Labenz

Yeah, I can't wait for that. All right.

Speaker 1

And topical ads to life.

Nathan Labenz

Fascinating stuff, guys. I really look forward to watching your progress. Thanks for the remedial education and a grand vision today. It's really extraordinary.

What a time to be alive. Vlad Tenev and Tudor Achim, co-founders of Harmonic, thank you both for being part of The Cognitive Revolution.

Vlad Tenev

Thanks for having me.

Tudor Achim

Pleasure to be with you.