[BidClub_]
The a16z Show · · 67 分钟

AI 公司为什么会想放慢脚步?

Ali GhodsiMartin CasadoSarah Wang

AI与软件技术企业经营
YouTube
TL;DR
  • Ali Ghodsi 的核心判断是,“现在的存在性风险接近于零”,而公开传播末日论的领导者是不负责任的。 他认为,向公众谈论“全人类都将被抹去”的场景,可能伤害那些并未深入技术细节的人群的心理健康;Martin Casado 那位乡村学校教师姐姐甚至发短信问,要不要“为 AI 世界末日准备好小木屋”,而围绕暂停 AI 的政治机器已经开始运转。
  • Ghodsi 给出了递归式自我改进真正危险的4项条件:下一代模型必须同时需要超线性减少的 GPU 和训练时间、具备更高智能,并且这3项改进还要重复发生。 现实情况恰恰相反:前沿模型训练所需算力门槛已升至约50亿—100亿美元,各实验室每年只做1—2次这样的训练,流程“非常脆弱”,而且已经有多次训练失败。“没有证据表明这4件事正在发生。”
  • 真正的运营风险在网络安全,而不是超级智能。 从 CVE 发布到漏洞被武器化的时间,已从2018—19年的2—3年缩短至2022年的8—9个月,如今更是“基本只需几小时”,因此自动化检测和威胁狩猎迫在眉睫,因为“绝大多数组织离做到这一点还很远”。Ghodsi 认为,数据/AI 与网络安全正在合流,Databricks 的 LakeWatch 检测产品正体现这一判断。
  • Casado 对公关策略的批评是:“pace 就像一个让末日论者不高兴、也让政策圈不高兴的诡异谷。” 在 Casado 看来,Dario 的文件“完全合理”,但“与放慢节奏毫无关系”;实验室本应像 Zuck 那样把重点放在保障前沿模型安全上。Ghodsi 的反驳是公地悲剧逻辑:“我不会单方面停下来,那我不就成了冤大头。”他预计行业自律和政府监管会相互渗透,但没有说监管已经不可避免。
  • 企业层面的判断是:前沿进展就算停下来,对大多数组织也无关紧要。 企业运行的是“非常、非常强化且高效的 Google 搜索”式聊天机器人,而不是 agent 集群;瓶颈在上下文而非智能——“我们只需要它从60%提升到70%”——解决方案是像 Google 搜索索引一样离线计算 ontology。Casado 的保留意见是,前沿冻结“会对实验室造成灾难”,因为按他的估计,智能的价格大约每6个月下降一个数量级。
  • 成本控制正在成为一个市场:Databricks 通过 Unity Gateway 预算、智能路由和 harness 套利压平了自身成本曲线,让 token 数持续上升而支出保持平稳——同一个模型换一套 harness,成本就可能相差2倍。 Casado 透露,一家已实现规模化的公司上周开始从前沿模型转向 GLM,并估计按美元计算,开源模型约占初创公司支出的5%,但按 token 计算占比超过60%。Wang 表示,一些初创公司的外部产品使用中,开源模型占比已接近90%。
  • Agent 正在成为新的数据库买家:Neon 和 Lakebase 上创建的数据库中,超过90%是 agent 而非人类创建的。 Neon 没有继续参与 DBA 时代的数据库战争,而是专注于一个新角色:亚秒级启动、TB 级克隆、分支能力和对 agent 友好的定价;这与 uv 和 ripgrep 所代表的轻量、快速模式相同。
  • 在收尾谈到 p(doom) 时,Ghodsi 给出的答案是“不到10%——不,接近于零”,Wang 表示认同。 Casado 则换了个框架:“没有 AI 时,我的 p(doom) 远高于有 AI 时。”
摘要 · 为研究而整理的核心内容

1. 领导者有责任不要吓坏公众

  • Ghodsi 开场就亮明立场:“除非有非常、非常充分的理由,否则领导者有责任不要无端吓坏人。”既然“现在的存在性风险接近于零”,他反问:“那为什么要把所有人都吓坏?”技术细节应留在研究者之间讨论,而不是在电视或 Twitter 上向数百万人广播“全人类有10%的概率被抹去”。
  • Casado 给出的证据是,恐慌已经失控:他的姐姐是一名住在亚利桑那州乡村的学校教师,周日给他发短信问:“我该为 AI 世界末日准备小木屋吗?水已经备好了。你什么时候过来?”
  • 其二阶成本在于,Casado 进门时,Elizabeth Warren 刚刚呼吁暂停所有 AI 开发,“紧跟在 Bernie 后面,而 Bernie 也在和 Bannon——Steve Bannon——合作”。Casado 说,“联邦体系现在已经开始运转”,这可能与传播这套信息的初衷背道而驰。

2. 政治两端都在升温

  • Ghodsi 坚称,这场博弈并非单边进行:商业一侧,那些希望 IPO 漂亮收场的人会说:“别搞砸我的 IPO……大家能不能都闭嘴,让我们把钱拿回来?”他们有资源、有关系,也能“动用关系”。另一边则有人在问:“我们怎么把它武器化?……把这个种下去,把这些讨论串炒起来。”
  • Casado 补充说,两党——“不包括 Trump 本人”——都同意应该在某种程度上约束 AI;甚至 Greg Abbott 也说过,德州不能再建数据中心。

3. Casado 认为“放慢节奏”是一次公关灾难

  • 他的完整批评是:行业自律很正常,强调“安全和保障很重要”并呼吁加强监督,本来会是合理的信息。但放慢节奏“与安全和保障是正交的。你完全可以慢慢制造一件武器”。他的结论是:“pace 就像一个让末日论者不高兴、也让政策圈不高兴的诡异谷。”Dario 的文件“完全合理”,但 Casado 说它“与放慢节奏毫无关系”。
  • Hugging Face–OpenAI 事件成了检验案例。Casado 的反应是:“老兄,先他妈的把你们的东西安全做好。”这属于经典的安全控制,而不是放慢节奏。Ghodsi 反驳说:“但这就是放慢节奏”:如果安全团队要监控数百万 GPU 小时的强化学习训练中的每一个 token,速度“会显著放慢”,这正是实验室主张建立共同护栏的原因。
  • Ghodsi 用公地悲剧来解释:“我不会单方面停下来,那我不就成了冤大头……为什么不先让你停?”因此实验室会说:“你能不能进来把我们停下来?”两人都认为 Zuck 的做法更务实:他聚焦于安全、保障和自律;而 Dario 以“我们需要放慢前沿进展”开篇,试图同时回应末日论者和政客,结果“某种程度上两边都没满足”。
  • 这一节最后落地成一个笑话:“我们现在学到的是,Martin 真的很讨厌 pace 这个词。我以后再也不会对你用这个词了。”

4. 递归式自我改进的4项测试——而前沿训练目前恰恰相反

  • Ghodsi 认为,真正值得担忧的递归式自我改进需要满足:下一代模型所需资源和 GPU 数量以超线性方式减少,训练时间缩短,智能水平提高,而且这个循环还要重复——“必须同时发生,而不是其中任意一项”。
  • Casado 认为,只要有一项条件不成立——例如资源需求保持不变——硬件上限最终就会让进程自行放慢。软件能够自我编写并不算数:“Databricks 90%多的软件都是 AI 写的。剩下那一点如果也由 AI 写,会有关系吗?不会。”
  • Ghodsi 说,人们把“我的价值是什么、这对我来说是不是吓人”与 Bostrom 在2014年理论推演的超级智能混为一谈。
  • Wang 的补充依据是,训练前沿模型所需的最低算力仍在上升;她估计约为50亿美元,Ghodsi 接着说:“50亿到100亿美元。”他补充称,每家实验室每年只进行1—2次这样的训练,需要更多资源和人力,流程更加脆弱,对大型数据中心和网络的要求极高,而且“已经有多次训练失败”。这与更快、更便宜、更聪明的递归式改进完全相反。

5. 试金石,以及没有出现的蠕虫末日

  • Casado 给出的黑箱测试是:面对动态、自适应系统,要相信数据,而不是相信“你亲眼看到的东西”。假设 Anthropic 两周后只剩12个人,却以越来越快的速度产出模型,“我们大概就该注意了”。Ghodsi 进一步限定说,这样的情况足够构成证据,但并非必要条件;更相关的测试是,实际负责预训练和后训练的团队是否在缩小,同时使用更少的 GPU。“事实并非如此。”
  • Casado 回忆,美国曾因担心 Saddam Hussein 用 PlayStation 做模拟而对其实施出口管制,但这些担忧后来都没有成为现实。Ghodsi 开玩笑说:“也许只是年纪大了之后失忆了。”他的反驳是,今天的规模不同:“我们现在把更多东西互联起来了”,agent 可能开始在系统之间“跳跃”,并且“有点像病毒一样”扩散。
  • Casado 最尖锐的经验性质疑是:在互联网发展到相当阶段时,蠕虫已经“瘫痪医院、摧毁关键基础设施,并造成数百亿美元的经济损失”,而当时的网络建设远不如今天。“我们还没有看到任何与蠕虫早期相称的事情。这个脱节该怎么解释?”Ghodsi 没有完全回答,只说:“我晚上睡得很好。”

6. 真正的风险在网络安全,这是一场自动化竞赛

  • 最关键的数据是:从 CVE 发布到漏洞被武器化,2018—19年需要2—3年,2022年需要8—9个月,而“现在……基本只需要几小时。漏洞会立即被武器化”。
  • Ghodsi 描述的机制是:SOC 分析师醒来后会面对数百封检测邮件,其中很多是误报,而“人类响应攻击的速度不够快”。组织需要自动化检测和威胁狩猎,让 agent 攻击自己的系统。银行和部分高度重视安全的组织已经在做,但“绝大多数组织离做到这一点还很远”。如果行业跟不上,可能出现“网站宕机、整个系统暂时停止”,造成经济损失,甚至导致人员受伤——但这不是存在性风险。
  • 通过 Databricks 的 LakeWatch 检测产品,市场判断也浮现出来:内部 agent 现在生成日志、轨迹和指纹的规模高出“许多个数量级”,因此数据/AI 与网络安全正在融合,两个市场“最终会合并”。

7. 两个问题被混为一谈,以及谁有资格检查实验室

  • Ghodsi 将问题拆开:Bostrom 式超级智能——一个能在几秒内写出新颖、经过同行评审的博士论文,或瞬间完成数千年思考的系统——不会只是工程问题,“那会非常具有存在性风险”。但“我们现在还没有走上那条路”。
  • 真正改变的是,人们现在可以运行大量能力足够强的 agent:“我们以前根本不能说,‘把1万个 agent 放进沙盒里运行1个月,让它们完成价值1亿美元的工资劳动。’现在我们能做到。”Ghodsi 认为,随之而来的网络安全风险很严重,但可以通过工程手段解决。
  • 谈到实验室检查员,核心是可信度。如果由 Yann LeCun——Ghodsi 认为他对风险持怀疑态度——检查实验室后确认“没什么可看的”,Ghodsi 说自己会感到放心;如果 LeCun 出来时明显受到震动、并且改变了看法,那同样会是一个有意义的信号。Ghodsi 倾向于由一组多元化的检查员负责,也不信任那些早已先入为主的人。
  • 对 Elon 提出的让各家实验室互相审计,Ghodsi 说:“如果拳击比赛在擂台上进行,拳手是不是应该互相当裁判?能行吗?不行。”他预计,只要另一家公司发布强大的模型,竞争对手就会喊“犯规”,因此更倾向于第三方。
  • Casado 提出,行业自我监管可能演化为政府监管,并以 FINRA 为例——这是一个与政府相连的自律组织。Ghodsi 说,如果公司一边宣称存在性风险,一边要求监管机构来监管自己,监管者就很难一直拒绝;“我不认为这种情况会持续太久。”

8. 四维棋问题:末日论与 IPO 配额

  • Wang 转述了 Elon 的犬儒式解读:“一方面你说全人类都会死。另一方面你又说,‘嘿,你想要多少 IPO 配额?’”Ghodsi 的解释是,真实恐惧、监管自利和既得利益可以同时存在——“人们总能在脑子里找到办法,让这些事情和谐地对齐”。
  • 他也承认,行业一直有一种营销噱头:“最新模型太厉害了,简直是我训练出来的。难以置信。它几乎吓到我了。”但网络安全威胁同时是真实存在的,所以“也许两者并不矛盾”。
  • 关于公司为什么要上市,Ghodsi 说,作为一家已实现规模化的私人公司负责人,他认为公司其实更愿意保持私有;它们需要资本,并把规模扩张规律和资本本身视为战略优势。

9. 企业不需要更聪明的模型,需要的是上下文

  • Ghodsi 说,大约从上一年的 Q3 或 Q4 开始,人们就一直说 AI 比身边大多数人更聪明;但当他问谁在使用数百或数千个“协同成群、互相谈判”的 agent 时,几乎没人举手。大多数企业运行的是一个聊天机器人——“非常、非常强化且高效的 Google 搜索”——外加编程业务,后者“我们可以讨论 ROI”。
  • 他的诊断是,模型已经足够聪明,只是“没有参加过每一场会议”,不知道每个人脑子里装着什么。把这些上下文接入后,“能释放出巨大的生产力提升”。“我们不需要一个能求解 Navier–Stokes 方程的更聪明模型……我们只需要它从60%提升到70%。”他的挑衅式结论是:如果前沿模型停止进步,“对绝大多数组织而言其实没关系”。
  • Casado 的保留意见是:“这会对实验室造成灾难,因为智能的价格正在渐近式下降——大概每6个月降到1/10之类的。”

10. 账本的另一面:没人听说过的应用案例

  • Ghodsi 用一组上行案例回应末日论饱和:Crisis Text Line 使用大语言模型识别有自残或自杀风险的青少年;Omnipod 使用 AI 学习糖尿病患者的胰岛素释放和血糖水平,并据此释放胰岛素;Zipline 的自动无人机最初为非洲难民运送血液,随后扩展到其他场景。
  • 更深一层的案例是 TEDDY,即 Transformer-Enhanced Drug Discovery。该项目与 Merck 合作完成并已发表,预测的是基因调控网络的响应,而不是下一个 token,帮助区分因果细胞和反应性细胞,并有望降低药物研发成本。
  • Novo Nordisk 在所有试验中使用 Genie,将获得洞察所需的时间——例如一项肥胖研究——从数周压缩到数分钟。“我们希望这些应用全部继续,我们不希望暂停它们。”

11. Ontology:把入职5年的员工编码成一张图

  • Ghodsi 通过一个最好的类比来定义 ontology:两个能力相同的员工,一个入职第1天,一个已经工作5年。后者“拥有一套关于组织如何运转的 ontology”。“不要看组织架构图。不要去问那个人。去问这个人。”要捕捉这些信息,就必须把一切数字化——包括转录会议内容,哪怕这会带来法律问题——再将其提炼成一张带权限意识、供 agent 使用的图谱。
  • 架构上的问题是,今天的 agent 循环会一次检查一个 MCP server 的资源,就像 Google 如果通过实时爬取网站来构建搜索,每次只爬一个网站、耗时10分钟:成本高、速度慢、质量低。相反,应当“一直在离线计算那张索引”,这是一个类似 PageRank、但因访问控制、权限和异构对象而更复杂的问题。Ghodsi 认为 Palantir 擅长提取隐性知识,而 Databricks 可以据此自动构建图谱。
  • Databricks 自身就是证明:公司有数百万个节点,而且按 Ghodsi 的说法,其 ontology 比任何客户的都大。他曾问销售运营团队 Fortune 500 的渗透率;对方当时无法在飞行途中登录 Genie,于是他发短信给 CFO,CFO“直接把 Genie 的截图原样复制粘贴给我”。如今这个动词已经在公司内部固定下来:“有人能不能 Genie 一下?”员工会在会议期间查询这套 ontology。与一年前相比,“公司已经彻底变了”。
  • Ghodsi 说,Databricks 的 FDE 团队帮助企业收集缺失的上下文、构建面向客户的 agent,并加入护栏。他以 Fox 的 Sports AI 为例:系统可以回答体育问题,但会拒绝把对话引向政治的尝试。

12. 从 token 最大化到价值最大化,以及转向开源

  • 这段旅程始于上一年的 Q4 左右:Ghodsi 亲自把 AI 编写的代码提交到生产环境,并用排行榜推动其他人照做——“如果 CEO 能把代码提交到生产环境……你也应该能做到”。到了2—3月,token 最大化“已经失控”。
  • 解决方案包括为个人和团队设置 Unity Gateway 预算约束、预测性成本分析、自动选择更便宜模型的智能路由,以及 Omnigent 的多路复用 harness。“harness 本身很重要”:同一个模型换一套 harness,“实际成本可以相差2倍”。结果是 token 数继续上升,而成本大致保持平稳。
  • Casado 观察到一个市场信号:“历史上第一次,上周有一家已经实现规模化的公司说,它们要从前沿模型转向 GLM。”他将此与更早的 DeepSeek 和 Kimi 时刻对比,后两者都没有对市场产生可感知的影响。
  • 谈到初创公司,Casado 估计,开源模型按美元计算约占支出的5%,但按 token 计算超过60%。Wang 表示,一些初创公司的外部产品使用中,开源模型占比已接近90%,而内部使用仍更多依赖前沿模型。
  • 一个正在形成的模式是,用 Fable 或 Astra 做架构设计和审计,用更便宜的模型负责实现。Ghodsi 的总结是,应该把小型开源模型与大型专家模型混合使用,有时在两者之间来回切换,而不是每个琐碎任务都调用最新、最昂贵的模型。

13. 后训练的评估陷阱、原生 agent 数据库与 p(doom)

  • 谈到2023年夏天收购 Mosaic 时,Ghodsi 说,如今再为客户做预训练“完全没有意义”,因为强大的预训练模型已经存在。但针对有特定任务的初创公司,在开源模型上进行强化学习后训练,可以降低成本、提升速度,并保留其 IP 控制权。
  • 企业不太可能迅速跟进,因为“做好评估很难”。Databricks 生成了评估集,并把它们放在产品最前面,但客户并不使用,于是公司把它们移到后端,变成可选功能。Ghodsi 将其比作测试驱动开发:“所有人都说这是正确做法,但实际上没人照做。”
  • 按 Ghodsi 的说法,Neon/Lakebase 赢得 agent 数据库竞赛的原因,是它们专注于一个新角色。数据库启动远不到1秒,TB 级数据库可以在不到1秒内克隆,分支是杀手级功能,agent 试验期间定价也不会失控。他将这种做法与 uv 和 ripgrep 相比:两者都是对 Unix 工具的轻量级重写,但速度快、对 agent 有用。“Neon 和 Lakebase 上创建的数据库中,超过90%实际上是 agent 创建的。”Casado 说,他已经开始把 Neon 作为标准数据库使用,并惊讶于加入更大公司后,产品仍然有了实质性改善。
  • 收尾时,Ghodsi 对 p(doom) 的答案是:“不到10%。不,接近于零。”Wang 表示同意。Casado 最后重新定义了这个问题:“没有 AI 时,我的 p(doom) 远高于有 AI 时。”
完整逐字稿
Ali Ghodsi

As a business leader, there’s a tragedy of the commons. If you want to stop, if you want to go slower, why don’t you go slower? I’m competing. I want to win.

Martin Casado

There are almost 2 camps. One believes that this is actually an engineering problem, and the other believes you have to slow it down.

Ali Ghodsi

Humans don’t respond fast enough to the attacks that are happening. You need to automate all of those, and most organizations are actually not close to doing that. Is RSI, or recursive self-improvement, that the labs are doing leading us there? That’s the big question.

Sarah Wang

Something Elon said was that this is some elaborate 4D chess, because on the one hand you’re saying all of humanity will die—

Ali Ghodsi

Mm-hmm.

Sarah Wang

On the other hand, you’re saying, “Hey, what do you want for your IPO allocation?”

Martin Casado

For the first time ever, a company at scale last week said that they’re moving from the frontier models to GLM. Do you think that’s a trend, or do you think that’s just a one-off anecdote?

Sarah Wang

Thank you for being here, Ali.

Ali Ghodsi

Super excited.

So we obviously want to get to Databricks, but there’s a broader conversation going on right now about AI. Dario’s weighed in, Jakob’s weighed in, and Elon’s weighed in, but we want to hear what Ali Ghodsi thinks.

In terms of the topic, broadly speaking, of pacing the frontier, what is your strongest agreement with what’s being put out there? Where do you disagree? And maybe where is there nuance that’s not being captured?

Ali Ghodsi

Yeah, happy to cover it. Martin and I argue a lot, so—

Martin Casado

So—

Ali Ghodsi

I’m sure that’s not going to take long.

Martin Casado

We’ll try and rein it in this time.

Ali Ghodsi

Yeah, try to stay calm.

1. Leaders Should Not Freak People

I do think, first and foremost, that leaders have a responsibility not to freak people out unnecessarily unless there’s a really, really good reason. There are always different people in society who are at different places in their mind-space. Talking about existential risks and scenarios where all of humanity is going to be wiped out, I think, is irresponsible. It can tip a lot of people over and cause mental health issues.

Martin Casado

Unless you have something that’s going to wipe people out.

Ali Ghodsi

Yeah. As I said, if there’s an actual reason for it, then that’s a different story. But I think that right now the existential risk is close to zero, so why freak everybody out? It’s not needed.

There are risks—we’ll get into that. That’s probably where we disagree. But first and foremost, I think leaders should not freak everyone out. If there are technical nuances in how we’re doing AI research, researchers can discuss that.

Sarah Wang

Yeah.

Ali Ghodsi

You don’t need to go on TV or blast millions of people on Twitter every time, saying, “Hey, I think there’s a 10% risk that all of humanity is going to be wiped out.” I don’t think that’s helpful for a lot of people. I think it causes a lot of harm for people who get stressed out and aren’t immersed in the nuances of all this and what it means. I don’t think we should do that. I don’t think it’s fruitful. It doesn’t really help anyone.

Martin Casado

I think this is very, very true for the general public. My sister, who’s a schoolteacher in rural Arizona, texted me on Sunday and said, “Martin, should I prepare the cabin for the AI apocalypse? I’ve got water set up. When are you showing up?” She’s kind of a prepper anyway. I said, “Hold on. We’re not there yet.”

Ali Ghodsi

Yeah.

Martin Casado

Clearly, this has spilled over to the populace, which I agree is unnecessary and has blowback.

I think there’s a second issue. I don’t know if you saw this while we were walking in here, but I was checking X, and Elizabeth Warren just talked about pausing all AI development. That, of course, is on the coattails of Bernie, who is also working with Bannon—Steve Bannon.

So, in addition to scaring people, the federal complex is now spinning up. I think that could actually be quite contrary to the goals of the message. There’s more than just public hysteria at stake here.

2. The Politics Behind AI Pacing

Ali Ghodsi

There’s a lot of politics going on, but I’m in all of these groups, and I see both sides. There’s heavy politics happening on both sides.

Martin Casado

Oh, yeah. This is—

Ali Ghodsi

This is happening on both sides.

Martin Casado

No, this is a—I think both parties, not including Trump himself, agree that AI should be constrained at some level, right? Even Greg Abbott said you can’t have data centers in Texas.

Ali Ghodsi

I’m talking about the other side of this argument as well. Let me give you an example.

Martin Casado

Right.

Ali Ghodsi

I’m not talking about politicians. I’m talking about the fact that there’s politics happening on both sides.

Martin Casado

Right.

Ali Ghodsi

There’s politics on the business side—people who want to see great IPOs and get returns on their investments.

Martin Casado

Right.

Ali Ghodsi

They’re saying, “Don’t mess up my IPO. Can everybody just shut up so that we can get our money back?”

There’s that. They have resources, and they’re using them. There’s politics on that side, and those people aren’t sitting quietly and doing nothing. They can pull strings, and they have connections.

On the other side, there are all the people saying, “Okay, how do we weaponize this? This is awesome. This guy tweeted this. Let’s weaponize this one. Let’s plant this. Let’s pump these threads.”

Martin Casado

No, 100%.

Ali Ghodsi

You know—

3. Pacing Is Not Safety

Martin Casado

Totally. Let’s talk about the specific leverage point that everybody is using, because I actually think this is a classic case of a PR misstep. It’s not just the doomy-gloomy type of stuff.

Here’s the PR misstep, I think: It’s not unusual for industries to try to regulate themselves. It’s just not. Saying, “Security and safety are important. They’re important with every tech epoch, and we want to have some oversight”—that was very, very sensible.

Ali Ghodsi

Mm-hmm.

Martin Casado

The problem is that it’s couched in this notion of pacing, and there are a number of issues with pacing. First off, it’s orthogonal to safety and security. You can slowly build a weapon; that’s not different from building a weapon. People don’t feel that it’s genuine because these companies have been at a dead run.

Sarah Wang

They’re still buying more compute to be even faster.

Martin Casado

No, I mean, they just haven’t done it historically. But it also feels like this kind of milquetoast capitulation to the pause people. It’s like, “Well, you say pause, I say pacing,” which is almost like pause, but it’s not pause.

They chose this kind of flag to follow around pacing. But if you actually read the document Dario wrote—did you read the document Dario wrote? It’s a totally sensible document.

Ali Ghodsi

I read it, yeah.

Martin Casado

It just has nothing to do with pacing, right? And so I honestly—

Ali Ghodsi

No, he does mention it. Look, I kind of disagree. There’s a tragedy of the commons. There’s this argument: “Hey, if you want to stop, if you want to go slower, why don’t you go slower? Why are you writing articles?” A lot of people are making that argument.

Martin Casado

Yeah.

Ali Ghodsi

As a business leader, I understand that. There’s a tragedy of the commons. I’m competing, I want to win, and if you’re running ahead, I’m going to—

Martin Casado

There’s also the market equilibrium, which suggests that pacing is probably impractical anyway.

Ali Ghodsi

Yeah.

Martin Casado

Yeah.

Ali Ghodsi

I’m just saying that it makes some sense for people to say, “If you guys don’t stop this, this tragedy of the commons is going to continue.” I’m not going to stop racing because there are IPOs at stake and there’s competition at stake.

Sarah Wang

Yeah.

Ali Ghodsi

There’s also some animosity between the people. I’m not going to stop unilaterally. I’ll be a sucker. Why don’t you stop first?

Martin Casado

Right.

Ali Ghodsi

So then they’re saying, “Can you come in and stop us?” But I think you could also make the argument that if you look at the Hugging Face–OpenAI incident, that…

By the way, I think these companies are great, and I think they are probably investing a lot of resources.

Martin Casado

No, I agree.

Ali Ghodsi

But it's very clear—

Sarah Wang

Yeah.

Ali Ghodsi

If you read what happened, they weren't monitoring every token coming out. They were just running these RL experiments and then, after the fact, coming in and checking, “Okay, what happened?” So they should have paced. They should have been much slower in that particular incident, right?

Martin Casado

I just don't want to quibble on syntax, but words matter with PR, right?

Ali Ghodsi

Yeah.

Martin Casado

So let's take the Hugging Face incident. When I read that, you know what my reaction was? It was not, “Oh, OpenAI should pace.”

Ali Ghodsi

Mm-hmm.

Martin Casado

It was like, “Dude, fucking secure your thing,” right? Do security controls like we always have done.

Ali Ghodsi

It is pacing, though. It is pacing.

Martin Casado

It's not pacing. Like—

Ali Ghodsi

It is.

Martin Casado

In the history of the internet, we had all of these things. We're like, “Let's pace the growth of the internet.” Let's do security, let's do control, let's do whatever.

Ali Ghodsi

It is pacing.

Martin Casado

I mean, you could call it pacing.

Ali Ghodsi

You should do that, but it is pacing in the sense that, look, I face this all the time. I have a legal department at Databricks, I have a security department at Databricks, and they are always like, “Hey, slow everything down for everything, not AI.” Literally every little thing. “Oh, you're going to go on a podcast. Well, what's the script for it? What are you going to say? Let's review that.”

Martin Casado

Yeah, go ahead.

Ali Ghodsi

And, you know, what's the legal position? You cannot say this, you can say that, you can say that. Everything you say has to be materially true. You cannot leave—

Martin Casado

Are these pacing people in this room with us right now?

Sarah Wang

Who wants to do that? Yeah.

Ali Ghodsi

So you're running an RL experiment, you're training the next model. Should the security team be there, run all their monitors, and look at every— I mean, millions of GPU-hours of tokens were produced, and these agents were running amok in the sandboxes. It would have slowed them down significantly if you had the security team sit there—

Martin Casado

Right.

Ali Ghodsi

—and look at all of the stuff.

Martin Casado

I agree. I've—

Ali Ghodsi

Now, they're saying, “Hey, if we do that, it'll slow us down, and I'm not sure the other side is doing that. So can you guys come in and slow us down? Just tell us. Put some guardrails around us. We'll happily follow the rules and do the secure thing.” Otherwise, it doesn't make sense, because we'll get our butts kicked.

Martin Casado

I just think nuanced, second-order words don't work when people are really afraid. You're like, “I'm going to pace, and therefore things like these don't happen.” I literally think we should have just been like, “Safety and security are paramount. We're going to put in these controls.” That's the important thing, and I do think that nuance actually got lost. If you look at the—

Sarah Wang

Well, kind of what Zuck said. Do you agree with what—

Ali Ghodsi

I thought what Zuck did was very pragmatic.

Sarah Wang

He was like, “Hey, we're going to pace ourselves. We're going to put in security—the reason we released this later is because of security.”

Martin Casado

Mm-hmm.

Ali Ghodsi

Well, listen, what I thought was so great about Zuck is that he was very focused on security, safety, and self-regulation. Dario's first five words or whatever are like, “We need to pace the frontier,” right? It just puts you in a very different mindset than—

Sarah Wang

Yeah.

Ali Ghodsi

What he could have said was, “We need to secure the frontier.” Fine. “We need safety of the frontier.” At some level, I think they were trying to optimize both for the doomers, which calls for a pause, and for politicians, and they kind of didn't satisfy either. The doomers are like—

Martin Casado

But the thing is, don't you think people are actually freaking out inside the labs? There are a lot of safety people that are genuinely freaking out. By the way, not all of them are EA people and so on. And then people are surprised, right?

Ali Ghodsi

Right.

Martin Casado

But here's the thing: using the word “pace” doesn't help either of them. I think it's literally like you're—

Sarah Wang

Yeah.

Martin Casado

You're trying to find this—pace is like the uncanny valley of making the doomer people unhappy and the policy people unhappy. The doomer people are like, “That's not a pause, this is pacing.”

Ali Ghodsi

Yeah.

Martin Casado

And everybody else is like, “Well, this isn't real. You're not going to do it anyways, and you're not focused on security.” So, independent of what we should do, which we should talk about—

Ali Ghodsi

Yeah.

Martin Casado

I just think the way it was presented was bad, and it didn't work. That's why we're having the blowback.

Ali Ghodsi

Yeah, some of these guys are not trained PR people, and yes, some of the stuff—

Martin Casado

Well, they have to—

Ali Ghodsi

I agree, I agree with—

Martin Casado

No, but they do. They have trained PR people.

Ali Ghodsi

I agree with the core premise that we shouldn't freak the public out. I think that—

Martin Casado

Or the government.

Ali Ghodsi

Existential risk right now is close to zero. But let's talk about the core thing, which is the fact that anyone who's doing big reinforcement learning runs and is giving it a reward function—so unleashing, saying, “Hey, here's 10,000 agents and here's $100 million. Let's put them in parallel and let them run on a gigantic cluster for a month or two. Try to solve anything.” It doesn't need to be a security thing. It could be: do anything. Solve this math puzzle. Really bad things can happen. Really bad things, meaning things get hacked. Cyber is the primary one.

Martin Casado

Yeah.

Sarah Wang

Yeah.

Ali Ghodsi

Right? That is real, right? I think making it, “Hey, this is an existential risk,” and so on, was a mistake. I think it's not good to scare the public that way. I think it's become something that everyone—not just your sister, everybody around the planet—is now talking about. I've had all kinds of people who never care about this stuff, and find this extremely boring, ping me and say, “What do you really actually think about this? This is really important in my count, so now I'm starting to worry about it.” Then it becomes a political issue, and we have elections here coming up. But there are elections all around the world, so you're going to see they're not going to sit still in other parts of the world either. I think it's our responsibility to talk about this in a balanced way and actually expose the risks. I think that superintelligence, that idea from that book, is very, very far away. I don't see any evidence that we're actually marching toward that or that it's going to happen.

Martin Casado

Same.

Ali Ghodsi

Apparently some people at the labs are freaked out that maybe there's progress toward that. I think it comes from RSI, recursive self-improvement—the models improving themselves. I would love to understand—have they seen something we don't know? There are kind of 4 criteria. If those 4 things are happening, I would love to understand them. The first is: does the next model require fewer resources, fewer GPUs to train, and superlinearly—not just by a tiny little bit?

Martin Casado

Mm.

Ali Ghodsi

The next model takes less time to train as well—that's the second condition.

Martin Casado

Mm.

Ali Ghodsi

Third, the accuracy of the model, the intelligence, is increasing. And fourth, we can do the first 3 again and again and again.

Martin Casado

All of those at the same time, right?

Ali Ghodsi

All at the same time, not just any of them. Yeah.

Martin Casado

Okay.

Ali Ghodsi

All 4.

Martin Casado

Okay.

Ali Ghodsi

If all 4 are happening, then you can imagine a way where you can—

Martin Casado

You know, because if any of them does not happen—for instance, if resources are constant—then that's okay, because we're going to run out of hardware, so then it will pace itself. We will not have enough hardware to do that.

Ali Ghodsi

Not enough GPUs.

Martin Casado

Yeah.

Ali Ghodsi

Right? Time, same thing. So if you just mean that the software is writing itself, we're already there today.

Martin Casado

Right.

Ali Ghodsi

Like, 90-some percent of the software in Databricks is written by AI. Does it matter if the last few percent is also written by AI? No, it doesn't matter really that much. But if you're getting these 4 conditions, then you might get a speedup where the next model, let's say, takes half the amount of time and half the resources, and it is more intelligent. And you keep doing that, then you might end up in a situation where—I don't know.

By the way, I don't even know if that necessarily leads you to superintelligence per se.

Martin Casado

No, because then that could still converge, yeah.

Ali Ghodsi

But it could, so then that would be more risky. It would be nice if they could share all that data and we could shed some light and transparency on that.

Martin Casado

I actually think it's a great breakdown that you have.

Ali Ghodsi

I don't think anyone is using that as a definition, actually, right?

Martin Casado

It's a great one.

Ali Ghodsi

I think there is a little bit of people freaking out about, “Oh my God, emergent behavior, now it's creating itself,” and so on. But I think, as I said, for a lot of people, their definition is just, “Hey, if I'm not even coding anymore and it's coding itself.” But I think they're conflating, “Hey, what's my value, and is it scary for me?” with, “Hey, that then means we'll get that superintelligence that in 2014 was theoretically hypothesized by Bostrom.”

Martin Casado

Yeah.

Ali Ghodsi

Yeah.

Sarah Wang

Yeah. Well, your point on compute, though, is a really good one that's missed in, I think, a lot of arguments on RSI, right? Because as far as we can tell, the minimum threshold for compute needed to train a good model just keeps going up.

Ali Ghodsi

Yeah.

Sarah Wang

It was 100 million, 1 billion; now it's probably like 5 billion.

Ali Ghodsi

To train a model now?

Sarah Wang

To train, like, a frontier model, right?

Ali Ghodsi

5 to 10 billion, yeah.

Sarah Wang

Right, right, exactly. Versus—

Ali Ghodsi

Million or billion?

Martin Casado

Billion.

Sarah Wang

Billion.

Ali Ghodsi

Billion.

Martin Casado

Billion, yeah.

Sarah Wang

Versus your second point.

Ali Ghodsi

Yes.

Martin Casado

It's very expensive.

Ali Ghodsi

Yeah, frontier models are very expensive.

Martin Casado

To replicate the frontier 6 months later is about 1/20 the cost.

Ali Ghodsi

No, I think Sarah has a great point, which is—this is a good argument against this whole thing. The next model—first of all, each of these labs does only 1 or 2 such runs a year, and they take—the opposite of the 4 criteria that I mentioned, right?

Sarah Wang

Right.

Ali Ghodsi

It's going to take more resources, involve more humans, and be even more brittle.

Sarah Wang

Right.

Ali Ghodsi

They have to build out the data centers. I mean, the labs are not necessarily doing that, but others have to build the data centers, and they have to be gigantic. They have to get the GPUs and networking right. They have to do the engineering to make sure that they can tolerate it, because every order of magnitude more GPUs you cram in there, you now have to worry about errors that you didn't have to worry about before. So you have to increase the robustness of the process. It's a very brittle process.

Sarah Wang

Yeah.

Ali Ghodsi

And if it fails, you've squandered so much money, so they're very, very careful with that run, and there have been multiple runs that have been botched.

Martin Casado

Yeah.

Ali Ghodsi

So it's the opposite of that: the next model is faster, cheaper, smarter, and there is recursive improvement. It's the opposite. It's taking longer, it's more brittle, there are more people, and it's harder to pull off. So I do think that is true.

Sarah Wang

Yeah.

Ali Ghodsi

With respect to RSI, with respect to actual cyber risks and things—

Martin Casado

Oh, it's very high. It's very high.

Ali Ghodsi

Getting hacked, we need to take it super seriously, yeah.

Martin Casado

So I'm going to have another litmus test here. I actually love your 4 criteria. I was literally just waiting to argue with it, but I actually think this is very good. So let me give you a black-box one. When you're dealing with these dynamic, adaptive systems, what are you going to believe? Are you going to believe the numbers or your lying eyes, right? So I think you kind of have to go to the numbers on these ones. What are the numbers to look at? I really think you should just basically—and maybe going public is the right way to do it—if these companies continue to grow, reduce the number of people and the amount of money that goes into them, then I would say something is definitely happening here. I do think you can actually black-box this and take a look. But none of those indicate that they're hiring like crazy, they're—

Ali Ghodsi

But that's not fair.

Martin Casado

They're burning money like crazy.

Ali Ghodsi

That's not fair, because companies are not necessarily efficient, right? So what if you have— I mean, OpenAI itself was doing a million different activities. A very small—

Martin Casado

Well—

Ali Ghodsi

A very small—

Martin Casado

I just have another— I agree. Just another litmus test. We have kind of 2 litmus tests. We have your litmus test, which I think is great, but then you would actually have to have a way to instrument it.

Ali Ghodsi

Yeah.

Martin Casado

And then we should have the black-box litmus test. I mean, listen, if Anthropic in 2 weeks is 12 people and they continue to grow and they're putting out models at an increasing rate, I think we should probably take notice of that.

Ali Ghodsi

That's a sufficient criterion, but it's not a necessary condition, right?

Martin Casado

Totally, yeah. Yeah.

Ali Ghodsi

So I'm just saying that there could be that. Really, the right way to do this, then, is to look at—okay, the pre-training and the post-training that's being done, that's really necessary.

Martin Casado

Yeah.

Ali Ghodsi

Because they have so many resources that they might be doing a lot of other stuff they don't need to do.

Martin Casado

Fair enough.

Ali Ghodsi

But they're doing it because they can just hire—

Martin Casado

Right.

Ali Ghodsi

The people, because they have infinite money—

Martin Casado

Right.

Ali Ghodsi

So really, the people who are training the next model—is that team tiny, tiny, and actually getting reduced? Are they doing less and less work, with AI doing the post-training, and then all just using fewer GPUs? That's not the case.

Martin Casado

We've had this argument many times as an industry before. I remember when we learned how to really cluster computers, because the mainframe was actually kind of limited by things like memory coherence. Remember that?

Ali Ghodsi

Yeah.

Martin Casado

You could only make it so big.

Ali Ghodsi

Yep.

Martin Casado

Then we went to the client-server era, and we didn't have that problem. Then we started creating supercomputers, which were basically just clustered computers.

Ali Ghodsi

At some point, the internet happened.

Martin Casado

And that was also roughly when GPUs started getting good. Do you remember that we would actually export-control PlayStations? Because we were worried that Saddam Hussein would use them to do simulations. The arguments were very similar: these things are getting infinitely powerful. We were using them to simulate nuclear weapons—which we were; I was.

Ali Ghodsi

Mm-hmm.

Martin Casado

This stuff has existential risk. Actually, they didn't use those words, but this has the potential for nuclear weapons or whatever, and we should stop it. None of that came to pass.

Ali Ghodsi

Mm-hmm.

Martin Casado

So I think a very reasonable discussion is: Is this time different?

Ali Ghodsi

Yeah.

Martin Casado

Yes or no? I don't have an answer to that. But I'm a VC.

Ali Ghodsi

Well—

Martin Casado

But, you know, you're a—

Ali Ghodsi

Yeah, I mean, look, I think I was—I'm not old enough to remember, so ignorance is bliss. So I can take this—

Martin Casado

Wait.

Ali Ghodsi

So, no, no. I don't recall—

Martin Casado

Wait. Huh, huh, huh.

Ali Ghodsi

PlayStations being illegal and Saddam Hussein being—

Martin Casado

What?

Ali Ghodsi

I just don't know. Or maybe I'm just saying—

Martin Casado

How old are you? This is like—

Ali Ghodsi

Maybe I'm just ignorant.

Martin Casado

This is like 1999.

Ali Ghodsi

Okay. Maybe I'm old and ignorant.

Martin Casado

What the fuck?

Ali Ghodsi

I mean, you know—

Martin Casado

Maybe they didn't care in Sweden. Is that the—

Ali Ghodsi

Maybe it's just amnesia from age—

Martin Casado

Is that the conclusion? Jesus.

Ali Ghodsi

But—

Martin Casado

Sweden—

Ali Ghodsi

Whatever it is—

Martin Casado

Sweden doesn't care about export controls in the US.

Ali Ghodsi

Yeah, whatever it is, I think it's at a different scale now, right? With AI and what we're doing, the pace of development and so on. They are freaking out at the frontier. I do think cyber is actually one of the biggest ones—

Martin Casado

Yeah.

Ali Ghodsi

That we're going to see, right?

Martin Casado

Sure.

Ali Ghodsi

Because there's just so much infrastructure on the planet.

By the way, way more than it was—whatever Saddam Hussein or Xbox or whatever it was—

Martin Casado

Yeah.

Ali Ghodsi

We've just interconnected way more things, and they're dependent. The planet just looks different today in terms of internet and technology dependency and interconnection than it did 30 years ago.

So I just want to make this point: there's so much infrastructure that's insecure. If you're going to unleash these agents, they're going to find loopholes, they're going to find exploits—

Martin Casado

Oh, they—

Ali Ghodsi

They're going to break in here and there.

Martin Casado

Yeah.

Ali Ghodsi

So this is a real risk.

Martin Casado

So why do you think—

Ali Ghodsi

And by the way—

Martin Casado

Sure.

Ali Ghodsi

This time it didn't do that, but you could imagine a scenario where it starts hopping. It takes resources and starts executing itself elsewhere, so it spreads like a virus a little bit.

Martin Casado

Yeah. So—

Ali Ghodsi

That's a real risk.

Martin Casado

So this is pure curiosity. I promise I'm not trying to be a foil here, but why do you think we just haven't seen very much, then? Again, I'm much older than you. I remember very well when the internet came out. By this point, we had literally taken out 10%, disabled hospitals, taken out critical infrastructure, and caused tens of billions of dollars in economic damage from worms. All of that had already happened.

And to your point, we had much less buildout. Less of the economy was on it. AI has so many people who want to find risks and threats. We're running so fast, so much money has been poured into it, and we haven't seen anything commensurate with the early days of worms. What is that disconnect?

Ali Ghodsi

Mm-hmm. Yep. Look, I do remember those days.

Martin Casado

But it's the same time.

Ali Ghodsi

Yeah. I would just say that I am sleeping well at night, and I don't think there's existential risk right now. I do think there's a lot of infrastructure that needs to be secured.

We have a product in the market, in the detection market, LakeWatch, that helps you do detections. The space is moving so fast because you used to have these SOC teams—security operations center people—who would look at what intrusions were happening, how you were being attacked, and so on. Now the humans just can't keep up.

Martin Casado

Yeah.

Ali Ghodsi

So the whole cybersecurity space is being transitioned to fully automated detection using agents. If we don't do that, I mean, we're rushing. The industry is rushing to do that super fast.

Martin Casado

Yeah.

Ali Ghodsi

If we don't do that—

Martin Casado

Yeah.

Ali Ghodsi

—you will start seeing those kinds of things: sites going down, whole systems stopping for a while. There will be consequences—not existential, but economic damage and people getting hurt could happen.

We just have to race very fast to do all of those things. Humans don't respond fast enough to the attacks that are happening, so you need to automate all of it. Most organizations are not close to doing that. The banks are doing it.

Martin Casado

Right.

Ali Ghodsi

Some of the people who are extremely security-conscious are doing it. But most of the industry today is running with old-school security operations centers and people who wake up every day to hundreds of detection emails that have fired. Many of them are false positives, so you can ignore them. But some of them are not, and they just don't have time to go through them.

You need to have automated threat hunting, where you're actually attacking your own systems automatically with agents and so on. It hasn't happened. If we just say, "Hey, this is like the internet in the early days," bad things are going to happen. So there is a race going on.

Martin Casado

You know, I was actually very surprised. Earlier this morning, I was on a call with a founder. I feel like you and I are pretty close, and we talk periodically. I feel like I know a fair bit about Databricks. This founder was basically saying, "Yeah, we're doing all of this observability, agent threat detection, and we're using Databricks." I didn't even know that you had this offering, quite frankly.

Just from an education standpoint, how extensive have you gotten in agentic AI observability, security, and safety?

Ali Ghodsi

Yeah. We gave a talk this year at RSA, actually with Ben Horowitz. The issue is that data and AI are blending with cybersecurity. These 2 markets are collapsing—

Martin Casado

I don't know. Because I think—

Ali Ghodsi

You know?

Martin Casado

At least—

Ali Ghodsi

The reason they're collapsing is that it used to be like, okay, you have data and AI—the kind of stuff Databricks and companies like that used to do. You have a bunch of data, you run AI and machine learning, and that just lives separately. Then you have the cyber world, where you want to detect if something bad is happening—if bad people are trying to hack you or doing things. You need to detect that, okay?

But now, on the data and AI side, we have agents running. Internally in the company, people have agents running, and the agents are also doing things with other people's agents. They're producing a lot of data.

Martin Casado

Yeah.

Ali Ghodsi

Logs, trails, fingerprints that are being left. Now you have these agents internally doing that. These worlds start merging more and more. All the data that's being produced needs to be analyzed, and the scale at which you need to do that is many, many orders of magnitude greater than it was just 1 or 2 years ago. Things have changed dramatically.

In 2018 and 2019, the time it would take from a CVE vulnerability being published until you saw it actually weaponized in the industry would be 2 or 3 years. That went down significantly by 2022, but it was still 8 or 9 months.

Martin Casado

Yeah, yeah.

Ali Ghodsi

So that's fine. You have 8 or 9 months from a vulnerability to weaponization. That was 2022.

Martin Casado

Yeah, yeah.

Ali Ghodsi

Now, if you look at this curve from 2022 until now, it's down to basically hours.

Martin Casado

Yeah, yeah.

Ali Ghodsi

It's down to basically no time. Things get immediately weaponized. So you need to do it in an automated way, with a data and AI platform approach. These markets—

Martin Casado

Yeah.

Ali Ghodsi

—are just going to collapse, actually.

Martin Casado

Yeah, yeah, yeah. I agree. So it's a very specific question. To pull back on the existential X-risk discussion, it feels like there are almost 2 camps.

Ali Ghodsi

Yes.

Martin Casado

There's 1 camp that believes this is actually an engineering problem, and companies like Databricks can solve it through products, engineering solutions, and services. As an industry, we just need to solve that problem.

Ali Ghodsi

Yeah.

Martin Casado

And there are others who believe there is no engineering solution. You have to slow it down. You have to use regulation. It's more like a nuclear weapon, et cetera.

Does this mean you believe it is an engineering problem, or are you not quite comfortable saying that yet? Because what it was actually like was, "Why would we even buy Databricks, man? Let's just put this stuff in a national lab."

Ali Ghodsi

Just pace it—that's the only solution. I know—

Martin Casado

But pause it—

Ali Ghodsi

Everybody just paces themselves a little bit, then we'll be fine. The bad guys do—

Martin Casado

By the way, there is no line of inquiry that ever gets to pacing, I don't think. I think it's like—

Ali Ghodsi

Yeah.

Martin Casado

—you pause it, or you—

Ali Ghodsi

Yeah.

Martin Casado

—solve it.

Ali Ghodsi

What we've learned here is that Martin really hates the word "pacing."

Martin Casado

I hate the word "pacing."

Ali Ghodsi

I will never use that word with you again. Clearly. Duly noted. But, yeah.

Martin Casado

Frustrated Martin.

Ali Ghodsi

Yeah. Is it an engineering problem that can be solved by engineers, or is there more to it? I actually think, which problem are we talking about? There are 2 separate problems that I think are being conflated.

Martin Casado

Yeah. Mm.

Ali Ghodsi

There is the superintelligence problem.

Martin Casado

Yeah.

Ali Ghodsi

You know? A lot of this comes from Bostrom's 2014 book, Superintelligence: Paths, Dangers, Strategies. If you look at the definitions, I think people don't have clear definitions of what superintelligence...

If you read his book, those definitions are kind of crazy. I think what he had in mind when he said “superintelligence” is AIs that—I don’t know what the examples were. Something like—

Martin Casado

No, everything could do anything.

Ali Ghodsi

Yeah. Something like they write a whole PhD thesis with novel, peer-reviewed material in a couple of seconds.

Martin Casado

Yeah.

Ali Ghodsi

And they can do millennia worth of thought instantaneously.

Martin Casado

Yeah.

Ali Ghodsi

This is the level of how fast they are and how intelligent they are. It’s like—

Martin Casado

They can learn.

Ali Ghodsi

Yeah. It’s just many, many, many orders of magnitude.

Martin Casado

Many.

Ali Ghodsi

Right? The scale of the problem is just completely different.

Martin Casado

Yeah.

Ali Ghodsi

So if such a thing exists, do I think that’s just an engineering problem to solve? No, I think that’s actually very—if such a thing would happen, that would be very existential.

Martin Casado

Of course.

Ali Ghodsi

And that’s what everybody agrees on. So I think that’s being mixed with the fact that now we have agents that are nowhere near that. There’s nothing like that, and we don’t have anything towards that path right now.

Martin Casado

Yeah.

Ali Ghodsi

But these agents are capable.

Martin Casado

Yeah.

Ali Ghodsi

You can do something with them that you could never do before in the history of mankind. I do think an inflection point has happened. Something has changed. We have good security researchers at Databricks, but I could never say, “Let’s get 10,000 of them in a sandbox for a month and have them do $100 million worth of salary-wage work.”

Martin Casado

Totally.

Ali Ghodsi

We can do that now. We just turn on a button and we can get 100,000 of them. Or mathematics: we can say, “Hey, we want to solve a conjecture.” Let’s get pretty good mathematicians, but let’s have 10,000 of them collaborate, and then you can make very fast progress. I think this leads to all these cyber risks. I think cyber is the major problem here.

Martin Casado

Yeah. Mm-hmm.

Ali Ghodsi

This, I think, you can solve with engineering. We are working on it, and many others are working on it. There are still risks, but they’re not existential. I think we should do it.

There is the superintelligence thing. That’s the thing that could write a novel PhD thesis or reason intuitively in 11-dimensional space physics instantaneously without writing anything down—something humans can’t do.

Martin Casado

Yeah.

Ali Ghodsi

That kind of superintelligence.

Martin Casado

Yeah.

Ali Ghodsi

The question is: Is RSI, or recursive self-improvement, that the labs are doing leading us there? Are we going to get there?

Martin Casado

Trying to do.

Ali Ghodsi

Trying to do. Is that what’s going to happen?

Martin Casado

Yeah.

Ali Ghodsi

And how fast is that going to happen? That’s the big question. They’ve suggested that we should have inspectors come in and look at what we’re doing, and that’s a good idea. Have them go in there and get the data. I would love to—the question is, who are the inspectors?

Martin Casado

Right.

Ali Ghodsi

Because you can stack that.

Martin Casado

Who do you think, Martin?

Ali Ghodsi

Yeah. There are a bunch of people on either camp. Actually, I wouldn’t care if they were the inspectors, but I would not be very impressed by what they say because they’ve already made up their minds even before they go in there.

Martin Casado

Right. Exactly.

Ali Ghodsi

But let’s say, as an example, if Yann LeCun, who was one of the inventors of this deep neural network technology—one of the pioneers—said, “Hey, there’s nothing to see here. There’s no risk,” paraphrasing him. “This is nonsense. Keep on going. Go fast, fast, fast.”

Martin Casado

None of us would believe it.

Ali Ghodsi

I’m putting words in his mouth. I mean, those aren’t his exact words. Now, if he was one of the inspectors and he went in there and had a look, then came out and said, “Hey, I’ve looked, and it’s just what I said. There’s nothing to see here. Just keep going,” I would feel very good about that.

Martin Casado

Yeah.

Ali Ghodsi

I would say, “Okay, I would feel very—” Or if he comes out and says, “Oh, my God,” and he’s wobbling and has changed his mind a little bit, that would also have a lot of interesting signals. I think it comes down to who we pick as inspectors, and I think it’s a good idea. Let’s have some of them and pick—

Martin Casado

What—

Ali Ghodsi

Pick a diverse set of people so that we can get different, nuanced points of view.

Martin Casado

What do you think about this kind of Elon Musk view, which is less third-party? It feels like there are 3 proposals. The OpenAI and Anthropic one is a third party.

Ali Ghodsi

Yeah.

Martin Casado

The Elon Musk one, as far as I can tell, is that the labs cross-check each other, like peer review in science.

Ali Ghodsi

Yeah.

Martin Casado

And then the Mark Zuckerberg one is police yourself, right?

Ali Ghodsi

Yeah.

Martin Casado

What do you think about this middle one?

Ali Ghodsi

That they should pace each other—evaluate each other? I mean—

Martin Casado

Like I evaluate you, you evaluate me.

Ali Ghodsi

I think if we have boxing matches in the ring, the boxers should just be the judges of each other. Would that work? No. They would scream “Foul!” all the time. “Foul, foul, foul!” The moment the other guy puts out a great model, it’s, “Big superintelligence risk.” Absolutely. They have not been responsible.

When vested interests are at play, there are IPO plans, and these 2 companies are so competitive—and they have this history between them—they’ll be very fair to each other, I’m sure. That’s why you need a third party, right? Why do we have judges in the world at all? Why do we have third parties at all? Why can’t people just figure things out between themselves?

Sarah Wang

So, I—

Ali Ghodsi

They should try. If they want to do it, they should try, but I’m skeptical that they wouldn’t just be biased in multiple ways to self-judge each other.

Sarah Wang

I think it’s a fair point.

So, I have to ask, Ali, do you think—something Elon also said, I think it was at the All-In Summit—he was like, “This is some elaborate 4D chess,” because on the one hand you’re saying all of humanity will die.

Ali Ghodsi

Mm-hmm.

Sarah Wang

On the other hand, you’re saying, “Hey, what do you want for your IPO allocation?” Right? That is probably a more cynical view, but how do you reconcile that? The dissonance gets a lot of people. How do you think that gets reconciled?

Ali Ghodsi

Look, I think all of these things get mixed. I think there are people who are freaked out, and I do think there are people saying, “Hey, if there was regulation that would pace us”—sorry to use the word—“that would be good for us.” Right? “That would be good for us.” But I also think that people have vested interests.

Sarah Wang

Yeah.

Ali Ghodsi

These things—and usually people figure out a way to get all of these things to align harmonically in their heads.

Sarah Wang

Mm.

Ali Ghodsi

Do I think there has been a tendency in the past, in general, to use marketing stunts by saying, “Oh, my God, this latest model is so good that I trained it. It’s unbelievable. It’s almost scaring me”? And then—

Sarah Wang

Mm-hmm.

Ali Ghodsi

—the whole world starts focusing on it? Yeah, there’s been that kind of marketing going on.

Sarah Wang

Yeah.

Ali Ghodsi

But at the same time, as I said, the time from CVE to actually—

Sarah Wang

Right.

Ali Ghodsi

—weaponize an exploit has been going down from years to minutes now, in just 3 or 4 years. So it’s real. The cyberattacks are real. But there’s also a great marketing ploy: whenever you train a new model, make lots of noise around how much of a crazy risk it is to the world. It helps you, right? Maybe they’re not in contradiction—these things.

Martin Casado

Maybe. And—

So, I mean, you and I are networking folks, and there’s a long history of forming third parties to help arbitrate things, right? Like the IETF or—

Ali Ghodsi

Yes.

Martin Casado

—the IEEE or even ICANN.

Ali Ghodsi

I see where this is going.

Martin Casado

No, no, no. My question to you is: I think this is actually a very sensible proposal that they have.

Ali Ghodsi

Mm-hmm.

Martin Casado

I actually agree with you. You probably want to make sure it’s independent, which it’s not clear right now. Whatever.

Ali Ghodsi

Yeah. And there are going to be a lot of arguments about who you put there, and everybody’s going to disagree.

Martin Casado

Right, right, right, right. But you said, “Why do we have judges?” The state stepping in is actually quite a different thing from industry self-policing.

Ali Ghodsi

Mm-hmm.

Martin Casado

So at what point do you think it makes sense to consider federal involvement? Or do you think now is the time to consider actual federal involvement, as opposed to more industry self-policing? These are just very different approaches.

Ali Ghodsi

They are different, but they kind of bleed into each other. For instance, FINRA is not completely independent. It is self-regulatory, but it’s linked to the government. I think these things will bleed over.

Martin Casado

Do you think that they evolve into—historically, they’ve—

Ali Ghodsi

Yeah, I mean—

Martin Casado

The industry self-polices, and then it evolves into regulation.

Ali Ghodsi

Look, if they’re saying there is existential risk, and they’re saying, “Come police us and regulate us,” I think it’s very hard for regulators to say, “No, we’re not going to do that.” So far, they’ve said that, but I think that’s not going to last very long.

Martin Casado

It’s so funny. Wasn’t it David Sacks who said, “I’ve never had a CEO ask us to regulate them”?

Ali Ghodsi

And the CEOs are saying, “I’ve never had a regulator say no to that.”

Martin Casado

They know. The reality is that the actual metapolitical machinery is already in motion, right?

Ali Ghodsi

Yeah.

Martin Casado

Everyone has a talking point. Obama has came out. It is a major issue. Do you think there’s a reality that it’s too late, that this will be a major issue in the midterms, and we’re actually going to have heavy-handed federal regulation? Is this all going to be paused, with Anthropic going into the DOE? Are we past that point, or do you think we can actually end up with sensible self-policing regulation?

Ali Ghodsi

I think we should strive toward doing the right thing.

Martin Casado

Yeah, of course.

Ali Ghodsi

I think there are still some degrees of freedom in how things evolve while there’s still time. And, yeah, you’re right that, largely, you have these companies where we’re pumping in so many billions of dollars—

Martin Casado

Yeah.

Ali Ghodsi

The way reinforcement learning works is that you give it a reward function that’s verifiable, like, “We’re going to solve this math problem,” or this narrow area of programming, and so on. We pour so much money into that, and you can get quite good results in that narrow kind of domain. That doesn’t mean that you’re getting that superintelligence.

Martin Casado

No, but you can even trick yourself into thinking that fewer inputs are giving you a better outcome just because you’re running so many experiences and you’ve thought about it so much, right?

Ali Ghodsi

Mm-hmm.

Martin Casado

But it’s actually very hard to do a closed experiment this way, given how many resources are going in.

Sarah Wang

Fundraising is going up astronomically, to your point.

Martin Casado

Yes, yes, and they burn the money.

Ali Ghodsi

That’s why these companies are going public, right?

Martin Casado

They burn the money.

Sarah Wang

Yeah.

Ali Ghodsi

They would stay private. I mean, as someone who runs a private company at scale, I think they would prefer to stay private otherwise. Why are they going public? Because they need the capital, and they consider the scaling laws and the capital to be a strategic advantage.

Martin Casado

Yeah.

Ali Ghodsi

But I would say, let’s go back to the 4 things that I listed. If those are true, would you want to know about it, and would that be worrisome? That could get out of hand. Now, there’s no evidence that those 4 are happening, but if that was actually where it was headed—

Martin Casado

Yeah, yeah, yeah. I actually think understanding any sort of self-propelling property in any system is important. We’ve done this in the past with dynamic systems. We’ve done this with compilers. We did this with all the research on nanotechnology. It’s been a common interest of ours.

Ali Ghodsi

Yeah.

Martin Casado

I don’t think it’s new as an area of interest. We should continue to have that interest. I just think the fear is that these particular systems are metaeconomic systems that are so complex that the risk is that you’re claiming you’re seeing it when you’re not seeing it.

Ali Ghodsi

Yeah.

Sarah Wang

Mm-hmm.

Martin Casado

And I think a lot of that’s happening right now.

Ali Ghodsi

Yeah.

Martin Casado

But, of course, if you see it, you want to know.

Ali Ghodsi

But it is fair to say that the labs are now focusing a lot on RSI, and that’s where they’re headed next. Maybe they’re just unjustifiably worried themselves, just like they were worried about GPT-2, right?

Martin Casado

Yeah.

Ali Ghodsi

They were like, “GPT-2 is world-ending,” and then it wasn’t. GPT-3 and GPT-4 came out, and it was like—

Martin Casado

I don’t want to quibble. A lot of the time when they say RSI, they’re actually talking about autocatalytic effects, and autocatalytic effects have been in our industry for a very, very long time.

Ali Ghodsi

Yeah.

Martin Casado

For example, there’s no way you can create a computer chip without a computer chip. You just cannot do it.

Ali Ghodsi

Yep. Anyone with a computer science degree would say you’re right. A compiler writes its—

Martin Casado

My own compiler—

Ali Ghodsi

—own compiler.

Martin Casado

Well, that becomes closer to RSI.

Ali Ghodsi

Yeah.

Martin Casado

The steam engine was autocatalytic, right?

Ali Ghodsi

Yeah.

Martin Casado

My full-time job is people coming out of labs and starting companies, and they all say RSI because everybody says RSI. Maybe 1% of those are actually RSI. They’re more like, “We use AI for data cleaning. We use AI for making GPU kernels.”

Ali Ghodsi

Let’s make the distinction. So you’re saying it’s basically a catalyst in the sense that they’re using AI—

Martin Casado

It’s—

Ali Ghodsi

—to speed things up.

Martin Casado

It’s autocatalytic, yes.

Ali Ghodsi

Yes.

Martin Casado

So I would say, of what we hear, the RSI—

Sarah Wang

Like a model trains another model, right?

Martin Casado

You’re using a model to build a GPU kernel.

Sarah Wang

Yeah.

Martin Casado

You’re using a model to do—

Sarah Wang

Yeah, exactly.

Martin Casado

—data cleaning.

Ali Ghodsi

Yeah, yeah.

Martin Casado

Just like I use a computer to design a computer.

Ali Ghodsi

Yeah, it’s a catalyst.

Sarah Wang

Right.

Martin Casado

The internet was autocatalytic because it allowed people to collaborate remotely. So I would say—this is anecdotal—90% of the calories are autocatalytic, which is 100% what you’d expect.

Ali Ghodsi

And that’s been going on for a while, though. I mean, that’s not even new.

Martin Casado

Yeah, yeah, but you would expect that.

Ali Ghodsi

But I would say that there is also now a focus on: can we get the model to train itself? This is kind of like the autoresearch that Karpathy did, but now they want to do that at scale.

Martin Casado

There are teams that do that.

Ali Ghodsi

There are teams doing it.

Martin Casado

It is not nearly as many calories as you would expect. I feel like I have a good sample of this because they all come and talk to us.

Ali Ghodsi

Great. Can we maybe have you be one of the inspectors? Can we get all that data and have all of us look at that data? Maybe there’s nothing to see here, you know? I personally don’t think there’s a very high probability that those 4 criteria are happening.

Sarah Wang

That’d be incredible.

Greg Brockman went on the pod this week and said—

Martin Casado

Yes.

Sarah Wang

—we’re in the AGI era.

Martin Casado

Yeah. So now I ask the same question after he said that, and now everybody’s saying we have AGI.

Sarah Wang

We feel validated.

Martin Casado

But people have, for a very long time, when I asked this question, said that AI is smarter than most of the people around me most of the time. That’s been the case since almost Q3 or Q4 of last year. Then I asked them, “How many of you have hundreds or thousands of agents that you are managing, that are coordinating with each other in swarms and negotiating and automating your life and everything around you? And if so, raise your hand.” Almost nobody raises their hand.

Ali Ghodsi

Of course, Martin has done that at home. But—

Sarah Wang

No, no, most enterprises are on Microsoft Copilot. That’s the extent of their AI—

Ali Ghodsi

Most—

Sarah Wang

—from what we’ve seen.

Ali Ghodsi

Most enterprises I talk to, when I ask this question, they’re like, “No, we don’t have any of that.” So we’re like, “Okay, what are you doing, then?” They’re using a chatbot, asking questions from a chatbot. That’s basically a very glorified, efficient Google search of the old days. The result is just a faster Google search.

Then coding is happening, so people are using it for coding, though we can discuss the ROI there. But there’s no agentic work that’s automated across the whole enterprise. That has just not happened.

So then why is that? I think the real reason is, if you actually look at it, the models are smart enough, but they just don’t have the context that exists inside any organization. They haven’t been in every meeting. They don’t know what’s in everybody’s heads. They don’t know all the processes.

There are always a couple of employees who know everything in every organization. You tap them on the shoulder, and everybody’s like, “Oh, my God, what would happen if he or she quits?” They don’t have that context. If you just fused that and gave that context to the AI models—to the frontier models today—I think there are so many productivity gains you could get for any organization on the planet.

For that, we actually don’t need smarter models. We don’t need a smarter model that can solve Navier–Stokes or mathematical conjectures, or do better on Humanity’s Last Exam. We need it to just go from 60% to 70%.

Martin Casado

Yeah.

Ali Ghodsi

None of that is needed. I think people are very upset about, “Oh, if we pause the frontier.” But actually, if the frontier doesn’t advance, it doesn’t matter for the vast majority of organizations on the planet. They’re just so far behind in the adoption curve of actually automating things—

Martin Casado

Yeah—

Ali Ghodsi

—and getting value out of this stuff.

Martin Casado

But it would be disastrous to the labs because the price of intelligence is dropping asymptotically. I think it’s going down by one-tenth every 6 months or something like that. So that would—

Ali Ghodsi

Oh—

Martin Casado

—dramatically change their businesses if you weren’t pushing the frontier.

Ali Ghodsi

Yeah. But this is what we should focus on, right? There are 2 sides.

Martin Casado

Oh, agree.

Ali Ghodsi

We discuss the costs a lot here. There’s a cost-benefit analysis that we should do on everything, right? We’ve discussed the costs a lot here: Is there an existential threat? Is there cyber risk? Are there things we should be worried about, and so on? That’s the cost side.

4. The Benefits Are Real

What’s the benefit? Now that this has become a public thing and the whole public cares about AI, they’re asking, “Well, what’s in it for me? What am I getting out of it?”

It seems like nothing. Why are we—

Sarah Wang

So how do they get there? What are some of the use cases you’ve seen to date—

Ali Ghodsi

Mm-hmm.

Sarah Wang

—that have maybe surprised you to the upside?

Ali Ghodsi

Yeah. First of all, there’s so much worry about existential risk and so on. I think a lot of people just don’t know what a cool use case is, where people are actually doing interesting things.

We have a lot of use cases that are fascinating. One that I like is Crisis Text Line. They actually use large language models with us to detect if teenagers want to self-harm or attempt suicide.

Martin Casado

Wow.

Sarah Wang

Interesting.

Ali Ghodsi

You know—

Martin Casado

That’s awesome.

Sarah Wang

Oh, wow, yeah.

Ali Ghodsi

That’s an awesome use case.

Martin Casado

Super cool.

Ali Ghodsi

It actually saves lives. That’s a great organization, and it’s doing amazing work.

Martin Casado

Wow.

Ali Ghodsi

Another one that’s kind of interesting is the Omnipod, which is for diabetes patients. They can put on the Omnipod, and it uses AI to learn your insulin release and your glucose levels and actually release insulin accordingly.

I don’t know if you remember, people used to stick themselves, right? But this now happens automatically. It’s self-learned AI for your body.

Sarah Wang

Wow.

Ali Ghodsi

It’s a cool use case. Zipline is another one. They’re doing awesome work.

Martin Casado

Oh, yeah.

Sarah Wang

Oh, yeah.

Ali Ghodsi

When they started, they had these drones that were completely automated and all AI-driven, everything from battery optimization to the routes and everything. They were delivering food in areas of need—

Martin Casado

Like blood—blood to refugees.

Ali Ghodsi

Blood to refugees.

Martin Casado

Real serious stuff, yeah.

Ali Ghodsi

Yeah, they started in Africa and then expanded elsewhere in the world.

Martin Casado

Yeah.

Ali Ghodsi

That’s also all AI. It’s an AI use case built on Databricks. That’s a cool one, but there are more advanced ones, too.

One that I kind of like, but that’s harder to explain, is a model that we built with Merck. It’s a transformer-based model called TEDDY: Transformer-Enhanced Drug Discovery—

Sarah Wang

Love that.

Ali Ghodsi

—as the name suggests. They actually published the research, so you can check it out.

It’s a model that, instead of predicting the next token in English, predicts how the gene regulatory network, the DRN, is going to respond. It can detect which cells are causal and which ones are just reactive—they’re just reacting.

Therefore, they can start using this in drug discovery and significantly reduce the cost of developing drugs that target specific diseases. That’s a super cool use case.

There are lots of these. Genie, which I mentioned—you have this ontology, and you can ask it any question. Novo Nordisk is using this.

Sarah Wang

Okay, cool.

Ali Ghodsi

They built this GLP-1 drug, but what Novo Nordisk is doing now is using it for all of the trials they’re running.

Sarah Wang

Oh, wow.

Ali Ghodsi

It can compress the time it takes to get insights—from weeks down to minutes—versus if you’re doing an obesity study or something.

There are a lot of amazing use cases for AI. We should not forget these upsides, too. We want all of these, and we do not want to pause these.

Sarah Wang

Yeah, exactly. You’re totally right. So—

Ali Ghodsi

Yeah.

Sarah Wang

—how do they—let’s say you map out the next 12 months—how do enterprises actually get value? You dropped the word “context,” but how do they operationalize that?

5. Enterprise AI Needs Context

Ali Ghodsi

It’s actually harder than most people believe. First and foremost, we have to make sure that we’ve digitized everything that’s happening in an organization. You cannot just wave a magic wand and make that happen.

Every meeting has to be transcribed.

Sarah Wang

Mm.

Ali Ghodsi

You have to be able to get all the context from all the meetings and everything that’s happening. All the digital content has to be fed to the AI. You have to build—we call it an ontology. We build that.

But first and foremost, you have to collect that. That itself is a problem in many organizations because legal teams will say, “Don’t record every call. Don’t record everything.” So you have to do that in a way where it’s—

Sarah Wang

Can you define “ontology” for everyone? I know Palantir says the word a lot, but it’s not like they own the word. What does that mean? For the people listening, how should they think about it?

Ali Ghodsi

Ontology just means that, in an organization, it’s the relationship between all the abstract concepts—all the goals, all the departments, all the people, and all the projects that are going on. What do they exactly mean, and what’s the relationship between the people, the resources, and what that company does?

It’s the difference between a person who is a new employee in the company and just started today and a person who has worked there for 5 years. Let’s say they’re equally skilled, have the same educational background, and are equally smart and hardworking.

One person is on her first day at work today. The other one has been there for 5 years.

Sarah Wang

Perfect.

Ali Ghodsi

What’s the difference between these 2 people? One has an ontology of how that organization works, who the people are, and how you get stuff done.

Don’t look at the org chart.

Sarah Wang

Mm.

Ali Ghodsi

Don’t go ask that person.

Sarah Wang

Yeah.

Ali Ghodsi

They will not get anything done. You go ask this person, and they’ll get it done for you. And that’s not how it works, so you don’t need to file that paperwork here.

And this project—this is what's going on. This is essential. So there's just a lot of ingrained knowledge that's sitting—

Sarah Wang

Yeah.

Ali Ghodsi

…in everybody's heads—who knows how an organization works. That's why people say, in startup land, “Hey, if you lose most of your people, that company can't recover from it.” It doesn't—

Sarah Wang

Mm-hmm.

Ali Ghodsi

You can't just replenish and hire new people.

Sarah Wang

Right.

Ali Ghodsi

The people are so essential. How do we get that context—that's the ontology—and give it to the AI? Part of that is that we just have to have the recording and all of that. But the second part is: how do you actually distill it down into a graph?

Sarah Wang

Mm.

Ali Ghodsi

Actually, a digital graph that you can then feed to the AI. So the way a lot of the agents work today—like Cloud Coder, any of them, Codex, Pi, OpenCode, or you can go through the whole slew of them—they have this loop, an agentic loop. It can reason, but then it goes and checks every resource one at a time. So it'll go to this MCP server for your question and try to see: Is the answer here? Is there another one? It synthesizes it and gives you an answer. But it's kind of slow.

I liken this to: if Google had built Google Search this way 25 years ago, we would've said, “Okay, we're going to get 10 blue links. We search for our key terms here.” But instead of giving you 10 blue links, it would've gone to 1 website, summarized with an LLM what it does, found a few hyperlinks, jumped in parallel to a few of them, read a few websites, done that for 10 minutes, and then given you its best 10 blue links it would find.

Well, that would be very expensive. It would cost a lot of money to do that every time you go on the web. 2, it would've taken a long time, and then you've got to wait 10 minutes. And 3, the quality would be bad, because you're actually only looking at a very small subset of everything that exists out there, right?

Speaker 2

Mm-hmm.

Ali Ghodsi

So how do they do it? They have an index, right? You never leave Google's servers. You search for it, and it hits the index. The reverse index immediately gets you the 10 blue links within less than 100 milliseconds.

We need to do the same thing for the AI. So the ontology is that. We need to compute that index offline all the time. It's almost like the PageRank algorithm that Google had invented back in the day, but it's more complicated. Google was just looking at a web where everybody can go on the web. Here, there are permissions.

Martin Casado

And the links existed, and—

Ali Ghodsi

Yeah. Here, there's permissions involved. The data I'm allowed to access might not be the data that you're allowed to access, so there's privacy.

Speaker 2

Right.

Ali Ghodsi

There's access control. Also, there are many different types of objects here that we're dealing with, not just websites. So the problem is a little bit harder. But it's manageable. You can actually do it.

I'm convinced you can do this and you can get massive productivity gains out of it, because we did it for Databricks.

Sarah Wang

Yeah, exactly.

Ali Ghodsi

We did it for ourselves.

Sarah Wang

Same part, yeah, yeah.

Ali Ghodsi

Yeah, and we're like, the company's just completely changed. It's not the way it was, I would say, 1 year ago, because of this.

Speaker 2

I mean, you've been dogfooding Databricks for Databricks forever. But maybe say more about the impact you've seen as an organization.

Ali Ghodsi

Once we got this ontology and started working on it, we actually had probably the largest—of all of our customers, we have the largest ontology. Our ontology is bigger for us than any of our customers' when they use us to build their ontology, because Databricks uses Databricks more than anyone else uses Databricks.

So it's like millions and millions of nodes in the graph, in the ontology graph that we have. What happens in an organization? You have a tree-structured organization, and information flows up and down the tree structure. If you can't make a decision, you escalate to your boss; maybe they can tiebreak, and it escalates up.

They need to get up to speed on what's happening, and they need to get all the context, and then they make decisions. Once decisions get made, you have to percolate them down in the organization. A lot of this can now be done by AI if you have an ontology.

Why? Because what happens in a meeting? In a meeting, you go through some analysis. Someone has done the analysis. They probably have a PowerPoint deck with some pretty graphs in it. That person who did the analysis is some smart person who used Excel and made some models.

Speaker 2

Right.

Ali Ghodsi

There are some numerical calculations. So a lot of that you can now just do with AI. The AI can do the analysis for you. It has all the context. It can present it in a way that you want. You can ask questions about it. Instead of having follow-up meetings, you can directly ask questions from the AI.

It's very similar—it's along the lines of what Jack Dorsey has said that you can do to the organization.

Speaker 2

Mm-hmm.

Ali Ghodsi

It's just a concrete way of implementing it.

Speaker 2

Yeah.

Ali Ghodsi

It's a game changer for us. Everybody's on their phones now in the meetings on Genie, and they're asking Genie questions. You can see, as soon as someone says something complicated, you see everybody go to their phone.

Sarah Wang

Can you share that finance anecdote you mentioned once in a board meeting?

Ali Ghodsi

Yeah, sure. Internal board meeting. It actually was needed for one of our presentations. I needed to know how many customers we have in the Fortune 500—

Sarah Wang

Uh-huh.

Ali Ghodsi

…that use us. What's our penetration of the Fortune 500? And I asked one of the people in sales ops because I thought she would have it. She texted me back and said, “Oh, sorry, I can't log in to Genie right now. I'm on a flight.”

And I said, “Well, if you're just going to log in to Genie, I can do that myself.” I asked you because I thought you had something authoritative that I don't have access to. So then I was a little bit angry, so I texted the CFO instead, Dave. I texted Dave and said, “Hey, do you know what our Fortune 500 penetration is?”

He just copy-pasted a screenshot of Genie back to me. That's it. So he also has that.

Sarah Wang

I love that so much.

Ali Ghodsi

So then I said, “Does anyone do anything novel here, or is everybody just going to Genie and asking the ontology for answers?”

Speaker 2

It's like, “Let me Genie that for you,” instead of “Let me Google it.”

Ali Ghodsi

Yeah. That's what everybody's doing. Now we just say, “Hey, can someone just Genie this?” “Can you just get it from the ontology?”

Speaker 2

Absolutely.

Ali Ghodsi

So I do think it's a game changer. But it's not just that you press a button and you have an ontology in an organization. And I think Palantir has done a great job of going to organizations, getting a lot of that tacit knowledge written down, and getting it into the organizations.

We automatically then take that and build the graph, and then we feed that graph into the agents—

Speaker 2

Mm.

Ali Ghodsi

…so that we can answer the question in a way that business leaders would like to see it, which is in graphs, in an analytical way, and in a way where you can interrogate that question and continue asking questions and getting answers, so you can make decisions. And then disseminating that information in the organization.

Speaker 2

Yeah, it's pretty amazing. You sort of bookmarked that developers are obviously using AI, with questionable value. I want to follow up with you on that because I feel like you guys were one of the earliest—

Ali Ghodsi

Mm-hmm.

Speaker 2

…and I don't want to use the word “token maxing” because it has such a negative connotation, but I think in terms of applauding people who can use—

Ali Ghodsi

Yeah.

Speaker 2

…AI to become more productive, you guys were—

Ali Ghodsi

Yeah.

Speaker 2

…at the forefront of that, right? And then, of course, there's this cycle of, “Oh, shoot, people are being wasteful. Now we need to value max.” What was your own journey on that?

Ali Ghodsi

Yeah.

Speaker 2

How do you guys think about value maxing, not token maxing? And then I'm going to throw in AI Gateway in this, right?

Ali Ghodsi

Yeah.

Speaker 2

Because I think the cost-management piece is actually getting more important, and you guys are helping people do that. But maybe tie that in, to the extent it's relevant.

6. AI Costs Need Control

Ali Ghodsi

Yeah. So around Q4 last year was when the models got really, really good, and we started noticing that they were actually starting to give us much better productivity. I started using the models myself to start committing code into production for Databricks—the actual—

I said, “I want to take it all the way to production.” So I did that and started pushing the organization: “Hey, everyone needs to do that. I have done it. Why are you not?”

If the CEO can commit code to production on a very sensitive data platform that has all these security requirements, you should be able to do that too—you being any manager or anyone in the organization.

So I started pushing everyone very hard, and we started making leaderboards in Q4. In January and February, when we kicked off the year, we were already in full swing. Everybody was using this stuff, and we were pushing and managing it. The whole token-maxing thing was already happening around the February-March period.

Sarah Wang

Right.

Ali Ghodsi

We were lucky to be maybe a few quarters ahead of folks, so we could see what was happening here, and it was getting out of hand. We already had a gateway called Unity Gateway. This gateway was being used to provide token capacity, so you could get OpenAI, Anthropic, Gemini, or Groq capacity.

Sarah Wang

Mm-hmm.

Ali Ghodsi

Any customer can come to us, and we'll provide them that capacity because we have relationships with those providers and can provide any open-source model. We started putting budget constraints in place and giving people warnings: “You have this much of your budget left. You're getting close to your ceiling.” We did that for each person and group, and then we started doing great analytics so we could predict exactly where the costs were going.

Sarah Wang

Mm.

Ali Ghodsi

Yeah, and then we added smart routers that could pick cheaper models if you were getting close to your budget or if you had simple questions. We started doing that. We also built a harness called Omnigent, which can multiplex between the different harnesses. It turns out the harness itself matters. If you use the same model but different harnesses, there's almost a 2X cost difference.

Sarah Wang

Right.

Ali Ghodsi

Even with the exact same model.

Sarah Wang

Mm.

Ali Ghodsi

With the same version but a different harness, you get a 2X difference in actual cost.

Sarah Wang

Interesting.

Ali Ghodsi

So if you can change the harness, you can get a lot of leverage on cost. We started using all of this. That way, we were able to bend the curve. Our costs for AI have basically been stagnant: the tokens have continued to go up, but the costs have stayed roughly the same.

Sarah Wang

Hmm.

Ali Ghodsi

That's been super, super important for us, and there's huge demand for this. I think every organization is going through this now.

Martin Casado

Yeah. For the first time ever, a company at scale said last week that it's moving from the frontier models to GLM. This is a large engineering organization. Do you see this? Do you think that's a trend, or do you think that's just a one-off anecdote? I've been hearing about it. I remember the first DeepSeek moment and NVIDIA stock, right? And then that turned out not to be real.

Ali Ghodsi

Yeah.

Martin Casado

Then there was the Kimi moment, followed by the next DeepSeek moment. None of it seems to have had an appreciable impact on the market. But now the anecdotes I have are pretty real, and it seems to be happening. I'd love your view.

Ali Ghodsi

I think people want both. They want the latest model that's super intelligent for difficult tasks where they get ROI.

Martin Casado

Yeah.

Ali Ghodsi

But then there are a lot of mundane, dumb things. People literally use their harness to rename files and whatnot.

Martin Casado

Yeah.

Ali Ghodsi

You're paying orders of magnitude more for that.

Sarah Wang

Yeah.

Ali Ghodsi

Please type that in yourself. Don't have the model do that. It's going to spin for 5 minutes, then it's going to rename the file for you and cost you cents.

Martin Casado

Sorry.

Ali Ghodsi

I'm just saying, type it in yourself.

Martin Casado

I'm just wondering: do you actually see market movement on that?

Ali Ghodsi

Yeah. No, people are moving on it, but the pattern is either to use this expert pattern, where you have a small, cheap, open-source model that uses an expert model—the big ones—or vice versa, or a way in which they can ping-pong them to each other. But people are also multiplexing harnesses and changing harnesses so they can control costs.

People have found, for instance, that Pi is very efficient as a harness.

Martin Casado

Yeah.

Ali Ghodsi

I think there's going to be a multitude of these. The models themselves are stochastic, as you said. Every time they give a different answer, and they're changing so much.

Martin Casado

Yeah.

Ali Ghodsi

There's a lot of experimentation happening, so I think we're going to get to a world where you're not always using the smartest model for everything, which has been the paradigm for the last couple of years.

Martin Casado

Yeah.

Ali Ghodsi

A new model comes out, it's super smart, and they use it for everything—even really simple, mundane tasks.

Martin Casado

Yeah. I'll tell you what I see. I see people using Fable and Astra for architecture.

Ali Ghodsi

Uh-huh.

Martin Casado

They use a cheap model for implementation, and then Fable or Astra for audit.

Ali Ghodsi

Yeah.

Martin Casado

That seems to be this emerging pattern.

Ali Ghodsi

Yeah. What are you guys seeing in the startups?

Martin Casado

Well, that's it. That's honestly the pattern.

Ali Ghodsi

Mm.

Martin Casado

How much open source?

Ali Ghodsi

Either.

Martin Casado

By token or by dollar?

Ali Ghodsi

Either.

Martin Casado

By dollar, open source is 5%. It's very little, but by token count it's over 60%.

Ali Ghodsi

Hmm.

Sarah Wang

Yeah. We talked to, say, Decagon or something like that. I think it's different for internal use versus external product use.

Ali Ghodsi

Mm-hmm.

Sarah Wang

For external product use, I think they're almost up to 90% open source.

Ali Ghodsi

Wow. Hmm.

Sarah Wang

Internally—and I don't want to say for Decagon in particular—a lot of them are like, “We don't care. We'll just use frontier. We're not thinking about cost control.” But as it gets bigger—

Ali Ghodsi

Yeah.

Sarah Wang

—you and I were in another board meeting where they actually did bring that down, just from a waste perspective.

Ali Ghodsi

Yeah.

Sarah Wang

So I definitely see that moving more toward open source on the product side. And actually, that's a—

Martin Casado

Yeah.

Sarah Wang

—related question, maybe around open source, but post-training specifically.

Ali Ghodsi

Yeah.

Sarah Wang

I feel like you were kind of early. I remember talking to you in 2023.

Ali Ghodsi

Yeah.

Sarah Wang

When did you buy Mosaic?

Ali Ghodsi

2023.

Sarah Wang

2023.

Ali Ghodsi

Summer.

Sarah Wang

Okay. So this vision that you had in 2023 kind of came true in 2026.

Ali Ghodsi

Mm-hmm.

Sarah Wang

I don't know if you guys would agree, right? That's sort of what we're hearing.

Ali Ghodsi

What vision?

Sarah Wang

Just: “Hey, we're going to actually—you're going to own your own intelligence.”

Ali Ghodsi

Oh, right. Yeah, yeah.

Sarah Wang

“You're going to—”

Ali Ghodsi

Mm-hmm.

Sarah Wang

“—post-train your own open-source models,” et cetera. That's definitely what the startups are doing.

Ali Ghodsi

Yeah.

Sarah Wang

I don't know if that's what the enterprises are doing yet—

Ali Ghodsi

No.

Sarah Wang

—I mean, do you feel like you were early to that?

Ali Ghodsi

Yeah. First of all, when we started, it was also, “Hey, we're also pre-training for you,” which doesn't make any sense.

Sarah Wang

Hmm.

Ali Ghodsi

You can—

Sarah Wang

Right, right.

Ali Ghodsi

There are very good pretrained models now that you can use.

Sarah Wang

Right, right.

Ali Ghodsi

But you can actually do post-training on the model, and you can do reinforcement learning.

Sarah Wang

Yeah.

Ali Ghodsi

We're actually doing it at scale. Many of those startups are actually customers, so we help them use RL, or reinforcement learning, environments where we can make the models very, very good at the specific tasks they're doing.

It makes a lot of sense for them to do that. If you have a startup and it's offering a product, and that product does something specific, it's not just general intelligence. It does something specific for you. It makes a lot of sense to take a really good open-source model and use reinforcement learning to make it really good at that specific task.

You can cut the cost down, make it really fast, and control your own IP. In that sense, it's possible. But large enterprises just need basic automation.

Sarah Wang

Mm.

Ali Ghodsi

It's just too much for them to do this right now.

Sarah Wang

Yeah.

Ali Ghodsi

I think one of the challenges—

Sarah Wang

See.

Ali Ghodsi

…is that you need good evals.

Sarah Wang

Totally.

Ali Ghodsi

And making good evals is hard. So while the startups can do that and they're motivated to do that, for other organizations, the easy button might be just to use a frontier model—

Sarah Wang

Mm-hmm.

Ali Ghodsi

…rather than having to create—

Sarah Wang

Yeah.

Ali Ghodsi

…your own evals. We even automatically generated evals for the customer in the product, and we had them front and center, but then people didn't want to use them. So we said, "Okay, let's move them to the back end so that they're optional."

Sarah Wang

Right.

Ali Ghodsi

And then they would never go to them. So I would say, in general, the—

Sarah Wang

Why? They just don't want to get into it? Is it too complicated?

Ali Ghodsi

I think you want—

Sarah Wang

Hmm. Yeah.

Ali Ghodsi

…quick—

Martin Casado

Yeah.

Sarah Wang

Yeah, that makes sense.

Ali Ghodsi

…reinforcement of, like, "Hey—

Sarah Wang

Yeah.

Ali Ghodsi

…there's a new model. I want to try this out. I want to get this problem solved." You don't have time to go do the scientific method of, "Let's make an eval. Let's have a great baseline." It's sort of like TDD, test-driven development.

Sarah Wang

Hmm.

Ali Ghodsi

In software engineering, did people actually do test-driven development? Very few did, right? Everyone said it was the right way to do it, but nobody actually did it in practice. So that's the same. That's kind of a little bit of the curse of training your own model: The evals are the hard part.

Sarah Wang

I know you have an FDE model at Databricks.

Ali Ghodsi

Yeah.

Sarah Wang

That's a very popular word—or acronym—right now. But to get these into enterprises that large, is a full FDE model required?

Ali Ghodsi

Mm-hmm.

Sarah Wang

Or how do you—

Ali Ghodsi

Yeah, I mean—

Sarah Wang

…how has that evolved, maybe?

Ali Ghodsi

Yeah, we've had these FDEs, and the demand for them has gone up significantly.

Sarah Wang

Mm.

Ali Ghodsi

A lot of it is, how do we build that ontology? The ontology is automatic, but if you're not collecting any information—

Sarah Wang

Yeah.

Ali Ghodsi

…you're not recording anything.

Sarah Wang

Right.

Ali Ghodsi

So that's one of the key things that we do. But also things like, "I want to build an agent. I want to put it on—I want it to be customer-facing, have really low latency, and have guardrails so people can't abuse it or ask it things that we don't want it to answer," and so on. So we can build that, like the Sports AI that Fox has. You can go chat with it about sports events, and you can try to ask it about politics, and it's very good at rejecting you and moving on to talking about sports instead. So the FDEs built that—

Sarah Wang

Wow.

Ali Ghodsi

…so we'll help organizations actually get started with AI.

Sarah Wang

Mm.

Ali Ghodsi

It is important.

Sarah Wang

Yeah.

Ali Ghodsi

Because many organizations simply don't have the in-house expertise to build this stuff. So they just need a little bit of help on the side, and then they get started.

Sarah Wang

Yeah. Makes sense. So this is more related to the agent side, but I saw recently that I think a third party—

Ali Ghodsi

Mm.

Sarah Wang

…a neutral third party, I think, did some tests that Lakebase or Neon was actually the Postgres database of choice for agents. I thought that was interesting, one, because it's exciting for Databricks, but two, I probably wouldn't have guessed that maybe a year ago.

Martin Casado

Yeah, it's a surprise for sure.

Sarah Wang

Yeah, it was a surprise, just because there are others out there that have great developer momentum as well. But it was pretty clearly number 1. So I'm curious: How did you guys crack this, and what makes you win across the agents? Because if you win the agents now, you win the market.

Ali Ghodsi

Yeah. I think a lot of credit should go to Neon and Nikita and the team. I think what they've done is they've just been obsessive about how to make the models—how to make the models pick and have agents favor Lakebase or Neon as a database. What do they do? The agents want to experiment. They're going off, they're trying to build a little bit of software. They need the database, so you need the database to come up quickly.

Sarah Wang

Mm-hmm.

Ali Ghodsi

They had this obsession that everything should take far less than a second. The database comes up in far less than a second. You can clone a gigantic database, like a terabyte database, in less than a second. So it's highly elastic, highly responsive, and then they built this killer feature called branching.

Sarah Wang

Mm.

Ali Ghodsi

Branching just lets you branch the database, and you can have many, many branches of the same database. They just made it very, very lightweight. We saw this with other things with agents, right? Like uv and ripgrep—basically reimplementations of a lot of the tools on Unix, making them really, really blazing fast and lightweight.

Sarah Wang

Mm.

Ali Ghodsi

And also sort of fail-safe for agents. They've just done this to a harder problem, which is databases.

Sarah Wang

Right.

Ali Ghodsi

So now you have a Postgres database, and the Postgres database has all these advantages: It's really fast, it's nimble, it's fail-safe. You can go back to snapshots. You can do those things. So I think that's why. It's just easier—

Sarah Wang

Yeah.

Ali Ghodsi

…for agents to use this. They also made sure it had a pricing model that was—

Sarah Wang

Yeah. Totally.

Ali Ghodsi

You don't want the cost to run up just because agents are building some software and experimenting.

Sarah Wang

Right.

Ali Ghodsi

You're okay paying for your database if it's for production use and lots of people are using it. But just to experiment—

Sarah Wang

Mm.

Ali Ghodsi

I think they were just obsessed. They weren't trying to win the database war or trying to be better than some other vendor. They were obsessed with, "How are we the best for the agents?"

Sarah Wang

Mm.

Ali Ghodsi

That's a new persona, because in databases, the obsession has been, how do we help DBAs? How do we help app developers? How do we help the people who are using the database?

Sarah Wang

Yeah, exactly.

Ali Ghodsi

They changed the game and said, "Hey, how do we focus on agents—

Sarah Wang

Uh-huh.

Ali Ghodsi

…and help agents get the best database they want?" And now over 90% of the databases that are created on Neon and Lakebase are actually created by agents.

Sarah Wang

Wow.

Ali Ghodsi

So it's not even humans. The numbers speak for themselves. They're—

Martin Casado

By the way, it's remarkable. I've started to use Neon as my standard database.

Ali Ghodsi

Mm.

Martin Casado

And it was bizarre to me, because normally when you enter a large company, things slow down. It's actually like the product has gotten materially better.

Ali Ghodsi

Yeah.

Martin Casado

Are they totally independent, or do they work with the rest of the company? How does that work?

Ali Ghodsi

No, it's a great team. I mean—

Martin Casado

Unbelievable.

Ali Ghodsi

…they work very closely together. We love databases and data, so—

Martin Casado

Amazing.

Ali Ghodsi

We live that. But no, the team does a great job of just making things super-fast, snappy—

Martin Casado

Super—

Ali Ghodsi

…and great for agents.

Martin Casado

Yeah. All right, Ali, what is your p(doom)?

Ali Ghodsi

Less than 10%. No, close to zero. What about yours?

Martin Casado

I don't know. My only answer is that my p(doom) without AI is much higher than my p(doom) with AI. How's that?

Ali Ghodsi

Wow, that's a lot. What about yourself?

Martin Casado

I went on a technical—

Sarah Wang

I would agree with Ali on this one.

Ali Ghodsi

Okay, cool.

Sarah Wang

Yeah.

Ali Ghodsi

Okay.