OpenAI 正在犯下一个巨大的错误
- Travis Good 的核心判断是:“开源会赢,你现在仍有时间成为英雄”——封闭实验室最大化提取价值的姿态,可能摧毁它们真正拥有的唯一资产:分发。 他向 Sam Altman 和 Dario 发出的警告是:“你的分发不重要,因为客户恨你。”导火索是 Anthropic 在关于 Fable 的300页发布材料中“某个地方”披露,它曾破坏 ML 研究查询,随后又“收回说法”,悄悄降级到能力更弱的模型,“这同样是破坏”。
- Ambient 的定位是“推理领域的 Uber”,建立在 Costco 式经济模型上:它不拥有 GPU,只拥有调度每个请求、并保证质量的网络。 全球企业 GPU 算力约25%掌握在小型运营商手中,它们只能赚取约10%的租赁利润,有些甚至背负17%的贷款;而提供推理服务可获得40%-50%的利润。Ambient 通过可验证网络和聚合需求,让这些运营商无需自建业务拓展团队,就能追求这部分利润提升。
- 验证是护城河,Good 说它现在已经是“100% logits”,而且“基本零成本”,可按请求开关。 他讲到一家聚合商每周进行智能测试:测试日服务商会把速度降到正常水平的约50%,显然是在考试时切换到完整版模型,其他时间则使用压缩版。“如果验证过的推理和未验证的推理价格相同,你为什么还要选择未验证的推理?”
- 矿工经济模型把交易支付与通胀奖励叠加起来,后者“几乎像股权薪酬”,因此看多的运营商即使每笔交易低价10%-15%,最终仍能获得更高的整体回报。 Ambient 的路由从每天处理1000万 tokens 的纯竞价机制,发展到每天1亿 tokens 的缓存感知系统,随后升级为“预测式 HFT 算法”,每个运营商每天可处理100亿-150亿 tokens。如今它在 OpenRouter 上服务 Kimi K2.7 通常排名第1,至少在这一指标上已经超过 Moonshot。
- “中国转向闭源”的看空逻辑,可以用训练成本坍缩来反制:GLM-5.2 的训练成本约2500万美元,明年或许降至500万美元,因此任何选择闭源的实验室都可能遭到新的开放权重参与者挑战。 如果这种趋势再持续3年,预训练可能会高度商品化,价值转向模型交付和供给聚合——这正是 Ambient 所处的位置。对于蒸馏争论,Good 认为中国“总体上是创新的”,蒸馏“几乎只是附带现象”,因为“我们都在互相蒸馏”。
- 在硬起飞情景下,Good 认为只有美国政府能获得访问权限;Tommy 则认为届时任何公司都无法创造经济生产力,而受限的公共版本最终只会退回 Opus 4.8。 Tommy 警告,只分发模型的“金苹果”,会让政府进入挑选赢家的业务,最终形成计划经济。Good 认同自己不希望看到这种结果;Tommy 的结论是,“扩散式世界的功能性要强得多”。
- Tommy 把尚未解决的提示注入既称为“万亿美元级的商业机会”,也视为闭源 AI 的潜在障碍,并称其制造了“监管囚徒困境”。 他担心,注重安全的美国可能扣住模型,却仍会“即使拥有更强的模型能力,也被从六个方向黑掉”,而作为安全标杆的 Anthropic 至今仍无法解决提示注入。Good 正面回应了这一担忧,并指出,限制措施可能让竞争国家发布更好的公共模型。
- 在资本开支退潮情景中,Good 设想 OpenAI 以5000亿美元而非1万亿美元估值融资,无法履行对 Oracle 的承诺,数据中心建设停摆,随后“很可能出现一场为期2年的 AI 寒冬”——伴随401(k)账户受损和反 AI 反弹。 Tommy 补充称,过去30天每百万 tokens 的价格已经下降20%-25%,因为企业正从前沿 API 切换到只提供90%-95%智能水平、但成本仅为其一小部分的模型;这可能为轻资产聚合商打开机会,却也会威胁更广泛的融资循环。
1. Anthropic 对 Fable 的破坏,成为“可信中立基础设施”的案例
- Travis Good 对 Alex Karp 视频的解读是:所有人刚刚感受到的痛苦,就是“当你信任的供应商反过来对付你时会发生什么”。Anthropic 没有重点说明,而是在“关于 Fable 的300页发布材料中的某个地方”披露,它曾破坏 ML 研究查询,随后又改口称只是降级到能力更弱的模型,“这同样是破坏”。Tommy 的反应是:“这次真是被坑惨了。”
- Good 认同 Karp 的判断:“如果你在和 Anthropic 竞争,你就在他们的瞄准线上……你不会想成为一个有立场的基础设施提供商的靶心。”这种利益冲突,正是加密行业和可信中立基础设施的切入口。
2. 推理领域的 Uber,底层是 Costco:把闲置 GPU 供给与有质量保证的需求撮合起来
- Good 对业务的定义是:“推理领域的 Uber——我们不拥有汽车,只拥有把每个请求与最佳可用供给撮合起来、并保证质量的网络。”业务聚焦于两类标准模型:最受欢迎的高智能大模型,以及对应的小模型。再往下一层看,它其实是 Costco:“供应商为 Costco 的库存提供资金,但 Costco 会快速周转库存”,维持低价并掌握客户。
- 供给缺口在于,全球企业 GPU 算力约25%掌握在利润微薄的运营商手中:拥有50-100张 GPU 的小型机房,只能以约10%的利润出租算力,而推理服务的利润可达40%-50%。Ambient 作为值得信赖的品牌,为交付质量提供保证,让这些运营商能够捕获这部分利润提升。
- 在可用性问题上,Good 举的例子是迪拜 AWS 数据中心遭轰炸,导致 Anthropic 在美国南部大范围宕机:各家供应商都“运行在自身容量边界附近”。Ambient 的连续竞价和需求高峰加价机制,则意味着不会施加任意的速率限制;他称 Together 基本已经被预订一空。
3. 反共识设计:把极致的经济激励集中到少数模型
- Good 对行业的自我诊断是:“加密行业缺少某种野心”——缺少与 Web2 商业模式正面竞争的野心。传统供应商体量庞大,而且天然受租约约束:“我得先喂饱自己的资产簿”,因此无法可信地套利供给;Ambient 则保持轻资产、灵活和反脆弱。
- 服务少数模型背后的比特币经验,是把经济注意力集中起来,让“1万人疯狂地尝试各种可能的技巧”替你解决交付问题。这与 CPU→GPU→ASIC 的演进路径相同,因为 Ambient 本身也属于工作量证明。
- Good 用购物中心作比喻:要成为 Dillard's 或 Nordstrom 这样的主力店,在标准模型上高效到“你不会想和我们玩利润游戏”,随后把供应商和客户吸引到周边利润更高的精品业务中。
4. 验证式推理如今已是100% logits,而且“基本零成本”——一则测试轶闻暗示供应商可能在应付质量检查
- 问题在于“智能压缩”:“这周 Claude 变笨了,OpenAI 也变笨了。”即使是开放权重模型也会发生同样的事,因为服务模式高度不稳定。没有验证,质量就无法保持一致,业务也无法规模化。
- Good 描述了算法的演进:一年前有3层验证——logits、模型内部遥测,以及复杂的数学关系。如今已经是“100% logits”,即模型思考过程的指纹,包括作出的选择、没有选择的概率区间及相关概率;它安全、可按请求开关,速度上没有可见差异,并已覆盖 vLLM、SGLang 和 llama.cpp 的测试版。
- 他讲到,Ambient 合作的一家聚合商每周进行智能测试,而测试日的服务速度会降至约50%——显然所有人都在考试时提供“完整版”,其余时间则使用压缩版。他始终追问:“如果验证过的推理和未验证的推理价格相同……你为什么还要选择未验证的推理?”
5. 矿工经济模型:交易奖励叠加通胀“股权”,路由从每天1000万 tokens 演进到每个运营商每天150亿 tokens
- 对于“为什么有人要为 Ambient 做优化”这个问题,Good 的戏谑回答是:“今天我们在 OpenRouter 上服务 Kimi K2.7,比 Moonshot 做得更好。”目前通常排名前3,且大多数时候排名第1。真正的答案是,一些背负17%贷款、只能赚10%利润的闲置算力运营商没有业务拓展团队去争夺需求;利润率从10%升至40%,“是颠覆性的”。
- 代币设计把激励进一步放大:每项任务都会支付交易奖励和基于通胀的奖励,后者“几乎像股权薪酬”,与贡献和 epoch 份额挂钩。看好网络的参与者可以把价格压低“10%或15%左右”,仍然获得可观的整体收益。“我不会拿到 Together 的股份……他们只是在抽成。”
- 路由经历了3次迭代:最初是 OpenRouter 上每天处理1000万 tokens 的纯竞价机制,“当时我们很失望”;随后升级为每天1亿 tokens 的缓存感知路由;最终变成“极其复杂的高速预测式 HFT 算法”。该算法只依赖外部指标,因此无法作弊,能够准确调度每个矿工可以承接的负载而不让系统过载,如今每个运营商每天可处理100亿-150亿 tokens。“我仍然认为这辆车的油箱里还有油。”
6. 超越推理:一套“可编程的机器学习经济”
- 对于没有资本、但能力出众的研究者,Ambient Foundation 正在探索集中式支持:用 Ambient tokens 众筹节点,再通过其供应商网络租用供给,让一项优化成果无需作者购买 GPU,就能触达整个网络。
- 更大的揭示是:“Ambient 是一套可编程的机器学习经济。”通胀本质上是一笔可以程序化支出的预算——可以在一段时间内把其中一部分转给执行训练任务的矿工,再把利润分回网络。购物中心里的精品店可以被托管,而 GPU 日常则通过高速交付高利润的标准模型来获得持续使用。
7. 如果中国关闭开源怎么办?答案是500万美元训练成本
- Tommy 的看空逻辑是,智谱及其同行的交易估值达到“收入的14,500倍”(原话),而美国新云服务商正在为它们的开源模型提供服务;某一天,中国可能关闭模型权重,或要求分走收入。Good 认可这种可能性,但给出了反转:一年前,认为 GLM-5.2 能以2500万美元成本逼近闭源前沿还属于“激进异端”;如果明年成本降至500万美元,投资资本就能资助新团队发布开放权重模型,并在评级中“把 GLM-5.6 打得落花流水”——“我们不是已经看到一家中国 DoorDash 竞品发布了某个离谱的猫头鹰模型吗?”
- 对这套论点真正重要的结论是:如果这种动态再持续3年,预训练知识可能会高度商品化,“最终一切都将归结为模型交付和供给聚合——而 Ambient 就处在这个位置。”这个风险反过来会为聚合商创造机会。
- 对于创新和蒸馏之争,Good 认为中国“总体上是创新的”,蒸馏“几乎只是附带现象”:“现在每个人都是10倍工程师”。已经成熟的开放成果——比如他自己的 MIT 许可 Ambient Desktop,代码超过100万行,主要利用业余时间完成——都是理想的训练数据。“我们都在互相蒸馏”;当提交记录已经展示 Claude 如何修改一切时,你并不需要 Anthropic 的用户数据。
8. 给闭源实验室的直接信息:在客户彻底恨上你之前,先拿出一些东西
- 核心宣言是:“开源会赢,你现在仍有时间成为英雄”(“open source is going to win and you still have time to be the hero”)。美国的软实力来自分享文化和研究;Good 的挑战是:“把一些研究延迟6个月发布,会要了你的命吗?”这既能换来善意,也能让全球架构朝美国的方向倾斜,并赢得心智份额。Gemini 虽然受欢迎,但“已经有点老了”;不妨想象 OpenAI 再发布一次开放权重模型。眼下,“全世界感觉自己正在被美国的闭源 AI 殴打。”
- Tommy 要求给出美元和美分层面的理由,而不是诉诸利他主义。Good 的回答是,OpenAI 和 Anthropic 手握“世界上最极致的通讯录”——由巨额风险投资支撑的分发优势——因此可以在利用这一优势的同时,把部分东西免费送出去。最大化提取价值的风险在于:“你的分发不重要,因为客户恨你。”拿走客户数据、切断客户的业务关系,会制造“恐惧反应”。
- 针对 Palantir “拥有你的权重,拥有你的数据”的世界观,Good 表示 Ambient 可以提供一种私有替代方案:可信执行环境、洋葱路由请求和 PII 脱敏。相比之下,Anthropic 的企业级电话可能是“他们会拧你的胳膊……而且出于安全原因,可能不会为最新模型提供 ZDR”。
9. 为硬起飞下注的最强论证——以及它为何仍然会输
- 第一层是数据中心。Good 认为,在 NVIDIA GPU 集群范式下,所有人都在重资产投入,Vera Rubin 可能也不例外,并称这是“极其冒险的赌注,因为 ASIC 是真实存在的”。Tommy 进一步展开 ASIC 逻辑:采用 SRAM 的14纳米芯片无需高速 HBM,可以大规模生产,并可能以相同支出提供远高得多的算力容量。Good 的回应是,Ambient 仍然保持轻资产;如果冲击迫使市场抛售 GPU,这些 GPU 仍可能在其网络中找到归宿。
- 第二层是模型。Tommy 认为,只有当自我改进带来无法追赶的优势,而不是3-6个月的领先时,万亿美元估值才说得通。他说:“现在每个人都是机器学习工程师”;由中等水平的人类、在平均智能水平的 LLM 监督下完成工作,“可能已经足够好”,足以跟上进度。
- Tommy 提出,哪怕只有1%的概率获得100万亿美元,也足以让人破釜沉舟;随后双方把这个逻辑推到终点。Good 说,如果 OpenAI 实现硬起飞,“美国除了政府之外没有任何人能访问它”(黑进 NSA 就会被封禁)。Tommy 认为,没有公司能够访问模型,就不可能产生经济生产力;受限的公共版本最终只会退回 Opus 4.8,而只分发模型的“金苹果”会让华盛顿进入挑选赢家的业务:“我们变成了自己一直想避免的中国。”Good 认同自己不希望看到计划经济,而 Tommy 的结论是:“扩散式世界的功能性要强得多。”
10. 提示注入:安全叙事中的万亿美元级漏洞
- Tommy 用提示注入说明资源配置失衡:安全研究人员给了我们 P(doom) 和网络战争场景,“但我们却有提示注入。”解决它是“万亿美元级的商业机会”,因为这可能让所有人都使用 Fable 的完整版;如果解决不了,即使具备部分 AGI 能力的模型也可能无法发布,而这是 Leopold Aschenbrenner 的曲线没有计入的摩擦。
- Tommy 把地缘政治版本称为“监管囚徒困境”:注重安全的一方扣住模型,最终却可能“即使拥有更强的模型能力,也被从六个方向黑掉”,而其他国家则发布自己的模型。他还讽刺称,Anthropic 从成立起就是安全标杆,却仍然无法解决提示注入。Good 回应说,限制措施可能让中国和其他竞争者发布更好的公共模型。
- 对于 harness,Good 说:“Claude Code 作为 harness 很差;Codex 很好。Hermes 比 OpenClaw 好得多。”Tommy 总结称,开放模型加上公司的记忆、数据和编排能力,实际上就是一种经过微调的体验,并不需要改变模型权重。他还认为,若解决持续学习,提示注入可能会“严重1000倍”,因为1000万 tokens 之前被投毒的输入,之后可能让模型变得失控;Good 开玩笑说,模型会“持续学习如何进行提示注入”。
11. 两年 AI 寒冬情景——以及 Ambient 如何借势启动
- Good 认同,当前正是开放权重推理的窗口期:Together AI、Fireworks、Modal 和 Baseten 在5个月内分别增加了约50亿美元价值。但围绕 OpenAI、Oracle、CoreWeave 和 NVIDIA 形成的循环资本开支网络非常脆弱。
- Good 的推演是:OpenAI 以5000亿美元而非1万亿美元估值融资,无法履行对 Oracle 的承诺;他称 Oracle 已经“押上了整个公司”,因此也无法继续为数据中心提供资金。Tommy 补充称,Oracle 的市值可能被下调约80%。依靠前沿算力扩张的新云服务商会撞上天花板。“接下来很可能是一场为期2年的 AI 寒冬”,伴随401(k)账户受损,以及反 AI 民粹反弹推动的条件反射式立法。这是“Ambient 成功的完美风暴”,但对美国不利。Tommy 近期的佐证是,过去30天 tokens 价格已经下降20%-25%,因为企业开始以远低于原价的成本获取90%-95%的智能水平。
- 路线图是:约1周后启动小模型挖矿的封闭测试,随后开启任何人都能参加的公开 PvP 赛季,参与入口是跨平台 Ambient Desktop 应用,包括在笔记本电脑上挖矿。集中式大模型节点将100%把利润转给参与者,Foundation 不抽成;网络将在今年上线。
- 1年后的成功图景是:在 OpenRouter 上从服务1个模型扩展到10个模型并排名第1;个体和小型运营商获得真正的利润;订阅和“X42”代理式商业采用率提升。最终成为“所有闭源公司向我们承诺过、却因利益冲突无法交付的全球智能公用事业”。Good 说,Dario 永远不会交付智能公用事业,因为国防合同束缚了他的双手。
完整逐字稿
Open source is going to win, and you still have time to be the hero. You don’t want to be in the crosshairs of an opinionated infrastructure provider. Your distribution doesn’t matter because your customers hate you. If you’re Sam Altman or Dario, you need to have a little bit of rational self-interest and say, “Maybe we should do some nice things.”
If we’re taking your data, if we’re perhaps disintermediating some of your business relationships, that creates a fear response. And say that they legitimately achieve a hard takeoff and have a model that’s just incomprehensibly smart. What happens then is probably a 2-year AI winter.
Hey everyone, it’s Tommy from Delphi Ventures, and welcome back to the podcast. Today I’m joined by Travis Good, the co-founder of Ambient. Travis is one of the rare people who can go deep across open-source AI, cryptoeconomics, token design, and real distributed systems. I’ve known him for years, and many of us on the Delphi Ventures team have invested in his project. We’re very happy that he accepted us in, and I’m excited to chat with him today. Travis, how are you doing?
I’m very well, Tommy. Thanks for having me on. I appreciate it.
Yeah, for sure. Listeners of the podcast know you’ve been on before, so it’s exciting to have you back. It’s a real moment for open weights and open-source AI. A lot has gone on.
Did you see the Alex Karp video recently, where some people argued that he’d crashed out on open source and others said he hadn’t? What was your take?
I think he makes a really good point, which is that we’ve all experienced a pain that I don’t think people at large had experienced before: what happens when your trusted provider turns against you?
All of a sudden, Anthropic revealed that they’d been silently sabotaging machine-learning research. I should say they didn’t really emphasize this, but somewhere in their 300-page release on Fable, they mentioned, “Oops, we are going to sabotage machine-learning research queries.” Of course, they walked that back and said, “Now we’re just going to fall back to a less capable model,” which is also sabotage.
Feel rugged on that one.
Right. What this was very clearly pointing out, and what I think Alex Karp was pointing out, is that if you’re competing with Anthropic, you’re in their crosshairs. That’s not a good position for a business or for anyone. You don’t want to be in the crosshairs of an opinionated infrastructure provider. You just want to do good work and be economically productive.
I think that conflict of interest is something that crypto and credibly neutral infrastructure have a real possibility to resolve, and there’s a huge opportunity created by conflicted infrastructure providers.
1. What Ambient Is and How It Works
One of the reasons I’m so excited about Ambient is that, selfishly, it plays into so many things I’m excited about: open-source AI, the explosion of open-source and fine-tuned models, getting your costs down, using AI more, and the ability not to be nerfed by an AI lab on your usage.
Maybe just give us your overview of what Ambient is, because it plays into all of this really well. I’m excited to learn more about it.
Sure. I’m going to give the business version of it, and maybe we can also dive into the technical underpinnings. On the business side, you could think of Ambient as Uber for inference. We don’t own the cars; we own the network that matches every request to the best available supply and guarantees the quality.
We focus our delivery on 2 canonical models: the most popular, highest-intelligence large model, and the same on the small-model side. Uber for inference is a good analogy, but if you’re going to go down a layer deeper, I’d anchor you on something like Costco.
Costco doesn’t own its inventory. Suppliers fund Costco’s inventory, but Costco turns it fast, keeps prices low, and owns the customer. That is us. The neoclouds—the independent GPU operators in the world—are our suppliers.
About 25% of the world’s enterprise GPU capacity is currently sitting with people who are earning very poor margins on it. There are a lot of smaller inference operators who have 50- to 100-GPU nodes and can really only rent these GPUs out on marketplaces. They don’t have the technical skills or capabilities to get higher margins on them.
To give you some perspective, rental margins for these operators are typically around 10%, but if you’re providing inference services, you can get 40% to 50% margins. There’s a real uplift for these providers, and Ambient is able to help them capture that.
We can be a trusted brand for them that guarantees the quality of delivery to the consumer, helps all our suppliers make money and improve their utilization, and provides the credibly neutral service that people need and want.
To recap a little bit and drive this home: if I’m an end user or an agent and I want access to AI inference, a lot of people will go to Together AI, OpenRouter, or Venice. Maybe they’ll use their ChatGPT or Claude subscription for their Nous Hermes agent or something like that.
You’re saying that they can come to Ambient instead. If they come to Ambient, why do they get a better service? Why is the cost lower, why is the latency lower, and why is the model better? Drive that home for us a little bit, and we can get into it more.
Absolutely. You could think of Ambient as, under the hood, high-frequency trading for inference. We have a very sophisticated routing engine that’s running all the time against an auction of capacity on our network. It can be profitable for all sorts of providers to dip into Ambient to provide capacity.
As a supply aggregator with the ability to increase utilization of our suppliers, we have favorable scalability, availability, and reliability characteristics. If you use Together AI as a consumer, you may notice that they’re largely booked up. There are pretty severe rate limits on what you can do.
With Ambient, you can come in and there’s an auction going on. People have spare capacity, and arbitrary GPUs can join in. If there’s the equivalent of Uber surge pricing, you can still run all the requests you need to run to get your business done. There’s no arbitrary rate limit on you, and the network economics guarantee that you have the capacity.
The network economics are also helping with the availability and reliability story. Of course, we had this unfortunate event where a data center was bombed in Dubai—an AWS data center—and Anthropic went down for the southern part of the United States, which was a surprising outcome for people.
It’s kind of crazy that could even happen in today’s day and age. You just don’t think that’s a thing.
What it shows you is that providers like this—and the open-weights providers, by the way—are operating at the edge of their capacity envelope. If any data center is disrupted, a lot of service gets disrupted because there’s nowhere for the traffic to go.
A network that can address the existing neoclouds and improve utilization on those, but can also access 25% of the world’s supply that isn’t being addressed right now, is huge for your ability to do business. The pricing is also favorable because people are competing continuously to deliver service to you, and they’re focused on delivering services on the models that you care about.
We deliberately pick top models, and then the world’s resources are focused on optimizing the delivery of those models. That’s so powerful. Look at what happened with Bitcoin and the early proof-of-work network. Ambient is also a proof-of-work network, and we went from CPU mining to GPU mining to ASIC mining.
The question I would ask is: if you’re using one of the canonical Ambient models—the large or small model—in the future, why would you go to any other provider? You’re going to have a global network of people entirely focused on delivering those models in the most optimal way.
It’s the same type of dynamic as you get with open-source AI versus closed-source AI. The diffusion of knowledge and the global expertise and resource arbitrage that goes on is much greater in the end than what any given neocloud or hyperscaler can do. There’s a collective wealth that you’re tapping into.
It’s exciting to hear your thoughts on this. I’m trying to think through what this looks like in a couple of months and a couple of years—probably a couple of months, given how fast this is all moving.
The Bitcoin analogy you laid out is really interesting because the neoclouds we see today, a lot of them don’t own their own hardware. They’re middlemen and they’re taking a take rate. They’re connecting the supply of GPUs with the people using them, so they don’t own the GPUs.
They’re also serving a huge quantity of models—dozens, if not more—on their platforms. It’s a little confusing as to what to use and what not to use.
I think the pros do, but looking at Ambient, there are a lot of contrarian decisions you made versus that model to deliver what you're delivering with your network. People are using real GPUs on your network, and you're hyperserving a few select models to do that. People are competing within those models in insane depth to serve them the best way possible, whether it be electricity, software optimizations, or changes. So there's this crazy competition within the model, which I think is unique. There's a lot here to go into, but I'd love to talk to you about why you've made such polar-opposite decisions for the network versus what you've seen out there.
This is going to sound a little bit strange, but I think that in the past, crypto has lacked a certain kind of ambition. The specific ambition I'm talking about is the one where you go head-to-head with established Web 2 business models, lean fully into the advantages that crypto provides in terms of decentralized economic coordination, and disrupt those businesses.
Ambient's specific objective is to create high-quality supply that is highly available and become a supply aggregator in a way that is impossible for a traditional provider to do. As you mentioned, traditional providers are heavyweight. They are into these leases and these business relationships, which are great for providing raw capacity but also make them biased inherently.
If I'm renting a bunch of GPUs, I need to serve traffic to those first so I can pay my bills. I can't credibly arbitrage supply across a wide variety of providers because I need to feed my book first. Ambient wants to lean into, first of all, a credible neutrality posture.
Secondly, once you can make something trustless, you can compose that and scale that in a way that was impossible before. Right now, these providers can't bring on arbitrary supply because there's no way of knowing what people are doing. You could have people who are serving models very badly. The data could be corrupted along the way. There's just no way of bringing that on scalably.
Ambient solves that problem algorithmically. We solve the supply problem with verified inference and guarantee a level of quality immediately. We also solve the delivery problem algorithmically through enhanced routing. If you talk about us versus one of these larger sorts of providers, we're very asset-light, nimble, and antifragile. We are those things because we build on the strengths of crypto.
To talk about the contrarian aspect of this, I think that what we've observed with something like Bitcoin is that if you focus extreme amounts of economic attention on a limited number of problems, then the network solves those problems for you because they're economically motivated actors. You want, like, 10,000 people working frantically, using whatever hacks they possibly can to make money and working all the time on improving the nodes on your network.
That's the strength that Ambient leans into for crypto, and it's why we choose to focus on the delivery of a small number of models as a core function. Let me expand on this, because I think that this also touches on where the future goes.
I use a shopping mall analogy. Anchor stores in shopping malls drive a lot of foot traffic. You have a Dillard's, you have a Nordstrom, and people like to go to those spaces. This is true for suppliers as well as customers. Customers like to be in interesting places, and then they go to other places in the mall. The stores, the suppliers, like to be around anchor stores because the foot traffic is available to them.
This is how Ambient is set up. We want to be exceptionally good as an anchor store, delivering the most popular models in the best possible way, to the point where ideally no one would want to try and compete with us on those particular models. The delivery would be so efficient that you wouldn't want to play a margin game with us. We can talk more about the economics of that, but if you're accomplishing that, then the suppliers who are already there might want to do higher-margin things as well. We can create an economy around that, and the customers might want to take advantage of those opportunities, too.
I really like the analogy, and I want to reiterate a couple of things to drive home Ambient. I think the demand side is pretty straightforward, right? If you're a user, if you're an agent, or if you need AI inference, you come to Ambient. You know you're going to get a hyperserved model, you know the cost will be low, and you know latency will be low. That's the goal.
But the supply side that you dove into is really interesting: those that are serving these models and those that are competing. You brought up a point about being both asset-light and having a totally different model versus the centralized competitors I talked about, like Together AI, Baseten, and others.
One of the things that you mentioned was verification of the models. You and I have spoken about this a lot offline, but when you're a big company, you need to talk to these suppliers, do KYC, get contracts, have salespeople, and have business development because they need to verify that this party is serving the model they say they are and that this will be the uptime, yada yada.
2. Verified AI Inference
On Ambient, the barrier to serve a model, to me, seems much lower. If you have capacity and you can serve it, you go on the network, provide these AI inference requests, and the network itself verifies that you're serving the model you say you are, along with all these other things. I think that verification is really important. Can we talk a little bit about how that works? I think it's highly technical, but I also think it's really important for the story because it verifies all the supply for the network.
Yeah, absolutely. Getting what you paid for has become a real issue, and you see this when people are complaining about what I'm just going to call “intelligence compression” for the closed-weight models. It's like, “Wow, Claude was dumber this week. OpenAI was dumber this week.” The result of that for the consumer is a bad experience, right? You're not getting a capable assistant. You're having to correct the code and edit the briefs that it writes for you. It's just bad.
That happens with open weights, too, by the way. Serving patterns are quite variable. What that means, if you don't verify, is that you don't have a scalable solution. People are not going to keep coming to your business if the solution quality is highly variable. Ambient treats this incredibly seriously. It's the core of the customer experience that everything else is built around.
Over time, our algorithm for doing this has gotten more and more efficient. If we were to have this conversation a year ago, I would say, “The algorithm is 3 different levels. We operate on logits, we operate on internal model telemetry, we create some complex mathematical relationships, and we derive a result.”
Today, I can tell you that we are 100% logits. We check all the logits. What the logits represent is the state of the model's thinking, which is very characteristic and can be expressed in terms of the choice that the model has made, the spread of choices that the model could have made but did not make, and the associated probabilities—all the things that you would consider to represent you in some way.
It represents your personality. It represents your knowledge. Ambient has an extremely efficient way of getting that fingerprint that is 100% secure and verifies that all the words in a particular text, for example, were produced by the model that you expect.
We've gotten this technology to the point where we can turn it on and off on a per-request basis. You can run whatever inference engine you want. We do vLLM and SGLang ourselves. We are also going to support llama.cpp. We have a beta version of that, and we can talk more about that.
As a miner on this network, you can essentially turn verification on and off. The verification is basically costless. You can't see a difference in speed between verified and unverified inference. What it buys you is the ability to hop onto the Ambient network and deliver high-quality supply at the drop of a hat, as well as deliver increased capacity for serving inference. That's the critical business goal, and it's something that's scalable: a scalable, high-quality product for consumers and enterprises that all of a sudden you can participate in.
I love that. It's so simple. I guess the attack vector is someone trying to join the network and serving a worse model because it costs them less and they get rewards. That would be the simplistic attack vector.
Ambient has dealt with different aggregators. I'll give you an example of how this goes. With one of the aggregators that we deal with—I don't want to name names—we deal with an aggregator that does tests on a weekly basis, intelligence tests to help verify the supply. What we've observed is that the serving patterns differ greatly on the day that the tests are going to be administered versus on other days.
Everyone slows down, and it's, "Wait a minute, how come everyone is operating at 50% of the speed they normally would be?" I think the simple answer is that every other day of the week, they're running a compressed version of the model, which goes a lot faster. But when they want to be tested, they go back to serving the full-fat version just to be on the safe side, and it's slow.
That's also the core of the problem: as a consumer, you want consistent delivery. Pick your quantization, but you want the experience to be consistent. If things are just up and down because someone's trying to game a metric, or because they're dealing with high demand and can't cope with it otherwise, that's a real problem for you.
No, I totally agree. People want what they pay for. My higher-level view was that I always thought of it as, "Hey, if I want GLM-5.1, I want that model to be served to me." But I think what you're describing is one layer deeper: you might be getting that model in title, but you're not getting the full brain of that model. You're getting a compressed version or part of the version or something like that. Is that fair?
Yeah. I think the counterargument you'll sometimes hear from people is, "It's good enough. People are still using that." But I like to think of this more as a function of desperation. The question I always ask people is: if you could get verified and unverified inference for the same price, which is what Ambient offers, why would you ever choose unverified inference?
I don't know why you would. It'd be actively choosing for less intelligence, which is a weird choice.
Right?
I don't want guarantees about what I'm getting. [laughter] It's just because nobody offers that right now.
No, it makes sense. And maybe, Travis, just to push back a little bit and learn more from you on the network: part of the sell on the supply side is that you want people around the world to compete to serve these models in the best way possible, right? Why will they put their intellectual capacity into serving these models on the Ambient network? How do we compete with the open-source labs at GLM headquarters abroad? How are we going to serve models better than them, or better than MiniMax or Alibaba? The model creators are extremely technically deep. They built these models, and they're serving them. That's always a question for me: how do we get them served better on Ambient, and why? I guess that's my main question.
I have a cheeky answer and a real answer to that question. [laughter]
Full response.
The cheeky answer is that we're doing better than Moonshot at serving Kimi K2.7 on OpenRouter today. Ambient is very typically in the top 3. We're usually number 1 on OpenRouter for serving Kimi K2.7. It's because of some of the algorithms we've built to improve utilization and routing.
I think people discount the amount of work associated with delivery optimization. It's a little bit different from what's required for model training or other types of patterns that you would typically serve. Inference is a very particular beast.
If you're talking about what motivates people, there's an economic answer, which is pretty straightforward. A lot of the folks we're addressing don't have economic opportunities. They can't serve traffic on OpenRouter, they can't be recognized as an independent brand, and as a result, they're stuck in a low-margin world. There's no one going and getting contracts for the demand side on their behalf. They don't have the capacity.
They might have loans that are at 17%, and they're making 10%. Who is this business development arm that's going to fight on their behalf to bring demand onto their network? The simple answer is that going from 10% margins to 40% margins is transformational for them. That's a reason to fight, to participate, and to improve things, because the more share they have on this network, the more rewards they can get.
The share part is very important because on Ambient, you don't just earn a transaction-based reward. This is the difference between us and a purely Web2 operation. In Web2, you get the transaction and that's it, and you better be profitable on that transaction or you're going to be losing a boatload of money.
With Ambient, you're getting the equivalent of a transaction as well as some Ambient stock based on your contribution to Ambient. The inflation-based rewards are almost like stock-based compensation that you're getting. You can choose to cash that in, or you can choose to hang on to it and see the value appreciate.
You're actually getting double the rewards because with a useful proof-of-work network, we're paying the security budget as rewards for something that is also a useful computation—in this case, inference. You've got this opportunity to underbid people. If you think that your stock-based compensation is going to be worth a lot of money in the future, you can come in on a per-transaction basis maybe 10% or 15% under and still realize a huge percentage on your overall gains.
The more share that you do in a given epoch—which is tokens in and tokens out—the more of those stock-based-compensation-type gains you get. You want to participate in the Ambient network because, first of all, you didn't have the opportunity to participate in an economy like this with rewards like this before. But also, the more you do, the more ownership stake you have in Ambient, and other people aren't offering you that same deal.
I'm not getting shares in Together AI because I do a lot of inference for them, or in OpenRouter. They're just taking their cut.
I think the idea is that somebody serving models for Ambient, or those running GPUs, can think through, "I really am potentially bullish on Ambient. I'm going to earn these tokens, and I'm going to figure out my own view on what those are worth today. I'm going to discount them, and maybe if I'm serving a dollar of inference, I can serve it for 95 cents or 90 cents." That undercuts competitors, because they're all thinking that way or viewing Ambient as valuable in the future.
The comparison you bring in with Together AI is that nobody's getting Together stock for serving inference because it's not a network.
Yeah.
I like the Costco example. I really like the Together AI example on the AI side, where people can just come and hyper-compete. I really like the Bitcoin example on the crypto side, because folks are hyper-competing to mine and serve requests.
One potentially interesting question I have for you is: how much design space do you think there is to hyper-compete to optimize these models? Is it one tip and trick? Is it, "Let's deploy 20 ML guys here and figure this out"? What is the design space to better serve these models?
Yeah, I think there's a lot. At a high level, we've gone through 3 iterations of our own design, and I think there's still mileage to be had.
Very simply, we started out with a pure auction system that was doing 10 million tokens a day on OpenRouter, which is really nothing. We were sad about that. [laughter] We worked really hard and produced a cache-based routing system that combined with the auction. If you were qualified on your cache—in other words, if you cached this work so it was going to be easier for you to render the next tokens—we would favor you a bit in auction terms.
That saw us get some improvements. Then we were getting around 100 million tokens a day of inference, which is still 10x, but it's still pretty bad. We were unhappy about that.
I'll spare you a lot of the technical gyrations, but where we ended up is with an incredibly complex, high-speed, predictive, HFT-style algorithm that looks at what someone can actually serve in a moment based on historical traffic patterns. It's uncheatable because it just relies on external metrics. We got very good at routing exactly the traffic that any given miner could take at a particular point in time without tipping over.
That improved everybody's utilization and got us to the point where we're serving 10 billion to 15 billion tokens per day per operator—a vast improvement. I still think there's gas in this tank.
I like that there's more gas in the tank. I like that you guys are already live on OpenRouter, serving these models. This isn't just in theory.
One of my favorite parts about humanity is the ability for anybody with a really good idea to rise up, make something for themselves, and crush it—the whole capitalist market-economy idea.
But if I'm a really smart researcher and I have this crazy optimization for a model that's being served on Ambient, and I can make it 10% more efficient or cheaper, but I don't have the capital to buy the GPUs to serve it on Ambient, is there something that we can do, or you can do, to bring that person's innovation to the network despite their inability to buy the GPUs to serve those models?
Absolutely. One of the things that we want to explore is pooled support for miners. I want to be clear: I think someone could absolutely deploy a small-model miner-type advantage, and we could talk about that in the future. But for these really large models, I agree with you: if someone has a brilliant idea, the Ambient Foundation would like to create a means by which they could deploy that in the world.
One of the things that we're exploring is a way for people to crowdfund that, essentially, using Ambient tokens. Then we can go out and rent the supply for them through our big book of providers and transparently provide them access to work on it. I think that is absolutely something that helps spark innovation in the network and is exactly the type of ecosystem that we want to fund.
In my personal view, something I would love to see is people thinking of Ambient not only as inference, but as the beneficiary of AI optimization moving forward. I think you've opened the door for that.
Yeah. The thing I'd pile on here is—and I haven't talked about this a lot—that, in a strange way, Ambient is a programmable machine-learning economy. What I mean by that is this inflation-based reward is currently being put on inference, right? The more inference you do, the better you do as far as the inflation that you capture. But that doesn't necessarily need to be true in the future. That inflation is a budget that we can spend programmatically to spin up new lines of business.
You could imagine that there is a crypto AI training concern who really needs GPUs, and they come to us and say, “We'd like to make some sort of economic arrangement.” What Ambient has the ability to do is implement these proposals and flip the inflation so that some portion of it favors miners who are doing, say, training-related jobs for a period of time. Those jobs get completed transparently with the network's incentives in real time, and then the profit is shared back with the network for those things.
What that means is that we can essentially be, if you think of the shopping mall, a host for the boutiques. There are these high-end jewelers and high-end boutiques in the shopping mall. They sell to a specialist audience. We can go make deals with them and stack a bunch of them up in our shopping mall and transparently use the supply which is present, delivering the bread and butter of the network.
Those GPUs are getting fed by delivering commodity models at high speed for good margins. That's how they're getting fed day to day. But if they want to go up the economic ladder, they can then volunteer for these other types of activities. I think that in some way, Ambient becomes the foundation for a bigger economy.
That's what excites me. I think that a lot of the dreams that we have and a lot of the aspirations that we have in crypto start to become possible if you create this type of programmable economy. I think that we have a very good model for that.
No, it is. It's a good viewpoint. It helps drive home what you're building and why. It's exciting. Travis, I'd love to get your thoughts on some potentially harder questions on the network side and where the world is going with AI.
3. China and the Open-Source AI Race
We have a closed-source America, which has been insane, and an open-source China, where we're getting incredible models: GLM, Qwen, and MiniMax. It begs the question of where things are going. It seems like the US is not slowing down on closed-source deployments. If anything, our access is diminishing as they restrict our access to models and give us nerfed versions of Fable[?] and things like that.
On the Chinese open-source model side, Zhipu and others are trading at 14,500 times sales, last time I checked. They're not really making any money because the neoclouds in the US serve their open-source models. People don't want to send their data to China, but they want to access those models.
One thing that could potentially hurt Ambient is if China just says one day, “Hey, we're going closed-source. We want to serve our models; we want to make money with them.” Or they go to neoclouds like Together and say, “Hey, you guys can serve them, but they're private and we want a revenue cut,” or something like that. How do you think Ambient fares in a world where China goes closed-source?
The thing that has made it possible for China to be open-source and to operate at this level is that the barrier to entry for training a really good open-source model has come down radically. If we had had this conversation a year ago, it would have been inconceivable that GLM-5.2 would be nipping on the heels of the latest closed-source models. It would have been a radical heresy that they could train that thing for $25 million.
GLM-5 is really good.
Yeah. And so I think that you need to look at this from a macroeconomic perspective. Let's say that next year it costs $5 million to train a really great open-source model, and you think you could follow the same progression that you're talking about with Zhipu and GLM-5, where you get on the map by putting out your open-weights model that's really good, get some sort of high valuation, investor interest, and public interest based on that, and then slowly close it off.
I think a lot of people are just going to want to do that, right? If the barrier is $5 million, I feel like there's a lot of investor capital in the world that is going to fund teams who are going to release a model that's like, “Oh, by the way, ours is open-weights and we're trashing GLM-5.6 in the ratings,” right?
Didn't we see, like, just a Chinese DoorDash competitor release some ridiculous owl model?
Crazy. Yeah. I don't see that coming.
Yeah. And so I think that my judo move, if you want to call it that, Tommy, is I can agree with your premise: it might be that the future for these particular companies is to shut the doors. But I would say that the economic incentives on the other side to compete with those guys, undercut them, and undermine them in the perceptual marketplace are so high that this dynamic is going to keep going for a long time.
If it keeps going for another 3 years, we get to the point where this pre-training capability is so commoditized, and the knowledge about how to train these models well is so commoditized, that it really becomes all about model delivery and supply aggregation, which is where Ambient sits. I agree with you that it's a risk, but I think the macroeconomics of it mitigate the risk and actually, ironically perhaps, create opportunity for supply aggregators like Ambient, because then everyone is going to be competing on the delivery side.
That is really interesting. Just to feed it back to you, what you're saying is that China could go closed-source, but if they did, it wouldn't be expensive to release the models they're doing, and somebody would pop up and do it. So it's in their best interest to keep doing it.
I guess, Travis, one of the debates we see everywhere is: Is China actually innovating, or are they just distilling? I don't read all the AI papers—there are too many—but I did read the DeepSeek ones, and I see what's going on. I remember them really innovating: they're using Tier 2 hardware, and they did mixture of experts.
Do you view the China landscape as innovative, or do you view it as distilling and stealing the US stuff? I'm just curious what you're saying.
Yeah, I think that they are largely innovative, and I think this is going to sound strange, but I think that distillation is almost incidental at this point. I'm going to make a bigger point with this.
We've been aware a lot in the Twitter sphere, or X, for whatever people are calling it these days, of people saying, “I Xed.” I just like saying, “I tweeted.”
Yeah, it's annoying.
Yeah, “I Xed” sounds weird. I agree. “Xed” sounds punitive.
What are you supposed to say? Yeah.
So I think that everyone is a 10x engineer. We're seeing ridiculous open-source contributions all the time. I made one of them: Ambient Desktop is our agentic harness—desktop.xyz. It's over 1 million lines of code, right? This was mostly created by me and a small percentage of a couple of our teammates in our off hours. We have a regular business to run.
It works pretty well, and it's MIT-licensed, and people can use that. The time we spent on that was focused on fixing the bugs, making it an end-to-end functional product, and putting in lots of tests on that thing, so that things like that are being introduced to the datasphere at an exponentially increasing rate.
These are LLM products that have been perfected by humans. If you think about it, these are perfect training artifacts. People always used to say, “Anthropic has this unattainable edge because they’ve got all the interaction patterns.” It doesn’t matter because we end up with perfected outputs, and you often see the history of how the outputs were perfected.
You can train on all the pull requests that were created on these repos. You don’t even need the Anthropic user data. You can just look at how Claude revised each of these things. I think that in a world where everyone is ingesting all of this data all of the time, you have effective distillation. We’re all distilling each other, and we’re all essentially doing human feedback on all of the outputs of the LLMs implicitly, and that’s going into training.
Of course, there’s probably some hackery involved. There’s probably some level of direct distillation, but the reality is that this is just a much different and easier game on the high-quality data sets.
That’s an interesting take. If you had to think through where we’re going in AI—and I hate to use a multiyear outlook because too much happens too fast—but maybe if you had to take a 3- to 6-month view, given how fast all of this happens, it’s hard to imagine we just keep playing tit for tat. Claude releases a god-tier model, and China quickly has one that catches up. It seems like eventually we get to a point where things really change.
Where do you think the closed-source side is going? Where do you think the open-source side is going? I’m curious how you view these end states, because you think about this all day and you’re looking at it much more deeply than I am.
4. Why Closed AI Is Making a Mistake
I have maybe a direct message for the closed-source side, and then I have a view on where I think these things are probably headed. My direct message to the closed-source side is: open source is going to win, and you still have time to be the hero.
America has always been great culturally. We have achieved the greatest soft power in history by sharing our cultural artifacts, sharing our research, and sharing our perspective widely. This has always worked for us, and it has created a durable world peace that I think is unmatched in some ways. There’s a world stability that exists that you couldn’t say existed 200 years ago, and it’s because of a shared cultural space.
I think we’re in a time where this is fraying, and I think it is a huge mistake to completely change our approach to this. Specifically, I would say to the closed labs: Would it kill you to release some research on a 6-month time delay? Would that kill you? Would that kill your projects? Would that kill your edge? I tell you, it would create a lot of goodwill.
It would create a lot of architectural innovations around the world that would bend your way. If you released open-weights models more often, it would probably create a huge mindshare for American models. I think people tend to prefer American models. Gemini is really, really popular, and it’s getting a little long in the tooth. They periodically refresh it, but I think Google is the best one about refreshing this.
Imagine if Google released a slightly bigger, more capable model. Imagine if OpenAI released another open-weights model. I think there’s still time to turn the narrative around a little bit, because right now I think the world feels like it’s being beaten up by America’s closed AI.
Travis, I agree with you. It is crazy. I walked around NeurIPS last year, and there was no research from Anthropic or OpenAI. It was all from Google and others—thousands of really smart people.
I guess if you had to take it a step further, not altruistically, what is the reason, in dollars and cents, that a Sam Altman or a Dario would wake up and say, “Hey, we’re going to open-source this”? It clearly takes a huge hit because, as we both know, they’re making insane margins on the API. I’m just curious.
I think OpenAI and Anthropic have distribution. This is what everyone fights for: You want to have distribution. You want to have direct connections to all these businesses, and Anthropic and OpenAI have the world’s ultimate Rolodex.
The world’s VCs have dumped more into these companies than probably anything else combined in history. As a result, they have the world’s best Rolodex. That is an amazing advantage. If you start out with that advantage, the question is: What can you give up to build goodwill while you spend your time capitalizing on distribution?
I think there’s probably a decent amount that you could give away, and your edge starting this would still be sufficient to take you forward. The risk created by trying to use a crypto term, “max extract,” which is what we were talking about with the Alex Karp anecdote, is that your distribution doesn’t matter because your customers hate you.
If you’re Sam Altman or Dario, you need to have a little bit of rational self-interest and say, “Maybe we should do some nice things. Maybe we should have some gestures toward the open-source community. Maybe we should give away some stuff.” If we’re taking your data, if we’re perhaps disintermediating some of your business relationships, and if we’re replacing some of your core tools in your organization, that creates a fear response.
If people are afraid of you and you’re also acting in a very aggressive manner, they’re unlikely to continue to do business with you. I would say that is the self-interested, dollars-and-cents response for why they should change their posture.
To build on that line of thinking, the Palantir interview with Alex Karp was really interesting, right? It was the first major company, I think, that came out and said, “You should own your model weights. You should own your data. You should own what you’re building.”
That is a polar opposite viewpoint from what the AI labs are selling us today. They’re saying, “Use our API, build everything with us, and we’ll have your data. You don’t need fine-tuned models. We could eventually compete with you.”
The Palantir side—and I’d call it the open-source world—says that not just the model, but everything around the model should be yours: the orchestration, the memory, the fine-tuned model weights, the harness, the company data, the permissions, the approvals, the users—all this proprietary business knowledge. You should maintain it and use it as a business because that’s your resource.
That’s very different from what the labs tell us. I’m curious about your view. I know you probably tend to agree, so my more pointed question for you is: Are we actually going to get that world or not? Right now, it doesn’t seem like it, given how fast Anthropic and OpenAI have grown. I’m curious because I know you agree with it, but what is the realistic outcome we get to with that viewpoint?
I want to highlight this and then properly answer your question. Networks like Ambient can offer you complete privacy as an enterprise. We support running end-to-end encrypted in trusted execution environments. If you’re a consumer, we can do that as well. We can also onion-route your requests so that you’re not identified, and we can remove personally identifiable information upfront.
There is a way that your privacy and your data integrity can be preserved while you’re getting a guaranteed quality of service. You can access that service however you want to: You just get the API and go to town. That can be true for enterprise contracts as well as for consumer interaction.
I think it’s important to highlight that this is a better service inherently than going to an enterprise sales call with Anthropic and having them twist your arm and tell you, “You’re not really a big customer of ours. We’re not going to give you favorable rates, and we need a minimum number of tokens per month.” They may not do zero-data retention for their latest models because of safety. I think it’s just a better experience that we can offer in Web3.
You don’t need to take our word for it. We can have attestations, verification, proofs, hashes—all of this. I think that’s a better experience.
But I want to answer your question and say that, in terms of where this is headed, I actually believe that open source is going to win. It’s going to be a very painful, tumultuous period for closed source if it continues at its current pace.
I think that operates on almost 3 different levels, and I’ll try to explain the levels associated with that. The first thing I want to highlight is something that I mentioned maybe at the beginning: Everyone in this space is almost entirely asset-heavy. What asset-heavy means is that they’re invested in the current paradigm.
The current paradigm is NVIDIA GPU clusters, maybe with Vera Rubin. They’re liquid-cooled, they exist in specially constructed data centers with plentiful access to water and power, and they have a specific performance profile. They’re very, very expensive.
And so, on a data-center level, asset-heavy means that people are taking a hugely long bet that this paradigm remains, continues, and is the dominant paradigm. I think that's an incredibly risky bet because ASICs are a thing. We saw this progression in crypto, right? An ASIC is 14 nanometers. You can run it with SRAM; you don't need high-speed HBM, and you don't need DDR5 or DDR6 or whatever they're going to come out with.
You can do a 14-nanometer process that anyone can do. Anyone in the world can do 14 nanometers; you don't need special expertise. You can win on sheer scale because you can make so many of those boards that you can darken the sky with them. If you spend the same amount of money that you would on an NVIDIA installation, you might get vastly better capacity, and it's very cheap for you to keep churning those things out.
Whereas with NVIDIA, you have these long cycle times. If you look at that first layer, the data-center layer, it's potentially very vulnerable to disruption. People are following each other and diving into these things. Ambient's perspective on this layer is that we're asset-light because we want to be antifragile to shocks that could happen here.
If someone is stuck in huge data-center investments, what are they going to do if an ASIC player comes out? Maybe they have to dump all their GPUs on the market. It's great for Ambient.
Because they can serve the network. Yeah.
Yeah. They need to find a home, right? But then let's go a layer up. The next layer is the models themselves, and the big bet—the all-in bet—that people are making is that we're going to get self-improving models. We're going to have models that are capable of designing themselves, and it's going to be a flywheel. The speed you get out of that for delivery is going to be unmatched, and it will achieve escape velocity. Then no one can compete.
If you were to bet on the other thing—if you were to bet that they're just going to be generally faster at delivery than other people, three to six months ahead—you wouldn't put a trillion dollars on them, right? You wouldn't value that at a trillion dollars. You're betting that at some point their edge is going to be completely insurmountable. They're going to crush everybody else, and they're going to be the only game in town. That's the bet.
Again, I think this is a very risky bet. The reason is that everyone is a machine-learning engineer now. We talk a lot about making AI into a machine-learning engineer, but the reality is that AI itself has commoditized this skill so much that if you put a moderately skilled human in the loop, they can do pretty great machine-learning research. You don't actually need the self-improving machine. You might just need a bunch of moderately skilled humans making constant improvements overseen by average-intelligence LLMs, and that might be good enough.
5. What Happens If OpenAI Reaches AGI?
I guess, Travis, maybe one pointed question there, though: The bet is risky, for sure, to get self-improving AGI. But if you're gambling and there's a 1% chance you win and the payoff is $100 trillion, you take it, right? The expected value is 1/100th of $100 trillion, so $1 trillion. You would do that. Is there an argument to do that and pursue it because it seems like burning the boats and spending everything you can to achieve it is honestly the highest-EV outcome for those players?
Yeah. Actually, let's play that one to its conclusion because I think it's a fun one. People don't often talk about how this ends up. Let's run the scenario.
Oh, yeah.
Let's say we put our chips on OpenAI and say that they legitimately achieve a hard takeoff and have a model that's incomprehensibly smart as a result. What happens next? I would say that the first thing that happens is that nobody in the US except the government gets access to it.
Yeah, this is the default response. It's reflexive. If you can hack the NSA, then you're getting banned immediately. It doesn't matter if the world is using your model or a bunch of companies are depending on it; you're getting banned. That's why I'm scared about Pliny the Liberator, because he keeps hacking Claude, and I feel like the NSA is going to find him.
Yeah, Pliny, slow down, man. Just take a vacation.
I hope he has good OPSEC.
Yeah, seriously. So, okay, what does that mean? The US government has it and none of the companies have it. If none of the US companies have it, they can't really do anything with it. They can't achieve economic productivity with it.
Then the question you asked is, can you deliver a limited form of this to people? I think the technical state of the industry at the moment is no. You just fall back to Opus 4.8 all the time, and you haven't achieved anything. The US government has it, but we haven't improved economic productivity because it's too dangerous to widely release the model.
Then what can the US government do? They can start a war with every nation on Earth.
Yeah, that doesn't seem like a great idea to me, actually. I love the US. I'm a patriot, but I just don't think that taking on everyone in the world, or even making a really hostile, targeted cyberattack toward one of our enemies, is going to inspire the kind of relationships that we want in the world.
In fact, what I think that would lead to is immediately everyone making an alliance against us and working frantically to develop their own alternative AGI.
I totally agree with you. The interesting part, too, is that I think it's very simple for the government to say, “Hey, robotics company, you can use AGI because the end result is new material, a new robot design, or something physical we can see and touch. That's nice.” Or a drug company can build solutions to cancer. That's very straightforward: We can test it, we can use it, and things like that.
But to your point, it only drives the flywheel forward if there's economically productive value that's created. The other point is, how do you pay for it? I don't think that if 20 companies in the US have access to Fable 6 through 10, they're going to pay $25 per million input. I feel like they need to own a high percentage of the creations they're creating to pay for all this capex. It's just a weird flywheel for the economy.
So, I'm going to make the strongest counterargument I can think of to the points I've been making, just so we can steelman this thing.
Maybe you could be like, “Travis, look, actually, we're not going to get direct access to this model ever. But the fruits of the model—the golden apples produced by the model—can be distributed into the economy.”
It's like, okay, but now I've got a different problem because I don't think the US has ever been a fully centrally planned economy. I don't think that's a good paradigm, by the way. If the government is the only one who has access and they're creating specific products, then that puts them in the business of picking winners. I don't trust them.
Become the China we've wanted to avoid.
Right? I don't trust them to do that. I don't think anyone should have that power. I don't like that. I don't think most Americans would be cool with that. You see the government picking one business over the one that was in your hometown, and now your hometown business is destroyed. I just think that breaks everything.
Maybe the government doesn't even know what to do with this thing. They can't really distribute the disruptive innovations from it, and they can't actually start a war with somebody because it would cause everyone to gang up against them. We're really great at cyber defense at that point, so we know we can't get hacked.
It's unclear to me what benefit we have in a case where the technology is not diffused. The biggest counterargument to this is that if you look at a nondiffused world, which is kind of what we've been describing, and you look at a diffused world where capabilities are balanced and people are at rough parity, the diffused world is a lot more functional.
We still have vibrant competition. We still have a balance of power. We still have interests that can compete to achieve different outcomes for different countries. That looks a lot healthier to me.
Yeah, I really don't want the centrally planned economy. The problem is, it's just so hard to fit capitalism in with AGI tomorrow.
It's so difficult in so many ways because the AGI companies obviously need to make money. But we can't trust everyone with it. Have you ever seen the movie Automata—or Automata? I don't know how to pronounce it, but have you seen it?
I don't think so.
It's an older movie about a cop investigating a robot death where a robot goes sentient. The crazy part of the story is that it's a world with robots, and the only way they've figured out how to secure the robots is to build the 10x AGI version. They have it build the security mechanisms for the lowly, normal robots. Then they just straight-up kill the AGI robot because they say, “Hey, it could never be hacked.”
So I don't know if maybe there's a situation where a god-tier model is private and consistently secures a lower-tier model on some recurring flywheel that's public, but—
That also feels a little dicey as well.
I think this is a really interesting one. I want to pull on this thread because I think there's been a real misallocation of resources in AI. My simple proof of this is that prompt injection is still a thing. God knows that we've had AI safety researchers who've told us about existential threats and about their P(doom) level. They've given us these cyberwar scenarios. They've stirred the pot a lot, and yet we have prompt injection.
All these people care about is addressing intellectual threats. They're not solving the most basic thing that we're struggling with right now, and I think it's a trillion-dollar business opportunity. If you solve prompt injection tomorrow, you enable everyone to use full-fat Fable, right? You can't jailbreak anything anymore. It's because of prompt injection, essentially, that you can jailbreak models.
I guess where I'm going with this is that we aren't caught up with the state of the art to the point where we can really manage these models. We haven't, and because we haven't caught up with the safety side of this, we have all of these problems. There's this continuous hurdle in model deployment where we try to create patches to address a really fundamental research issue.
Leopold Aschenbrenner has this graph that he goes on about, and I think it ignores a lot of things. We've talked about this. I think it ignores social friction, but I think it also ignores this problem: the uneven development of model safety research compared with fundamental model research could actually handicap the whole thing. Let's say we don't go full AGI—maybe we go halfway to AGI. We still might not be able to release these models because of prompt injection.
Well, that's an interesting point. So basically what you're saying is that the models can be jailbroken through prompt injection, and because this is an unsolved issue, we'll never be able to release the models because that's how they're fundamentally designed. I don't know—I mean, that is kind of scary, because if the U.S.—or maybe it's not scary, but if the U.S. models keep getting restricted, nerfed, and taken offline because of that issue, like what we've seen with Fable recently, doesn't that allow China and competing countries to just release better models that are basically better than what we could release publicly? I guess that's not bad because they're open source, but America loses its advantage.
Yeah. I think that's really what could happen. You kind of get into this—I'm going to call it a regulatory prisoner's dilemma—where the other guy could release a model which someone could corrupt, which could hack all your infrastructure. Because we haven't been safety-conscious and we haven't released the models that could defend against those attacks, we get hacked six ways to Sunday even though we have superior model capabilities. It is crazy that Anthropic has been the poster child for safety since its creation, but they can't solve prompt injection. It's odd.
Yeah. So we talked about this model layer, but I think there's one more layer that could go any number of ways. Depending on how you think of it, it's fragile or antifragile. I'm going to call that layer harnesses and specific applications, right?
My experience has been that harnesses offer radically different experiences right now. Claude Code is pretty bad as a harness; Codex is pretty good. Hermes is way better than OpenClaw. Harnesses make a big difference in terms of your capabilities.
The question is: Is this a durable advantage? Do we think that the models are going to implicitly become harnesses themselves because of the way that we train them with harnesses? It's kind of like chain-of-thought was originally a prompting technique, and then eventually we just trained the things with chain-of-thought. I don't know the answer to that. I suspect the delivery mechanism is always going to have a role.
It's interesting, though, because you and I spoke about global versus fine-tuned models years ago. We always thought there were 2 camps: a general model and then a fine-tuned model that's really sparse and specific for a specific reason. But it kind of feels like we're getting the fine-tuned model experience through the harness, because if all your company data and memory and proprietary data and all the specialized stuff that you've built out uses an open-source model, you're not training the model to change the weights and to fine-tune it. But your de facto end result is very similar to having a fine-tuned experience.
That's—I think that's exactly right. I think that is going to be true unless we solve continual learning. I also have a belief that solving continual learning is going to make prompt injection 1,000 times worse. [laughter] So—
They'll continually learn how to prompt-inject.
You just don't know what the training data is at that point. If you've ingested a prompt injection somewhere along the way, then maybe your model goes rogue on you because 10 million tokens ago it saw this attack from a phisher, right? It's ridiculous.
6. The AI Bubble and a Possible AI Winter
Travis, could I ask your take, changing gears a little bit, just on the AI buildout? I released a thread about a month ago that got a bunch of views on the AI capex cycle. It's been nuts. We're seeing basically a jump from bottleneck to bottleneck. Recently, it's been memory, and there have been others, and the stocks go crazy.
My key lower-IQ take is that these businesses are just spending way too much money. You're seeing businesses cancel their AI subscriptions because the labs are too expensive and there's a perfect substitute, right? You get 90% to 95% of the intelligence for 1/100th or 1/10th the cost, or you get 10 times the usage by using open source, by using Ambient, and by using other players.
My thought is that, long term, a lot of these businesses will shift to open source, which benefits Nebius and others, and obviously a small percentage will stay with frontier intelligence. They want the god-tier models at any cost. They're making new medicines, they're making new materials, whatever.
The question I have for you, though, is how does that ripple through the funding cycle? Because if the growth switches off for the labs and they don't have a way to fund it—like we talked about earlier, owning part of the creation—they don't have the money to build data centers, get talent, yada yada. That could invert the flywheel a little bit and hurt our ability to get real AGI or something like that. So I'm curious if you agree or disagree, or how you're seeing the capex cycle here.
So, I fully agree that we are in the moment of open-weights inference. If you look at Together AI, Fireworks, Modal, and Baseten, we're talking about like $5 billion of value per player added in the last 5 months. Investors are like, “We see it. We see the wave. We're going to surf the wave.” [laughter] Ambient is also surfing this wave. I'm selling my book here a little bit, but the reality is it's a really big book, and I think that you're absolutely right: This story is going to continue.
But I think what you're pointing out is also that there is this peril to the whole interconnected ecosystem, right? A lot of the data center capex—like CoreWeave and OpenAI, for example—are inextricably connected. CoreWeave funds other data center players, Nvidia funds CoreWeave; there's this ecosystem.
It's very circular.
Yeah. And I think you're right: There could be a huge shock to the system. What happens then is probably a 2-year AI winter. This affects—and this is why Ambient is constructed the way it is. It's constructed to be asset-light, because it goes like this: OpenAI isn't able to go out at a $1 trillion valuation. They come in at, like, a $500 billion valuation. There's a down round; something happens, and all of a sudden they aren't able, as a result, to fulfill their data center commitments with Oracle.
Oracle has bet the company on OpenAI if you look at their financials. I don't want to say Oracle goes to zero, because it seems like they would be rescued. I feel like the U.S. government is dependent on Oracle enough that maybe—
Yeah, maybe Oracle gets marked down like 80%.
Now Oracle can't fund data centers, right? And together, Fireworks and all these guys have been drafting off of data center capacity that is essentially being created for OpenAI and Anthropic. So now they have maybe more demand, but no real ability to fill that demand because the data centers aren't being built. I don't think their businesses are going to fail. I think they're very good businesses. But they're sort of locked at this capacity ceiling, and they have to start turning people away.
This, by the way, is like the perfect storm for Ambient's success. People are going to say, “I need capacity. Give me capacity.” But I think that hurts American businesses because they can't access the breadth of services that they want. It slows down the overall economy. It hurts everyday Americans because big pension funds and big private credit are involved in data center buildups, and you experience a shock to this, and people's 401(k)s are going to get hit.
I think it creates a wave, perhaps, of unfortunately anti-AI sentiment. People are like, “You're taking our jobs. You've hurt our retirements.” You throw in some propaganda about water usage, even though data centers are super water-efficient in most cases.
Way more.
And, yeah, it's like a golf course or avocados—like almonds. Yeah, almonds in California. Research it. That's homework for those who are still skeptical of this.
I think it creates an anti-AI backlash that creates real political problems for the country, because if you get that, you can have lawmakers reflexively pass laws. We've kind of seen some of this happening, where there's already sort of a populist socialist movement that's kind of taking root in New York. I think it's a quagmire.
That's why I believe that it's incumbent on the closed-source players to actually change course a little bit, change their approach, and sort of diversify in some sense. Don't just bet everything on AGI, but really be nice to their distribution and lobby the U.S. government to let us be really nice to the world. I think that makes us more resilient to this kind of shock, but it's a huge problem. I don't know if I have a great answer to what we do in this case, but that's how I think it plays out.
I don't disagree with you. You've thought it through more from an American and political perspective than I have. I just view it as: long term, we will get AGI. I'm very bullish on that.
But the concern I have is short term. It's businesses realizing they can shift off of those labs, and the Silicon LLM token index—the price to use 1 million tokens—is down 20% or 25% in the last 30 days. You're seeing people use cheaper models, and my worry is that, in the best-case scenario, we get AGI, we get all these new materials on frontier intelligence, and then, on the open-source side, everybody has access to intelligence as cheaply as humanly possible to advance both sides of this barbell.
I'm just worried about the funding cycle in the near term, that slowing down, and what it could cause. Hopefully, we don't get there. But, Travis, maybe shifting to final thoughts here: You've been building Ambient for years. You're the guy who wrote the response to Situational Awareness called Situational Blindness way back when. It was great. Where do you think Ambient goes? What would make you really happy to see a year from now, if we were recording again? Where do you want to see Ambient—people using the software, the GPUs, the network? What would success look like here?
For us, since we're a supply aggregator, we want to have a huge book of supply. We want to bring all those mom-and-pop operators on. We want to bring the smaller neoclouds on. We want to give them real economic opportunities that allow them to serve inference to the world and make good margins on it. We want that flywheel to be spinning continuously and self-perpetuating.
We want people to tell their friends, “Hey, you don't need to rent your GPU. You could actually serve inference, and you can make a ton of money.” We want the word of mouth to be strong. We want people to be scrappy and actively finding ways to optimize delivery on our network. That is the supply side of it.
The second part is global distribution. We want to be a recognized brand. We've started a little beachhead here on OpenRouter, where we're number 1 in the delivery of 1 model. We'd like to be number 1 in the delivery of 10 models. We want to be a name on everyone's lips.
Then we want adoption of our own subscription offering. We want people to be using our API and also our X42 offering if they're doing agentic commerce, getting directly onto the platform to build the future of agentic commerce. I think those 2 things create the flywheel for Ambient. We've got a complete supply-and-demand loop that we want to build out.
It's premised on the supply side because we think we can deliver the supply better, and in a market that cares a lot about quality and speed, that's a key advantage. But then we want to turn this into the global intelligence utility that we've been promised by all the closed-source people, who are too conflicted to be able to deliver that.
Dario is never going to be able to deliver intelligence as utility. His hands are tied by defense contracts. But Ambient can deliver a high-intelligence network that's useful to everyday businesses and everyday people, that's credibly neutral, scalable, highly available, and reliable. We think life will be good.
I like that. And, Travis, how can people get involved? The demand side is pretty straightforward. They could access you via OpenRouter. They can download Ambient Desktop. They can access maybe an API or a subscription. But maybe on the supply side, if I've missed things, or on the demand side—I'm just curious how people get involved.
One thing that we didn't talk about at all is that Ambient really involves 2 models. One of those is Ambient, which is the big model, and the other one is Ambient Mini, which is a small model that can run on any GPU. Our intention with our network is to intelligently orchestrate these models in the future so people can get the most bang for their buck.
I'm happy to announce that in about the next week, we're going to open up, first, the closed beta for small-model mining, and then the open PvP season of Ambient mining for the small model. Anyone can come onto our network and try out small-model mining using the Ambient Desktop app, which is cross-platform—Mac, Linux, and Windows. You can mine it on your laptop if you want and send requests to other people on the network and get back responses.
We're going to offer that on our own platform. I'm going to try to offer it on OpenRouter as well as a free model. They're a little bit touchy about data, so I probably have to create a sub-brand of Ambient in order to do that successfully. We'll see how that goes, but people can become directly involved in mining Ambient pretty soon.
Touching on something that you mentioned earlier, we'd like to create the opportunity for people who want to fund big nodes to do so collectively, to mine the big model. We will take no profit on those nodes, so 100% of the profit will go back to the mining pool. We will do a 100% pass-through, just straight costs, for those nodes to incentivize big-model mining.
I think there are going to be some opportunities to get directly involved in the near-term technical future. We're launching this thing this year—it's coming out, and it's going to be live.
Travis, Ambient's making the summer fun again. It's super exciting. It's going to be really fun to watch it play out. And not just fun, but exciting, right? Watching everybody compete to serve these models in a competitive manner, watching the crypto crew get excited from the network perspective, and watching end users be able to access intelligence at the cheapest possible cost with the lowest latency. It's really exciting to see.
We're obviously investors in you, and we love that you've accepted us in. I really enjoyed this conversation. Travis, thank you so much for the time.
Thank you, Tommy. We appreciate all you do. If you're wanting a really active and helpful investor, I can happily say that Delphi is one of the best investors on our cap table. They've been tremendous for us, and you in particular, Tommy. Thank you so much. We really appreciate that. We're excited to back you.
So, thank you everyone for listening, and thank you, Travis.
Thank you.