OpenRouter 会以 $10BN 卖给 Stripe 吗?
- Alex Atallah 不会确认有关 OpenRouter 以100亿美元出售给 Stripe 的传闻:「我无法评论,但无论发生什么,我们都会执行这一愿景。」 他将 OpenRouter 定义为生态基础设施,确保「一家垄断企业不会接管」模型访问,同时提前回应5.5%的按量付费抽成批评:企业版采用承诺支出模式、不收取费用;自带密钥会彻底消除费用;自助式商业版也即将推出。
- AI 将是一个「巨型市场」、科技史上最大的市场,甚至可能是人类历史上最大的市场,但不会由单一模型赢下。 即便每家企业都训练专有模型——Fireworks 的 Lin 将其概括为「掌握自己的智能」——博弈论也会把企业重新拉回生态:竞争对手和模型实验室会持续发布用你没有的数据训练的模型,因此企业有动力不断尝试;这也是 Atallah 认为专业化内部模型「肯定」利好 OpenRouter,而非利空的原因。
- 平台上刚刚上演了一个近乎完美的杰文斯悖论:GPT 5.6 Luna 的价格在2周内下降10倍(OpenAI 先降5倍,再与 OpenRouter 协同降2倍),使用量增长13倍,随后恢复此前的增长速度。 Luna 的代币量已经超过 GLM,成为「极其漫长时间以来」首个进入 OpenRouter 排名前三至五位的 OpenAI 模型。他承认 Harry 提出的限制:「我们可能确实低估了前沿模型」,因为 OpenRouter 的数据偏向相信多模型路线的用户。
- 美国企业通常比对中国模型更担心前沿模型,因为「对数据政策究竟如何」存在更多困惑,而且无法在自己的机器上或通过自选供应商运行这些模型。 按照他的判断,Claude Design 是有意进行团队锁定——对 Anthropic 来说「并不是一大笔收入」,但能让「设计团队真正关心 Anthropic 的模型」——因此,薄层市场包装的应用最容易暴露在短期威胁之下。
- 美国在开放权重模型上「仍然非常、非常落后」:GLM 5.2 是「非常大的一步」,Kimi 则是 Moonshot 追上这一步,OpenRouter 7月上线了70个模型,约「每10小时一个模型」。 Harry 担心,围绕 DeepSeek 的国家冠军效应会在12个月内进一步拉大鸿沟,这一判断基本没有受到反驳;Atallah 的方案是蒸馏中国模型——「Sonnet 是 Opus 的部分蒸馏版本」——并把算力导向美国的新兴模型实验室,而 Harry 警告芯片优势不会长久:「如果他们4周能造出一座桥,我认为他们6个月就能造出一块芯片。」
- 推理服务商层可能抵抗商品化,因为市场受到供给约束,而 NVIDIA 希望其保持分散——「NVIDIA 的首要任务之一就是不能出现客户集中」。 一个代币并不等于另一个代币:Moonshot 自己的基准测试显示,为 Kimi K3 提供服务的各家供应商「数字差异很大」;OpenRouter 的路由器会在质量、价格和速度发生变化时重新分配流量,「每5分钟」调整一次。
- 尽管 OpenRouter 试图把切换成本降到接近于零,模型忠诚度依然真实存在,其驱动力包括「我的应用能运行,我不想把它弄坏」、既有模型持续降价,以及个人评测——「很好,反正我本来就喜欢 communique 2.6」。 Harness 并非靠捆绑留住用户,而是作为非模型实验室掌握用户关系的体验层继续存在;包括 Jeff Dean 新项目、Cognition 和 Cursor 在内的 agent 实验室,都有明确动力发布自己的模型。
- 中国开放模型正在快速推进,但 Atallah 也在追问其能力如何与中国的防火墙和审查制度相互作用。 OpenRouter 的应对方式是提供安全访问:下架被认为不安全的模型,并提供企业可以一键应用于全部推理流量的提示词注入防护和个人身份信息脱敏。
- 一个尚未充分讨论的组织变化是,AI 使用让员工成本从固定值变成动态值。 Atallah 建议将生产率与推理成本放进一个「值得庆祝」象限和一个「需要担忧」象限,同时让员工更多影响自己产生的成本。他最期待罕见病研究,以及通过众包改善城市和乡村生活质量,例如找出美国或英国所有的铅水管。
1. 创立时的意外:推理服务商击败了超大规模云厂商
- OpenSea 留给 Atallah 的经验是:NFT 在2020年10月爆发后,平台出现「服务器要融化了」,他的执念变成不能成为「套用到加密行业的 Twitter 故障鲸鱼」,也就是通过压力测试承受10倍于已见需求的流量。这套纪律后来直接迁移到 AI 领域,因为「所有公司,尤其是 Anthropic,都经历了不可预测的增长」。
- OpenRouter 创业时的核心判断,意外地错在了一个利好的方向:Atallah 原以为开放权重模型的托管层可能由一家超大规模云厂商垄断,结果却出现了独立推理服务商组成的生态。「你多久会听到有人把 GLM 跑在超大规模云厂商上?从来没有。」Fireworks、Together 及同业在托管模型和处理边缘情况上「快得多」。
- 早期的 OpenRouter 甚至不展示服务商,只显示「Provider One」和「Provider Fallback」,因为当时根本不确定这一层是否会成为一个市场。事实证明,正常运行时间「不会神奇地由市场供给侧解决」。
2. 他为何认为推理层可能抵抗商品化:NVIDIA 需要异质化
- 对于利润率压缩的看空逻辑,Atallah 的回答是结构性的:市场「受到极其严重的供给约束」,而且这种状态可能会持续一段时间;与此同时,NVIDIA 希望市场保持分散——「NVIDIA 的首要任务之一就是不能出现客户集中……他们希望市场保持异质化」。
- 他认同 Lin 对 Gavin Baker「一个代币就是一个代币」说法的修正。Moonshot 最近对所有服务 Kimi K3 的供应商进行基准测试,即便使用静态且广为人知的测试集,结果也「差异很大」。OpenRouter 看到的结果还在持续变化——「这些模型……非常情绪化,也非常不确定」。
- 路由器的作用,就是让这种差异可以被交易:当某个服务商能让一个代币跑得更远时,「它会立刻获得更多流量……对于大型模型,每5分钟都会出现重大变化」。当被追问最喜欢哪家服务商时,他保持中立,但更看好定制硬件玩家,以及可移植的 LoRAs/「cartridges」;后者有望让微调后更换基础模型的成本降到「几百美元,也可能只要几十美元」。
3. 「人类历史上最大的市场」:多模型未来不可避免
- Harry 提出挑战:如果每家公司都拥有一个用自身数据训练的专业模型,OpenRouter 还有什么意义?Atallah 回答「不,我不同意」——使命是「增加 AI 的神经多样性」。即便存在一个假设中完美的模型,也会有一个用不同数据训练的神经多样性竞争者出现,进而产生同时使用两者的需求。「创造力无法验证……同时使用两个模型,更有可能得到有创意的想法。」
- 博弈论的下一步是:其他人也在用新获得的数据训练模型,因此无论目标是利润还是增长,「你都有动力去使用生态创造的东西」。你的内部模型「永远不可能赢下整个市场」。
- 他的结论是:这将成为「巨型市场、科技史上最大的市场,甚至可能是人类历史上最大的市场」。因此,企业很可能会构建带有品牌属性的专业模型——「你的品牌是护城河的重要组成部分,而模型会成为承载品牌的一种方式」——同时把其他工作路由到整个生态中。
4. 路由器成为时尚:为参与而做,还是为胜利而做
- 对于 Ramp、Merge 等公司纷纷推出路由产品,Atallah 的评价是:「很多公司做路由器,只是因为这很时髦……你是在为了参与而做,还是为了存在而做,而不是为了赢。」他的定位则是:「我100%专注于打造最好的路由器、网关和 LLM 市场……这对我们不是一条支线任务。」
- 第二个不那么显眼的反对理由是:一个没有开放完整市场的网关,会「降低所有用户的杠杆」。能访问每一个模型,就能接触每一项创新;把这些切断,「就是让公司里的所有员工失去他们所需要的东西」。
- Harry 反驳说,企业在规模较小时喜欢5.5%的抽成,但一旦这成为真正的成本项目,就会自己搭建系统。Atallah 部分承认责任在己方:「有些确实是我们的错」,因为定价不够透明;企业版采用承诺支出、不收取费用,自带推理服务也会取消这笔费用,自助式商业版即将推出。
- 未来3年的收入逻辑取决于市场是否继续「每年增长10至15倍,甚至可能更多」。只要增长持续,收入仍将主要来自 OpenRouter 最擅长的业务:为「持续低估自身推理需求」的客户提供计划外推理容量、故障转移和正常运行时间。如果增长放缓,他认为大型 SMB SaaS 将成为 OpenRouter 的重要增长来源。
5. 一个近乎完美的杰文斯案例,以及数据上的坦诚限制
- 对话中最干净的案例是:GPT 5.6 Luna 的价格先由 OpenAI 降低5倍,再与 OpenRouter 协同降低2倍,2周内合计下降10倍;使用量增长13倍,短暂持平后恢复此前的增长速度。这是一个「近乎完美的杰文斯悖论故事」:当时 DeepSeek 以有竞争力的价格上线,GLM 也很便宜,而 Luna 现在已经超过 GLM,成为「极其漫长时间以来」首个按代币量进入 OpenRouter 前三至五位的 OpenAI 模型。
- Harry 质疑这一结果是否具有代表性:OpenRouter 可能只占代币总量的1.5%至2%,而且怀疑者会认为其排名被中国模型主导。Atallah 没有回避:「我们确实偏向相信我们这套论点的人……我们可能确实低估了前沿模型」,但他认为随着多模型论点扩散,数据会越来越具代表性。
6. 企业更害怕前沿模型,而不是中国模型
- 针对 Karp 所说企业对前沿模型实验室「感到恐惧」,Atallah 确实观察到企业对 Claude Design 和 Figma 的紧张情绪。如果模型实验室不在意某个市场,应用包装创业公司当然可以生存;但模型实验室的动机在于锁定团队:Claude Design「可能不是一笔重要收入」,但具有战略意义,因为它能让目标客户「再多一个真正想坚持使用 Anthropic 的团队」。
- 对于 Figma 被 Claude Design 挤压的具体问题,他并不信服:设计师确实试用了 Claude Design,但「到目前为止,我还没听到持续使用的后续案例」,而且 Figma 的业绩「相当不错」。Harry 说:「这就是为什么你不想成为上市公司……业绩很好,Figma 却下跌。可怜的 Dylan。」
- 他确认了双方的明显不对称:企业通常比面对中国模型更害怕前沿模型——「对数据政策、提示词究竟发生了什么,存在更多困惑」——因为企业无法在自己的机器上运行这些模型,也无法自由选择服务商。Harry 感叹:「身处这样一个世界,真是太奇怪了。」
- 对于把流量路由给 Moonshot 或 Alibaba 的 Qwen 是否意味着承担责任,他说:「我不能假装知道它们内部发生了什么。」OpenRouter 的答案是安全访问:下架被认为不安全的模型,并提供一键式提示词注入防护和个人身份信息脱敏,因为「你不能仅仅因为互联网上有一些坏东西,就在公司里封禁整个互联网」。
7. 美国在开放权重模型上「非常、非常落后」,差距可能继续扩大
- 模型数量的扩张令人震惊:「7月,我们上线了70个模型,约每10小时一个模型。」下一波来自 agent 实验室:Jeff Dean 目前正在 Google 内部启动一个 agent lab,Cognition 和 Cursor 已经有模型,Lovable 尚未公开发布;这些实验室都有「非常明确的动力」,通过自己的 agent 分发模型。
- 按他的判断,质量排序是这样的:「GLM 5.2 是开放权重模型迈出的一大步,Kimi 则有点像 Moonshot 追上了这一步。」Kimi 还不具备前沿模型级别的网络安全能力,在长周期任务上也落后,但它「是一位很好的作家」;与此同时,前沿模型随着编码能力增强出现「声音退化」——「咚、咚、咚,这才是问题所在」。
- Harry 判断,未来12个月鸿沟会进一步扩大:DeepSeek 成为中国的「国家冠军」,监管和政策让位于发展,资金也会被提供给它;而一家正在筹集数十亿美元的美国开源实验室,其商业模式「存在疑问」。Atallah 则追问另一边的问题:「DeepSeek 能走到防火墙之外多远?如果防火墙对中国很重要……总会有些东西发生变化。」Harry 转述朋友 Jason Lankan 的报告:DeepSeek 甚至无法弄清国内 Starbucks 什么时候开门——「比我们强得不可思议。国内则烂透了」。
- 如果由他担任美国开放权重生态的假想主管,他会先蒸馏中国模型——这是行业惯例,「Sonnet 是 Opus 的部分蒸馏版本」;同时利用 RL rollout 检查输出是否符合对齐要求,并通过 NVIDIA、Google 的 TPU、Amazon 的 Trainium 以及他所谓的 neo-chips,解决新兴模型实验室的算力获取问题。Harry 提醒,算力优势是会过期的:「如果他们4周能造出一座桥,我认为他们6个月就能造出一块芯片。」
8. 真正留住用户的是可运行的应用、个人评测和 harness,而不只是记忆
- 尽管 OpenRouter 试图把切换成本降到「接近于零」,其流失数据仍显示用户存在真实忠诚度,根源有3个:「我的应用能运行,我不想把它弄坏」;既有模型会随时间降价,而新模型会把价格曲线重新抬高;以及个人评测——如果随机测试失败,用户会说:「很好,反正我本来就喜欢 communique 2.6,继续吧。」
- 记忆确实是一种留存机制,但竞争焦点在于哪一层拥有记忆:模型、应用、推理服务商还是路由器。「不可能由某一层捕获所有有价值的记忆,因为应用掌握了大量模型实验室没有的重要上下文。」
- Harry 质疑:「harness 和应用有什么区别?感觉这只是文字游戏。」Atallah 给出了实质性回答:harness 具备可组合性并基于 Unix,相比通过登录和虚拟浏览器编排一个应用,拥有「极其、极其、极其少的未知未知」。随着模型变强,harness 仍会生存,因为很多 harness 正在删除最终「变成负担」的系统提示词垃圾;对非模型实验室而言,harness 是掌握用户关系的方式。
- Harry 自己的行为也说明了工具层的价值:他通过 Arena 提示模型,用过 Pergamom 和 Kimi,也觉得 Meta 的 Muse 很出色——「我没有忠诚度,给我看结果」。Atallah 认同由此延伸出的「编排器+子 agent」架构:由一个编排器模型把确定性任务分发给低成本开放权重子 agent,是「所有人都需要探索的优秀架构」。对于 Meta,他的态度建设性但尖锐:「人们现在还不太知道该拿 Muse Spark 做什么。」
9. Stripe 的不置可否,以及把员工 AI 成本视为动态变量
- 对于传闻中的 Stripe 以100亿美元收购 OpenRouter,Atallah 回应:「我无法评论,但无论发生什么,我们都会执行这一愿景。」他将 OpenRouter 描述为生态关键基础设施,目标是提供「安全的 AI 访问,让一家垄断企业不会接管」。对于个人可能获得的巨额财富,他避而不谈,表示资本应投向「不太适合风险投资」解决的问题,包括一个尚未公开的 AI 赋能研究资助项目。
- 快问快答中,他称赞 Poolside 的模型,并把一家新兴美国实验室视为自己认为被低估的标的:这家实验室正在构建有意思的编码模型,体量小但效率很高,还有实用的访问工具。未来3年新兴模型实验室的死亡率不是70%——「如果把整合计算在内,我会说是50」。对于 Dario 的末日论调,他说:「我个人欣赏 Anthropic 的偏执……如果没有人极度偏执,就没有人提供那种声音。」
- 他从使用数据中看到的一个尚未得到充分讨论的现象是:员工成本现在变成了动态变量。他建议管理者把生产率与推理支出画在一张图上,划出「值得庆祝」象限和「需要担忧」象限,后者的特征是「他们的 AI 使用量高得离谱」;同时将部分路由选择下放给员工,因为「他们可以控制自己产生多少成本」。
- 他最后表示,最令他兴奋的是罕见病研究,以及通过众包提升城市或乡村生活质量,例如找出美国或英国的每一条铅水管。AI 可以让拥有好想法的人获得更大杠杆,去解决那些此前只能被放弃的问题。
It's going to be the biggest market in tech ever. A lot of companies are making routers because it's fashionable. The model apps have several incentives to go after you eventually. In July, we launched 70 models—about 1 model every 10 hours. America is still very, very behind, but GLM-5.2 was a really big step for open-weight models.
There are reports that you are selling to Stripe for $10 billion. Is that going to happen?
This is 20VC with me, Harry Stebbings. Now, the only thing that I really care about anymore is providing the best, most relevant interviews at the right time to you. So today, we have Alex Atallah, co-founder and CEO of OpenRouter, the gateway to the world of LLMs. They reportedly have had offers from Stripe for $10 billion. They've raised at a valuation of over a billion and a half. They are the market leader, and this interview could not come at a more prescient time. It'll be very interesting to see whether the company chooses to stay private or sell to Stripe. We shall see. But this interview was recorded before, so we will check back in a couple of weeks. This was an incredible show, and it was, think, awesome to have Alex in the studio.
Alex, I am so excited for this, dude. I have wanted to make this one happen for a while. I've heard so many things from Matt at Menlo. I've stalked the shit out of you speaking to Anjani, even your roommate, before this show. So thank you for joining me, dude.
Thank you. It's great to be here.
Now, I want to start a little bit pre-OpenRouter and start with OpenSea. It was a pretty incredible journey. What did you take with you to OpenRouter, having seen all that you saw with OpenSea?
Yeah. OpenSea started as the first NFT marketplace. Similar to OpenRouter, it was very small for a long time. We kept the team very small until roughly the Series A, or a little bit afterward.
This was before AI. Right after NFTs started blowing up in October 2020, we were like, “Oh my goodness, we are understaffed. The servers are melting.” Our search index was exploding, and we had a couple of big outages. It was tough to keep the site up. It was like, “Oh my God, we're going to become the Twitter Fail Whale,” but applied to crypto.
My biggest goal was to make sure we weren't the Twitter Fail Whale for crypto. It took a little while to create the team and get the platform and infrastructure under control. We had to make sure we could predictably scale—in other words, do load testing to help the site sustain 10 times the load even when we weren't seeing that load.
With crypto, you just don't know. There were these moments when we would get incredible traffic spikes, and it was very dependent on the content and the community. So I built a lot of infrastructure and took those scaling responsibilities to OpenRouter.
I spent a lot of time thinking, “Okay, how do we make something that's going to basically be always up and that people can really count on from an infrastructure point of view, even when there are huge surges in tumultuous markets?” That has been very helpful for AI, of course, because all companies—especially Anthropic—have seen unpredictable growth, and we have as well.
We've had a couple of bumps, but overall it's been significantly better. OpenSea just drilled that into me in a way that I could take productively to OpenRouter.
Can I ask you, when you go back to the founding thesis of the company, what has happened in the ecosystem and in the model landscape that you did not expect to happen?
One thing we did not expect was that an ecosystem of companies would emerge to host and serve the open-weight models. Early on, it wasn't clear that the market wasn't going to be a monopoly where just the 3 hyperscalers served all the open-weight models and startups were really far behind.
In reality, how often do you hear people running GLM on a hyperscaler? Never. They're using inference providers like Fireworks and Together, and there are big lists of providers doing the best job of hosting all the open-weight models.
In the early days, we had what I think we called Provider One and Provider Fallback. We didn't show which providers were actually doing the hosting because we weren't really a marketplace. We were an exploration tool for finding and discovering new LLMs, and we wanted to build a marketplace of model labs. But we weren't sure whether the inference-provider layer would actually be a marketplace.
It turned out that those companies were doing a much better job than the hyperscalers. They were much faster to host the models and figure out the edge cases involved in hosting them, and uptime was going to be a constant problem. It wasn't magically going to get solved by the supply side of the market.
A lot of people suggest that the inference-provider layer is a commoditizable element that will be removed, or that its margins will be competed away over time. What would you say to that theory?
Right now, we're in a massively supply-constrained market, and it's likely going to be supply-constrained for a while. All the inference providers are pretty much constantly short.
You might ask, “Okay, GPUs are really, really beneficial. Why doesn't Google, Amazon, or Azure run around, buy up all the GPUs, and take all these inference providers out of business?” The people making the GPUs don't want that. One of NVIDIA's top priorities is not having customer concentration.
They want lots of customers to have separate allocations of GPUs. They want the heterogeneity of the market and competition on the compute layer, and this is good for the ecosystem. Users also want this. It's good for NVIDIA and for end users because it allows these inference providers to come up with new innovations in how to serve the models better.
Even with a single model like Kimi K3, Moonshot posted a benchmark showing all the inference providers and how well they were serving Kimi K3. The numbers were pretty different for benchmarks that are really static and well known.
We post this continuously. We're always benchmarking all of the models on all of the inference providers—all the open-weight providers—and finding really different results constantly. The results change over time. These models are very emotional. They're very nondeterministic.
I had Lin on the show from Fireworks, and I said that Gavin Baker had told me, “A token is a token is a token.” She corrected me and said that a token is not a token, because one provider can make a token go so much further than another provider.
It's like how you get to the store: you can drive around the whole block, or you can drive straight to the store. Tokens can be made more efficient and go further, and that's the job of the provider.
Yeah, I agree with that. In some ways, we're providing a service to help people discover providers. Ultimately, when one provider is making a token go further, we spend an enormous amount of time on our central router technology so that provider immediately gets more traffic.
As soon as we detect that there's a quality improvement, a speed-up, or a price reduction, that provider immediately starts getting more traffic. This happens 24/7.
Every 5 minutes, there are big changes for the big models, and so it actually makes the experience better.
You can only invest in 1 inference provider. Which one do you invest in?
I probably have to stay neutral on this. I really like the inference providers that are doing custom hardware and very low-level optimizations. I like providers that are also trying to figure out how to make customization easier. Today, you fine-tune models and create this new, fully independent model from the base model. Many inference providers are creating these LoRAs—some call them cartridges—that are much more portable, potentially between models. We might see a future where, when you do a fine-tune and want to change the base model layer, it only costs maybe a few hundred dollars, maybe a few dozen dollars, to change it.
It's okay. I understood that Fireworks is your favorite. It's okay, I get it. Mine too. My question is, when Lin was on the show, she was like, “Oh, you don't want to rent your intelligence; you want to own it,” and we're going to see companies have specialized models trained on their own data and proprietary to them. In a world where every company has specialized models that are really fine-tuned to them and their preferences, is that good for an OpenRouter business or not?
Oh, definitely.
Why? Because you'd stick with 1 model—
Mm-hmm.
—that's yours, proprietary, trained on your data, and not be open to the diaspora of models that is available.
No, I disagree. I think our mission from the very beginning has been to increase neurodiversity in AI for the whole ecosystem, and we really believe that a multi-model future is inevitable. Let's say, hypothetically, there's 1 model that fulfills all of your desires, either within your company or as a consumer. More and more people start using that model, and then someone decides, “You know what? I'm going to create a neurodivergent model. I'm going to create a model that's a little bit different, that talks a little differently, that has ideas the first model could never have come up with because it's trained on completely different data.”
Then it creates inevitable demand to use both models. Creativity is not verifiable. You can't really put an easy number on creative ideas, and when you use 2 models together, you're more likely to get creative ideas than if you just use 1. It's just a fact: if that other model was trained in a different way on a different data set or has made a big update, consolidation on 1 model just doesn't make any sense to me.
I totally get you. So you'll have companies that have a core workflow, or their core, which is their own specialized model, and then they'll use a plethora of other models. They'll use OpenRouter for that other model selection.
Yes, and I think that when companies make their own model trained on their own data, the ecosystem around them is all doing the same thing. You have to play out the game theory for these things a little bit. If everybody is doing this as well, and all the model labs are constantly creating new models using new data that they've acquired or bought from other companies—potentially data that's valuable to you—what is in your best interest? It's to go and try out those other models and see if you can be more productive with them, if you can merge them together to get better state-of-the-art performance, or if you can reduce your costs using these other models.
Whether your goal is to improve your margins or grow your company, you are incentivized to use what the ecosystem creates. The model that you made, you're going to have to continuously improve it to keep up, and it's never going to win the whole market. So this is going to be a massive market, the biggest market in tech ever, and probably the biggest market in human history. No one's going to win all of it.
You're not going to build a model that wins all of it, so you might as well build a model that's known to specialize in something very useful and important to your company and your business, and be known for that specialty. I think a lot of enterprises are going to move in that direction: make their own models, make their own branded intelligence. Your brand is a big part of your moat, and that model will be a way your brand carries around.
You mentioned the immense time that you spend on the routing technology that you have. A lot of people are thinking that we're seeing the commoditization of the routing technology. You're seeing Ramp release products like this. I mentioned earlier that Merge, a company we invest in, has released its product. Several companies are releasing routing technology that is similar or claiming to be similar. Are we seeing the commoditization of this layer?
A lot of companies are making routers because it's fashionable. They're seeing growth happen here, or they're making gateways, at least. There are 2 issues with that. First, it immediately puts you in the mindset of copying instead of winning something. You're playing to play, or you're playing to exist, rather than playing to win. Maybe you're just trying to serve your existing customer base and you want to see some AI growth happen.
I think it immediately puts that gateway many, many months behind the companies that are fully focused on it. I am 100% focused on building the best router and gateway and LLM marketplace, and it shows in our product, in the benchmarks that we create internally, and in how we see ourselves compared to the competition. This is not a side quest for us, as it may be for some other companies.
The other problem is that it reduces the leverage of all of your users. I really deeply believe in giving users and developers more leverage. Fundamentally, giving them access to more models is about giving them more leverage over all the innovations that happen in AI. You want to be able to access them all. You want to reduce your dependency on any individual one.
If you build on top of a router or a gateway that doesn't give you access to the full market, full flexibility, or full customizability, it doesn't give you the full leverage of the whole ecosystem. Then you're being cut out. You're cutting out all your employees at your company from things they need. OpenRouter is fundamentally about giving people more choice because that gives them more leverage.
You do that at a price, at a 5.5% take.
That was sort of our pay-as-you-go plan. We then added an enterprise plan with a totally different pricing model, and it's been very successful so far. It's based on committed spend, and there are no fees on that committed spend.
Because that was going to be my question. Ultimately, companies will love it while they're small, and then as you scale, you're like, “Shit, this is really freaking expensive. I'll just build my own routing tech now because it's become such a significant part of my cost base.”
I figured some of those companies just haven't realized that we have an enterprise plan, and some of that is our fault for not having a better, more detailed pricing model. We're soon going to introduce a business self-serve plan that also just makes it make a lot more sense.
If you bring your own inference to OpenRouter—if you bring your own keys—that fee goes away. For inference that we are providing you, when you go into OpenRouter's capacity and you're not on our enterprise plan, that's when that fee comes in. We need to be able to predict demand a little bit, so that's why we do the committed spend.
What would be the main revenue line of OpenRouter in 3 years' time?
I think it's going to depend on the economy in so many ways. If the overall AI market keeps growing the way it's been growing over the next 4 years, at 10–15× every year, or potentially more, that's a lot of growth. Under that scenario, I would expect people to continue to underestimate how much inference they're going to need. Our revenue is going to be dominated by the same things that dominate it today, which is helping people with unplanned inference capacity, both enterprises and startups.
That's what OpenRouter is best at. When you need to try models that you weren't expecting to need to try, or when you're using more inference than you thought you were going to use on particular models, we make sure that isn't going to be an issue for your company by providing the best failover and best uptime. This is really, really a good thing to do when the market is continuously underestimating its inference needs and growing at this rate.
If this growth rate continues over the next 4 years, it's going to be a wild amount of growth, and the economy has some limits to it. I can see major SMB SaaS growing for us if growth does not keep going 10–15× per year.
We've seen token prices fall 90%, give or take, in 18 months. Is the reduction of token prices helpful or hurtful to your business? Because obviously, you have a take on spend. If they come down and spend is more efficient, seemingly it's bad for your business. You have a shrinking pie to take from.
Well, a lot of people talk about the Jevons paradox: when prices go down by 10x, usage increases by more than 10x. But no one has really done a great job modeling it. We do have a lot of spot stories that confirm it.
For example, GPT-5.6 Luna on OpenRouter. OpenAI cut prices by 5x and then, in coordination with us, by another 2x. So, in total, the price of Luna has dropped 10x on OpenRouter over the last 2 weeks. Guess how much usage has grown?
13x. So it's a close-to-perfect Jevons paradox story, where you drop prices 10x and usage grows by more than 10x—just a bit more. The usage is pretty stable. It grew, flattened out at 13x, and then it's been growing at the same rate it was growing before it hit the 13x multiple.
So that's pretty interesting, and it's a pretty low-variance story. There are few other confounding variables in it, and it also happened in the middle of DeepSeek launching and having a really good price, and GLM having a really good price. Now Luna is being used more than GLM on OpenRouter.
GLM used to be one of the top 3 or 4 models by token volume, and now Luna is past it. This is the first time OpenAI has had a model on our platform in the top 3 to 5 models by token volume in an extremely long time. So it was a really big and interesting move.
How reflective of the market are your token volumes? It's about—I may get this wrong—maybe 1.5% to 2% of token volumes. How reflective are they? Because a lot of people, when I say, "Oh, the top 5 models when I look at OpenRouter are all Chinese. What does that mean?" they'll go, "Oh, well, Harry, no offense to OpenRouter, but it's not reflective of the market. Most people who use frontier models don't go through that. They use frontier APIs, and so it's not counted." To what extent are your rankings reflective of true token usage?
Yeah. That's a really good question. We try to estimate how they're off by surveying people sometimes or looking at surveys other people have done. We definitely have a bias toward people who believe our thesis, which is that the future is multi-model, and companies who want multiple models.
There are still companies out there—I very rarely run into them now—that are just like, "Oh, yeah, we're an OpenAI shop. We only do OpenAI models." So we're not going to see any of those companies, and I think those companies are primarily focused on the hyperscalers, OpenAI, Anthropic, and Gemini. We probably undercount the frontier models.
But I think, over time, our thesis is becoming more and more common in other companies. The moment that they're like, "Oh, yeah, we need to use other models," our data becomes more representative. As we scale up, the data becomes more representative in general. So my hope is that it just becomes better and better data over time.
Alex Karp said on CNBC, in his rather wonderfully energetic way, that companies are terrified of working with frontier model providers.
Mm-hmm.
Do you think they are?
I haven't seen what he talked about there when I talk to our customers. But there's definitely a little skittishness, particularly when Claude Design came out around Figma. That part I did see, and I do think that there are real concerns.
Figma's very different, but if a startup is only building a go-to-market wrapper around intelligence—"Hey, we're a company that brings AI to this market and does so by doing the right integrations and customizing the system prompt"—you're going to be fine if the model labs don't care about that market, which there will be many markets like that.
But the model labs have several incentives to go after you eventually. One is getting multiple teams within companies they do care about to be dependent on them. This is my theory behind why Claude Design was strategic. While it's not a massive amount of revenue for Anthropic, probably not a significant amount of revenue, it does get the design team to really care about Anthropic models.
The companies they want now have another team that really wants to stick to Anthropic. That team strategy can make you compete with the model labs. I think companies that find themselves saying, "Oh, we're building a product for a team," where that team has now become strategic for the model labs—for companies they actually care about—are where I see the most near-term threat.
Do you think Claude Design will have a meaningful impact on the Figma business? I speak to many founders today who are, bluntly, switching from Figma to Claude Design, and it's cannibalizing their Figma usage. Do you think that will happen?
I saw a lot of designers try out Claude Design, including our own, but so far I haven't heard the repeat story. I don't know. Honestly, I haven't talked to very many designers about this. I certainly haven't heard a lot of chatter about Claude Design, and if you just look at the numbers for Figma, they're quite good. They had incredible earnings.
Oh, right.
Yeah, that's crazy.
This is why you don't want to be public, dude. You see great numbers—Figma down. I'm like, "Poor Dylan." What?
Yeah.
Do you know what I mean?
Yeah.
It's like, really? Come on. We were talking about the different models that we have on offer—
Mm.
—and whether companies are willing to work with frontier models. The rate of model development feels immense. Do you think we will see the same rate of model development continue over the next 1, 2, 3 years?
Frontier model development or just general model development?
General models—
General models.
—both frontier—
Yeah.
—and open.
Yeah.
Just because, I mean, every single day there are 2, 3, 4 new models.
In July, we launched 70 models. It's about 1 model every 10 hours. There are some agent labs starting, too, that are all going to probably make models eventually. Jeff Dean is starting an agent lab right now from Google.
The companies that are known for making agents have a very clear incentive to create their own models and distribute them through the agent, and we haven't even seen the start of that. We've seen the start of it, but we haven't seen it really pick up. Cognition has a model. Cursor has a model. Does Lovable have a model yet? I don't think so.
Not publicly.
Yeah. So the agent labs are going to develop models, I think. There's pressure from both the GPU makers, like NVIDIA, to create more competition and diversity in the space, plus pressure from us and from investors who just want to try new things that could improve intelligence in some neurodivergent way. Those are strong incentives.
I think they're enough to incentivize more founders to make Neo labs. If American open-weight models pick up steam, then it gives these Neo labs a base to train on that's not Chinese, which will then probably create more American Neo labs.
Do you think we should be concerned by the rate and quality of Chinese open models?
We should. We're behind. America is very, very behind still. I think things are picking up. We have Poolside, Thinking Machines, and Arcee.
Do you feel a sense of responsibility for that? What I mean by that is, you are a routing business, and you could route a company to a Chinese model that, who knows, people are worried about backdoors and CCP involvement. You could be the deliverer of that to those models. Do you feel a sense of responsibility for that?
We do feel a responsibility to have safe access for all of these models. Customer trust is our paramount goal. If one of these models is unsafe to use—generally considered unsafe—we pull it from the platform.
If there's a way to use it in an unsafe way—I mean, there's a way to use all the models in an unsafe way—then we believe in using technology to make it safe and working with the model labs themselves to figure out how they're doing it on their side, so that we can be state-of-the-art or better.
We spend an enormous amount of time making sure that our practices match the best practices we're seeing coming out of the labs, or better. Because we're a way of exploring all the models and finding them for the first time, we're a good focal point for deploying safety measures across your whole company.
For example, we have prompt-injection protection. You can just turn it on and immediately flag prompts that look like prompt injection.
We have PII redaction. We have a couple of different things that you can automatically turn on with a click and get an added safety layer on top of all of your inference. We build that so enterprises feel like they can safely deploy new models and their employees can try them out.
I think of the models a little bit like the internet. You can't just ban the internet at your company because there are some bad things on the internet. You can create guardrails, and you should. You need to use AI to build the best possible guardrails that you can.
I'm with you, but do you think you actually know what's going on within Moonshot or Alibaba with Qwen? These are incredibly secretive organizations in the depths of China.
I can't pretend I know what's going on inside them. As a U.S. company, we're going to follow the best practices of what happens in the U.S. to make sure that we're not doing something irresponsible.
What do you think U.S. companies are more nervous about: frontier models or Chinese models?
I think they're more nervous about frontier models, usually in part because there's much more confusion around the data policy—about what's actually happening to the prompts that I'm sending, where they're being stored, and how they're being looked at. You can't run them on your own machine or with a provider of your choice.
That immediately creates all of this uncertainty in a lot of enterprises, and it's uncertainty that they can also pattern-match. It's very similar to running on their own infrastructure versus running in their VPC and knowing who can see the data.
How extraordinary is that, though? They're more nervous about U.S. companies headquartered in Silicon Valley, where you can see, touch, and feel the headquarters and the leaders. It's such a strange world to be in.
Yeah, it is very strange, especially with frontier models having the biggest cyber posture right now.
What do you make of every company posturing, “Ha-ha, we hacked someone”? First you had OpenAI, then you had Anthropic, and then you had Zuck coming out saying, “I didn't want to miss the party. We did, too.”
Yeah. Well, I think they have to talk about it. The right thing to do is to reveal when there's been a cyber incident involving your model. Covering it up doesn't work. It's not going to work in the long term, and it certainly looks like they're all bragging about it.
But really, if you were in their position and something happened with one of the models, and you had to make the choice about whether to publish it or not, I think the right thing to do is to publish it regardless of how people are going to spin it. So I doubt that they're actually thinking of the felony bench [?] or whatever it's called.
How significant was the latest Kimi model, which got so much attention? Was it as significant as everyone thought?
It's quite good. It's not cyber-capable in the same way the frontier models are, and on long-horizon tasks, I think it's still a bit behind the frontier models. But GLM-5.2 was a really big step for open-weight models. Kimi was kind of Moonshot getting up to that step. That's a little bit how I see it.
Kimi is also a very good writer. The voice and tone are both pretty good, whereas some of the frontier models have voice degradation that happens when they get better at coding, especially. Then it's like, “Oh my God, I can't read this output anymore.”
The output sounds like, “Three of the four arguments you made are right, and one is a turning point. Dun, dun, dun, dun, dun, here's the rub.” Sometimes it's just impossible to read what they're saying, and this stuff is fixable. But Kimi, I think, has always had pretty interesting writing.
In 12 months, will the chasm between U.S. open source and Chinese open source be bigger or smaller than it is today? My fear is that it'll be—
Bigger.
—bigger because when you have DeepSeek, it becomes the national champion in China. Xi Jinping is going, “This is our AI horse. I will concentrate all of my money and efforts behind this, and I will supplement this ecosystem to the end. This is the winner.”
When you see another Moonshot come out, suddenly all regulation gets moved aside, all policy gets pushed aside, and all funding becomes available. Everything is allowed. You are free to run. These guys are unfettered in their ability to do whatever they want to get to their end goal.
Whereas OpenAI, Anthropic, and all the other providers in the U.S., especially open source—fuck, you go and try raising billions of dollars for a U.S. open-source model. It's tough, actually. Not impossible at all, but tougher. The business model is questionable, AI research is super expensive, and you're competing against OpenAI and Anthropic.
I think the comparative landscapes they sit in mean that the Chinese open-source providers are just inherently advantaged, sadly.
They have very good researchers, and I think Americans underestimate that a lot. I do think they're going to be concerned about the cyber posture of their models, and they do seem very concerned about censoring the models and censoring the information that the models can provide to people.
While today people complain about American models censoring more due to cyber, I'm not sure that's always going to hold. As the Chinese models grow in importance for China, what are they going to do? Are they going to drop the Great Firewall? Are they going to give up on putting the firewall around the models?
I don't know that much about China, but it does seem strange that they don't seem to care more about what the models are capable of. I've never seen anyone do a profile of what you can do with DeepSeek that you can't do with the internet in China that's available to you within the border—what information you can access.
I've never seen anyone do a real deep dive. How far past the firewall does DeepSeek go? If the firewall matters to China, if it's going to matter in 10 years, something's going to change.
What's interesting is that the abilities of the Chinese models outside of China are immense.
Yeah.
The abilities of the Chinese models inside China are actually relatively limited.
Oh, the guardrails they put on.
The guardrails are—
Yeah, yeah.
—incredibly stringent and prohibitive. It's ironic that they are incredibly superior to us. Shit, domestically, they're terrible.
Interesting.
I literally just had—
Yeah.
—my dear friend Jason Lankan, who runs SAS, come back and say, “Couldn't figure out what time Starbucks opened on DeepSeek.” It wasn't on offer. It would say, “Not allowed.” Wild.
Wow.
Very basic, rudimentary requests. We're speaking about all of these different models.
Yeah, yeah.
The thing I think about is loyalty. You have this incredible seat in the ecosystem where you can see everything. Do we see any developer loyalty today with models?
Honestly, we do see some. We try to make switching costs close to zero so that when new models come out, people can try them out really easily. But we also measure retention and churn from all the models.
We share this data with model labs when they ask for it, so they can know, “For my model that just came out, which models drove traffic to it? For those users, when they leave, which models are they leaving to?” We'll make this more and more available to the world soon.
We do notice in the churn data that there are developers who continuously stick to models even when there are better models out there for their use cases. I think it's a combination of a couple of root factors.
One is: my app works, and I don't want to break it. If the support bot starts saying something weird that I didn't expect, why add more headache? I've already done all this optimization, and I've already put all these guardrails around it.
Another is that new models aren't necessarily going to make your pricing better. In fact, in general, what happens is that the price of current models goes down over time, especially when new advancements in the labs happen. You'll see intelligence jump, but the price curve also jumps, and then starts going down over time.
So it's not necessarily the most price-effective thing to do to shift over to the newest model, even for open weights. The third reason is just fundamental trust in the outputs. If I'm using a model to do my work and I like the way it talks, I probably have some eval—a personal eval.
A lot of people have these personal evals that are just random tests that they give the models. If the random test doesn't look really good on the new model, they'll just be like, “Good, I liked communique 2.6 anyway. Keep going.”
People thought before that memory would be the retentive mechanism. OpenAI has all of my previous queries and prompts. It knows that I live in London, I do podcasting, and so on.
Yeah, yeah.
That will make it a better model for me moving forward. Is memory no longer a retentive mechanism?
Memory is really interesting.
I've always thought it was a retentive mechanism, and the question is where it lives. Is it gonna live with the model? Is it gonna live with the inference provider? Is it gonna live with the app? Is it gonna live with the infrastructure provider—the router?
My guess is that all of those layers are going to try to own memory in different ways. There are gonna be advantages to sticking your memory in each layer. If you stick it with the app, then the memory has the most app-related context and is model-agnostic. If you stick it with the model, the memory might perform the best on personalized benchmarks and perhaps have the best ultimate intelligence. I think the model labs are gonna work on memory.
The ultimate thing might be: is there a good combination? Can I use memory in the model and memory at the infrastructure layer or the app layer at the same time? Is that gonna confuse the model? We don't know yet.
I do think that it's impossible for one layer to capture all valuable memory because the apps own so much important context that the model labs don't have. The model labs, in order to get this to work, will have to incentivize the apps to give them that context instead of—
Speaking of apps and the models there—Claw Code, Cursor, Bundle, Model, and Harness—is the router absorbed into the agent framework before it ever has the chance to be independent when you have the agent and the harness together?
The harnesses are pretty interesting because, in our early days, one of our early bets was that most apps were underestimating the desire for users to choose the model. Most apps in the very early days, in 2023 and 2024, didn't even make it clear which model was being used under the hood. They were like, "People are not gonna care about that. They just want AI." One of our strong convictions then was that people were going to want to use particular models. They were going to care about who they were talking to.
It's like, I wanna know which employees I'm talking to when I'm trying to solve a problem, and models will be kind of like that. That has played out. In Notion, you can choose the model that you talk to, even though you'd think an app like that might want to obscure it completely.
A similar thing happened with harnesses, where, particularly with developers, they started to build an affinity to different harnesses. That's because it's a user experience. I think that is my favorite argument for why harnesses are gonna stick around.
It's not that they're being bundled with the models, because, in fact, as models get better, they get more resourceful, and the junk that gets thrown into the system prompt just becomes a handicap. Anthropic published a good article about this, where they showed that they got rid of stuff from the system prompt, and suddenly fewer contradictions showed up later on with user prompts and the model performed better. We're seeing a lot of the harnesses right now deleting code in order to perform better with the latest frontier models.
I don't think that means that harnesses are bad. In fact, I think we'll see more harnesses come up in the future because it's a way of building a user experience on top of models. It's a way for developers who are not model labs to own a user relationship, and that is just going to be incredibly valuable for the economy to have that layer.
I'm gonna get killed for this. What's the difference between a harness and an app? It feels like this word wank, with everyone talking about harnesses and the harness, and I'm like, "Is that not an app?"
The nice thing about harnesses compared to apps is that they're more composable. I can have a harness call another harness. I can have a harness spin up another harness in a sandbox in the cloud.
Is that not what APIs did for apps?
Yes, but it's much more reliable and deterministic and easy for users to grok with a harness because the harnesses are Unix-based, and the models are so well-trained on Unix and Bash commands.
If I'm telling a harness to go orchestrate an app in the cloud, it's gonna be like, "Oh, boy, does this app—how do you log into this app? Do I need your password? Do I need to fire up a virtual browser? It's gonna be pretty slow. I'll figure it out. Okay, I fired up a browser, and now I need your password, and I'm gonna try to find the input where to put it in, and apparently there's probably an API in this app somewhere. I need to look up the docs to figure it out, and okay, now I've got the API."
There are so many unknown unknowns when you're composing around an app. There are very, very, very, very few unknown unknowns when you're composing around a harness. I think it just gives developers more flexibility, and flexibility that they can inspect.
With API calls, you're just seeing a whole bunch of code flying around the screen. With a harness, I can jump into the harness, look at what's going on, and talk in English about it. So it's much more user-friendly.
We've seen Meta and Muse really be a focus for Zuck. We've seen Alex Wang front and center much more. Were you impressed by what Meta delivered with Muse?
They've been doing a good job, yeah. It takes a while to set up a whole new model lab from scratch, and I'm sure there's a lot of organizational debt to deal with.
Do you think they will be a serious challenger?
I do. I think they have the resources. I think there are some competitive things they can do around the model that help people in ways that the model labs are not as interested in doing, like having a social network and a focus on people.
It's something for the brand that maybe Groq and SpaceX AI have too. But they do need to find their niche. I think people don't quite know what to do with Muse Spark yet—when to use it, when to go for it, or what its core advantage is.
They just released a coding harness. They're trying to be a generally capable model right now. I expect that in the future they're going to be like, "Look, we are way better at this thing," and that's gonna be a really important moment for them.
I see. I have to say, I was impressed by it, actually. Do you know what I use now? Maybe plug in one of our mutual friends, Anastasios, in Arena. It's so weird. I'll put my prompt in Arena, and then, obviously, it comes back with a load of different model options.
Yeah.
I used one the other day—Pergamom?
Pergamom.
Yeah. It was Kimi and Pergamom.
Yeah.
It offers you 4 different options, and it takes me to models that I would never have used before. Actually, Muse has come up a couple of times as being pretty impressive. I love that as a discovery mechanism for models that I would never have used. I would never go to Kimi—
Yeah, yeah.
Honestly, dude, I'd just go to ChatGPT. It's really interesting. It basically, though, goes to the point of the model layer—
Uh-huh.
—just becoming a utility layer.
What do you mean by that?
Actually, I have no loyalty to them.
Oh.
I have no affiliation with the brand. I go to Arena, and I wanna see what you got for me. Show me the results. I don't care if it's Kimi—
Yeah.
—or Muse or Claude or Sonnet or... Do you know what I mean? I just wanna see the options you got, and I'll pick the best from that. I'd rather run 4 in parallel.
Do you buy this whole idea that we're gonna have one frontier model run 4 open models, and the frontier model might be 160 IQ points and the open models might be 120 IQ points, but that will be a model infrastructure or structure that we'll work with?
I totally think that that is a great architecture that everybody needs to explore. We've been helping lots of developers do this. You have sub-agents. We have a sub-agent server tool that we tune to be really, really good at using models generally.
Then you have an orchestrator model that calls out to the sub-agents when it wants particular tasks to get done. These sub-agents are very, very low-cost, and they're focused on deterministic tasks. This is what open-weight models are generally really good at compared to frontier models.
When you have a deterministic task where you know the shape of the output, you know the type of problem that you're working on, and it's a type of problem that has been solved, like classifying some text, for example, then you should definitely use a low-cost model from OpenRouter. Then have the orchestrator model read the results and continue working on the unknown, nondeterministic task that it was set out to do.
I wanna create an open American ecosystem, yeah? More amazing open American models. I make you head of this program. What would you do to encourage and incentivize the open U.S. ecosystem to compete more vociferously with the Chinese?
I think I would spend time talking to the current American labs a little bit more to figure out what distilling the Chinese models looks like for them and how effective it is. You can probably get pretty far distilling the Chinese models.
The nice thing about the open-weight models and the Chinese models is that they allow distillation, and most of them do. That means that you can take the outputs of these models to do reinforcement learning on top of the model that you’re building. This is a very important and common practice in AI that all labs do. Also, when you distill, you see the output, so you can inspect it to make sure that it’s aligned.
If there’s anything about the open-weight models that you’re worried isn’t aligned with the voice or constitution of the model you’re creating, you have a much better shot at catching it when you’re doing these RL rollouts. The other thing I would try to figure out is the compute question. Compute is a huge advantage that I think we still have relative to China, and these neo-labs need a shot. There needs to be an easier way to get compute to the right talent in all countries, but especially if we’re trying to create a competitive American neo-lab system.
NVIDIA’s been doing a good job of this, but there’s Google, there are TPUs, and there’s Trainium from Amazon. I would work with all of the hardware companies, and also like the neo chips to help with compute.
I don’t think we’ll have that compute advantage for long. I think you see DeepSeek and ByteDance both aggressively pursuing their own chips now. The export controls mean that they have to, and this is the number one problem for Xi Jinping in his race to win the AI war. If they build a bridge in 4 weeks, I think they’ll manage a chip in 6 months.
Yeah. Staying ahead in the chip war is critical for America.
Is distillation wrong?
I mean, distillation is a technique to build models.
But people view it with cynicism and shade.
Yeah.
They’re just distilled models.
It’s a technique to build models. The closed-weight model labs distill models too. Sonnet is a partially distilled version of Opus, and this is how you make smaller models out of bigger models. It’s an important way to teach your model new things when you find something useful in the ecosystem.
We do think that labs have a right to say it’s not allowed in their terms of service. A company can cut off access to someone who is trying to build a competitive model. If you’re just trying to build a smaller model that’s really focused on doing one specific thing and isn’t competitive, most of the frontier labs don’t prohibit that, to my knowledge.
There are going to be markets for companies that allow it and companies that don’t, and we make sure that we help both companies uphold their terms of service.
I have to ask you one question before we do a quick-fire round. I’m going to get killed if I don’t ask it. There are reports that you are selling to Stripe for $10 billion. Is that going to happen?
I can’t comment, but whatever happens, we’re going to execute on the vision. What we’re doing is critical for the ecosystem, and we believe for safe access to AI, where one monopoly doesn’t take over, where we have a vibrant ecosystem of models that everyone can explore, and when new providers, new server tools, and new inference-adjacent technology come online, there’s a really easy way to discover it and connect it with all of your existing AI.
I always think in these situations that my response would be, “Well, I own 22% of the company. $10 billion, $2.2 billion. Ooh, now I’m a venture capitalist.” But is it hard not to think like that?
I don’t really think about it.
Do you not?
I don’t spend a lot personally. What I do with personal capital, I really want to help people work on problems that don’t lend themselves very well to venture capital. They’re sort of falling into this gray area of problems that people need to solve but are really tough to fund because they don’t come with a business model attached.
I think there are very cool things to do now in the nonprofit space because you can use AI to review way more data than you ever could before. I’m not quite ready to talk about it publicly yet, but I do want to do something that helps researchers work on those problems and get grants to do it.
One really cool example I think of this is David Fialkow, who’s one of the founders of General Catalyst. He basically finds incredible stories that won’t get funded for movies and funds them to shine a light on them because he thinks they’re very important.
The Dissident obviously told the story of Khashoggi. And then Icarus, which is the story of Russian doping. These were films that would not get funded had it not been for his funding because they were politically sensitive and charged. He’s like, “I’m going to enable the stories of these forbidden tales.”
Yeah. It’s kind of like that. I love that stuff.
He’s great.
He’s fucking awesome. Anyway, are you ready for a quick-fire round?
Sure.
Okay. What is the most underrated model on OpenRouter today?
Ooh, good one. Poolside’s models are great. For my quick-fire answer, I’d probably say a new American lab building interesting coding models that are small and highly effective, and building a lot of useful tools for accessing them. Good team.
70% of neo labs will die in the next 3 years. Agree or disagree?
Disagree. 70% seems very high. Of neo labs, there aren't that many neo labs. If getting acquired by one of the model labs counts as dying, I do think there’ll probably be some potential consolidation.
Consolidation.
But if you include the consolidation, I’d say 50%.
Do you think Dario should be less negative and more positive as a voice in AI?
I think it’s important to have somebody who’s very paranoid about the future and how things are going to shake up. I personally appreciate Anthropic’s paranoia. Obviously, there are areas where I want other model labs not to feel like they’re just being pushed off the table, but I’m a big believer in neurodiversity, and Anthropic is a part of the neurodiversity map that really matters.
If no one is being extremely paranoid, then no one is offering that voice, so I appreciate that they’re doing it.
What’s the craziest thing that you see from your seat, on top of everyone’s usage, that you don’t think people talk about enough?
A lot of companies are obviously worried about cost management and freaking out about the amount of inference they’re spending, and they don’t know how to think about it. It’s a whole new way of doing business and thinking about your OpEx.
The old way of thinking about how much you give your employees is that you give them a salary, and you kind of forget about it. Someone knows what everyone’s making, but it’s a static number that gets readjusted on a quarterly basis, maybe after performance reviews.
Really, your employees all cost totally different, dynamic amounts now. I think a lot of companies are putting it on them to do routing, and I think in the future there’s a good chance that it will get pushed down to the employee level. Your employees should figure out which tools and models to use that are best for their tasks, and then we should figure out how much you’re costing due to the choices that you make as an employee.
Your cost as an employee is going to be a dynamic number, and it’s going to be dependent on how effectively that employee is using expensive and cheap models to do their job. I advise companies to still do their normal management work, like having their managers assess how effective and productive employees are.
But also line it up with how much their employees cost, and then come up with a quadrant of celebration. These employees are doing a good job, and they’re pretty price-effective or cost-effective. And then a quadrant of concern: these employees are maybe doing a so-so job, and, whoa, they are not cost-effective at all. Their AI usage is off the charts. Then you address the quadrant of concern.
I don’t think people talk about how to think about employee cost in the age of AI. It really should be a dynamic number, not a static thing that only a few people know about.
Wonderful. But can you imagine going to someone and saying, “I’m sorry, you were worth $100,000 last month. Now you’re worth $50,000”? I think it would make planning impersonal.
Well, they’re in control of how much they cost. That’s the great thing. All employees are in control of how much they cost and can influence that. Now you get to think, “Okay, how good am I as an employee, and how efficient am I being as well?”
Final one.
Yeah.
When you look at the landscape today, there are so many things to be excited about. What are you singly most excited about?
Two things come to mind.
One is rare disease research, which I think is one of those things that has been intelligence-bottlenecked, or really just the inference bottleneck. It involves trying out lots of ideas and seeing if they work.
The other is crowdsourcing productive urban-life improvements. For example, imagine if someone was curious about finding every lead pipe in America or every lead pipe in the UK and had an approach to it. But they really need to make it mature and stress-test it. Now you can use AI to do that.
We just might solve some weird problems that everyone has given up on because you need a crazy idea to come from somewhere. Brilliant ideas are evenly distributed all over the world. They can come from anywhere, and now you just give them leverage to actually work.
So I'm excited about broad urban or rural quality-of-life improvements that we'll be able to make.
Nice, dude. I've wanted to do this one for a while. I'm so glad we could do it in person as well. I was worried that we were going to have to do it remotely. It is so much nicer to do it in person. You've been fantastic, so thank you so much for doing it with me.
Likewise. This was great.