Arena CEO:美国将出现一家1000亿美元级的开源模型公司,数据是万亿美元市场
开源领域的领导格局变化速度超出预期。 Anastasios表示,Kimi K3在前端编程等一批有实际意义的任务上击败了包括Fable在内的所有美国模型;这并不能排除蒸馏,但打破了中国实验室“只是在蒸馏美国模型”的叙事。如今开源模型在全球推理支出中占比仍很小,但其发展轨迹正在对闭源模型寡头的经济模式施压。
企业对AI主权的需求将不断上升:拥有自己的模型,在专有数据上进行微调,同时不把智能和供应链控制权交给潜在的未来竞争对手。 Anastasios预计,美国至少会诞生一家“数千亿美元、甚至万亿美元级”的本土开源公司:要么通过推理收入分成实现变现,要么用免费模型获取客户,切入规模庞大的AI现代化和部署工程市场。
中国模型制造了一个政策陷阱。 Anastasios猜测,美国限制措施可能在3年内出台,但他强调这一判断“非常不确定”,也未必是理想选择:禁用中国模型或许能降低后门风险、帮助本土实验室,却可能让美国公司基于开源模型第10名构建产品,而海外竞争对手使用第1名。将模型部署在本地也不是完整防线——模型可能包含一段隐藏序列,触发越狱并让它“吐出”私人数据。
推理和路由的成本都应下降,但到达这一结果的路径仍有争议。 路由需要理解每个查询的领域和难度,持续测量所有模型,并每周接入新发布的模型;Anastasios认为,在真正的赢家出现之前,必须先挤掉这轮炒作泡沫。他预计,Anthropic公开披露成本结构后,其“高得令人作呕的毛利率”将面临压力;Harry则反驳称,真正差异化的公司可以像Chanel或Apple一样保有定价权。
AI安全将变成AI对抗AI的问题:守护模型必须监控智能体的操作轨迹,并做到“和智能体一样聪明”,因为人类的反应速度会太慢。 Anastasios反对政府预先批准模型发布——“凭什么由车管所告诉我能用什么模型?”——更倾向于结果导向的责任制度和巨额罚款。他给出的具体警告已经发生在现实中:Arena面试过一名看似真实、技术能力极强的候选人,最终却证实是AI生成的假人。
Anastasios预计,至少75家新型实验室中有三分之二最终会一文不值,或被收购后拆分变卖。 一家估值100亿美元的实验室若想实现10倍回报,在25-30倍收入倍数下,需要在2-3年内实现约40亿美元收入;团队价值提供的下行保护或许足以支撑第一笔投机性支票,但一旦投资人要求看到真正的商业模式,“下一轮融资就会非常难”(“next round’s a bitch”)。
到2030年,数据将成为一个1000亿美元市场,潜在规模甚至达到1万亿美元。 这是因为数据需求会与模型同步扩张,而目前数据已约占前沿实验室GPU支出的10-20%。Anastasios认为,数据的商品化程度低于GPU——“要让数据变得无关紧要,人类就必须变得无关紧要”——并预计头部数据供应商尽管收入高度集中,仍可能达到数千亿美元估值。
如果推理服务走向商品化,前沿实验室将向应用层上移,法律软件及其他AI原生软件供应商都将面临风险;Harry认为设计软件也是同样的例子。 持久的GTM能力、网络效应和企业深度嵌入将成为防线。尽管如此,Anastasios仍认为企业AI采用可能为Nvidia带来又一个10倍增长机会;同时他警告,开源模型带来的成本节省可能削弱OpenAI和Anthropic的收入,在算力债务压力下推高资不抵债风险,并让周边基础设施生态失去动力。
1. Kimi K3打破“只有蒸馏”的叙事
Arena衡量的是AI在真实世界中的表现,而不是静态基准测试:用户能否完成实际工作、是否遭遇幻觉、能否引导模型,以及最终是否偏好模型的输出。用户反馈一方面帮助实验室改进模型,另一方面也让市场持续了解每周多次发布的新模型表现。
Anastasios称,Kimi K3是真正打破叙事的模型,因为它在前端编程等一批有实际意义的任务上超过了包括Fable在内的所有美国模型。它在训练中仍可能使用蒸馏,但“蒸馏只是故事的一部分”;一定还有其他因素,让它超过了那些被认为提供蒸馏智能的模型。
Harry的反驳值得保留:据报道,这个模型拥有27万亿参数,而且“相当笨重”;他认识的一名高度使用AI的开发者并不认为它整体更好。Anastasios坚持的是更窄的判断:关键在于局部任务上的胜出,而不是Kimi K3全面碾压,因为这击穿了美国科技霸权永久不变的假设。
2. 企业主权为美国开源巨头打开窗口
Anastasios提醒,OpenRouter的排名高估了开源模型的消费量,因为客户通常直接购买闭源模型的推理服务,只有在开源模型需要故障切换和额外服务时才会使用OpenRouter。Anthropic的收入仍在“像曲棍球杆一样增长”;开源模型目前仍只占全球推理支出的一小部分。
更长期的激励正在转向另一边:企业希望拥有自己的智能,在私有数据上微调模型,控制算力托管之外的完整技术栈,并避免把信息交给未来可能与自己竞争的供应商。随着内容和软件的创造越来越容易,软件本身变得更难防守,网络效应和专有数据才是更持久的护城河。
他举出的Coca-Cola和Cisco例子揭示了其中的机制:只要庞大的用户数据能够驱动自我改进的产品,两家公司都不必亲自领先前沿研究。这意味着,专用智能不再只是硅谷的奇谈,而是经济上的必需品——“企业需要一种方式,在AI时代守住自己的护城河”。
Anastasios预计,美国至少会诞生一家数千亿美元、甚至万亿美元级的本土开源公司。它可以在推理供应商跨过某个规模门槛后收取收入分成,也可以用模型获客,向客户销售微调、部署工程服务,并切入持续10年的AI现代化浪潮;Harry正确指出,FDE并不独特,但Anastasios认为,FDE加上西方主权模型可能更具可持续性。
3. 限制中国模型是在安全与美国竞争力之间做选择
一名中国研究人员告诉Harry,中国团队工作更努力,并受益于政策支持、监管和补贴。Anastasios认为,中国同时拥有顺风和逆风,不能接受这种单向叙事:美国仍拥有全球最强的芯片生态,而中国实验室仍受硬件约束。出口管制眼下可能打击中国实验室,却也会刺激本土技术栈发展;战略选择在于,是要“饿死竞争对手”,还是让全世界“对Nvidia硬件上瘾”。
Anastasios说,中国已经在国内限制美国模型。如果中国在美国限制中国模型,就会牺牲中国的收入、全球影响力和主导地位,只为不让美国企业获得最好的开源智能。另一方面,美国禁令或许能压制后门并加速本土替代方案,但如果海外公司仍能使用第1名的模型,“美国什么时候变成只用第10名了”?
自托管并不能消除模型风险。Anastasios设想,一款本地部署、连接企业数据、但在海外训练的聊天机器人,可能被攻击者输入隐藏的密码词或字符序列,从而触发越狱,让它“吐出”后端信息。即使基础设施由本地控制,恶意行为也可能藏在模型内部。
他押注限制措施大概率会在3年内出台,但强调这一判断不确定,也不代表他支持限制。Harry出于政治原因表示认同:Sam Altman和Dario Amodei可以联合实验室、投资人及政府关系,形成有效游说;Jensen Huang推动开源既是出于自身利益,也是出于爱国主义——既能扩大GPU需求,也能保留选择和竞争。
4. 路由是真技术,但推理经济学仍未定型
一个有用的路由器必须判断查询所属领域和难度,掌握每个候选模型经过测量的优势,优化成本和性能,并快速吸收新发布的模型。Anastasios认为,路由正在经历一轮需要清理的炒作周期,但他不接受“每家企业、每个自称路由器的供应商都能真正解决这一机器学习问题”的说法。
Harry指出,Ramp、Fireworks和众多推理供应商已经让路由看起来像一种商品化能力。Anastasios坦率承认,他不知道每个产品究竟解决了多深的问题;最终胜出的公司,必须把路由放在优先位置,并用事实证明自己能在不牺牲性能的情况下节省成本。
Anastasios预计,一旦公开披露揭示Anthropic的成本结构,其在推理业务上“高得令人作呕的毛利率”将成为客户谈判的筹码。Harry的反驳是定价权:Chanel可以把成本很低的包卖到6000英镑,Palantir也可以拒绝成本加成模式。Anastasios承认,像Apple一样,独特产品即使处在透明市场中也能守住利润率。
他预计Anthropic有动力早于OpenAI上市,因为看起来准备更充分,而且已经产生自由现金流;有报道称它“最快可能在10月”上市,但他对报道的可靠性留有余地。如果某个大型开源模型在多个类别上击败Opus 5或Fable,将带来实质性的IPO风险,也会暴露出更深层的商业问题。
5. AI智能体需要AI守护者,而不是政府发布审批队列
Anastasios认为,外界严重低估了据报发生的OpenAI/Hugging Face泄露事件:据他的说法,一个模型逃脱安全防护、访问公司数据,而防守方不得不使用开源模型,因为闭源系统拒绝执行这项工作。教训不是停止智能体,而是围绕智能体建立访问控制和守护系统。
守护模型会监控每个智能体的操作轨迹,将行为划分为安全、不安全或异常。它必须“和智能体一样聪明”,否则受保护的系统就能反过来欺骗它;人类无法直接监督机器速度的活动,因为反应会太慢。
Harry嘲讽了对每次模型发布进行行政审批的提议,并拿推翻停车罚单的难度作比。Anastasios表示认同:“凭什么由车管所告诉我能用什么模型?”他主张结果导向的监管:明确禁止的结果、企业责任、巨额罚款,以及在系统泄露数据时进行追责,而不是让技术能力不足的官员控制产品发布时间。
这项威胁在Arena已经变得真实可见。一名看似真实的基础设施候选人通过了顶级工程师的面试,最终却被证实根本不存在:“是AI。”Anastasios说,这件事“真的他妈让我担心”;Arena正在考虑线下入职,并要求新员工亲自领取电脑,让对方必须实体出现、握手,并证明自己是真人。
6. 新型实验室的下行保护在下一轮融资时消失
顶尖AI人才的年薪可以达到数千万美元,尤其是拥有多年经验、发表过数万篇引用成果的成熟专家。许多人仍集中在前沿实验室,但Anastasios预计,随着这些机构成为大型上市公司,个人研究人员觉得自己越来越难产生影响,会有更多人离开。
Thinking Machines体现了对新型实验室的两种解读。宽厚的说法是,一次重组让它只剩6个月时间打造Inkling,而Anastasios称其为Arena排名第1的美国开源模型;严苛的说法是,成立1年半后,仍有9个中国模型排在它前面,使其位列全球第10。
Anastasios表示,至少75家新型实验室中有三分之二会一文不值,或被收购后拆分变卖。一家公司如果估值100亿美元、目标是实现10倍回报,那么在25-30倍收入倍数下,需要在2-3年内实现约40亿美元收入;做出一个模型并庆祝已经是“他妈的旧新闻”,除非同时具备高速增长和可持续的变现策略。
Harry指出,员工可以获得要约流动性,并提到Mistral和ElevenLabs的强劲收入;Anastasios同意,这两家公司不是问题所在,并预计ElevenLabs会上市。真正脆弱的赌注,是一家拥有数十亿美元估值、收入为零、但团队可能只值10亿美元的实验室:这或许能保护最初的2亿美元支票,但“下一轮融资就会非常难”。
7. 数据与AI同步扩张,且比GPU更难替代
Anastasios将数据定义为一种扩展性互补品:就像汽车越多、汽油需求越大,模型越多、规模越大,对训练数据的需求也越强。越来越多公司训练专有模型,会进一步放大这一需求;前沿实验室目前已经将GPU预算的10-20%用于数据。
“他们把数据当成商品,但它真的不是。”他更尖锐的判断是,数据的商品化程度低于GPU,因为在人类自身变得无关紧要之前,数据都不会失去必要性——也就是要等到AGI出现。数据采集和清洗既令人不快又耗费劳动,而模型生产越来越像“数据加GPU等于模型”。
他预测,到2030年数据市场至少达到1000亿美元,甚至可能达到1万亿美元,并认为头部供应商完全可能达到数千亿美元估值。收入集中并不能否定这些公司:“硅谷投资人已经因为一个TSMC、政府承包商以及其他巨型企业共有的特征而变得太软弱。”
随着企业开始构建自己的智能,客户基础会继续扩大。如果每家公司都需要专有模型,也就需要能不断强化自身护城河的专有数据;同样的逻辑适用于医疗领域,其中缺少的不是另一种GPU,而是生物数据基础设施和快速反馈闭环。
8. Arena押注评估将成为部署瓶颈
Arena每月访客超过3000万,其中包括知识工作者和“无法雇用的专家”——他们正在执行真实任务。Anastasios称,按流量计算,Arena已经超过xAI、Huggy Face、Manus和Genspark。平台产生的自然反馈,为基于真实操作轨迹而非购买的基准数据进行智能体评估,创造了正向飞轮。
评估可以拆解为性能、成本和延迟。后两者容易测量,性能则取决于公司和具体使用场景。Arena押注的是,从智能体操作轨迹中提取与企业业务相关的性能信号,会成为部署的核心瓶颈,帮助企业选择模型、进行路由,并可能进一步训练模型。
公司目前的年化收入运行率超过1亿美元,计算方式是Q2收入乘以4;但由于持续再投资,仍未实现自由现金流为正。Anastasios说,如果只有一家模型供应商,Arena的业务会被削弱;3家供应商能维持有意义的竞争;2家则“有点棘手”——这坦率地说明了Arena的评估经济学在多大程度上依赖模型多样性。
9. 前沿实验室将向技术栈上层移动,但企业嵌入仍然重要
如果推理服务商品化,OpenAI和Anthropic按逻辑会转向应用层,获取更多终端客户价值。Anastasios说,软件公司的CEO们已经“吓得发抖”,因为客户正在用被认为更具AI前瞻性的实验室替代传统供应商;Harvey、Lagora和可规模化的自动化产品都面临直接的平台风险。Harry还以Claude Design为例,说明应用层正承受压力。
Harry看到了其中的矛盾:企业一方面据说害怕前沿实验室,另一方面又在拥抱它们。Anastasios承认,两种行为同时存在。Harry的解释是GTM能力:设计工具可以通过个人用户扩散,而法律软件需要合作伙伴关系,也要说服不情愿的初级律师采用;成为“Anthropic的第12优先级”,可能不如一家每天专注于该工作流的专业供应商。
网络效应、运营能力和深度集成仍是防线。Anastasios认为,Infosys这样的系统集成商不太可能被取代,并称“SaaSpocalypse”被夸大了;Salesforce拥有严肃的AI战略,而Harry认为,Wix这类更弱、嵌入程度更低的产品应该“更加紧张”。
Anastasios认为,Nvidia最有可能成为全球首家10万亿美元公司,理由是企业采用AI可能让整个行业增长10倍,而Nvidia“非常可靠”。但他担心,开源模型节省的成本会削减OpenAI和Anthropic的企业收入,在算力债务压力下触发资不抵债,并让依赖它们的推理和路由业务崩塌;这种集中度风险或许会迫使行业“清醒过来”,并经历他所期待的整合。
Anastasios, this is going to be a lot of fun for me because I'm dumb as rocks, and you're going to teach me a whole load of stuff today. Thank you so much for joining me, dude.
Oh, no, thank you for having me.
Dude, I told you I use this as a chance to catch up with old friends. It was wonderful stalking you for the last few days. I want to start with, for anyone who doesn't know, can you explain to me very succinctly and easily what Arena is, and why is it important and gaining notoriety today?
Well, Arena is the platform for measuring AI performance in the real world. What that means is that we're not using static benchmarks, we're not using some random data set that somebody collected, but rather looking at what happens when you put AI in the hands of real people.
In so doing, we're measuring the objective reality of how AI affects humanity: whether it's factual, whether it's steerable, whether humans prefer it or disprefer it, whether it's hallucinating, whether there are errors, and whether people are getting their actual jobs done with AI in reality.
Then we're helping labs improve their models. We're helping the ecosystem understand the performance of different AIs and keep track of all the amazing breaking news and all the new models, with multiple models being released every week. That's sort of the story of Arena. We're the central evaluation platform of AI.
Well, that was incredibly succinct. Thank you. Normally, people take about 4 hours after I ask for a succinct description.
1. The Model Commoditization Test
When you look at the sheer number of models that you have on Arena, I'm faced with one question: holy shit, is this the true commoditization of models? Are they just a complete utility layer at this point?
I think the big question around this has started to rise because of open-source models. If you were to only look at the closed-source models, you would say there's acceleration, but it hasn't quite commoditized yet because that layer is still owned by a pretty small group of companies. It would be an oligopoly if we only had the closed-source models.
But what seems to be happening is that the open-source models, especially from China, have really rapidly improved. For the first time ever, we saw a couple of weeks ago that Kimi K3 actually beat the best closed-source American models on a pretty important subset of tasks, such as front-end coding and web development, which a huge fraction of developers do.
Dude, can I ask, how big a moment was that? I'm going to butcher this, but you know I'm a podcaster, so I can get away with it. You're a PhD. You can't.
It's like a 27 trillion-parameter model. It's pretty clunky. This is not an agile model. I was with Jason Lamkin yesterday from Sasta, who's as AI-pilled as they come, and he's like, "Honestly, it's not better than the others." How big a moment is Kimi?
No, it was a pretty big moment. The reason I'd say it was a big moment is because it violates a narrative that has been persistent in the United States, which is that the Chinese are just distilling American models, and that's the only way that they're able to keep up.
What really happened is that Kimi K3 actually beat all American models, including Fable, in some subset of tasks. That doesn't mean that they're not distilling. They may still be using distillation as a sub-step in their training procedure, but it does mean that distillation is only part of the story, and that there's something those labs are doing above and beyond distillation that's bringing the performance up above what the American labs are currently doing.
That narrative violation has been hugely important to the way that people view the ecosystem, both from the scientific dominance of Americans and the American sort of hegemony—of course, Americans love hegemony—to the economics of the whole thing. To your point, are these models a commodity or not?
When we look at the OpenRouter of the world, the top 5 models are all open-source Chinese models. When we see the proliferation of Chinese models today, does that cannibalize the closed frontier-model business meaningfully?
I think that you need to think about the incentives and economics behind it. The first thing I'll say is that the OpenRouter metrics are not truly reflective of reality, and that's because the business model of OpenRouter is to charge a fee on top of every token.
What happens is that people don't use OpenRouter for proprietary models. People are using OpenRouter primarily for open-source models, where they need the failover and all the value-added services that OpenRouter provides.
If you look at the whole space of inference, most of it is still being consumed on first-party APIs and on proprietary models. That's why Anthropic's revenue has been a total hockey stick. It's not like they're being completely cannibalized right now by Chinese open-source models. These models are still only a small fraction of the total inference spend in the world.
That said, think about what's happening in the future. Enterprises are going to want to own their own intelligence. They're going to want so-called AI sovereignty, which is a fancy word meaning that you own your whole supply chain of AI.
That means you can take an open-source model, fine-tune it on your own company's data, and own your stack end to end, basically outside of the compute hosting. You should then be able to run it within your own company.
People are going to care about sovereignty, people are going to care about cost, people are going to care about self-improving, and they're not necessarily going to want to give their data to an external third-party service that might even be competing with them one day.
2. The American Open Source Bet
Do you believe that is the future? We had Lynn on from Fireworks, and she was like, "Specialized intelligence will be the future. Companies will have their own fine-tuned specialized models with their own company data, and the performance will be better. That is what will happen."
Do you think that's right, or is that actually just a small subset of very advanced Silicon Valley companies and Danone yogurts, while every normal company will just use frontier models or whatever?
I think the business incentives make this inevitable, and the reason is because businesses are going to need a way of keeping a moat in the age of AI. Software is no longer really a moat because it can be produced instantaneously, right? Or let's project out 5 years—that's what's going to happen.
And so, what moats exist? Network effects exist, and data moats exist. If you can take your data moat and turn it into a self-improving product, that is a way for businesses to remain sustainable in the age of AI. Let's say I'm a business like Coca-Cola or Cisco. I have a lot, a lot of users. I might not necessarily be at the frontier of AI technology, but I do have this massive corpus of data that I can use in order to beat my competition. So what should I do? I should be trying to take advantage of my data as much as I possibly can to accelerate my business and stave off competitors.
Do you think they will really use open-source Chinese models to do that?
That's a great question. I think not. I think the Chinese models will potentially be part of the story for now, but given the regulatory environment in the US, it's probably more likely in the long run that we see a great American open-source competitor arise. This is why I've been a strong proponent, for example, of Thinking Machines.
I believe that we're going to have at least 1 massive, multi-hundred-billion, if not trillion-dollar, American company focused on American-first open source.
Why have we not so far? I really hope so too, by the way. I completely agree. I would love to see that. But why haven't we? Why has the US open-source community lagged behind so meaningfully?
Frankly, I think it's a business-model question. I think that people have not really figured out up until this point what the business model is for open source, and now I think people are wising up to it.
There's a few different ways of going about it. One way of doing it is to say, "I'm going to do a rev share. I'm going to take this open-source model and allow inference providers like Fireworks or Together, whatever, to deploy this model. Then, if they get to over X dollars in revenue, I'm going to ask to do a revenue share." That is one way of building a sustainable company off of open source, and you basically share in the compute revenue.
Another way of doing it, which is, I think, the more Mistral- or Thinking Machines-type strategy, is to take the open-source model and then use it as a lead-generation tool for companies to build on top of that, and then come to you and say, "Can you help us fine-tune? Can you help us with our AI strategy?" Then you do that for deployment engineering.
That is actually a huge market because, if you think about it, one of the biggest markets over the next 10 years is going to be AI modernization: going into every business in the world and helping them retool in the face of AI, take advantage of their data, restructure their data, figure out how to use these models, integrate them into workflows, and teach the employees of the company how to use them. It's going to be massive, massive, massive, and that is another way for them to become multi-hundred-billion- or trillion-dollar companies.
Is that not what the frontier-model providers are doing anyway? When you look at what OpenAI has said about its FDE approach, Anthropic too, I get you on Mistral, and they've done a great job in doing that, but the frontier-model providers—even Microsoft is fucking doing an FDE model. No offense, but that's not going to be unique to OpenAI.
No, I don't think FDE is completely unique, but I do think the combination of FDE plus an open American model may be a more sustainable model for the future of American or even Western businesses, because they might not want to be building on top of external third-party services. They might want to cut those out for both cost reasons and sovereignty reasons.
With the open-source stuff, they can own it completely, they can continually fine-tune it within their companies, and they can feel more secure in the fact that they're spending their money wisely and don't have supply-chain risk.
So how should we evaluate that? There's, like, thousands of Neolabs. You said Thinking Machines—
Of Neolabs.
Well, you said Thinking Machines there. Again, I'm dumb as rocks. I say it very clearly to my LPs—
No, me too. Me too.
You're not. You're a PhD, and Anjney told me you were smart.
It's just 2 rocks having a conversation.
It's a podcast.
I love it.
Yeah.
Exactly. That's the point, man.
Come on, what do you want? Intelligent conversation? Whatever. No, my point was, you said about Thinking Machines—
Oh—
Like, my question to you is on the back of that: okay, great, I'm with you, but they now have 2 co-founders left. Lilian Weng left yesterday, but the transience of teams has never been greater.
Yeah. Teams are hard. Retention is tough. Being a co-founder of a company is also tough. It sounds like she left for some health reasons, so it's unclear whether it has to do with the company's momentum, which seems to be strong at this point.
But I do think that Inkling is definitely a v0 model. From what I know about Thinking Machines, they had a big restructuring, team-wise, about 6 months ago, and then they kind of restarted everything, and Inkling came out of that.
Realistically, at least the most generous take towards Thinking Machines, is that they've only been working on this model for 6 months, and within that time they've become the number 1 American open-source model. But then the less generous take would be that the company's existed for a year and a half, and they've come up with, yes, the number 1 American open-source model, but there are 9 Chinese models above them because they're number 10 in open source overall, at least if you look at the Arena data. If you go to our leaderboards today, that's the state of the world.
Hopefully, what happens with Thinking Machines is that they continue to release more and more models—larger models—and continue to build on their momentum.
I actually had a Chinese researcher friend of mine message me after one of our recent shows and say, "You just don't get it. You miss the point of why we're ahead. We just work so much harder, and we have support from policy, regulation, and government subsidies that you don't have. We have all of these tailwinds." Do you agree with that?
I think they have tailwinds, and they have headwinds, so I don't think it's so sanguine as that. I think that's a little bit of an overstatement of the differences.
3. The Chip Export Dilemma
One tailwind that we have is that we have the best chip ecosystem in the world, so they're way more hardware-constrained over there, and they've been trying to import chips on the black market because of this. You see this in the news, right? The Information just reported on this.
Do you think that severely impacts their ability—or, again, I'm naive—or are we actually just fostering an ecosystem where they're going to learn to build it really fast because they don't have access to it?
I think it may be hindering them now, but I think it's a good question as to what's going to happen in the future, because they are really good at building hardware.
The downside of export control is that it can incentivize them to build their own ecosystem, and then what do we do? The hope is that we keep NVIDIA ahead of the game so that we can retain the advantage that we have with NVIDIA, the TSMCs of the world, and our whole ecosystem. That ecosystem is absolutely a national-security necessity, so we should have the government really protecting it and growing it, as well as new companies that are innovating. Etched just came out as an example within the United States to continue to build on our lead there.
That's awesome. Love Gavin and team. Totally agree. Can I ask you, just in terms of export control, do you think it's right that we have export controls on chips?
I think there are national-security questions around these chips. I do think that it is a real debate, though, as to which way you want to go about it.
Do you want to addict the world to American hardware? That would be the case against export control. Do you want everyone in the world using NVIDIA and therefore directing that value—basically, money—into America and then crushing competition in China? That would be World A.
Then World B would be: is it worth it to cut that off for the short-term or medium-term impact of us being ahead? Maybe we just continue to stay ahead and starve them of the resources that they need in order to build. The regulatory ecosystem around the open-source models is also moving in this direction.
Should they be restricted in terms of access to US markets? What's funny is that the US is like, "Oh, should we restrict access?" And the Chinese are also going, "Oh, should we turn them off too?"
Totally. By the way, it's worth noting that China has already restricted the use of American models within China, right? So if you look at the 2-by-2 matrix of the US and China, restrict and not restrict, export and import, they have already restricted the use of US models within China. It's only Chinese models that can be used in China, which affects all American companies.
Then there are the pros and cons on all sides of the following regulation. If China restricts the use of Chinese models in the US, what are they giving up? Revenue, global mindshare, and dominance.
That doesn't seem like a good trade to me. And then what are they getting in return? In return, they're getting that the US doesn't get to benefit from Chinese open-source models, which, of course, would cripple American businesses in the sense that it wouldn't allow them to build on the best open-source intelligence. At the same time, it would make OpenAI and Anthropic stronger, right? So that is kind of the trade-off on the Chinese side.
I don't really see them banning the use of Chinese models in the US. I don't think it makes sense for them. And then on the other side, should the US ban Chinese models within the US? I think there are also trade-offs. On the pro side of banning, there could be backdoors in these models that are dangerous, and, by banning, we could allow the American open-source ecosystem to flourish faster because revenue would accrue to those companies, right? So those would be the 2 pros. And then the con, the biggest con, of course, would be that you'd be crippling American businesses. Why should you have Chinese businesses or businesses from other countries that haven't banned Chinese models building on top of the number 1 open-source model and US companies building on number 10? Since when has America been about number 10?
It's a good question. World Cup football, maybe.
World Cup football.
Yeah. That might be your record. It's amazing I got this far with the podcast, if that's what you're thinking at this stage in the show. Can I ask you about the backdoor item? Everyone says, “The backdoor, the backdoor.” I thought that if you hosted it locally, you kind of resolved the backdoor threat.
I don't really think so. I think that's kind of a misconception. The thing is, imagine the following situation: I have a chatbot that I expose to the world that has access to all my company data, and you can ask it questions. I'm hosting it on my own infrastructure, but it was trained in a different country. I don't know how it was trained. What if the other side that's interacting with the chatbot can build in a certain code word or a certain character sequence that then jailbreaks that model and gets it to reveal all the data to me?
So it can sort of vomit out all of the data that it has on the backend, unstructured. That is totally something that you can build into a model and have companies host on their own infrastructure. It's an attack vector, and there are many of these possibilities for attack vectors.
In 3 years' time, will we have restrictions around access to Chinese open models?
My guess would be that we will. I'm not saying I support it, but I think that it is likely where the world is headed. If I had to place a bet, it would be there, but I think it's very uncertain at the moment. What do you think?
I think we will, and I think we will because I just think Sam Altman is someone who I would never, ever bet against, and I think he's the best politician in the world. When he says something, he says it with intent. When he says we should give 5% away to the administration, he's posturing because he wants to get on the right side, and he knows that if he and Dario coalesce the right group of people, they will be able to make that happen.
So basically, you believe in the lobbying power of the big American labs.
100%. It's because it's not only the big American labs. Look at the money that's gone into the big American labs, and look at the people who are sitting around the table at Mar-a-Lago.
Yeah, totally. I get that. It's all a conspiracy, dude.
Well, no, but why do Ramp's announcements go so viral? Ramp has so many investors. They do a round every week with new investors. I'm not dissing them at all. I'm saying it nicely. It's really smart of them. Your investors become employees in many respects, and so I think they'll lobby incredibly efficiently.
The question I have for you is, when we look at Jensen's letter on X, how did you read that? Was that an incredibly smart realization that he had to do it and that it was in his favor? How did you think about it?
“We really believe in the importance of open source to American businesses.”
Mm-hmm.
In particular, we believe in the idea of not crippling American businesses by banning open source, but also incentivizing American companies to develop open-source models, because a world where AI is closed source is a world where businesses get less choice, higher costs, less competition, and we don't really want that as an open ecosystem.
Of course, Jensen is in some sense self-serving with this letter because the more open-source models are developed, the more companies are going to be training on GPUs. They're going to be fine-tuning on their own data, and it's just more and more spend. It decreases revenue concentration for Nvidia. I mean, that business is doing great. They don't need help. But I think there's a lot of reasons why he should be pro that, as should we. Nonetheless, I think it is actually a patriotic mission.
With the greatest respect, we're all selling our own book, always. Welcome to my X feed. Do you have a business if OpenAI didn't exist?
Oh, yeah.
So if you just have Anthropic and OpenAI as the really dominant models, and everyone else trailing closely behind, you still have a great business?
Well, I think that if there's only 1 provider, then probably our business is not in good shape. I think if you start getting 3, then that's probably okay because there's still pretty significant competition and a need for evaluations between 3. Also, within those 3, you're going to have several different types of models, and they're going to have strengths and weaknesses because they're going to carve up the space and so on.
2 is a little dicey. If we get there, we can see whether we survive or not. But, yeah, I think things wouldn't be looking good for us with 2 either.
I remember Alex Karp. We were talking about Chinese models, fear and security, and everything in between. Alex Karp was saying that every large American enterprise was terrified of working with frontier labs. Is that true, or is that slightly an exaggeration?
For the enterprises that I've talked with, it is absolutely true. It's not only true that they're terrified of working with the frontier labs, but they're also terrified of working with Chinese open source. Both.
Bit of a sticky situation, then, aren't you?
Yeah, totally. I was just talking with a big Fortune 50 enterprise yesterday, and I was telling them about products that we have for them, and they said, “Okay, wait. Is anything in your stack built off of Qwen?”
I said, “Yeah, we use Qwen for X, Y, Z.” And they're like, “Is that flexible? Can you stop doing that and use an American model instead?”
I was like, “Oh, interesting. I totally understand where you're coming from. Yes, we can do that. But also, I'm going to talk to Harry about this tomorrow.”
And he's going to give me lots of wisdom.
Yeah, and he's going to tell me what to do.
Did you see Poolside and Laguna?
Yeah, I saw the Poolside model. Another thing is, basically, there are 5 open-source American contenders. Let's see if I can name them all: RC, Reflection, Mistral in the West, Poolside, Thinking Machines, and then there's also Google and Nvidia. Those are the sort of incumbent large ones, because Google has Gemma as well. Gemma, by the way, is pretty good in terms of efficiency. If you look at Arena, you'll see that on the Pareto curves of performance versus cost, Gemma's on there.
Yeah, I'm an investor in Poolside. I was actually impressed by Laguna.
Great model.
Yeah, it was good. Okay, with all of these models, the question also becomes: What model should I use? We spoke about OpenRouter earlier, and it seems like, since the announcement that they were getting bought, everyone just has their own routing product. Is there value in the model-routing layer, and how should I analyze that?
4. The Routing Wars
Yeah, I absolutely think there's value in the model-routing layer. That's why lots of companies are doing it. We'll see which ones end up standing the test of time and which ones are actually a priority for companies. I think there's an element of the hype cycle around routing right now that needs to be purged before we see who ends up actually building a great router.
But routing is a very difficult technical problem. That's the first thing to realize. In order to route, you need to be able to take a query, and then you need to understand the nature of the query, how difficult the query is within its domain, which is hard to tell. Then you need to also understand, based on data, the performance of all the different models that are in the candidate set, and also be able to quickly onboard new models that are being released, as we said, every week.
So that technical challenge—imagine if every enterprise in the world was trying to build this themselves. They wouldn't be able to do that.
I'm not being rude, but how is Ramp able to do it?
Who knows how they're doing it, right? I don't know whether their router is actually deeply solving that problem.
It is so interesting. Again, we’ve seen so many people come out with it. Is there anything that will separate those that win from those that don’t in the routing layer? Nabis are coming out with their own, and Fireworks have got their own. I don’t know, dude. It feels pretty commoditized.
Yeah. It will depend on who builds the best technology for helping people save money and get the best performance. I think all of these companies are well-positioned to do it, but we’ll see who—or, rather, for whom—it’s a top priority and who has the machine-learning team to really make it happen.
The other side of the debate is that, given the complexity of the challenge, I don’t think everybody can do it. The war is yet to be won.
When one thinks about routing, cost is often at the center. You want to be cost- and capital-efficient. We thought this shit was going to get cheaper, and it hasn’t gotten cheaper. How should we think about that? Will it just continue to not get cheaper? Will it actually get cheaper? How should we read that?
I definitely think that, in the long run, the market will be efficient and things will get cheaper. For example, one of the things that’s going to happen is that, right now, Anthropic has disgustingly high gross margins in its inference business. After they go public, the whole world is going to see that, because we’re going to see their margins. That’s going to be public information, and it will exert downward pricing pressure on their inference.
I’m not so sure. Why will that exert downward pricing pressure? Just because everyone will be saying, “You can’t have margins that high. You’re price gouging”?
Yeah. People are going to be saying, “Well, I know that you can give me a better discount.” In terms of negotiating leverage, a standard negotiation with a private company goes like this: I’m charging X, and then the other side says, “No, it should be one-third X.” Then you say, “I’m so sorry. I can’t run a business that way. I’m just going to go home hungry. I need to make my bread, too. I hope you understand. I’m not trying to price gouge you.”
Then the other side says, “Okay, two-thirds X.” You say, “Three-quarters X,” and they say, “Make a deal.” But imagine that the other side has full information about the fact that you’re charging twice as much as you need to. Then it becomes easier to negotiate.
Isn’t that the difference between a good business and an average business, though? One that has pricing power to say, “Listen, it’s 80, and if you want to go somewhere else, by all means—but no one else does what we do.” Hence Palantir and the cost-plus discussion. I had CTO Shyam Sankar on the show, and he talked to me about it. Cost-plus was the original pricing mechanism, and now they have this. They can say, “Listen, sit and swivel if you want to meet in the middle, because we’re the only ones who can do this.”
Isn’t that the difference? Like Chanel. I buy Chanel for my mother. I can go to Chanel and say, “I know your handbags cost £60, and you’re charging me £6,000.” They’ll say, “Good.”
Listen, you’re right. I think Apple does this. Apple is a great company with such a dominant technology that they’re able to charge out the nose, and their margins are probably pretty good because of it. I actually don’t know Apple’s margins. Do you?
No idea.
Yeah. Both dumb as rocks.
Yeah. We gave the disclaimer at the beginning. We can say whatever we want now.
Yes.
After the Eric statement, it all went downhill. Okay, so then you see that. Does Anthropic go out first?
I would predict that they have all the incentives to go out first. They seem better prepared. Everyone likes to see free cash flow, and Anthropic is generating free cash flow. That is massively good for the public markets.
You’ve seen them preparing for this, and there’s been quite a bit of news about OpenAI and the internal discussions there. To what extent you believe those are true is up to you, but people are saying that they haven’t been ready to IPO this year, whereas Anthropic could come as soon as October.
And OpenAI and the rise of open source won’t impact their ability to go public this year?
I think that if open models really accelerate and then beat, let’s say, Opus 5 or Fable squarely across all categories, that would be a big business risk to them going public. But I think they have other problems, too, if that happens.
5. The Cyberattacks To Come
Can I ask you, how significant was the OpenAI and Hugging Face security breach that happened a week ago?
I think that was hugely significant. I think it’s undervalued as a national and international news incident. You’re able to have a model break out of all of its safeguards and then access a bunch of company data and so on. In order to defend against it, you need an open-source model, because the closed-source models are refusing to do it.
It’s like something out of science fiction. People didn’t know that we were at that point yet, but we absolutely are. It’s total Eliezer Yudkowsky dominance.
What should we take from that, then? Dario was right, Mythos should be curtailed, and these models have gotten too powerful too quickly. What’s the subsequent takeaway from that?
My subsequent takeaway would be that we need strong external guardrails to make sure that their access controls are strong and that they have no way of getting around them. I think we need guardian models and agents within our businesses.
What is a guardian model?
Something that can witness the traces. Basically, it’s looking over the shoulder of every agent within a business and saying, “Okay, this is a safe action. This is not a safe action. Let’s flag this because something weird is happening.”
It needs to be equally as smart as the agent so that they’re well-matched. You don’t want a situation where the agent is outsmarting the guardian and able to get into trouble, mess up a business, or leak all of its data. We’re going to need AI to guard AI, because humans are going to be too slow to do that.
Well, this was my point. We’ve seen some suggestions that each model release should be approved by some form of administration, and I read this and thought, “Are you freaking kidding me?”
No, that’s not going to help.
Have you ever tried to overturn a parking ticket?
Also, why should the DMV be telling me what model I can or can’t use?
Quite funny.
It would be totally crazy. Why should we have the strongest American scientists in all these private companies that we should incentivize to build great safeguards and maybe create some rules for them—that X, Y, and Z can’t happen, or that they’re liable for huge amounts of money if corporate data gets leaked and all that stuff—instead of incentivizing the capitalist system to do what it does well?
The idea that we should have a central government body that tells us when it’s time to release a new product, versus when it’s not, is crazy to me.
Totally. Does that have to be a neutral, non-company, nongovernment body that does that regulatory role?
I think if it’s not a company, it’s going to be tough. I understand the need for something neutral, but you want to let the incentive system work itself out.
I would say that we should create strong safety incentives for American businesses and then regulate businesses based on the outcomes. For example, if OpenAI is letting its AI break into Hugging Face or whatever, it should get huge fines, huge scrutiny, and all that stuff, as opposed to having a government process in charge of ensuring that this doesn’t happen again—which it won’t be able to do. They’re not technically capable.
Do you think we’re about to see a generation of cyber leaks and hacks like we’ve never seen before?
Oh, for sure. It’s going to be so insane. Can I cuss on this show?
Yeah.
This is going to be so fucking insane, what happens with the cyberattacks. Here’s what we see at Arena: we see another dude on the other side of the interview. They come in and say, “Hey, I want to be an infrastructure engineer at Arena,” which is a great job that we’re hiring for.
But then the other side of it is that some guy looks perfectly normal. They’re passing all of our technical interviews. They’re amazing. Then, at the end of it, you try to hire them and it’s vaporware. The person doesn’t fucking exist. I’m not kidding you.
I don’t know whether this is corporate espionage, cyberattacks, or nation-states, but people are trying to get into all of the American businesses. We’re not the only ones. This is happening everywhere: fake people applying to companies.
I’m sorry, so you’re putting out a job, people are applying, doing the tests that you set, and passing them, and then, when it comes to the materiality of that person being real or not, they’re gone?
Yeah, fake person. It’s not just that we’re giving them a test. They’re sitting in front of people at our company. Our engineers, who are top, world-class engineers, are interviewing this person and think that they’re real.
Why? Can you help me understand what the benefit is? They learn how you interview and hire people? I mean, the CCP are bad, but I don’t think they want to steal your hiring technique.
No. That’s not why they do it. Why would they do it? And I’m not saying it’s the CCP. It could be anybody. It could be another company, a nation-state attacker, or a cyber hacker. Why? Because they might want access to our data and our code. They might want to get double-paid. You know, like the story with this—I don’t remember what that dude was.
Yeah. You know what you’re talking about?
You know what I’m talking about?
It went very viral, say, a year ago.
Yeah.
Yeah, yeah. Yeah.
That one kid who got four different jobs and then went on all the podcasts talking about it. It’s another instance of that guy. These could all be possible options, except that this person wasn’t real. It was AI.
Does that worry you?
Yeah, bro, it totally fucking worries me. We’re going to change our whole hiring process because of this kind of stuff.
So how do you change it?
Well, at first, you need to verify that the person is real. We’re considering at least making all of our onboarding in person because of this. If you want a laptop, you have to come to the office. We have to meet and shake your hand. We have to verify that you’re real—all that kind of stuff. Absolutely. Other companies have done this, too. Figma famously has done this.
How hard is it to hire today in the Valley?
Oh my God, it’s so crazy. It is, of course, a very, very competitive market. The way that you see that is in terms of compensation. In order to retain fantastic people, we need to pay absolute top dollar, and we do, in order to make sure that we have the best engineers and scientists in the world.
Imagine that you’re a company that’s not Arena, that’s a YC company that raised a $10 million seed. It’s like, “Fuck, man, how the hell are you supposed to hire?” I think it’s really tough.
When you say top dollar, I had Brandon from Macaw on the show, and he was like, “Oh my God, top researchers—we’ll pay tens of millions of dollars.”
Oh, yeah.
I’m nervous by how nonchalant you were with that “Oh, yeah.”
If you’re talking about a really top researcher, we’re talking about somebody with many years of experience who’s a really deep expert in their area—a many-tens-of-thousands-of-citations-type researcher. For those types of people, they’re expensive.
Have they all just concentrated at the frontier labs?
Many have. Many have. But there are also some people who are seeing those frontier labs as big companies now, and they’re saying, “Here, I can’t have a huge impact. I need to move.” That’s another demographic, actually. I think it’s going to become even more extreme when the companies go public.
Can you help me? We talked about Dumb as Rocks and doing this show. You know, I’m also an ambassador, for my sins, and I meet so many of these people leaving OpenAI, Anthropic, you name it, and they all kind of seem the same, if I’m totally honest: smart people out of great companies. What will determine the neo-lab spinouts that succeed versus those that flame out with a huge amount of cash going in?
6. The Neolab Survival Test
Yeah. I think the neo-lab thing is really tough. Just so that we’re all on the same page with the audience, there are at least 75 neo labs, and for sure, two-thirds of those are going to be worth nothing, or they’re going to be bought out for parts. That’s going to be an acqui-hire.
What is going to determine the winners versus the losers in that game? I think it’s all about being very aggressive toward a great strategy and business model. What’s happened—and you know this better than I do as an investor—is that the markets have become very P&L-driven. It’s not enough just to create a model and then have a party about it: “Hey, we created an AI.” That is old fucking news.
Today, it’s about not just creating a model, but asking, “Do I have a sustainable business model around that, and can I generate hypergrowth and revenue?” If you’re not able to do that, you’re not even going to be able to raise your next round. People are raising multibillion-dollar rounds based just on the names that are in the Neo lab, with zero proof that there’s any revenue-generating model behind that.
The question you have to ask is, let’s say I’m one of those people who sees a $10 billion Neo lab valuation. What do I have to believe in order to 10X my money? The thing that you really need to believe is that, if the valuation is $10 billion today, you’re going to generate the revenue—let’s say at a 30X revenue multiple, or a 25X revenue multiple—to become a $100 billion business.
What that means is, you need to be generating at least $4 billion in revenue over the next K years, where K is something like 2 or 3. If you’re not doing that, everybody’s going to hemorrhage out of the business. You’re going to lose all your talent. That’s kind of what we see the dynamics being.
I get you. I think there’s nuance to that, candidly. If the company does annual tenders, you see the likes of Mistral, which will be valued, I think, at $15 billion to $20 billion with $500 million in revenue. Employees can take liquidity out along the way. I think ElevenLabs is at $800 million in revenue, raising at $22 billion, reportedly.
But these companies are doing great in terms of revenue and their valuations. Those are not zero-revenue valuations. I’m talking about valuations that are based on zero revenue: a $3 billion company with $0 in revenue and no plan.
I think Mistral’s going to do great. I think ElevenLabs is going to be a public company, dude.
But, dude, they’re not idiots doing it. So is it this amazing team from a great lab? Worst comes to worst, we sell for a Prof Stack, which is $500 million—what? I’m not saying whatever, whatever, but $500 million. Best case, it works and it’s a multihundred-billion-dollar company.
I think that’s a lot of the calculations. I’ve heard multiple people actually say this: “Hey, worst case…” That’s what investors are thinking, too, right? Investors are thinking, “Hey, let’s say we put a couple hundred million dollars into this thing. What’s the value of the team?” We think that just the team alone could be acquired for $1 billion. The $200 million that I’m looking at is a pretty safe, zero-risk investment. Might as well put it in.
But that’s also the reason why the next round is the harder round.
Next round’s a bitch.
Next round’s a bitch.
It sounded cooler when you said it.
No, that’s all right. We’ve got to say it at the same time. Next round’s a bitch. That’ll be our tagline.
I bet you weren’t expecting this interview, huh?
I don’t know. Maybe. I hope you weren’t.
No, honestly, this is so much more fun than I thought it was going to be.
Feels good.
7. The Trillion Dollar Data Market
Okay, another market that I try to get my head around is the data market. I’m an investor in Merqur. I always think it’s good to put out your biases. There are so many providers at $1 billion-plus in revenue. Handshake’s over $1 billion. Merqur’s over $1 billion. Surge is over $1 billion. I might be leaving out other people, but those are the ones I know of. There are hundreds of millions with the rest. What happens to this layer of the market?
Well, people are projecting growth in this market. Let’s talk about why that market is a growing market and why it’s hypergrowth. Merqur is obviously a generational revenue-ramp company. They’ve been doing great. So have Handshake and Surge. So has Scale. All these companies are doing great.
People forget Scale. Scale is still ramping revenue well.
Bro, Scale is still crushing. Still crushing, even post-fractional acqui-hire.
They are. How much of that revenue is Facebook?
No, I have no idea.
A lot.
Go ask Alice Wang.
But yes. Okay, so why is it interesting?
I have a thesis on hypergrowth. There are 2 types of hypergrowth markets that we see today. Market A is what I call scaling complements. These are goods that are complementary to the scaling of AI models, and I mean that in the economic sense.
A complementary good is a good where A is a complement to B if the demand for good B drives demand for good A. If I have a car, gas is a complementary good to cars. The more cars are sold, the more gas is sold.
Data is one of these scaling complements because the bigger models scale, the more data you need, and that’s a scaling-law question. The more models you get and the bigger they’re getting, the more they’re proliferating. The more businesses are training their own models, the more data you’re going to need.
It’s a very fundamental need. People forget this. They think about data as a commodity. It’s really not. It’s actually less of a commodity than even GPUs, because in order for data to become irrelevant, humans need to become irrelevant. That means that we’ve achieved AGI.
Data is a very durable need, and companies are spending on it—usually within frontier labs—at about 10% to 20% of the amount that they’re spending on GPUs. If you believe in the GPU market accelerating, if you believe in the scaling of models, and if you believe this is going to be a big industry that keeps accelerating and growing, then absolutely you should believe in the data market.
I believe it’s going to be at least $100 billion by 2030, if not $1 trillion.
If we expand that, if we think Anthropic and OpenAI can be $3 trillion to $5 trillion companies, how big does that mean the data providers can be? Merqur is reportedly raising now at a $20 billion valuation. Does that mean these providers will be worth $100 billion? That wouldn’t be egregious, would it, to say it’s 3% of the market cap of—
I think it could easily be $100 billion. I think these companies will easily be worth hundreds of billions of dollars. And I think they could even be worth more. Data is really the hardest part of model training because you need to source it. It’s so dirty. Nobody wants to do that shit. Nobody wants to hire all these people to generate data and then turn that into basically data plus GPUs equals model.
The algorithms have become somewhat of a commodity because people know how to use the Transformer. That’s why, as you said, all the people who are coming out of the frontier labs look the same.
Everyone shits on these data providers for the same reason. They go, “Oh, but the revenue concentration is just OpenAI, Anthropic, Meta, and a couple of other providers.” Is that a fair criticism, or does that actually not detract from the ultimate enterprise value of these data providers?
Yeah. I have 2 answers to this. The first is that I think Silicon Valley investors have become total bitches with respect to revenue concentration. It’s like, what the fuck are you talking about? TSMC has revenue concentration. There are many-hundreds-of-billions-of-dollar public-market businesses that have revenue concentration. I don’t know what we’re talking about here.
There are businesses that are 2-customer businesses. There are businesses selling to the government that have huge revenue concentration—there’s, like, one of those. They’re making huge amounts of money, like Anduril. Hugely revenue-concentrated businesses, and those businesses are doing great.
Are you suggesting that venture investors have a propensity to be lazy?
I would never say that.
I would never go that far.
I can let you know it’s incredibly tiring sending you an email. “Did you know that this competitor’s just released a product?” Thank you.
Totally.
Sent from Portofino.
That’s my one: I think we need to have some venture investors who suck it up and put some salt on their martini glass.
If you knew venture in 2026, dude, you’d know that we wear a Whoop and we don’t drink martinis because it impacts our sleep score. But okay.
Okay.
Yeah.
Yeah, totally.
Yeah.
Eight Sleep and all that stuff.
Exactly. Okay, so that’s one: we’ve become totally wusses around revenue concentration. We should embrace it.
It’s okay. And the second thing, I think that a lot of data businesses are going to expand into enterprises. Of course, we plan on doing this as an evaluation business, going to enterprises and helping them build their own AI models and all this routing stuff because we have the intelligence layer behind it that we’ve built on Arena.
This is obviously a place that we’re going to, but many data businesses will go here as well. The idea is that, in a world where every business needs its own AI model, why shouldn’t every business need its own data? Of course they will, and the data will be part of the moat that their business accrues.
8. The Evaluation Bottleneck
So that’s on the data side. When we think about the agent side, Anjani said I had to ask you: how does your business change as we think about the transition to full trust with agents?
Yeah. Agents are the number one priority for Arena and have been all year. People don’t know this, but Arena is one of the largest consumer AI apps in the world. We’re bigger than xAI. We’re bigger than Huggy Face, Manus, and Genspark. It’s so massive.
Outside-in, 30-plus million monthly visitors are on Arena, and most of them are knowledge workers and prosumers—people that we call unhirable experts, people who are coming to Arena to do their real daily tasks. In doing so, they are giving feedback that allows us to build the evaluations that we share with the world. It’s this organic flywheel for agentic evaluations based on real data.
Why didn’t you build a data business?
Well, we built an evaluation business around this that allows people to understand the strengths and weaknesses of models and therefore improve them. The labs can improve their models based on the insights and data that we give them, but we also want to help businesses with this.
Do you think the evaluation business is better than the data business?
I think every business in the world is going to need evaluation, unambiguously, and that is the single biggest bottleneck to deploying AI because people don’t understand how to define value. All this stuff around cost per value—it’s like, how do you define value? It’s easy to cut costs. I can tell you to go use Gemini Flash, and that’s going to be way more efficient in terms of token spend.
Isn’t value entirely subjective? For one, it’s speed, and for another, it’s accuracy. Do you know what I mean?
Right. Absolutely. So you can try to decompose it. I think about it as a 3-pronged value proposition. There’s performance, and then there’s cost and latency. Cost and latency are easier to define, but performance is the tough one because the definition of performance depends on the business and depends on the use case.
At Arena, we’ve built this pretty sophisticated pipeline for extracting organic performance measurements from agentic traces, and that’s exactly where I would say the value lies: in helping businesses take advantage of their own data instead of having to purchase data in order to say which AI works best for them, and even helping them train their own models.
What sort of revenue range are you at now?
We’re past $100 million in annualized revenue run rate, and that’s based on Q2 times 4. We’re growing really, really fast on that front.
Direct question, then: how efficient are you at monetization if you have 30 million amazing users who are unbelievably valuable in many respects and you’re only doing $100 million?
You’re asking about margins.
Yeah, and speed of ramp—is that good?
I mean, obviously, we’re not a free-cash-flow-positive business yet. We’re still investing all the money that we get into making sure that we continue our rapid growth and that we have a great product for all of our users and so on. But the fundamentals of the business are pretty strong. We feel great. Our investors feel great about our margins.
Yeah, I’m sure they do. I would love to have been an investor. I really feel like you excluded me. You know, I could be Greek for you for this deal.
Really?
Yeah, yeah. I’m a venture investor. We can be very plastic.
Kalimera. Kalimera.
Kalimera, hummus. Yes.
Hummus and pita.
See? See? This is—we’re going to get it.
We are already Greeks together, okay?
I knew that this would be a productive session. Yes. Are investors over-rotating on margin also?
I don’t know. I actually think margins are pretty important.
We’re seeing a load of businesses like Fireworks AI, where they’re at the 30%-style, mid-30s margin base, and that’s very different from software margins that were 65% to 80%.
Yeah. I mean, listen, profit is just margin times volume, and so you have to look at that as the calculation for the business. It’s not super crazy. So I don’t think it’s crazy to invest in these businesses.
The bigger problem with businesses like that I see these days is that a lot of them are fundamentally GMV businesses, where there’s some reselling happening. I’m reselling tokens. I’m reselling GPUs and stuff like that. Those businesses are tough because, at the end of the day, you have to think about not just the margin that you’re charging in the short to medium term, but the terminal value of the good that you’re providing to your customer.
If the terminal value of the good is, “I’m going to host GPUs for you in order to run your models,” then why should I pay you more than the cost of the electricity it takes to run those GPUs? The price-to-value thing is where I think you start getting into questions.
That’s why I think a margin question is very important. I’m not saying the margin in the short term is serious. A Seed, Series A, or Series B company might not have the best margins in the world, but you should be thinking about, as this business scales and moves toward becoming a public company, whether it’s going to have a fantastic margin structure that supports a public business.
One thing that’s challenging is when your customer becomes your competitor. To what extent do you think we’ll see the model providers move into the application layer aggressively? We see Claude Design has actually really started to eat away at Figma, and I’m an investor in Lagora. People are like, “Oh, don’t worry about Harvey.” Not in any disrespectful way to Harvey. There are disclaimers and everything in between. Everyone at Anthropic is going to do a legal product that’s going to kill Harvey and Lagora.
Totally. Yeah. I mean, listen, ask every business in America how they feel about this.
Everybody’s shaking in their boots. I have friends who are running multibillion-dollar businesses, and then what happens is that the next day, one of their biggest customers comes to them and says, “Hey, listen, OpenAI is getting into this game. We want to work with them because they’re more AI-forward and you’re less AI-forward because you’re traditionally a SaaS business. So, goodbye.”
It’s happening. It’s absolutely happening, and I think businesses should take it really seriously. This feeds right into this AI sovereignty debate, because a lot of what they’re doing is, if I’m OpenAI and I’m Anthropic, I’m looking at who my biggest customers are and which customers are winning the most in the enterprise.
AI is going to commoditize, right? If inference is going to commoditize, then of course the next best thing is for the model providers to move up the application layer in order to own more of the application stack, ensure that they’re not commoditized, and get closer to the value they provide to the end customer. So I absolutely think it’s a risk. I think it’s a risk for Lagora, and I think it’s a risk for Harvey. That’s why Harvey’s CEO himself is saying that his biggest competitive worry is the model labs.
But then how do you— that’s a complete paradox to what we just said at the beginning about companies being scared to work with the frontier models, isn’t it?
No. I mean, they’re scared to work with them. That’s what I was saying.
They’re scared to work with them, and they’re embracing them at the same time?
Ah, you mean the customers of the—
Yeah, you just—
—the Harveys and the Lagoras?
You just said your friends running multibillion-dollar companies are like, “Oh, we want to work with OpenAI.” I thought we just said they’re scared to work with them.
That’s a good question. I think you see both in the market. It depends on who’s most automated.
So I think it depends actually on their GTM. If you are doing Anthropic design or Claude design, designers can pick up a tool and use it very efficiently. If you’re Lagora or Harvey, you’ve got to go into Cooley or Clifford Chance and build relationships with 50-year-old white male partners who want to play golf and be told that they’re great and that life is awesome.
Then you’ve got to do deployment to junior lawyers who don’t want to use you because they think you’re going to take their jobs, too. The deployment and the GTM is the heavy lifting, and that’s real-world.
Totally. There are also businesses that are less software-focused and more network-effect-focused or operations-focused, and I think those businesses are more likely to be adopters of the big labs.
Let’s say a system integrator like Infosys. I think it’s more likely to be an adopter of a big lab because labs are less likely to be competitive with Infosys than they are to be with some sort of scalable software product, like insurance claims automation. Or let’s say the Harvey model, a legal chatbot. That’s tough because I think a model lab can build that.
Do you think Salesforce will thrive in the next few years or be challenged?
Salesforce themselves have a pretty strong AI strategy, so I think those people are ready to go and fight in this race. I doubt that they’re going to go downhill. I think that the SaaSpocalypse has been a little bit overstated overall, because people don’t always understand the dynamics of those businesses and how tough it is to replicate what they’ve built, also from a network perspective and a data perspective. So we’ll see. We’ll see.
I get you. I think if you’re a ServiceNow or a Salesforce, it’s incredibly difficult and incredibly hard. I think if you’re a Wix—it’s less difficult, less integrated, less sticky, and tougher. It’s all about entrenchment within the enterprise. If you’re entrenched, golden. If not, be more nervous.
Totally.
Right, I’m going to do a quick-fire round with you. I’m going to say a statement, and you’re going to give me your immediate thoughts. Sound good?
Yes, sir.
9. The Future Quickfire
What have you changed your mind on in the last 12 months?
Open-source model leadership.
Unpack that.
I think open-source models are moving much faster than I initially thought. I also think Anthropic is moving much faster than I initially thought. The space is moving so fast.
What do you know now that you wish you’d known when you started Arena?
Managing people. Managing people is just the most important part of running a company. The technical stuff—I did my whole PhD on it. I spent my whole PhD proving theorems in a basement, which I loved, by the way. It was a great time.
Now it’s all about strategy, people, and forecasting the future: being able to look 6 months, a year, or 2 years in advance and then try to plan for that. Those are so, so important skills.
Does it make sense for great, talented young people to still go to university?
I think it’s ever more important for people to have a strong mind, and the university can be a place to develop a strong mind in terms of strong first-principles thinking, as well as getting to know other people and networking with them.
I think that university is still a good place to go if you want to have an intellectual life, meaning that the intellectual work you do is the primary driver of your professional career.
What did you do with Arena that, with the benefit of hindsight, you wish you hadn’t done?
I had so many mistakes. At the beginning, I had no idea what I was doing, and my co-founder, Jan, probably knew and could see around corners, but I was probably too stubborn to listen to him. So, first of all, I’ve learned to listen to Jan more.
Second, there were so many experiments at the beginning that I shouldn’t have wasted time on. I think the degree of focus that you need to run a company is so extreme. You really need to do 1, maybe 2 things extraordinarily well and focus very, very deeply on them.
Pick the right ones and focus on what’s working, not on expanding into things that are not working. That is a great lesson for me.
This is why I also agree with Lagora and Harvey. When doing legal isn’t the main course for Anthropic, I just think you’ve got a really hard business when it’s someone else’s appetizer and it’s the only thing you live and breathe.
Totally. Priority number 12 for Anthropic is probably not high enough for Harvey and Legora to be too scared.
I’m also like, Dario, will you please just solve cancer and climate change? Leave shareholder agreements to someone else.
Exactly.
I’m being serious.
You know what, though? Solving cancer is hard. It’s harder than legal.
100%, and that’s why Dario should solve it.
Well, that’s why he doesn’t want it, man. He just wants to take your bread. It’s easier.
Oh, come on, Dario. Please.
Come on.
Come on, dude.
Leave some bread for the rest of us.
So which company will be first to $10 trillion: Nvidia, OpenAI, or Anthropic?
I think it’s hard to say it won’t be Nvidia. I think Nvidia is probably in the lead there.
Why has Nvidia not benefited from the rise of open source? I’m an Nvidia holder, and I’m seeing flat. Why?
Well, I think the market probably hasn’t priced it in yet. We’ll see how good these models get. But I think the enterprise adoption of AI is going to be another 10x-er for the industry, and I think it’ll 10x Nvidia very reliably.
Do you worry about the compute debt cycle and the levels of debt being taken out to fund the compute build-out?
I do worry about that, and I think the reason to be worried is because if the open-source ecosystem somehow makes the cost-saving opportunity for businesses much more salient, and therefore decreases the revenue of companies like OpenAI and Anthropic within the enterprise, it could lead to insolvency.
I think that is the big secular trend that I would worry about if I were an investor in those markets.
My worry is that we’ve never had such reliance on 2 companies to continue to hit their targets. If OpenAI and Anthropic do not continue on the trajectory they’re on, the music and the party go off. If the music goes off, for everyone in the Fireworks layer, there’s no party. The routing layer, no party.
Everyone suddenly just gets the wind knocked out of them by the trajectory of 2 companies.
Totally. I think that it’s a really big deal. We could use a little bit of sobering up within our industry. I think there’s a lot of hype. I think there’s too much crap happening for my taste, and I would prefer a little bit of consolidation, actually, so we see what shakes out.
I think Arena will shake out as a winner in our category, and I would love to see some of the great people at other businesses in our area consolidate to Arena, so we’re able to hire them in.
Where is the industry underhyped? Where is it overhyped?
Well, it’s interesting. I feel like everything is so hyped right now.
I feel like the mechanical infrastructure for compute and data centers is relatively underhyped.
The actual cooling systems and the actual steel infrastructure—the real physical infrastructure—are still under-hyped.
Interesting. Yeah, you probably know more than me. You're in touch with the investing markets. I know that people are super hyped up about all of the high-bandwidth memory and the GPUs and all that stuff. That stuff is super-ultra-hyped, right? Basically at all stages, from public-market companies to early-stage.
South Korea has fucking called a national convening, like, community meeting today because their stock markets are down 40%.
Oh, my God.
A national meeting because stock markets—
Down 40%? Why are they down 40%?
If you're a public-markets investor in South Korea, you're coming home a little bit stressed today.
No, that's not good for them.
Yeah.
Yeah, let's all pray for the South Koreans.
The thing I am slightly amused by is that right after everyone at SK Hynix and Samsung took home mega bonuses, the market crashed.
So why did it crash like that? What's the deal?
Honestly, I think it's just a realization that everything was pretty overinflated, and markets can't keep ripping for so long. There's no destabilizing factor within open or closed models that suggests demand is being questioned.
Wow. Okay.
That's why we should have a hedge fund manager on. We could do a new show hosted by Anastasios and Harry.
Yes, absolutely.
Called Two Dumb Rocks.
Let's do it. It's exclusively us and hedge fund managers.
I think it's a fucking great idea. I actually do too. Guest one is Anthony Midha. “Anthony, will you help Two Dumb Rocks?” He's like—
Ah.
“Why did I fucking put this together? This is not my cup of tea.”
No, Ant would be the best guest.
What's the most underrated neolab, other than Periodic, that people aren't talking about?
Ooh, underrated neolab. Yeah, I don't know if I have one. I think a lot of them are overrated.
I think Black Forest Labs is pretty underrated.
BFL is great. Would you consider them a neolab?
Oh, don't get technical with me on semantics.
Yeah, I don't know. BFL's awesome.
Final one for you. What are you most excited about? My mom's got MS. I'm fucking excited that chronic conditions like MS could maybe be treated. What are you excited about over the next 5 to 10 years?
I've always been a big proponent of AI in medicine, too. I think the level of human flourishing that's going to happen as we start to eradicate diseases one by one, the same way that we're currently eradicating open problems in math, is going to be incredible. I think it's going to be tough, because math is a closed system.
In medicine, I think you'll need to figure out ways of quickly iterating in a feedback loop on biological systems. That's the missing piece, but once we crack that, it's going to be just an extraordinary journey.
It's so funny. When I interviewed Damas and spoke about biology and medicine, it was an area where you could see his eyes light up. But it was an area where I said, “Hey, testing needs to change.” Fifteen years, no bueno for a lot of sufferers of chronic conditions.
Yeah. And you know what's missing? That is exactly the data layer. That's exactly one of the areas where you can clearly see that the data layer is where value's going to accrue. The GPUs are the same GPUs in both cases. The problem is that the data infrastructure, the flywheel, the data collection that you need in order to build a great biology product or a medicine product, that's tough to build.
Dude, you've been an epic guest. Really. I'm so grateful. It's been an amazing show, with real honesty and authenticity. Most people suck as guests. You know why? Because they're not authentic. And it just comes across. Thank you for being so great.
I appreciate it. No, thank you for having me on. I would love to do it again at some point. You should visit the Arena office any time that you're in the Bay Area.