四位CEO论AI未来:CoreWeave、Perplexity、Mistral与IREN
Michael Intrator × Aravind Srinivas × Arthur Mensch × Daniel Roberts
CoreWeave的核心判断是,AI算力到了集群规模就会脱离商品化,而合同现金流的持续时间远超市场担心的芯片周期。 Michael Intrator称“16个月左右GPU就过时”的说法是“无稽之谈”:公司平均合同期为5年,设备按6年折旧,并预计实际使用寿命会更长;他还表示,A100价格在过去一年上涨。只有当数据中心的电力转投其他基础设施、能够赚取更高毛利时,硬件才会真正过时。
CoreWeave的增长引擎是结构化融资,而不是押注无合同覆盖的GPU。 每个独立的“盒子”都包含客户合同、GPU和数据中心协议;客户付款先覆盖机房、电力、利息和本金,剩余现金才回到CoreWeave。这一结构支持公司在18个月内完成350亿美元融资,在一份5年合同签订后的2½年内覆盖全部成本,并帮助公司将资本成本降低600个基点。
基础设施需求仍高于实体供给,但瓶颈如今横跨内存、电力、光学器件、网络和施工,而不只是NVIDIA的芯片配额。 Intrator称,4年的需求已经“压垮全球计算产能”;IREN的Daniel Roberts则表示,“全球数据中心里没有闲置GPU”。长期合同保护CoreWeave免受需求真空冲击;IREN的优势在于4.5 GW电力,以及通过提前8年锁定土地和电力建立的先发优势。
Perplexity押注,价值会从某一个前沿模型迁移到中立的协调层,由它负责选择和调度不同模型。 Aravind Srinivas称公司是“瑞士”:GPT、Gemini、Claude、Kimi、Nemotron和Qwen可以各自专长,Perplexity负责自动路由任务;Model Council则解释多个模型在哪些地方一致、分歧。公司服务数千万消费者,称企业收入增速更快;尽管整体尚未盈利,但“每一分钱”收入都实现了正毛利率。
Perplexity的路线图,是把AI从回答问题的输入框变成电脑,最终变成操作系统。 Personal Computer将以Mac mini作为私有本地运行环境,在获得许可的情况下把前沿模型调用或长时间任务交给与用户隔离的服务器;Srinivas称其为“给小白用的OpenClaw”(“Open Claw for dummies”)。界面从目标而非程序出发,Linux、文件、模型、连接器和沙箱则位于编排层之下。
短期内真正被颠覆的将是定制软件和自动化后台工作,而不是某一个万能聊天机器人。 Perplexity Computer已经产出了公司董事会备忘录、合作伙伴演示文稿和媒体简报;Srinivas更长期的目标,是让AI替一家小企业投放广告、处理客服、收款和开发功能,而企业主可以“在Napa喝着葡萄酒”。他强调,这一未来“还没有到来”,也承认就业会出现阶段性替代,同时认为AI可能把主动权交还给那些不喜欢传统工作的人。
Mistral的企业路线结合了开放模型、人类专家信号和严格的执行控制。 Arthur Mensch表示,合成数据可以帮助小模型预热和压缩,“但最终仍然需要人类信号”;因此,Mistral会把可移植的训练工具和博士级工程师部署进客户基础设施,让数据无需回流。对于生产环境中的智能体,OpenClaw式自主性若没有确定性闸门、沙箱、基于角色的权限控制,以及防止薪酬等受限数据在组织内泄露的“上下文引擎”,仍然不够。
IREN的传统电力资产已经成为稀缺的AI资产,一份97亿美元的Microsoft合同只消耗其5%的容量。 其位于得州的750 MW旗舰项目属于4.5 GW电力组合;公司自成立以来一直使用100%可再生能源,包括不列颠哥伦比亚省水电,以及西得州风电和太阳能。Roberts的需求模型基于杰文斯悖论:如果10倍算力把图像生成时间从几分钟缩短到5-10秒,用户会生成更多图像,而不是把效率收益存起来。
1. CoreWeave先买下算力期权,再出售AI云
Intrator回溯称,CoreWeave起源于一家天然气算法对冲基金,其团队在业务空档期开始关注加密货币。他们放弃Bitcoin ASIC,是因为芯片设计商很可能拥有运营优势;但GPU既能挖Ethereum,也能服务其他工作负载,因此公司从一开始就把“算力当作期权”。
公司成立于2017年,前后挖矿约3年;对冲基金式的风险纪律帮助其穿越了多轮加密寒冬。随后,CoreWeave沿着工作负载栈从CGI渲染扩展到批处理计算和医学研究,寻找比本身高度波动的加密需求更稳定的用途。
真正决定性的学费发生在2020-21年:CoreWeave买入A100,并向EleutherAI捐出算力,同时学习神经网络基础设施。免费算力意味着,研究人员“即便我们一开始做得不够好,也不太会真的对我们生气”;当这些志愿者回到本职工作后,他们要求继续使用同样的基础设施,CoreWeave由此找到了商业切入口。
2. 算力到了集群规模,就不再是商品
在ChatGPT出现之前,规模定律已经改变了CoreWeave的定位。Intrator的区分很直接:“任何人都能运行一块GPU,但你能否运行足够大的集群,训练出一个足以改变世界的模型?”用他的话说,“算力在规模化后会脱离商品化”。
CoreWeave刻意占据“位于NVIDIA GPU之上、模型之下”的一层,为单一的定制化工作负载整合软件、运营和可观测性。Intrator将其与AWS对比:传统云计算把一个问题解决得很好,而加速AI需要一套全新的解决方案。
EleutherAI之后,公司与Mustafa参与的Inflection合作,随后客户扩展至超大规模云厂商、OpenAI及其他基础模型公司。训练业务带来了最初的收入;如今,产品化AI正在把推理从组织边缘推向核心工作流。
3. 推理实现AI变现,也延长每一代GPU的生命周期
Intrator称,推理是“AI投资的变现”:用户提出问题,或要求模型采取行动,正是模型能力跨入现实价值的环节。因此,他把CoreWeave基础设施上的推理量,视为衡量整个生态健康度的直接指标。
CoreWeave称其率先完成H100、H200和GB200的大规模商业部署,并将GB300确定为下一代架构。芯片的生命周期并不是训练一次就丢弃:最前沿的芯片先用于训练模型,再转入实验,随后长期服务于推理,“非常、非常长时间”都能产生价值。
主持人的追问引出了公开市场的看空逻辑:架构快速迭代必然缩短资产寿命。Intrator否定了这一前提:市场对最新芯片的需求增长,并不会消灭实验、渲染、小模型和推理等长尾需求,而这些工作负载仍可由旧架构承载。
4. 5年合同击穿GPU短生命周期论
Intrator称,折旧争议是持有空头仓位的交易员炒作的“无稽之谈”。CoreWeave客户签订5年或6年合同,平均期限为5年;声称设备16个月就会过时的说法,“与现场事实完全对不上”。
CoreWeave按6年折旧,并预计GPU运行时间会超过6年。更值得注意的是,Intrator表示,A100/Ampere价格在过去一年上涨,因为新公司和新工作流吸收了那些既无法证明最新H100级系统经济性、也无法获得其供应的算力。
主持人用iPhone类比说明转售逻辑:iPhone 12对原用户而言可能已经过时,但在另一个市场仍保有可观价值。同理,老GPU可以转给不追求最前沿性能的客户、迁移到另一座数据中心或其他地区,而不会在经济意义上凭空消失。
Intrator真正的过时标准是机会成本:当数据中心受限的电力,通过替换后的基础设施能够赚取更高利润时,旧设备才算过时。在此之前,老旧基础设施仍然“极其赚钱”;CoreWeave当时仍能以高于前一年的价格出售可用的旧算力。
5. “盒子”把客户合同变成更便宜的基础设施债务
CoreWeave先取得一份高信用质量客户合同,Intrator以Microsoft为例,然后把合同、采购的GPU和数据中心协议一起放进名为“盒子”的融资结构。客户向盒子付款,而不是直接向CoreWeave付款。
现金瀑布的第一顺位是支付数据中心费用和电费,随后支付利息和本金,剩余现金才回到CoreWeave。因此,贷款方承保的是有合同支撑的客户应收款和实物抵押品,而不是对未来现货需求持续高涨的笼统承诺。
Intrator表示,这一结构让一家当时并不知名的公司在18个月内融资350亿美元。在一份5年合同签订后的2½年内,盒子就已经覆盖包括本金和利息在内的全部成本,同时开始为CoreWeave创造回报。
每个盒子都是独立交易,因此单笔交易失败时不会轻易形成传染。随着交易持续兑现,CoreWeave在2年内将资本成本降低了600个基点;但这要求公司保持纪律,Intrator会拒绝1年的GPU订单,因为期限短到不足以摊销建设成本。
6. 稀缺性已经从GPU扩散到整条供应链
Intrator称,需求连续4年保持“势不可挡”,并持续压过全球算力供给。但CoreWeave仍然会为技术、战争或其他冲击导致的需求真空做建模;风险管理并不关心具体原因,5年合同和高质量交易对手才是保护措施。
GPU只是一个节流阀。电源机架、内存、存储、网络和光学器件都可能限制部署;Intrator指出当前瓶颈是内存——AI需求飙升与晶圆厂投资错位碰撞,而后者可能早在2023年就应该启动。
主持人将由此形成的繁荣—萧条周期同时视为破坏性和创造性力量:它会“清理灌木丛”,奖励幸存者,并留下基础设施。他把光纤过度建设与YouTube联系起来:免费上传内容受益于存储和带宽成本大幅下降。
主持人还引用了OpenAI CFO Sarah Friar的对比:GPT-3问世时,100万tokens的成本为32美元出头,如今已经降至9美分。Intrator认同,单位成本下降说明资本市场、资本主义和工程技术如何推动竞争,降低曾让人类创造力“被困住”的门槛,并为可能达到80亿人的群体提供从想法出发创造的工具。
7. Perplexity把答案准确性变成一台电脑
Srinivas将Perplexity Ask、Comet和Computer描述为一条信任度递进路径:接入互联网提升答案准确性;获得完整浏览器权限提升任务准确性;如今获得完整电脑权限后,AI可以执行用户原本需要手动完成的同一套数字工作。
他的标志性比喻是一支管弦乐队:数百个专业模型是乐器,子智能体是演奏者,Perplexity是指挥。编码、写作、图像、视频和音频能力只有在组合成“你演奏的音乐”——也就是最终交付成果——时才真正重要。
Personal Computer会让Perplexity与一台充当本地服务器的Mac mini同步。涉及私人数据的编排可以留在本地硬件上;获得许可的前沿模型调用或复杂的长时间任务,则可以转移到服务器端的一台电脑,并且只有该用户能够访问。
Srinivas将打包后的体验称为“给小白用的OpenClaw”(“Open Claw for dummies”):一个可执行文件,无需管理API密钥,无需为约100项服务分别计费,也无需手动穿越复杂的权限迷宫。战略选择是端到端整合,而不是把每个组件都暴露给用户。
8. 本地AI变成家电,AI则变成操作系统
主持人称,他曾在Mac Studio上本地运行Kimi 2.5,免费获得大约80%的前沿模型能力;他还提到一台配备750 GB内存的Dell-NVIDIA工作站。问题在于,一台1万美元的桌面电脑能否替代部分每月500美元的云账单,同时提升隐私性。
Srinivas预计,本地模型最初会作为处理报税、照片、邮件、日历和笔记等隐私敏感任务的子智能体。他拒绝把问题简单化为本地与云端的二选一:Google Workspace数据本来就存储在服务器上,而手机用户大多只关心任务是否完成,并不在意获得许可的机器究竟是哪一台。
长期来看,他预计家用AI服务器会像“买一台冰箱”或买一个互联网调制解调器一样普及,并负责协调家庭传感器。它的操作系统从目标而非指令开始;Linux、文件、模型、连接器和代码沙箱都位于AI抽象层之下,由AI决定如何执行。
9. 模型专业化是Perplexity的“瑞士”护城河
Perplexity拥有数千万月活用户,企业收入增速快于消费者业务。Enterprise Pro的价格为每用户每月40美元;Enterprise Max为400美元,超出包含的Computer额度后按使用量收费。Srinivas称,Max客户合计节省了超过1亿美元。
他的经济模型解释非常具体:Perplexity把订阅、模型路由、检索和搜索结合起来,而不是不加区分地转售最大上下文tokens,因此每一种收入来源都拥有正毛利率。但公司本身“目前仍未盈利”。
面对收购传闻以及资金预算远大于自己的竞争对手,这家约400人的公司给出的答案是中立性:“我们就像瑞士。”Kimi、Nemotron和Qwen可以与GPT、Gemini或Claude一起在底层运行,任何一个模型家族胜出,都不会让Perplexity的应用层变成沉没资产。
专业化进一步巩固了这一定位。Srinivas称,Perplexity的iOS工程师喜欢使用Codex,后端工程师则喜欢Claude Code;产品会自动为每个提示词路由模型,同时保留手动选择。Model Council更进一步,会比较多个答案,指出其中准确的一致点、分歧点和细微差别。
10. 更快交付,指向按需生成的软件
Srinivas称,速度是Perplexity的工作模式,前提是质量和信任能够得到保留。在Perplexity内部,即使不是工程师的员工,如今也会让Computer Slack机器人修复Bug。
Computer已经准备了Perplexity的董事会备忘录,一次性完成合作伙伴演示文稿,并取代了一场过去需要员工投入的传播简报。随着产品记忆能力增强,这种杠杆效应还会累积:它可以调取过往会议、演示文稿和合作关系,而不必每次从头重建上下文。
Srinivas将长期编排变得可靠归功于沙箱、终端、文件、子智能体、技能和命令行工具。模型只需加载必要上下文,之后即可丢弃,因此上下文窗口限制不再是决定性约束。
Srinivas称,他曾让智能体下载每一期All-In,提取其中提到的上市公司,绘制提及频率随时间变化的图表,分析股价影响和情绪,并提供链接回相关音频的时间戳。主持人另行描述了自己在被“Claude-pilled”之后使用一个智能体:它主动提出搭建定制CRM,而不只是导出电子表格。
11. 自动化企业可行,但还不是开箱即用
Srinivas区分了“一人公司创造10亿美元价值”的口号与真正新增的GDP:用AI替代研究人员,不等于创造10亿美元价值。他的目标更小也更具体——帮助一个人买一台Mac mini,经营一家年产出数十万或数百万美元的企业。
设想中的Computer会运行Instagram和Google广告,接入SEM与SEO工具,寻找用户,通过Stripe向用户收费,发布功能,并通过Intercom处理客服,而企业主则“在Napa喝着葡萄酒”。他的限定条件非常明确:“所有人都认为AI已经走到这一步了。其实还没有。”
Srinivas给自己的前景判断是70-80%高度乐观,20%担心就业会快速流失。他接受阶段性就业替代,但认为很多人并不喜欢传统工作;理想终局是让个人拥有主动权、所有权并创办小企业,前提是人们保持主动、坚韧和解决问题的能力。
12. 付费浏览器访问补上智能体缺失的代理层
Computer已经集成在Perplexity应用内,iOS版Comet则把浏览器体验带到了移动端。浏览器在战略上仍然不可替代:在网站围绕API和命令行工具完成组织之前,智能体仍需打开标签页、填写表单、点击控件和上传文件。原生浏览器控制能力让Perplexity可以自动化处理这些剩余的网页工作。
主持人追问了Reddit、LinkedIn等网站逐渐抵制自动化活动的问题,并提出付费、认证式访问。他设想的是有边界的行为——总结用户已订阅的内容,或找出7名员工,同时设定明确配额,不允许发帖或投票——而不是无限制抓取。
Srinivas不愿讨论具体谈判,但认可官方API和基于用户选择的“双赢”模式。主持人进一步将这一模式延伸至付费新闻订阅:智能体访问可以为出版商带来增量收入,也能让订阅更有用、更具黏性。
13. Mistral用开放模型把企业IP带进模型
Mensch宣布,Mistral将与NVIDIA共同训练下一代前沿模型,延续约18个月前与Mistral NeMo启动的合作。目标是打造领先的开源模型,再由Forge针对工程、物理、科学、金融服务和政府语种进行专业化。
Mistral虽然来自欧洲,但业务并不局限于欧洲:Mensch称,公司25%的业务和25%的研究人员位于美国,同时在法国、英国和新加坡设有团队。欧洲拥有更多制造业场景,也有一批在AI采用上落后后、希望实现跨越式追赶的企业。
通用模型对于编排仍不可或缺,但企业拥有数十年积累的知识产权和物理系统信号,封闭模型既无法深入吸收,也无法透明地利用这些资产。开放模型允许企业添加新参数、构建定制化工具链,并部署在任意云端、客户硬件或边缘设备上。
Forge把模型定制与前置部署工程结合起来,Studio则负责构建端到端智能体。Mensch主张的价值在于更低成本和更强控制力,尤其适用于ASML这类企业:其专有数据可以让模型在关键工业流程中获得独特能力。
14. 人类信号与确定性控制,把企业智能体和演示Demo区分开
Mistral的平台可以移植到客户自己的基础设施中,因此“不会有数据流回Mistral”。博士级工程师会与领域专家合作,处理图像扫描、缺陷检测等任务,再把足够的知识转移给客户,让客户能够自行重新训练,而无需永久依赖Mistral。
合成数据适合用于模型预热和压缩:大型教师模型可以高效生成训练材料,供更小的模型学习。但Mensch的边界判断很明确——“最终仍然需要人类信号”,即便专家反馈的获取成本很高。
OpenClaw展示了个人层面的自主性,但Mensch认为它缺少企业级基础能力。以HSBC的KYC流程为例,确定性闸门必须始终执行、保持可观测,并且能够向管理层作出保证;生产环境因此需要控制平面、沙箱和严格执行的访问控制。
主持人用薪酬数据举例,说明如果智能体对Gmail和Slack拥有root权限会产生什么风险。Mensch的回答是“上下文引擎”:将数据和元数据映射到角色,对未经授权的工程师薪酬查询直接拒绝;一旦这些信息流发生改变,管理层级和客户服务组织也可能需要重新设计。
15. IREN借Bitcoin为转向AI融资
Roberts称,IREN成立时的判断是,指数级增长的数字需求最终会与物理世界发生碰撞。Bitcoin挖矿是公司用来启动土地、电力和数据中心建设的第一种现金流工作负载,直到“更高、更好的用途”出现。
2020年与Dell签署的谅解备忘录曾是一次虚假的曙光:当时引入客户和算力还为时过早,IREN于是回到Bitcoin。大约在访谈前2年,需求开始变得切实可见,此后每月加速;公司如今正在“把所有Bitcoin换成AI芯片”。
Microsoft在上一年年底签署了一份97亿美元合同,但Roberts称该合同只消耗IREN容量的5%。公司自行开发站点,负责土地、许可、电网接入和建设;其得州旗舰项目为750 MW,整体电力组合达到4.5 GW,按Roberts的比较,接近整个湾区一年的用电量。
由于IREN在8年前就开始锁定土地和电力,Roberts称公司如今的约束已经变成“算力上线所需的时间”。施工劳动力、散热、内存和供应链形成“永久打地鼠”局面:数千人试图把偏远地区的实体基础设施转化为可运行的集群。
16. 电力地理与杰文斯悖论决定IREN下一轮建设周期
IREN称公司自成立以来一直使用100%可再生能源,包括不列颠哥伦比亚省水电,以及西得州风电和太阳能。西得州的机会在于地理错配:当地拥有45-50 GW可再生能源,但通往达拉斯和休斯敦的输电容量只有12 GW,因此算力被部署到过剩电力所在地,再以数字商品形式输出。
主持人原本预计电池会平滑间歇性供电,但Roberts表示IREN不需要电池:一旦拿到稀缺的电网接入,公用事业公司就能保证全天候、每周7天的可靠电力。在当地,IREN以每20英里为一档向外扩展招聘,累计发放了100万美元社区补助,并与大学和职业院校合作;Roberts称,主持人估计技术工种薪资为150,000-300,000美元,其中低端水平大体正确。
Roberts称需求“势不可挡”,并明确表示:“全球数据中心里没有闲置GPU。”效率提升反而会强化消费:如果10倍算力把图像生成从几分钟缩短到5-10秒,用户会生成更多图像——这是杰文斯悖论,而不是需求被摧毁。
短期来看,NVIDIA仍是最稳妥的路线图,原因在于其生态和标准;但定制芯片已经开始争取数据中心容量。核电可能扩大供给,但大概率需要10年甚至更久;网络则是眼前的问题,IREN西得州站点往返达拉斯的延迟为6 ms。太空数据中心仍受发射成本、辐射和工程难度限制。
One of the great companies of the AI era is, of course, CoreWeave. They’re building massive infrastructure for these hyperscalers, and in some ways, Michael Intrator, welcome to the program. You’re the original hyperscaler. You guys got in very early and secured your—I don’t know which GPUs you wound up getting—but you were very early to this trend. How did you get to it so early, and how did you build out this first, I guess at the time, neocloud?
We didn’t really start it as a neocloud. I was running an algorithmic hedge fund focused on natural gas, and when you build an algorithmic hedge fund, once the algorithms are built, you’re really just monitoring them, testing different pieces, and doing all that. But there’s also a lot of downtime, and we got super interested in crypto.
We’re pretty nerdy. We dig under the hood, and we started to get interested in the security layer. We looked at Bitcoin and Bitcoin mining, and we didn’t like it. We thought there was some brilliant engineer who built the ASIC, and they were probably going to be better at running it than we were.
We really began to focus on GPUs, mostly because you could mine Ethereum with them, but you could also do all these other things. Right from the start, we looked at compute as an option to deploy our computing power to different use cases.
We began the company in 2017 and spent the first 3 years mining crypto. We went through a couple of crypto winters, and because we had come from a hedge fund, we had real chops in risk management and how we thought about capital, risk exposure, allocation, and all of that. We were really careful around that right from the start.
We weathered crypto winter really well and began to scale the company. We immediately started to look for other use cases for this compute because crypto was pretty volatile.
Yeah, and crypto was a question mark at that time.
Absolutely. Bitcoin was speculative, and there were many other speculative projects. The only other people using this type of hardware were quants and medical researchers.
A good way to think about it is the progression of products that we started to work on. First was crypto, but we immediately moved from crypto to CGI rendering. We built projects that allowed folks who were trying to animate and render images—the things that make movies cool—to do that.
Then we moved to batch computing and started to look at medical research and different ways of using compute to drive science. We just kept moving up the stack in terms of the complexity of how GPUs could be used.
Ultimately, around 2020 or 2021, we started to figure out how GPUs could be used for neural networks. That wasn’t something we knew how to do, so we went out and bought a bunch of A100s and donated them to a group that was working on Luther AI. They were working on an open-source project, and the thought was that since these guys were taking the GPU compute as a donation, they couldn’t really get pissed at us if we weren’t very good at it initially.
They couldn’t complain about the SLA.
They kept telling us, “We need more of this. You’ve got to work on this.” That began to give us an understanding of what was necessary to run scalable, parallelized computing.
I feel like buying those initial GPUs was the tuition we paid to learn how to run this business. One of the interesting things is that all of those guys went back to their day jobs because they were all volunteers working on this. They were like-minded scientists.
When they got to their day jobs, they were all like, “I want that infrastructure. It’s built the right way. That’s the way research is going to want to use it.” That launched our business.
It was an amazing story. You went from crypto to these researchers, into academia and deep research. What’s the next card to turn over in the poker game?
What became very clear to us very early on was that the scaling laws were going to drive this. Remember, this was back in 2020 and 2021, before the ChatGPT moment occurred.
We began to understand that compute decommoditizes at scale. Anybody can run a GPU, but can you run a cluster that’s large enough to train a model that can change the world? That’s a different question.
We began to think about how to scale up our delivery of this computing to larger and larger clients. That was the next card to turn: thinking about how there was a component of this that would lean into our ability to access the capital necessary to deliver our solution to the broadest possible audience, to the most sophisticated consumers of this compute.
The next card was thinking about it as a business rather than as an engineering project: how to deliver the infrastructure and the software, and really everything in between. When you’re thinking about what we do, we live above the NVIDIA GPUs but below the models.
Everything in there—the software, the integration of software and operations, observability, and all the things you need to build a cloud that’s purpose-built for this one specific use case—is what we focus on. We don’t do everything. We really focus on one use case.
Web servers are different. You’ve got AWS. They do a great job. It’s a great solution. It was a brilliant solution to solve a problem.
We just looked at it and said, “There’s a new problem. Let’s look at this problem and try to come up with a solution to deliver compute that solves it.”
When did the language models start dialing and calling you for capacity?
Our first language model was really EleutherAI. Our first large commercial customer was Inflection. We worked with Mustafa and Inflection, and then we diversified from there into the hyperscalers and OpenAI, across the foundation-model landscape.
We just kept scaling, with the belief that the decommoditization of compute and the ability to deliver a solution were going to matter. The solution is building supercomputers that can change the world. That’s really what we began to focus on.
That led into training, and now the world has gone through this moment where we’ve moved from research into the productization of this. It’s beginning to work its way in from the fringe of organizations into the core of what they do.
You can see that every day in the amount of inference compute being driven through our infrastructure layer, which is massive.
They’re consuming it, not just building models, but deploying and utilizing them.
I always think of inference as the monetization of the investment in artificial intelligence.
When we see our compute being used to stand up the massive scale of inference hitting our compute every day, inference is when people ask the model a question and it comes back with an answer. That’s an inference. Or when you ask the model a question and then ask it to go do something, that’s inference.
That’s where you have the opportunity to really drive value outside of the model itself and into the real world. That’s exciting for us. That’s what we like to watch, and what I like to watch in terms of gauging the health of the business.
What chips are those?
We are the tip of the spear in bringing the new architectures out of NVIDIA into commercial production at scale. We were the first ones to bring the H100s at scale. We were the first ones to bring the H200s at scale, the first ones with the GB200s, and now you’ve got the GB300s.
One of the things that’s amazing and really fascinating for us is that people are using the bleeding-edge GPUs to train models as the new architectures come out. Then they take those GPUs and move them into different experiments. Over time, they move them into inference, and they continue to use them in inference for a very, very long time.
What is the shelf life of an H100 right now? That’s been a big debate, I think, for your company and for Microsoft. I guess Michael Berry—you must have known him when you were a quant—has been saying, “Oh, my God, the whole industry is falling. The sky is falling.”
We all know in the industry that people don’t just throw this hardware away. They find uses for it. The market finds its own use for technology. So, what’s the reality of the lifespan of these things?
My take on the GPU depreciation debate is that it’s nonsense. It’s a debate being brought to the forefront by some traders who have a short position in the stock, and they’re trying to talk it down.
Look, here's what we know. When we buy infrastructure, we're a success-based company, right? We're a small company on a relative basis compared to the enormous companies that we're competing with. Our clients come to us and buy compute for 5 years, for 6 years. Our average contract is 5 years. So any commentary by anyone, either inside or outside of the industry, that this stuff becomes obsolete in 16 months or whatever nonsense they're spewing, it doesn't in any way match up with the facts on the ground. The fact on the ground is they're buying it for 5 years, right? My approach to this has always been: if people are willing to pay me for it, it still has value.
Correct.
Pretty simple way of approaching it. We use a 6-year depreciation. We believe that the GPUs will last in excess of 6 years, but we felt like that was a fair and reasonable approach to a technology cycle that's moving at this velocity. The A100s, the Ampere GPUs—this year, the price has appreciated through the year. Now, why is that? I think it's because, as more installed capacity becomes available, you have new companies that come into existence with new use cases and different-size models. They're trying to build new commercial ventures that maybe have been blocked out of the H100s and never had an opportunity to run on them.
To make a very simple example for the audience, when you trade in your iPhone after 3 or 4 years, you're like, “Who's going to use an iPhone 12?” And it's like, “Have you been to South America or Africa, where you go to the store and buy an iPhone 12 or the Pixel 7, and it costs $50? That's still got great life left in it.”
Absolutely. Yeah. Well, and so, look, we find these amazing use cases: new companies that have come into existence or existing companies that have integrated new models into their workflow and are able to use the Ampere GPUs. And so they keep buying any GPUs that we have available. Once again, the concept that a GPU is no longer relevant or commercially viable after 16 or 18 months or 2 years—
Yeah, as far as it goes, it just doesn't make any sense.
It goes as far as it goes. I think sometimes people get caught up in Moore's law or in just how fast our industry is growing, and that there's so much at stake that big companies are demanding the most recent products. That doesn't mean that the lifespan has gotten shorter; it means the opportunity and the surface area of the opportunity have gotten much larger.
The industry has gotten so much attention for the unprecedented scale of capital that's coming to bear on this. Because of that, there tends to be an incredible focus on the companies that are building on the most advanced chipsets. The truth of the matter is that even within those companies, they have a long tail of useful life to provide inference horsepower, work on other experiments, and do less bleeding-edge activity that still needs to be done. Rendering comes to mind as well. We're making images on Nano Banana. There will be a use for it.
There is a moment in time where maybe the compute-to-power ratio doesn't make sense. My expectation is that obsolescence will be defined by the moment in time when the power in the data center, for me, will be able to be repurposed for a higher margin than the existing infrastructure provides. As I said, I fully expect this infrastructure to last in excess of 6 years, but the standard in the space has really been 5 years, with 1 exception, which is Amazon, at 6 years. That seems like the right schedule. I'm not making it up. That's what everybody's using.
The energy cost is the opportunity cost because we need that space. There's a better reward here, and that hardware might get resold to somebody else who wants it—a hobbyist or something.
It could be sent someplace else where they have more capacity and can repurpose it there. I kind of feel like we'll deal with that part of the business when we get there.
What I know right now is that it is extraordinarily profitable. It's very accretive to my company to continue to keep the infrastructure that's been up and running and that's been on these long-term contracts. As it rolls off, as it's been in use for 5 years and becomes available, I'm still able to sell it at a higher price than it was at a year ago.
There's competition now. When you were buying these from Jensen back in the day, you could buy them and have them shipped, I would assume, within 30 days or less. Nowadays, what's the wait like, even for you, a loyal old customer? Is there a bit of a battle? Is there politics to who gets the servers? You see some very big names talking about how they have to get an allocation. Is it still a little bit crazy? What's it like to be in that category, having to buy something everybody wants?
Look, I think of it as an affirmation of the business that we're in, right? The fact that we are attracting competitors means that the business is healthy and that there's a lot of people trying to deliver this service. The need for this infrastructure, and the need to integrate the infrastructure into the software layers to deliver it to artificial intelligence—whether at the model level, the inference level, the application level, or whatever level of the 5-layer cake that Jensen's focused on—is growing. The fact that there are more people coming into this doesn't discourage me.
As far as getting access to the GPUs, we show up like everybody else with a PO: we'd like to buy, and we're ready to pay.
What's the wait time like? Is it just really competitive or not? Because I talked to Jensen about it. I said, “How do you manage all these big egos, names, and companies trying to buy stuff?” And he said, “Well, they order it, and we give it to them in the order in which they order it.” Is it really like that?
It really is, right? He doesn't want to be in the position of playing favorites. That just seems like a bad place to be with your clients.
Or auctioning them off.
Yeah, that would be crazy. I'm not sure that would be good for the long-term business.
No.
Yeah, so our approach is—
Get some sovereigns coming in and saying, “I'll pay double.” They do that with Ferraris, too, sometimes.
[Laughter]
Because these are the Ferraris of computing, right?
By the way, they are. Yeah, they're the Bugattis. Our approach is to work with clients across the entire space to find opportunities with really interesting companies that can fit into our contracting requirements, where we're going to be able to go out and structure the debt that we require in order to build infrastructure at this scale.
How does all that debt work? That is something that you guys specialize in. Corporate debt—I'm in the venture business. People are like, “Why should I be in venture when corporate debt pays so well?” Corporate paper is so huge. I'm curious how this fits in and what interest rate people are paying on $1 billion in infrastructure. What do they pay on that?
Yeah, so CoreWeave has really been the innovator around a lot of the financing engines that have come to bear on this. We did the first GPU-backed loans. I think it's important—or I'm going to try to explain this in a way people can understand.
What we do is go out and find a client. Let's use Microsoft—you brought them up before, right? Microsoft comes to us and says, “We'd like to buy something.” We say, “Okay, great. We're going to sign a contract.”
Once I have a contract in hand, what I do is create something. It's not a particularly creative name. It's called “the box.” What I do with the box is take my contract with Microsoft and put it in the box. I go to Jensen and buy the GPUs, and put them in the box. I take my data center contract and put it in the box. Now the box governs cash flow. It has a waterfall of cash flow that comes into it and goes out of it.
The way it works is that I build the compute and deliver it to Microsoft, and they pay the box. They don't pay me. It goes into the box, and the first thing it does is pay the data center. It pays the power bill. It pays the interest and the principal. Then whatever's left flows back to us, right?
It is an incredibly well-structured, time-tested, pressure-tested vehicle to borrow money against client paper and all of the other collateral around the deal. That's why CoreWeave, a company that many people haven't ever heard of, was able to go out and raise $35 billion in 18 months to build infrastructure at scale.
What's important to understand is that the economics in this box are such that within 2 and 1/2 years of a 5-year deal, we've paid for everything. The principal has been paid off, and the interest has been paid off. The return into the box is such that we're able to generate returns to our company at the box level, right?
That gives the most sophisticated lenders in the world—whether it's banks, private equity funds, or whoever—confidence that they're going to be able to achieve the 1 rule of lending: “Give me my money back.” Yes, and so it's better when that happens.
So they look at this box and they're like, “Wow, we're really confident we're going to get our money back.”
And maybe they want 10 boxes.
That's correct.
And if any one box goes upside down, you can deal with it, and it's not as acute.
That's correct. They don't cross-pollinate; they don't cause a contagion across the boxes. One, and number two, as you do this and as you show the lenders how this financing tool and how this financing mechanism works, what they do is they continue to lend you money at progressively lower rates.
When you think about our cost of capital over the last 2 years, we have dropped our cost of capital by 600 basis points. Wow. It is enormous, right? You're seeing a company that is driving its cost of capital down toward where the hyperscalers borrow, which will enable us to be competitive with them over time. We have been extremely militant and diligent about feeding, watering, and caring for those boxes so that we continue to have access to the capital markets in a way that allows us to build and drive our business.
That means you have to say no. You have to say no to maybe some people who want to be in the box?
Yeah, we look at some deals and we're just like—they want to buy GPUs for a year, and I look at it and say, "That's not a deal that I can do because it's too short for me to amortize the expenses." And so I won't do that.
They can go to another provider who maybe wants to take that risk on, who has extra capacity.
Absolutely, but our business is really built around the risk management of being able to get to scale because, in my mind, during this period of disequilibrium—during this period where there were not enough GPUs in the world to provide the compute for all of the different use cases in artificial intelligence—the part that's important for me and for my company is to get enormously large so we can drive down our cost of capital, so that we have information flow coming in from all different parts of the market.
The large language models, high-speed trading, search—all of these things are feeding information back in to us that is letting us know what the next product we need to build is, where they need help scaling, or what type of compute they need. All of that information flow is incredibly valuable to us.
What can you tell us about demand? There have been reports of, "Hey, maybe the Oracle star base thing with OpenAI has been downsized, or maybe not." And then other folks—Microsoft is going big, Google's going big, Meta is going big—and those people obviously have massive cash flow. Apple seems to be MIA. They don't seem to want to play.
You've named a lot of really big companies with really big balance sheets that have the capacity to drive a lot of demand. Look, I have been truly steadfast in this for years now. For 4 years, the depth of the demand for the service we provide has been relentless and overwhelms the global capacity of the world to deliver enough compute to enable all of the demand for artificial intelligence to be satisfied. We have been relentless about that.
Sounds like Knicks tickets during the Patrick Ewing era. They got up to 50,000 people on the waitlist. So if magically the waitlist went away, if the constraint went away, and we just had a large amount of GPUs available, a lot of energy available, and a lot of data center available, how much capacity would just all of a sudden come out of the system? So, what would be deployed, I should say?
Remember how we build our business through this box. It's a 5-year box. If we had an air pocket—if demand were suddenly to disappear because of a technology breakthrough, because of a war, anything—the why, from a risk-management perspective, does not matter. You have to prepare your company for what happens if it happens.
Yeah. And so, by entering into these long-term contracts, by entering into contracts with counterparties that have large balance sheets, you are—or we are—protecting ourselves and our lenders.
Yeah, so that we are confident and they are confident, because you can see how confident they are by the rate that they're charging us continuing to decline, that they're ultimately going to get their money back, and that is the one rule of lending.
Yeah. And so, in terms of the capacity, if you were unconstrained and NVIDIA's Jensen says, "Order as many as you want," what would happen?
It is also important to understand the constraints aren't just GPUs.
Right. Right. Electricity.
It's power shelves, it's memory, it's storage, it's networking, it's optics—all of the things. There are various throttles.
Memory is the throttle right now, right?
Oh, yeah, it is. Oh, yeah, it is.
Why? How did memory become the throttle? Memory has historically been a cyclical business, right? We have seen these waves of demand driving up the cost for memory, and then it collapses, and then it drives it up. It's a very boom-and-bust business. It's cyclical in its nature because the fabs are so capital-intensive that people invest in the fabs, build a ton of capacity, and then overbuild if there's any type of downturn. We've seen that cycle again and again.
What's happening right now is the confluence of 2 things, right? One is, with all the demand for artificial intelligence and the corresponding demand for compute and the ancillary services around the GPU, the demand is through the roof. That's number one.
Number two is that there was probably an investment cycle that needed to happen back in 2023 that would have brought on the necessary fab capacity to be able to serve the demand. It's impossible to predict what just happened. Just with energy, it's impossible to predict what just happened, and now people are chasing energy. The data centers are going where the energy is. It's not based on real estate; it's based on where there's some wind.
Many times, when you have a capital-intensive business like building fabs, you will get this boom-and-bust cycle. Just like in energy, they overbuild. And then, you know, fiber. Yeah, I mean, there are a lot of examples of that.
In some ways, when you look at that, it's a beautiful aspect of capitalism: we're able to have a boom-and-bust cycle, and we're able to weather it, right? If you think about capitalism from first principles, something like that happens and we have too much fiber, it creates an opportunity for Google to buy it all up or the next person.
Listen, the boom-and-bust cycle does a lot of things. It clears out the underbrush. The strongest companies will be able to survive and take advantage of that, and it sows the seeds of future business. The other thing that it does is put that infrastructure into the ground. You put the fiber into the ground, which became the backbone of how we watch movies every day, how we communicate, and how we hop on a Zoom. COVID and all of these things were based on that infrastructure that was available to be consumed.
People don't recognize this fact. The premise of YouTube, from the founders whom I knew, Chad Hurley and his other partner, was that they basically had the realization that storage was coming down so quickly that they could offer free, unlimited uploads, and bandwidth was coming down, so they didn't have to charge people for sharing a video online.
Before that, if your video went viral, people were going to have their minds blown, but your server would turn off and it would say, "This person needs to pay their bill." They were getting charged for carriage by the megabit going out.
The business models change and evolve, and, like you said, Moore's law—and certainly NVIDIA's Jensen will talk about the fact that what is going on within accelerated compute dwarfs Moore's law. All of that is going to lead to more opportunity to build more companies that are going to do things like YouTube did, which has really changed the world.
The concept—I don't know if it was a million hours being uploaded every hour or minute—but at some point, Susan Wojcicki, rest in peace, said to me, "How much was being uploaded every minute?" It made no logical sense until she realized, "Well, there are 2 or 3 billion people on the service, and 10 bips upload." It's like, okay, 1 in 1,000 people upload. It's a big denominator.
I was sitting on a panel with Sarah Friar, CFO of OpenAI, and she every once in a while really puts out interesting information. She was talking about the cost of a million tokens when GPT-3 came out, and it was $32 and change. Now a million tokens costs 9 cents.
Right, and so you just see the incredible power of how the capital markets, how capitalism, are fueling engineering and fueling competition.
It becomes recursive now, too. I mean, these models—if you say to the model, "Hey, make yourself more efficient, spend less money, and lower the cost of tokens," it'll be like, "Okay, captain." I don't know if you saw Karpathy's recursive thing last weekend, but now civilians who've never worked in a language model or done computer science are like, "I'm going to try to do something recursive this weekend."
You know, it's one of the things that I talked to the other founders about. When you think about some of the things that AI does, right, it's lowering the barrier to operations.
So if you have a good idea or a great idea, you can open up your model and tell your model—you can vibe-code it, you can do all kinds of different things—and create things that never existed before. That's amazing, right? That's bringing down this incredible barrier that kept human creativity contained, and now, all of a sudden, there's this whole new vector of medical research or different approaches to baseball cards or whatever you want.
If you've got a great idea, if you've got a new creative idea, that's the valuable kernel right now that allows you to build new things and create new things. I just think that's incredibly exciting. You're bringing the minds of 8 billion people a tool that allows them to overcome what was insurmountable forever for humanity.
Yeah. It's a bright new future, Michael. I appreciate you sharing the information with us and the vision. I am really delighted to have Aravind Srinivas on the program.
Thank you for having me. It's so great.
I want to go through 3 stages in which I fell in love with your product. The first phase was that I could go in and pick my language model if I wanted to use OpenAI, Claude, or whatever it was. That was a real unlock for me.
On the sidebar, I noticed you had done essentially what Yahoo did in the early days: finance, sports. When I pulled my Knicks game up, it gave me a live version of that. When I pulled my stocks up, it summarized the news in real time, and I was like, "Wow, this execution's great."
I made you my front door to different models, and it made it easier for me to check them. Then you came out with the Comet browser, and I was like, "Holy cow, I can give this a series of instructions: Go to my LinkedIn, find everybody from this company, and put them into a Google Sheet." Boom, you were the first out of the gate with that.
Then, just in the last couple of weeks, I'd been Claude-pilled and using Open Claude, but you came out with Computer. I started using Computer, and boy, it's good. It's a really strong start, allowing me to do repetitive tasks very similar, in some ways, to Claude Cowork, or basically an engineer or developer using it.
So, are these the evolution of the company, and should I think about it that way? How do you look at Perplexity now? You have a very loyal fan base. You're making a lot of money. I don't know if you disclose it, but I think it's hundreds of millions to billions. You can tell us.
What is Perplexity in the face of Claude having a great run, OpenAI still doing strong, Grok doing very well, and Gemini coming on strong? There's like 6 or 7 of you, and you just happen to be one of my top 2 right now. Thank you. So, tell me.
First of all, thank you. Thank you so much. Perplexity has always been built for people who are always looking for the extra edge—the curious people. So, it's very natural that you are one of our power users.
One common theme for us for the last 3 and a half years is accuracy. Perplexity wants to be the company that's building the most accurate AI. When you want to give somebody answers, accuracy is very essential for building trust, because only then is the user going to ask the next set of questions.
It turns out it was a great idea to give AI access to the internet to be accurate. So that's the Perplexity Ask product. It turns out it's a great idea for AI to have full access to a browser so that it can be accurate when you task it to go do something that you would do yourself on a browser. Agentic browsing: Comet.
Now, the last phase is that it turns out it's a great idea for AI to be given full access to a computer, so that it can do whatever you do on a computer on its own, essentially becoming the computer itself.
It's an orchestra of everything AI can do today—every single capability each individual AI model has, be it GPT, Claude, Gemini, or anything else. An orchestra of all those capabilities. That's what Perplexity Computer is.
All these sub-agents that are running inside Computer are the musicians. The models are essentially the instruments. There are hundreds of models out there, each having its own specialization. Some are good at coding, some are good at writing, some are good at multimodal visual synthesis, image generation, video generation, or audio.
But what matters is the end output, the music you play. That's the work AI gets done for you, and that's what Perplexity Computer is. The AI itself is the computer now.
It still lives inside of a browser. Have you considered giving it desktop root access? That feels like the next place this is going, but that comes with a lot of security issues and a lot of trust issues.
As you mentioned, trust is paramount. Getting the right answer is what builds it, but also not getting hacked and not having it delete your files. So, how do you think about root access to my Windows machine? Obviously, iOS won't let you, but with an Android phone, it would let you. Do you have that in the works?
Yes. We announced something called Personal Computer—Perplexity Personal Computer. That's essentially going to take all the trust and reliability and the server-side execution of Perplexity Computer, but synchronize it with your local computer so that you can use it from your phone.
We're going to do this with the Mac mini, where you synchronize your computer with the Mac mini so that it becomes your local server. All the agent orchestration that has to do with your local private data will run on that local orchestration loop—that runtime—with the Mac mini.
Not on your servers, not on Anthropic's.
Exactly. It could still ping frontier models if it needs to, with your permission, but it will be orchestrating everything on your local hardware.
If it needs to run on the server-side hardware, if you don't want very complicated, long-running tasks to be running on your local hardware, you can delegate it to run on your server-side computer, which is again only accessible to you and you alone.
That way, we're going to bring this perfect, trustworthy hybrid between local and server-side.
And you'll make it easy to do. It'll just be abstracted. You install 1 executable, and boom, it's done.
It's like OpenClaw for dummies. Nobody needs to learn how to use it. Nobody needs to manage API keys. Nobody needs to manage separate billing across 100 different services, or figure out what you can give access to and what you can't access. We take care of that.
So it's the Steve Jobs way of doing it: end-to-end integration.
And how do you think about local models? I've started running Kimmy 2.5 on a Mac Studio. It's not as good as Claude, Gemini, or Grok, but you can probably do about 80% there for free.
Yeah, essentially.
Do you have one of those? Have you started testing on your local Mac Studio? I assume you have a Mac Studio and you're doing this yourself?
Yeah. I don't know if you saw Dell and NVIDIA announce a giant workstation. Is it a 3800? Something like that, with 750 GB of RAM.
Something like that, with 750 GB of RAM. So, what do you think about the desktop going back to workstation-server status?
I think it's very promising. My prediction is it'll initially start off as a sub-agent. Whatever you need to go—your tax returns, your personal photos, your emails, your calendar, all that stuff, those local apps, your personal notes, very personal notes—you can make sure that the models that access those tokens will be running on your local hardware if you want to, if you're that privacy-conscious.
More complicated stuff that accesses your data is already on the server side. For example, your Google Calendar—
Your Gmail.
This is personal data still, but an AI runtime can access that through your connector—your Google Calendar connector, your Google Workspace connector. That could run on the server side because, anyway, the data's on the servers. It's not even lying on your device.
So that sort of hybrid orchestration is where we are headed. I don't think it's a dichotomy between fully local versus fully server-side. It's all about choice.
Anyway, when you're on your phone, you don't actually care which server that workload's running from, because it's not going to be able to run on your phone anyway. The chips need to exist on a Mac Studio or a Mac mini, or on the new Dell that's coming out.
I really think the idea of spending $10,000 on a powerful desktop will appeal to people if it lowers their $500-a-month cloud bill.
Yes. This is an incredible savings, plus you get the benefit—
Yes, of privacy and not educating the language models on your personal data.
Yes. And it's going to be like you're buying a refrigerator, your internet modem. The cost for these will eventually go down, but it's not going to feel like you're wasting your money.
Every home has a lot of other sensors that run your home. They'll also be part of this orchestration loop. That's where it gets exciting, because now you can just dictate something to your phone, and that can control your entire home.
That's the dream that everybody has, and that entire orchestration loop can run on your local hardware, no problem.
And I'm curious what you think of the operating system. What's eventually going to be the operating system of this workstation?
AI is the operating system. Earlier, in the traditional operating system, you executed programmatically. Now you start with objectives, not specific instructions. You come up with a high-level objective: "Go build this website for me that takes all the transcripts of the All-In podcast and tracks the stock price just before the podcast and after."
Yeah, and charted for the max 7. Yeah, and charted over time. You can—so that's the objective.
But individually, it's running a file system, a code sandbox, and access to the internet. It's got its own HTML tools. So I think that's basically where models, systems, files, and connectors are all coming together. You would think of that as an OS.
Mhm. Except you're operating at an abstraction above that, where you're thinking in terms of objectives. Does it need to eventually become its own operating system in your mind?
It could be. People could think about it as, “Yeah, I have my Perplexity Computer running all the time.” Essentially, it runs on Linux machines right now. Every server-side computer is a Linux machine. Mark Zuckerberg recently tweeted, right after our release, “Turns out Linux computers were the right idea. Desktop Linux computers are finally going to work.”
Linux machines are stable and customizable, and you're not at the mercy of Apple's desire to contain the experience or Microsoft's attack surface for hackers. You build something rock-solid, and it does feel like Linux might actually become the eventual winner. It may not need to have a front end. You could access the Linux machine on your phone, running iOS or Android. It doesn't matter. The actual valuable runtime is running on Linux on the server.
You've done great as a consumer company. A lot of love there. Now I'm starting to see corporations engage with it. In fact, you'll be happy to know this: Last week, I took 2 people in my back office and said, “Stop working on OpenClaw. Your job is to do the back-office automation at our venture firm, using only Perplexity.” They said, “Perplexity Computer?” I said, “It will.” I'm going to see Aravind, so I'll talk to you about that. We need a really strong Slack connector.
It's already out.
It is? Okay, great. At first, we were sending reports in, but it wasn't interactive. That's perfect. So now you've got your company going in 2 different directions: this incredible consumer run you have. How many people are using the product every month?
Several tens of millions. Our computer exists as a Slack bot right now that you can add to your Slack workspace on the enterprise plan. Our entire company works like that. People are talking more to the computer on Slack than to other people.
That's the first volley. We were sending reports in, but it wasn't interactive. So now you've got your company going in 2 different directions. This incredible consumer run you have—how many people are using the product every month?
So tens of millions of people. That's very similar to the trajectory of Google's and Yahoo's consumer businesses. Now you've got corporate. How are you doing on the corporate side? Thousands of companies?
It's a growing business for us. It's growing faster than the consumer business in revenue. Things like Computer unlock entirely new possibilities. For example, we've saved more than $100 million for our Enterprise Max customers, who are on the highest tier of enterprise.
Explain that. What does it cost? $200 a month per person?
There are 2 tiers. One is Enterprise Pro, which is $40 a month, and there's Enterprise Max, which is $400 a month. On Computer, after you run out of your credits, you pay for the tokens. You pay for the usage.
Are you making money on the $400-a-month, $5,000-a-year one, or at this point are people going so crazy that—
One thing Perplexity has, unlike certain other wrapper companies, is that every dollar of revenue we make has positive gross margins. We’re not just selling tokens. Most of our revenue is recurring because people are paying a subscription fee. Because we route through multiple different models, we're very efficient in terms of how we spend on tokens. Because we have all this advantage with RAG, orchestration, and search, we don't actually need to blow up the context window of the models.
As a result, we have positive gross margins on all the revenue we make. Every single penny we make, we make profits on. Overall, the company is still not yet profitable, but we're working toward that.
You've had the opportunity to exit. There were a lot of rumors that Apple and other people were saying, “Hey, this is a great team.” How many people are on the team now?
About 400.
You've got a very coveted team. You obviously understand consumer, you obviously understand business, and it's a product-driven organization. Reports are that you declined the offers. But the world's getting hypercompetitive here. How do you keep up as a 400-person organization when you've got Sam Altman over here raising $100 billion, Elon putting data centers in space and merging with SpaceX and Twitter, Google with unlimited resources, and Amazon getting into the game? Gemini is a very strong product, and Google is really good at consumer.
I think we'd all agree that Facebook and Meta haven't figured it out yet, except maybe for serving us better ads. They haven't figured out the consumer case yet, but they'll copy it. They always do. How do you look at the playing field? The degree of difficulty—this isn't playing checkers. This is like playing against the 10 best chess players in the world. That's what you have to do every day. How do you think about it long-term? An independent company? Do you think you'll need to join forces at some point? And why didn't you take the deal? The deals that you were offered were incredible.
One advantage we have that all these companies you mentioned don't have is multi-model orchestration. We're like Switzerland. We don't have to have 1 horse in the race. If GPT wins, Gemini wins, Claude wins, or Llama wins, it doesn't matter to us. Even open-source models can win. No problem.
And you have them on the service? You have DeepSeek and Kimi?
We have Kimi, we have Nemotron, and we have a lot of usage of Qwen.
Alibaba Qwen?
Yes, silently under the hood. For us, the advantage of being able to take the best in each model and give the user the orchestration of everything they can do is something I don't think any of the companies you mentioned can do.
Right. Nor would they.
Nor would they. It makes no sense for them. It would be an admission that all the data centers and capex they've built out mean they still couldn't produce the best model themselves.
Dario Amodei, the CEO of Anthropic, said recently in an interview that models are specializing. Toward the beginning of last year, people thought models were going to commoditize, but toward the end of last year, models started specializing. Even within coding, Claude Code and Codex have very different capabilities. Our iOS engineers love using Codex. Our back-end engineers love using Claude Code.
Even within a specialization like coding, models have their own unique specialties. There are many other use cases outside coding where different models are good at different things. That means the orchestra conductor, which has no single horse in the race, can win by providing a very unique value and service to the customer that each of these amazing companies cannot.
So you're buying tokens wholesale from them and then charging customers for it?
We're going to take care of all that orchestration, so you don't have to manage tokens across different models.
I authenticate a couple of my different accounts—my Pro accounts—into Perplexity, but I don't have enough knowledge to know if you're abstracting that and people can just search across them as part of their Perplexity subscription.
No, we're not bundling subscriptions from other AIs. We just ping the models directly. What you get from us is Perplexity orchestration. When models are specializing, there's a bigger value in the one who knows how to build a great harness that can take the best from each model.
Does it auto-route today, or do you still have the dropdown? Somebody has to pick?
It definitely auto-routes to the best model for each prompt, but we also give users the flexibility to pick whatever model they want.
I've seen a bunch of startups hack this together. What do you think of doing the same query across multiple models?
We built a thing called Model Council.
Model Council, yeah. I saw Jensen Huang say in one of his interviews that he puts the same prompt into 5 different AIs and sees what each of them says. Everybody does that.
But then you still have to apply your biological compute to read every answer and figure out where they differ. It's like talking to 5 different doctors and trying to figure it out.
Exactly. It's dumb.
Model Council is a feature we built where it will not just give you the answers from each model. It will tell you exactly where they agree, where they disagree, and where the nuances are.
And that's in the interface? Model Council? I didn't know it was there. You release products at a pretty great cadence. Where did you learn that, and what's your philosophy of shipping product?
Our philosophy is that speed is our mode. One of the things that big companies cannot do is move at the speed we do and serve customers at the speed and quality we do. It's very hard to maintain quality, speed, and trust at the same time. Apple takes a long time to ship anything.
Right.
Because they're very worried about people not trusting them. And so some companies are bureaucratic, and they just take forever to ship something. They don't maintain what they ship. They may make a big deal about an event, but nobody even knows how to go and use that feature. They get abandoned.
Exactly. So, Perplexity has those advantages of being very small.
Mhm. Which is honestly one of the reasons why we built Computer, because now even non-engineers are shipping code here by just pinging a Slack bot and asking it to fix bugs.
So, this iteration has just been exponential. The moment I became Claude-pilled was when I was working with it and I was like, “Hey, I want to build my network. I know these 20 people in Japan. I had dinner with them during my recent trip. I want to know who they know, so check out LinkedIn and other things to see who they're associated with, and make me a mind map of it.”
“On the next trip, I want to meet with the next circle of those connections.” So I started asking. It said, “Okay, I got the results.” I was like, “Great.” It said, “Where do you want me to put them?” And I was like, “Well, where can you put them?”
It said, “I could put it in a Google Sheet. I could put it in a Notion table. I could put it here. I could give you a PDF. I could give you a CSV file, or I could write you a CRM.” And I was like, “Yeah, sure. Make me a CRM system.” And it made a CRM system.
Yeah. And I think maybe 1 out of 1,000 people working with AI have had that experience. Maybe it's 1 in 10,000, where your agent says, “I'll make you bespoke software.” Have you had that yet? Do you see that as a part of Computer, where when a person needs a spreadsheet, you don't launch Excel or Google Sheets—you just pop up a spreadsheet?
Yeah. Well, we have a board meeting tomorrow.
Okay, I'll come. [Laughter] So, pitch it to the board.
Sure. Computer made the memo. And we had a partner meeting to pitch a partnership idea. Earlier, we would have a design team do the whole deck. Computer did it in one shot.
I had a press briefing with a bunch of journalists. My comms person would usually give me a memo about what to say. Computer did it in one shot.
That's brutal.
It's crazy, and the context is so good because the memory is getting better. So, it knows that journalist from the last time. It knows the board meeting. It has all the previous decks.
When did that happen?
I think it happened with Opus or Phi.
Uh-huh.
Anthropic's Opus or Phi—that was the inflection point when models started being amazingly good at orchestration, reasoning, and tool calls. Claude Code brought in this new idea in AI that everything can happen inside a sandbox, a console, or a terminal with access to tools, where tools are just command-line tools. They don't even need to have a graphical user interface.
So, when you did that, and when you organized around files, subagents, skills, and CLIs, the model became very good at handling the context. The context window no longer became a problem. It just put whatever was necessary into the context whenever it wanted to and dumped it away when it wanted to. That made it suddenly so good at doing very long orchestration tasks.
Yeah, it's pretty crazy. I have every episode of This Week in Startups, all the transcripts, and then all of All-In.
One of the tasks I did, by the way—I can send it to you. I asked it, “I want you to download every All-In podcast.”
Yeah, since the beginning.
Yes. I want you to take a mention of all the public companies they mentioned during each episode. I want you to have a histogram of the counts, and I also want you to chart it across time. Then I want you to analyze the impact on the stock price and the sentiment of what we said.
Exactly. And it did. It clearly said—
Moving stocks?
It was about Google's stock going up. Yes, prior to that, you guys were talking a lot about Google.
And I said, “I made a bet publicly on the thing. I said, ‘I am buying a bunch of Google because I believe, even though they're behind, it's because they're too precious.’” You were mentioning a company that might be too precious at times and doesn't release. I was like, “That's that company. They need to release more.” And I told Sergey, “Give us the good stuff.”
Yeah. He started giving us the good stuff.
It literally gives you the timestamps of every single mention, and then I can click on it and actually hear exactly—
Yeah. Sweet. That's when I was like, “Damn. This would have been a week-long project.”
It would have been 10 hours a week of a researcher. I'm experiencing the same thing. When I do research notes, I've created my own mega-prompt, and it will tell me where you worked before, who's in your circle, who your competitors are, who your friends are, and so on.
Then it will find old podcasts. That's one of my secrets if you're an interviewer watching. I try to find what the person was talking about 5 years ago, 10 years ago, and then over 10 years ago. I've gone into interviews now with Michael Dell and talked about things he was talking about in the '90s.
It finds me some ancient stuff. You would pay a researcher or a producer $70,000 or $80,000 a year to do this, and they would have done a third of the job in 10 times longer.
Yeah. It's really gotten weird just in the last 6 months.
What do you think the next 6 months looks like?
I think the dream is to help businesses run as autonomously as possible. Everybody talks about how AI is going to create this 1-person, $1 billion company. Some people say it's already happened because people pay researchers like $1 billion. But it's not truly moving the GDP by $1 billion. It's not truly creating new value.
So, the best way to do that is to actually help a small business—people who would otherwise drive Ubers for extra passive income—buy a Mac mini, set up Perplexity Computer, and run their business on that, or run it on the server. It doesn't matter. They can actually make real money—hundreds of thousands or even millions a year—and grow it.
Have Computer run your ad campaigns on Instagram or Google. Integrate with SEM and SEO tools, find new users, integrate with Stripe, charge them, ship new features, have your own Intercom integration for customer support, and have all of this working while you can be sipping wine in Napa.
That's the dream. It feels awesome to say. Everybody thinks AI is already there. It's not there yet. Someone has to do that hard work. That's what we want to do.
It's a great vision, because when I watched startups 20 years ago, there were so many checkboxes they had to do. I have to find an office space. I have to put up a bunch of servers. I have to hire an HR firm. I have to hire a PR person. All this stuff.
Now I talk to young founders who have a 3-person team. They've come out of a16z, my Launch Accelerator program, or Y Combinator. And I'm like, “Okay, you raised $500,000. You raised $1 million. Who are you hiring?”
They're like, “I don't know if we need to hire anybody.” I'm like, “If you could hire somebody, who would you hire?” They're like, “Well, I do my own HR. I have this partner.” I'm like, “How are you doing hiring, anyway?”
They're like, “Well, I put out an ad, and then it sorts and ranks the candidates. It emails the top 10, asks them a bunch of questions, and then I meet with the last 2.” I'm like, “That's what a recruiter did.”
The entire recruiting job has been abstracted, and a tool like Computer is going to make that even faster.
There's work to do. A lot of connectors. A lot of specific workflows. People don't want to learn how to write essay-long prompts. It needs to be so quick, fast, and autonomous. You just set it up and it's done. You have an idea, you can turn it into a business, and start making money.
It's an incredible future, and it feels like it's right here. How do you think about job displacement? You're actually making the tool that enables people to be solo entrepreneurs and get to $1 million in revenue, but it's also the same tool that doesn't require them to hire. We've had this debate a million times on the podcast. Do you have moments where you're like, “Oh my God, this is really terrifying?”
Yeah. A lot of people are going to lose their jobs really fast.
Yeah. And then, oh my God, you can learn any skill you want, and all the things that were hard are now easy.
Yeah. I go back and forth. I'm 70–80% super positive about this, but 20% of the time I'm a little worried. Where do you sit?
I mean, America has always been about entrepreneurship, right? We've been about trying to build new things, discover new things, and go explore.
I think Henry Ford came and built factories and brought in jobs and things like that, and put people into a box. But the reality is that most people don't enjoy their jobs. They're doing it for—
They hate them. Exactly. So, there is suddenly a new possibility, a new opportunity to use these tools, learn them, and start your own mini-business.
If it pays for your needs for a year or multiple years, lets you have a high-quality life and good work-life balance, and gives you a true feeling of agency, ownership, and passion to get your ideas out there, I think that—even if there is temporary job displacement to deal with—that sort of glorious future is what we should look forward to.
I think you're exactly right. If there will be some displacement, then there's also going to be so many opportunities opening up, and it requires the individual to not be passive.
They have to be rugged individualists. They have to be resilient. Yeah, and they have to be resourceful. I think once you start playing with these tools, that's what happens.
Exactly. You all of a sudden feel—it brings out the best in you if you truly are in a good space. Yeah. And then today, Comet for iOS is out.
Yeah. I'm a Comet super fan. I required everybody to put this on. You were nice enough; I emailed you. I was like, “Can you send me some licenses?” You sent me a bunch of licenses, and I said, “Everybody, put this on.”
Because it was $300 a month when you first came out with the Comet browser. Now it's free, I think, for all users. Highly recommend it. Highly recommend getting a Pro account. It's only $20 a month to get into Perplexity, which is a joke. You can get on board for nothing—less than a dollar a day.
But what does iOS allow me to do? And how does it connect to Computer? That's another thing I'm having trouble with. Claude Code and Computer—there's not a good enough integration with this mobile device yet.
Yeah. Computer is already on the Perplexity app, so you can just toggle to Computer and start using it. Comet's uniqueness in Perplexity, for the company and the strategy, is the fact that you can control the browser.
The browser also becomes a tool for Computer, just like your Google Workspace and all these other things. Until the whole world is organized around CLIs and tools, there's still a lot of tasks we have to do manually on the web, on the browser: open tabs full of forms, click on things, upload stuff.
All that stuff, if you want to automate, you need a browser. You need an AI that can natively control the browser. So, that is Comet. That's why, no matter how many other tools in the market exist, like OpenClaw or Claude Cowork, executing tasks on a browser on the server side, along with all the other things, is something uniquely Perplexity can do.
Yeah, my dream is that you'll create an Android app that roots my Android phone. You just take over and see everything. One of the blockers I have now is that some of the websites have gotten a little persnickety.
I don't want to mention too many, but Reddit and LinkedIn. I'm a great Reddit user. I'm a great LinkedIn supporter. But sometimes I need to get my InMail from LinkedIn, and I just need to find 7 people at a company.
Is there going to be a solution between the LinkedIns and Reddits of the world and Claude and Perplexity? How is that negotiation going? You don't have to speak about any specific ones unless you want to. It feels like there's got to be a solution, and I'm willing to pay for it as a user. I'm willing to pay Reddit to allow my bot to show up and behave properly.
Yeah. Well, I cannot speak about any particular company, but we are happy to work with anyone. With Comet, our idea is to give people the flexibility to set things up on their own. Any official APIs that anyone's willing to offer, we're always happy to put that as part of Computer.
Here's what I think should happen. Let me see if you agree. This is for Steve Huffman at Reddit. I go on Reddit and get a Pro account for $20 a month. When I do that, I can authenticate whatever tool I want to do a series of well-behaved things a certain number of times a day.
It's not unlimited. I'm not going to scrape the whole site, but I would like to let Perplexity or Computer go and tell me, “Hey, what are people saying on the This Week in Startups and All-In subreddits? Summarize it for me so I get the customer feedback.”
I would literally name my agent and say, “It won't post on my behalf. It won't vote on my behalf. I just need it to do a couple of little read-only things.” This would be an easy solution.
Or with LinkedIn, I already pay them like $50 a month. They should just let the $50-a-month account work with Computer.
Yeah, absolutely. This is for Satya Nadella: Let LinkedIn work with Perplexity and the other players, and we'll pay you extra. Perfect.
It's a revenue stream. Don't you think API access for our customers is a revenue stream?
I think so. Fundamentally, giving users a choice and setting it up as a win-win for both the business and the user is where the world should head. I would say the same thing applies to any website in the world. If you want an AI to use it on your behalf, it should be okay, because that's what the user wants.
I have a paid New York Times subscription. Let me go in there and do 100 searches a day, a week, a month—whatever they choose. That would make the subscription that much stickier.
Exactly. All right, Aravind, love the product. Anybody at home, it's just tremendous. Go learn Computer and get the Comet browser. It has changed my business for the last 2 years. Love the product, and we'll have you back soon when you launch your operating system and come up with your own server and desktop server, but business is the focus. Yes. Great seeing you.
We have an amazing guest. Arthur Mensch is here, the CEO of Mistral AI. How are you doing, sir?
Great. Thank you for having me.
You're here at NVIDIA's big conference, and there's a big announcement. You're going to be working with NVIDIA to build models and open-source them. What is the big announcement here?
We're announcing that we're going to be training the next generation of frontier models with NVIDIA. It's something that we've done before with NVIDIA with Mistral NeMo, something we did around 18 months ago.
The point for us is really to be able to produce the best open-source models out there so that we can use those assets to specialize them through products that we do for our customers, like Forge, which helps us customize the models for the enterprises we work with in engineering, physics, and science, and make them better at certain languages when we work with governments, et cetera.
Mistral is obviously based in France. You're the leading AI company there. What's it like running the company and building a large language model in Europe? Obviously, there's regulation and all kinds of considerations around privacy. The French are known for protecting privacy. In the United States, we're known for taking it away.
How is the landscape there, and what do you have to deal with there that maybe you wouldn't have to deal with in America? What are the pros and the cons?
Let's say, first, 25% of our business is in the US, and 25% of our researchers are actually here. I spend a lot of time here, as well as in France, the UK, and Singapore, where we are.
Of course, they're different markets—markets where language is a topic, where manufacturing is a bigger piece of the pie than it is here. I'd say our strength has also been to work with European companies that are a bit lagging behind and want to adopt the technology to leap forward.
We've been able to do that through forward-deployed engineering engagements, through our Forge product, or through our Studio product, which allows you to deploy agents that do end-to-end automation.
On top of that, the thing that we announced today, like Forge, is something that's actually being used today with customers in the US because they come to us with needs for post-training, for making models more capable, and specifically good at financial services. What's happening is that we have this product, and we can bring the models to specialize them as well.
Your belief is that specialized, verticalized models—health care, finance, engineering, and different verticals—will win the day, or that a global model will win the day that does everything?
You need general-purpose models to do the orchestration part, et cetera. But at some point, enterprises sit on a lot of intellectual property and a lot of signals coming from physical systems, factories, and tools. It's actually not trivial to connect those systems, to connect that data to models that are closed-source.
If you have open models, you can add new parameters. You can make a lot of deeper things that you cannot do with closed models. You can also—and that's something that we do—we not only work on the model side, but also on the orchestration side.
We sit with subject-matter experts to understand their needs, and we build business applications that are fully bespoke to their needs by modifying the models, but also modifying the harness on top.
So, we believe that eventually, building an open-source technology is a way to save costs and have better control, because you can see the thing on every cloud that you want, on your hardware if you want, and deploy it on the edge if you want.
Eventually, from a customization perspective and from leveraging your decades of IP that you've been accruing in financial services and heavy manufacturing, companies like ASML, for instance, benefit from working with us because we take their data and build models that are specifically good for them and their purposes.
This training data uses experts to come in and refine a model. Most people don't know this business that well, but this has become a very large part of the industry. Obviously, Scale AI was doing it. They went to Facebook and lost a lot of the customer base who didn't want to send their data, I guess, over to Meta.
We're investors in a company called micro1 that's doing pretty well in this space. There are other folks doing it. Explain to the audience what you're doing specifically for companies, how this training works in a verticalized way, and how you silo that data. If you're working with one customer in aerospace or fintech, they might have a need set, but they may not want that training to go to a competitor.
I can give you a few examples. I think overall, the data segregation is super important. The way we have solved that is through a portable platform. Our technology is a set of services, a set of training tools, and a set of data-processing tools that I can take and put on the infrastructure of my customers.
Suddenly, from an IT perspective, when we talk to the CIOs, they realize that, from a security perspective, the flow of data doesn't go anywhere. There's no data flow coming back to Mistral because everything stays there. The way we then use that technology that has been deployed is that we're going to be working with the teams that are doing image scanning and defect detection with ASML, for instance.
We're going to send forward-deployed engineers and scientists. They're all PhDs, and they know how to train models. They spend some time with the subject-matter experts who can explain how an image is being detected, how you detect defects, and so on. Based on that, we're going to work out what kind of data needs to be used to train the models that are going to solve the task itself.
We send the technology and, typically, a few scientists because you do need that expertise transfer and knowledge transfer between our teams and the vertical experts. Then we make sure that eventually our team no longer needs to be there to retrain the models, get more data access, and so on. That combination of data segregation, expertise transfer, and knowledge transfer is the one thing that makes us quite unique and allows us to serve the most critical use cases and processes in industries that need to take their data and put it into models for them to work.
It seems that once we've exhausted the entire open web—what was available legally, on the gray market, and so on—I wouldn't have you comment on that controversy, but we've kind of exhausted what's in the open crawl, yeah?
We have.
And it's time to either make synthetic data or use experts. Do you believe in synthetic data, and where does that work and where does it fail?
We use synthetic data as a way to warm up the models. It's a way to be quite efficient at the beginning. If you have a large model and you want to train a small model, you will use your large model to process and produce a lot of synthetic data at the beginning.
Eventually, though, you do need to have human signal. Human signal is always a bit costly to acquire because you need to talk to the experts, and they need to give feedback to the machines. At the beginning, synthetic data allows you to do the compression, to further compress the models.
At the end, you do need to go and get data that is produced by humans. So, it's mostly an efficient way of training models, with bigger models used as teachers for smaller models, but it's not enough. You also need human signal.
Arthur, we've seen an incredible explosion. We're sitting here 52 days after OpenClaw, the year of our Lord, 52 days in. When you first saw OpenClaw and saw the reaction of hackers, founders, and startup CEOs—the amount of energy, and seeing it race to the top of GitHub with the most stars and likes, along with all these contributors—what did that say to you as an executive in the space who's been grinding on this for many years? What did that OpenClaw moment mean?
It resonates a lot with what we are doing with our customers because, pretty quickly, enterprises realized that if they wanted to make some gains with artificial intelligence and generative AI, they would need to automate full processes. To automate a full process as an enterprise, you can use OpenClaw, but it's actually not really enough because you have data problems and governance problems. You can't observe the process that's running, and in many cases you can't control it.
When you run a KYC process, if you're HSBC, for instance, one of our customers, you will want to have deterministic gates that are always going to do the same thing, in a way that is observable and allows you to guarantee to the CEO that it's always going to go through these gates. That's not something OpenClaw is providing because it doesn't have the kind of primitives that you need to work on collective productivity, observable productivity, and mission-critical systems.
On the other hand, the autonomy it gives and brings to people who are just individuals hacking things together is also a way to show enterprises that if you set up the right control plane and the right sandboxes, connect to the right data sources, and make sure that your access controls are well respected, then you can actually unleash the power of agents doing things for your employees. That can work on the platform; otherwise, you will not be at ease when you're sleeping.
It is definitely something you have to be thoughtful about. When I installed it, I gave my agent root access to my Google Docs, my G Suite, my Notion, my Zoom, and my Calendar—everything.
Then I realized, "Wow, with my enterprise edition of Gmail, I can essentially summarize every conversation going on in Gmail for my entire 21-person investment company, and then correlate it with every conversation in Slack." Then I realized, "Oh my gosh, there are compensation discussions going on. There is a person on a preferred performance improvement plan, or something like that."
I have to make sure nobody else can access this because the power comes from giving it access to data, but with great power comes great responsibility. I think people are learning that in real time.
It's a big problem because enterprise data is not a single thing that you want to put into a single system that's going to be accessible by everyone. You need to have this layer that actually understands what is in the data. You need to have a semantic of what can actually be exposed to HR or what can be exposed to engineering.
Typically, compensation is one of these things. You want to make sure that compensation data does not flow back to all of the enterprise because you're going to have a lot of problems if that's the case.
What you actually need, and which is hard to do, is what we call a context engine: a mapping of where the data sits that comes with a certain amount of metadata telling you that this data is not accessible to this part of the company. If someone in engineering is asking for something related to compensation, the system is going to tell you, "Look, you actually can't access that data."
That's hard. It's actually hard. You need to rethink entirely the way your IT systems are being connected. At some point, you also need to think about your management because your information flow is completely different today if you're connecting agents together with your data sources than it used to be.
Suddenly, maybe you don't need that manager whose only purpose was to take information from the bottom and put the information at the top. There are so many problems to solve. You need the right primitives, you need sandboxes, and you need role-based access control and these kinds of things. You have changes to make.
You need to rethink your entire customer-service department because suddenly you don't need that much transfer of information operated by humans.
All right, you have to go—you've got a flight to catch. It is so great to see you, Arthur. Continued success with Mistral.
I'm really lucky to have Daniel Roberts here. He's the co-CEO and co-founder, along with his brother, of IREN. They're a publicly traded company. They started in BTC.
Thanks. Pleasure to be here.
You and your brother started in Sydney 7 or 8 years ago, and you got in early on Bitcoin. All these Bitcoin miners wanted to have data centers, huh?
That's directionally right. The thesis we saw was this explosion of the digital world and the growth in the online world, and at some point the real world was going to struggle. So, we set about to build out large-scale data centers.
Yes, the first use case was Bitcoin mining, but as we said to our seed investors, use that to bootstrap the platform, generate cash flow, and layer in higher and better use cases over time as they emerge. Here we are today with AI. We are swapping out all the Bitcoin for AI chips.
When did you first start seeing the demand in the company shift from, "Hey, Bitcoin miners, we need some H100s, whatever it is," to, "Hey, we're this nonprofit OpenAI. Hey, we're this research lab. We need some AI compute"? When did that start hitting?
Look, we had a bit of a false dawn, I would say, back in 2020. We signed an MOU with Dell to start bringing on customers and compute, but in hindsight it was too early. So, we went back to Bitcoin and kept bootstrapping the platform.
I would say about 2 years ago, and month by month, the demand just continues to escalate.
And you were in so early that when you were looking at data-center space in the United States, you were one of 1 or 2 or 3 people looking at the space.
They were trying to sell you on space, yeah?
Yeah, so we actually developed the data centers ourselves. We go and find the land, get the permits, and apply for grid connections, and we were doing it at a scale that just amazed people at the time. Our flagship Texas site is 750 MW. Four years ago, that was unheard of. In the middle of the desert, we're building these big data centers, and the traditional data center industry was going, “What are you guys doing?” We said, “We believe in the future of digitization, high-performance computing,” and obviously, today, it's paying dividends.
I don't think anybody could have predicted, when ChatGPT came out, OpenClaw recently as a turning point, and then Microsoft, Google, and everybody embracing this. That's your big partner, Microsoft.
Yes, Microsoft is one of our early partners. We signed a $9.7 billion contract with them late last year. But, as I was explaining to you before the show, that's 5% of our capacity. So, things are busy at the moment.
And when you do these build-outs, the big conversation today is no longer the number of GPUs we're putting in; it's just power. Power is the constraint today, yeah?
For many in the industry, it is. But for us, because we started eight years ago tying up all this land and power, it isn't. We've got 4.5 GW. For context, that's almost as much power annually as the Bay Area uses in its entirety each year. It's huge. So, for us, the hurdle, or the constraint, is really time to compute, and that's emerging across the industry as well.
And time to compute means tradespeople coming to West Texas, living in a trailer that you set up, and then breaking ground on a data center, building foundations, and building water-cooling systems. This is hard manual labor going on, yeah?
Exactly. And this is the whole real-world challenge of responding to these digital exponential demand curves. They're unconstrained by the real world in terms of their appetite, and it just compounds. You need thousands of people out in these locations that haven't supported it. You put stress on supply chains. We're seeing what's happening with memory—every aspect of it. So, it's permanent whack-a-mole, permanently putting out fires to try and bring this compute online.
And you get to spend time there. What's it like when you set up a town or bring 1,000 or 2,000 people to what's pretty much a remote small town? I'm assuming that when you bring 1,000 people, there might only be 500 living there right now. What are those towns like? It sounds to me like something out of the gold-mining era, when people first went and were prospectors. It's a prospecting town?
Pretty much. I mean, the barbecue's great. That was the drawcard. Then, apart from that, we've always had a policy of hiring local and supporting the local community. This year, we're hitting $1 million in community grants cumulatively. That's things like local playgrounds and supporting the fire departments. But we will hire locally. Once we can't find that trade locally, we'll expand the radius by 20 miles and hire out of that, and so on and so on.
That's very thoughtful. These folks are coming—say, an electrician or a construction worker—and they've built houses or maybe corporate offices. Now they come for a tour of duty here, and the salaries go up massively, but they have to leave their family for a 3-month tour or something?
Yes and no, because typically, where we locate is where there's heavy electrical infrastructure. Where there's heavy electrical infrastructure is typically where old manufacturing and industry have closed down. So, we go in, leverage that sunk CapEx, rehire and retrain local workforces, and bring a new industry to town in these data centers.
Has that workforce now been completely depleted? Do we need to train another generation, a younger generation, to really embrace the trades?
100%. We're partnering with universities and trade colleges, absolutely.
And you go to a trade school, you go to a college, and people are getting degrees in philosophy and English literature. They're going $50K a year into debt, $200K a year into debt. What's the starting salary for a tradesperson working on a data center, doing electrical or construction?
Exactly.
What's the ballpark range?
Look, I won't talk specifics, but they are going up. The price is going up. It depends on the level, but yes, there is a rush for good tradespeople.
I'm hearing $150K to $300K. Am I in the ballpark?
At the lower end, directionally, you're right.
It's incredible when you think about it. There's concern about, “Hey, AI is taking jobs,” and then, on this other side of the ledger, we can't find enough talent to service it. Talk to me about energy sources and how you think about that. President Trump, Chris Wright, and the administration started with, “Hey, clean, beautiful coal.” Year 2, they're like, “All sources matter. Nuclear.” Obviously, natural gas is plentiful in that area, and we've got a lot of oil. People don't know this about Texas: in the United States, it's the number 1 source of solar installations. Talk to us about energy.
Our philosophy has been sustainability from day 1. We've used 100% renewable energy since inception.
What? 100%? Wait, how is that possible?
We use hydro in British Columbia, and we use wind and solar in West Texas. In West Texas, there's around 45 to 50 GW of wind and solar. The transmission line to export that down to the load centers in Dallas and Houston is 12 GW. So, you locate to the closest source of low-cost excess renewable energy, monetize it into this digital commodity, and export it at the speed of light as a token.
Great arbitrage. The wind is producing a lot, but it's harder to get the power from those areas where people are willing to put it up. People don't understand how big West Texas is. It's an incredible amount of land. And you're coming from Australia, where, on the west side, people also don't understand exactly how much pure undeveloped land there is.
So much land. And the issue is distance. You've got to spend billions of dollars on this transmission-connection infrastructure to move that power to where people actually want it. You can build wind farms and solar farms, but if you build them in the desert and no one can use them, then what's the point? So, the whole opportunity for our industry is to go to the source of that power and monetize it. The data centers follow the wind turbines and the solar installations.
How do you think about batteries? Are you able to put those online? Obviously, you're going to have periods where it's not a windy day. In Texas, we have very few days when it's overcast, so that problem's pretty much solved, but you're going to have 50 days where the sun's not beating down. How do you deal with the demand and soften that duck curve?
We don't need to.
Huh.
The utility does that on our behalf. This is why these grid connections are so scarce, so hard to get, and so highly valued, because once you get that grid connection, the utility underwrites all of that variability. They guarantee you 24/7 reliable power.
Got it. So, on their side, they're figuring it out. If something goes down, they could fall back, even though you're 100% committed to renewables. If they needed to fall back to gas or whatever, they have that ability out there, so you have that as a backup. There's a lot of talk, or a debate, about whether we're getting ahead of our skis and whether people are slowing down. There was some talk about the OpenAI project maybe downscaling a little bit. Is OpenAI a partner as well, or—
I can't comment.
Can't comment. Okay, so we'll read into that whatever we want. Are there pockets where people are saying, “Hey, let's slow down,” or is it still gangbusters?
It's at the upper end of the spectrum. It's gangbusters. We cannot meet demand. That's why the whole industry now is around time to compute. There are no idle GPUs in the world sitting in a data center.
What's your take on what happens when software makes things more efficient? This was a big discussion from Jensen himself during his 2.5-hour keynote yesterday. We're sitting here Wednesday; I think he did his keynote on Tuesday. He was talking about, “Hey, software is going to make it 50 times more efficient and lower the cost of tokens 50x.” Then you have transport also contributing to that. When do you think the curve goes from parabolic to simply growing at a ridiculous level? Is there a slowdown coming, or how are you planning for the future?
Look, I think it's actually the opposite. I think it feeds on itself. I'll give you 1 example. You go into ChatGPT today and generate an image. You hit enter on the prompt, and it's like the dial-up internet days. It takes minutes, and you're like, “I better get this prompt right.” Finally, 2 minutes later, it comes. Now, I'll give you an example: if we 10x the amount of compute available—which is an enormous task from where we are today—and those images take 5 to 10 seconds, are we going to generate more or fewer images? Many more. This is Jevons paradox. This is the theory of induced traffic. You build a couple more lanes, and people start to think, “Well, maybe the distance from Bondi Beach to the central business district in Sydney would be an acceptable commute.”
Love the analogy. Yeah. What do you think about, or what are you seeing? We're here at NVIDIA. Obviously, they make the leading-edge chips. They just bought Rocks, so now you've got 2 of the leading-edge chips coming out of the same company. But custom silicon is becoming a big discussion. Has that started to land in the data centers yet? Obviously, Google—I don't know if they're a customer you can tell us about—but they're making custom silicon.
Amazon is making custom silicon. Meta is making custom silicon. Talk to me about that revolution. Is it actually making it to the data centers yet?
Look, to various degrees, it is. They're promoting their products. They're trying to tie up data center capacity. So, yes, there's multiple silicon looking for homes. I think it's fair to say NVIDIA has a massive head start.
The ecosystem they've incubated and the standards that they're setting mean I would say the safest pathway to build out at scale early is to follow the NVIDIA roadmaps. But absolutely, over time, we are seeing these chips emerge.
And in terms of desktop computing, I know there was a survey announcement that Dell and NVIDIA are making a really powerful desktop—750 GB of RAM and a lot of power. You're going to be able to run some local models, open source, with OpenClaw and open-source models coming from Kimi and a bunch of the models out of China. And in the hacker group—which I think you started in, like I did, probably around a similar time—people are starting to get really obsessed with having a $10,000 or $20,000 desktop setup and running this locally. What do you think of that trend? I'm curious.
Yeah, the breakthroughs we're seeing in software—the way it's distributing power to every man and every woman in every house, and their ability to code and use products like OpenClaw—the generation of demand and appetite for compute at the local level, all the way through to these mega data centers, it's absolutely real. And as we see the emergence of agents using more and more compute, and as we see autonomous vehicles and other automation and robotics, it's absolutely going to compound.
And what about nuclear? The Trump administration really seemed to flip the switch on a growing belief that, hey, wait, nuclear's pretty great. It's clean. It's the original renewable, in a way. These new modular reactors have nothing to do with Chernobyl, Fukushima, or Three Mile Island. They're much safer. They're a completely different architecture.
Have those started to land yet? And since you correctly followed that trend in the great state of Texas, where I'm from, are you following nuclear?
I think you have to. I think the reality is it's going to take a decade, a bit longer, by the time big projects can come into commissioning. But now is the time to start that conversation, put in place policies, mobilize capital, and start that ball rolling.
Do you have a data center going up near nuclear?
No, not at the moment.
But you're actively tracking that activity?
Yes. Yeah, this seems pretty inevitable.
Feels like it. And if that happens, what impact does it have on your industry? Obviously, it's happening in China, and people always point to the Bitcoin miners—they were like the canary in the coal mine—near the hydro dams and near the nuclear plants where there was excess capacity. What impact do you think this has if you could actually have small modular reactors next to data centers?
Well, I think it just opens up the market and enhances the US's competitive advantage in this space. AI is inevitable. Robotics is inevitable. The reality is the correlation between human progress and energy consumption is really, really high over a very long time period.
So, if we can find a way to unlock new generation—clean generation, as nuclear—and locate that more at the source and enable more compute on a distributed basis, all those use cases we just discussed become easier, more fluid, faster. Then you get that positive flywheel around Jevons paradox and demand.
Talk to me about the architecture today of Ethernet and data moving between data centers and within data centers. That backbone is going through a paradigm shift as well, yeah?
Yeah, it is. Jensen coins the term, “The data center is the new computer.” You need to step back and say, “Right, this big building is essentially the old desktop PC we had under our desk at home.” You go, “Right, how does that work?”
All the cabling, the latency, the number of hops between each GPU, how they talk to each other, and the fabric around InfiniBand and Ethernet—it’s absolutely critical because every millisecond matters in terms of the performance of that cluster.
And where do you think—or what do you think of Elon’s vision? It’s obviously a longer-term vision of putting data centers in space, and there are a couple of other people working on it as well.
I mean, it’s very hard to argue with Elon. He’s been very right on a number of things for a very long time. I think sitting here today, it feels exceptionally difficult, given the cost of moving things to space and the challenges around radiation. There’s a huge amount of engineering challenges, but that’s never scared Elon before.
He’s uniquely qualified, and he’s inevitably right, but sometimes he’s late. He might be late to the party, he might show up at dessert, but generally he nails it.
How much of an issue is getting the data out of the data center to consumers today? Is that not something people are worried about when you’re building something out in West Texas? All that data and fiber—all that’s been taken care of, or does that become a gating issue at some point?
So, this was one of the big myths that we had to bust when we started this business, because everyone said, “Data centers must be located close to population centers, metropolitan areas. Latency is really important.” And we said, “Yeah, that’s right. Latency is important.”
But the reality is, in the US, Texas especially, there is fiber everywhere underneath the ground. Lots and lots and lots of it. And when you look at latency from our site, in the middle of the desert in West Texas, down to Dallas, the big carrier hotel—
Yeah. 6 ms round-trip latency. What’s 6 ms? There are 1,000 milliseconds in a second.
Yeah, we’re talking six. It’s effectively adjacent. It’s not even—yeah, it’s definitely not material.
Listen, continued success, and you’re hiring a lot of people.
Yeah, I think we’ve got 129 job advertisements up at the moment.
The company’s doing fantastic. Thanks for spending some time with us here at GTC.
Thanks, Jason. Appreciate it.