开源 AI 的真实图景|Token 成本将下降 10 倍、使用量将爆发 100 倍|Lin Qiao
- Lin Qiao 的创始论点直接挑战 AGI 至上主义:如果你认为智能是数据的导数,世界上绝大多数数据都是私有的、“锁在企业内部”,永远不会被分享——因此前沿将是专业化的私有智能,终局是“每个应用、每个用例对应数百万个专业模型”,而不是一个统治一切的模型。 Fireworks 每天处理的 Token 已超过 40 万亿,其中大多数来自定制模型,而不是现成模型。
- 针对 Harry 提出的 OpenAI 和 Anthropic 是否估值过高的问题,Lin 将前沿实验室重新定义为“电力线路”——不可或缺的基础设施,但不会取代其上构建的一切;当开源和闭源模型都已“跨过质量门槛”,且开放权重模型用少量专有数据调优后就能在你的评测上击败通用模型时,电力线路究竟是不是好生意? 她留下一个投资者问题:“我不希望看到的是,只有一家公司拥有智能。”
- Token 成本之所以尚未下降,是因为供应链约束;但竞争将在未来 3 年带来 10 倍降本,推动使用量增长 100 倍。 Fireworks 自身的 Token 数量到明年底可能增长“20 至 100 倍”,“我们还处在 S 曲线爆发的非常早期”。Harry 推断资本开支泡沫论“很荒谬”,她表示认同,但提醒真正的瓶颈在 Jensen 五层蛋糕的底部:能源、芯片,“我们受限于小部件——晶体管”。
- “扩张至破产”是推动企业转向开放权重模型的机制:PMF 与可持续业务已经脱钩,因为推理不像 SaaS 里的 CPU 那样是商品化资源——拥有巨大流量的在位企业无法让 AI 功能通过 CFO 这一关,因此必须拥有并调优自己的模型。 她对未来 3 年的非共识判断是:“每家公司都必须拥有自己的智能,这是必需品,不是可选项。”
- 在应用公司是否应训练自有模型的问题上,Cursor 开创了调优先例,如今几乎所有编码公司都在调优自己的模型;工具调用框架需要与驱动它的模型共同训练。 Harry 透露,Fireworks CTO Dima 曾在 Cursor 驻场数月;Lin 则介绍了跨 5 至 6 个数据中心区域、运行在分散 GPU 上的解耦式 RL 基础设施,而不是 InfiniBand 超级集群。
- Fireworks 的 AR 达到 8 亿美元,预计到年底“至少翻倍”,而团队只有 200 人。 30% 至 40% 的毛利率(SaaS 为 80%)被定义为超高速增长阶段的选择,而非新常态——“约束会拖慢创新”。明确的边界是:“我们绝不会进入应用层”;数据中心“始终在考虑范围内,问题只是时机”;芯片则被排除,因为工作负载变化太快,硬件折旧模型已经失灵——“仅一家供应商在 1 年内就有 3 个 SKU”。
- 面对 OpenRouter 排名前六的中国开源模型,Lin 给出的答案很务实:无论开源还是闭源,每个模型都要加护栏,因为每个提供商都会把自己的判断和品味注入训练过程。 如果中国限制访问,短期内“影响会很大”,但美国“能够靠自己构建那套开源系统,而且我们应该这么做”(Nvidia 正在供应链缺口之际训练一款可能名为 Nemotron 的模型)。
1. 创始论点:世界上大多数数据永远不会接触前沿模型
- Lin 回应 Harry 时给出的核心论点是:“如果你认为智能是数据的导数,那么绝大多数数据实际上并未用于训练通用智能模型。”训练语料主要来自公开互联网和标签,相对于世界全部数据而言“只是非常小的一部分”;而绝大多数数据“是私有的、锁在应用里、锁在企业里——它永远不会与其他任何人分享,因为这是公司的专有 IP”。Fireworks 要做的,就是把这些数据激活:“智能的前沿其实是私有智能、专业化智能。”
- 这份信念有其背景:她拥有分布式系统博士学位,曾在 LinkedIn 工作;到 2015 年,她已经拿出商业计划和联合创始人名单,但因为“我觉得自己在人和组织上的能力还不够”而暂缓。她加入 Facebook 时“暗中打算学 1 年或 2 年就离开”,结果一待就是 7 年,最终在 48 岁创办 Fireworks。Eric Vishria 原本坚持不招大科技公司高管担任董事,但他的顾问提醒他:“你见过多少大科技公司高管最后成功?非常少。”
- Harry 对这家公司的判断也很明确:一次 15 分钟的会面后,他开出了一张 1000 万美元的支票——“这是我 10 年投资生涯中最容易做出的投资决定之一。”
2. AGI 至上主义 vs “一支机器人军队”
- Harry 开场最尖锐的追问是:激活企业私有数据,不正是 Anthropic 围绕 Claude co-work 的论点吗?Lin 的重新定义是,Anthropic “完全相信 AGI”,其定义就是一个模型解决所有问题,而这也意味着它不会进行专业化。她的反驳不是技术性的,而是文明层面的:不同地区的价值观、政策和品味都不同;“如果未来世界要由一个标准统治,由一家公司规定品味,我们就会把自己变成一支机器人军队”。
- 她反复引用的一句话来自 Jensen Huang 在 GTC 主题演讲后的表述:“不存在所谓专业化的通用公司”(There's no specialized general company)。每家公司都建立在独特的信念之上,这些信念被写入产品设计、数据和对用户意图的理解中,“另一家置身局外的公司无法学习,也无法保证拥有这些东西”。
- 那么,Dario、Sam、Larry 和 Sergey 为什么都认为 AGI 不可避免?Lin 的回答是,他们“正在建造电力线路,用来输送一种非常强大的智能来源”——它至关重要,就像电力让她心爱的咖啡机得以运行;“但这条电力线路会取代我们所做的一切吗?我不这么认为。” Harry 把问题翻译成投资者语言,并故意不下结论:“电力线路是好生意吗?”谈到政府持股(Sam 向政府提出 5% 的方案),她提到 PG&E 的先例,最后强调:“我不希望看到的是,只有一家公司拥有智能。”
3. 开源模型跨过质量门槛——“扩张至破产”把所有人推向开源
- 最初押注在开源模型之上,是在它们“几乎还处于婴儿期”时做出的决定,根源来自 PyTorch 生态:“开放带来了控制权。”这份押注兑现了两次:开源和闭源模型都已“跨过质量门槛”,能够解决真实问题;开源模型也变得易于引导——只要加入少量公司的独特数据,就能“朝着你的评测指标不断爬升”,而且“很多时候最终结果是:用你的数据解决你的独特问题时,你会胜过通用模型”。Fireworks 自身已经在招聘、金融和内部编码代理中运行开源模型。
- 她对经济驱动力的标志性表述是:“你听说过‘扩张至破产’吗?”在 SaaS 时代,产品市场契合度和可持续业务基本等价,因为 CPU 是商品;现在两者已经分离。真正拥有 PMF 的初创公司“可能一路扩张到破产”;对于那些已经积累了 10 年流量的数字原生在位企业,情况更糟,因为它们的 CFO 无法为向全部流量推出 AI 功能找到合理性。替代方案是:“掌控自己的开放权重模型。”
- Harry 反问,Sam Altman 新发布、价格大幅下降的模型,难道不能解决这个问题?她承认“有可能”,但开放权重模型“没有采购成本”,而前沿实验室必须收回研发投入;“你就是无法定制那些通用模型……而使用开源模型,你拥有完全的控制权”。当用户达到数十亿规模时,“即便只降本 5%,意义也很大——更不用说我们过去看到的 5 倍至 10 倍降幅”。
4. 中国模型、主权与数百万模型的未来
- 面对“如今 OpenRouter 排名前六的模型都是中国模型”这一国家安全问题,Lin 拒绝接受这种提法:所有模型都应加护栏,无论开源还是闭源,因为任何提供商“都会把自己的判断、自己的品味注入模型训练过程;你无法保证它与自己的判断一致”。随后她给出本期最重要的判断:“这可能令人害怕,但我认为事实就是如此。未来将有数百万个专业模型——每个应用、每个用例对应一个模型。”
- 如果中国限制开源模型访问(上周已有相关报道),短期内“影响会很大”;但开源生态会吸引大量参与者,“就人才密度和资源而言,我确实相信美国能够靠自己构建那套开源系统,而且我们应该这么做”。Nvidia 正在供应链缺口之际训练一款可能名为 Nemotron 的模型:缺少美国本土开源模型,本质上是一个供应链问题。
- 谈到主权 AI,Harry 提到 Fable 曾被政府短暂禁用 19 天,成为欧洲的警钟。Lin 延续电力线路的比喻:“每个国家都应该拥有自己的电力线路……你不希望任何一个人切断你的供给。那将是极其可怕的时刻。”每家公司也同样如此。
5. 应用公司是否应训练自己的模型?Harvey、Legora 与 Cursor 的先例
- Harry 先披露了自己持有 Legora 的仓位,再提出问题:Harvey 已决定训练自己的模型,Legora 没有;1 年前,不自建模型的公司看起来是正确的,“现在看来它们错了”。Lin 解释说,应用开发周期已经从“数十名非常强的产品工程师和 PM、持续多个季度”缩短到“1 个人、几周时间”,因此实现本身不再是护城河,竞争已经转移到其他地方。
- Harry 的反驳值得保留:企业法律服务依赖多年关系型销售和定制化部署,“不像 11 Labs 那样拿起来就能用”。Lin 承认法律行业格外严苛——“律师通常更保守……法律对错误完全不容忍,这也是律师收费高的原因”;但她坚持认为,两家公司都拥有专有的工作流知识,而决定调用哪些工具的编排框架“需要与驱动它的模型共同训练”。
- 证据来自编码领域:“Cursor 可能是最早调优自有模型的公司之一”;“如今几乎所有编码公司都在调优自己的模型”。Harvey 与 Legora 的差异,或许只是时间问题。
6. Cursor 内部:运行在分散 GPU 上的分布式 RL
- Harry 表示,Fireworks CTO Dima 曾在 Cursor 驻场数月,帮助搭建 RL 基础设施。两家公司都对资本十分审慎,因此把强化学习拆成 trainer 和 rollout 两部分;它们没有采用超大规模云厂商使用的 InfiniBand 架构——一次连接 10,000 至 100,000 颗芯片,“极其昂贵,也非常难找”——而是“完全分布在全球 5 至 6 个数据中心区域,调动分散的 GPU”。难点在于跨区域同步新权重,避免奖励信号过时:“如果太陈旧,偏离就会太大。”
- 她用采用曲线解释这段合作的重要性:早期采用者是黑客,“他们拥有来自前沿实验室的研究人员,希望控制每一件事”;而后期大众市场只需要很少的控制权。Fireworks 瞄准后者,但与先行者合作,可以了解“要走到那一步需要什么”。
- SpaceX 收购 Cursor 后,集中度风险如何?她坦率回答:“所有人都在担心——整个行业都在担心。”整个行业都由少数达到逃逸速度的应用塑造,“所有模型公司都集中在 Cursor 上”。此后,“去年是编码之年,今年是 co-work 之年”——通用深度研究之外,法律、金融、客服、招聘和医疗等领域都在扩张,消费公司也开始重新思考如何用生成式 AI 重构推荐系统。
7. 每天 40 万亿 Token——以及资本开支泡沫论为何站不住脚
- 数据很直接:“我们现在每天处理超过 40 万亿 Token”,其中大多数来自定制模型。到明年底?“增长 20 至 100 倍都有可能……我们还处在 S 曲线爆发的非常早期。”Harry 因此得出结论,资本开支泡沫的说法很荒谬;她表示认同,并把注意力转向 Jensen 的五层蛋糕:“我们受限于底部几层”——能源、芯片,以及从未为 100 倍扩张设计过的制造产线。“我们受限于小部件——晶体管。”
- Harry 追问,Token 成本为什么还没有下降?Lin 的回答是:“供应链约束——但我们生活在自由经济中。”短缺会吸引竞争,竞争会压低成本。她的预测是:“未来 3 年成本下降 10 倍,而这 10 倍的降本将带来 100 倍的使用量。”Harry 提供的需求锚点是:一位投资者称,Salesforce 在 Anthropic 和 Claude Code 上的支出约占开发者薪资的 3.8%。
- 她坚持强调其中的细节:“不是所有 Token 都一样。”评估时应看每项任务的 Token 经济性:如果一个模型便宜 2 倍,但输出冗长 2 倍,那么完成同一任务的成本并没有变化。随着质量提升,“精确将成为优化的一部分”。
- 另一个尚未得到充分讨论的瓶颈是:“我们今天还没有为 10 万亿参数模型设计出优秀的系统。”解决这个问题,需要从模型、服务平台一路协同设计到芯片系统;她认为,这正是基础设施仍有创新空间的地方。
8. 这门生意:质量溢价、30% 至 40% 毛利率,以及技术栈的边界
- 对于 Together 更便宜的说法,她回应:“我们可能不是在比较同类产品。”Fireworks 的大多数流量来自定制模型,并以质量优先进行优化,甚至做到“零 KLD”:训练系统和推理系统在比特层面等价,“我们不会损失任何一比特准确率”。否则,“你为什么要用打折后的质量来支付训练投入?”Fireworks 采用一套配置服务一个场景的部署方式,并由一支应用机器学习团队(FDE 式)支持,目前正在开发自动化部署的代理。
- 对于相较 SaaS 80% 毛利率而言只有 30% 至 40% 毛利率的问题,她说:“我不认为这是新常态。”这反映的是超高速增长阶段。“利润率优化本质上是约束问题……而约束会拖慢创新。”只有确定系统未来要扩张 1000 倍后,才值得把它优化到极致。最糟糕的情况是,把利润率一路推高,然后停止增长。边界也很明确:“我们绝不会进入应用层”;数据中心“始终可以考虑,但问题只是时机”。
- 芯片则是完全不同的问题:“我知道造芯片极其困难。”Meta 的 MTIA 大约从 2018 年开始,服务于排名和推荐工作负载;只有在“工作负载稳定下来之后”才会流片,而如今的 AI 工作负载“非常、非常动态”。生成式 AI 之前的加速器创业公司多少是误打误撞:那些大量使用 SRAM 的设计,恰好适合内存需求旺盛的模型。她还提到 Nvidia 最近收购 Groq——后者高度依赖 SRAM 的芯片,天然适合与浮点运算密集型 GPU 搭配:前者负责 prefill,后者负责生成。这种异构数据中心设计让她确实很感兴趣。
- 折旧模型正在打破自建与采购之间的算术:过去硬件按 6 年折旧,而产品 3 年迭代一代;“现在仅一家供应商在 1 年内就有 3 个 SKU”。模型的性能高峰每周都在刷新,“更新的模型通常在最新硬件上运行得最好”。3 年之后,“你还想让 9 代以前的硬件运行 3 年前的模型吗?这很值得怀疑。”
9. 8 亿美元 AR、翻倍目标、George Hu 与 Jensen 的运营体系
- 业务轨迹是:AR 达到 8 亿美元,“我们认为到年底至少可以翻倍”,团队规模为 200 人,1 年前只有 50 人。最能说明客户增长的故事是 Cursor:签约时 Fireworks 的收入还只有“个位数百万美元”,就在 2 年前;随后 Cursor 在 2 年内增长了 100 倍至 1000 倍,选择 Fireworks 作为平台研发合作伙伴,自己专注于产品。Harry 用 Slack 的案例描述那个时代:18 个月从 100 万美元增长到 1000 万美元,曾经是风投界的黄金标准。
- 关于聘请前 Salesforce 总裁 George Hu,她 1 年前曾对他说:“我们可能对你来说还太小。”过去 1 年里,他先帮助面试高管,随后随着业务增长进入关键阶段而加入。她筛选人的标准并不是能力,而是“他们是否真正具备极致的主人翁意识”;“我们不是把人装进一个个盒子,再把盒子堆成一座塔。”
- Jensen 的启发来自他总能在 1 分钟内回复邮件:“领导力就是判断力,不是特权。”在高速变化的行业里,不能等信息逐层传递;“不知道究竟发生了什么,却拥有做判断的职位,这就是糟糕的领导力。”她今年也改变了一个看法,不再担心快速扩招会扼杀敏捷性。她承认自己犯过的错误是迟迟没有推进营销:“营销不是流程,而是教育……是清晰度。”
- 她对未来 3 年的判断是:“每家公司都必须拥有自己的智能,这是必需品,不是可选项。”这与每家公司都拥有自己的软件栈是同一套逻辑。而行业下一阶段将从“Token 极限优化”转向“ROI 极限优化”——前者只是阶段性现象,后者关乎如何真正经营一家企业。
What I don't want to see is only 1 company owning intelligence. That doesn't make sense to me. I think last year was the year of coding, and this year is the year of co-work.
I do think the cost of tokens will go down drastically. There will be a 10x cost reduction in the next 3 years, and this 10x cost reduction will drive 100x usage. We absolutely are not going to move into the application layer. It's very unclear to us whether we will move down into data centers and so on. That could always be on the table, but the question is—
Ready to go, Lin? I am so excited for this. I heard so many great things. I just got off the phone with your co-founder, Dima. I spoke to Alfred Lin, Sonia, Matt Miller, and many more, so thank you for joining me.
Thanks for having me.
Now, I heard that Eric Vishria has a rule: don't invest in big tech executives. But he broke that rule with you, which is very special.
I think so, too. A funny story: after we decided to shake hands, he called me and said he had talked with one of his advisers. His adviser questioned him: “Hey, how many big tech executives have you seen being successful in starting a company?” Very few.
He told me that, and I was surprised. I said, “Are we breaking our handshake now?” He said, “No.” Since then, we have worked very closely with each other.
Eric is one of the best. You also started the company when you were 48.
Yeah.
1. Why Starting a Company at 48 Was an Advantage
That's quite late. Can I ask you how you reflect on being a 48-year-old founder when we glorify starting a company when you're pretty much 15 these days?
I didn't think deeply about that. I always wanted to have a tech business myself. I actually wanted to start a business in 2015. I'm a first-generation immigrant and came to the US in 2000. I did my PhD in distributed systems and computer science, especially focused on databases.
Databases are a very complex system to build, with a lot of different objectives to optimize for. I pretty much touched every single aspect of processing data. Then I moved to LinkedIn to further drill down and build systems and products to be used to drive real impact.
At that time, I felt I was ready to start a company. I knew all the tech, I knew what product to build, I had a business proposal, and I had a list of people I wanted to start a company with. I spent time thinking about it, and I paused because I didn't think I had the skill set around people to build a company.
It's not just about the product. It's not just about the tech. It's actually about people. I decided I wanted to go to a place where I could learn the most about people, and the best company at that time was Facebook. It was a rising star in Silicon Valley.
Secretly, I was planning to learn for 1 or 2 years, leave, and go back to do my own business. I stayed there for 7 years.
2. The AI Layer Everyone Is Overlooking
With Fireworks, you saw something in inference that the world was not focused on. The world was focused on training. I think it's helpful for people to understand the stack, because beneath you, there are obviously chip providers—the NVIDIAs of the world—and above you, you've got the model providers, while you sit in between. Why is that a valuable part of the stack and not a commodity?
That's a really good question. Why bother with specialized intelligence? Why not just use generalized intelligence and worry about fewer things? You could just build on top of an API provided by frontier labs. Wouldn't that be much easier?
The argument is the following: if you think intelligence is a derivative of data, then the majority of the world's data is actually not used for training a general-intelligence model. The training data is coming from the public internet and labeled data. The public internet is a very small corpus of data compared with the world's data.
The majority of the world's data is actually private, locked inside applications and locked inside enterprises. It will never get shared with anyone else because it's a company's proprietary IP. So then the space becomes very interesting, because the majority of data is not being activated to derive any intelligence.
That's where we believe our role is: to activate that data. We believe the future frontier of intelligence is actually private intelligence and specialized intelligence. That's where Fireworks came from. From the beginning, we have been focusing on driving the value.
I have so many questions to ask you. I totally understand you in terms of the value in private data within some of these largest companies. Is that not the premise of Anthropic's enterprise business, though, with Claude for Work and with a lot of the adjacencies that they're building? Would Dario not say that that's exactly what we're going after?
That's interesting, because I view Anthropic as a company that fully believes in AGI. The definition of AGI is that there's 1 model that can solve all problems in the best way. To me, that's the definition of AGI.
That means you do not need to specialize. That 1 model should be able to solve all the problems. If it's so intelligent and has so much knowledge of every part of the businesses and every part of the jobs, and it can fulfill all of them, then why do you need to bother specializing?
That itself is a validation that we're living in a world that's not ruled by 1 principle. We are living in a fully diversified world. Give you 1 example: different regions will have different value systems and different policies. They'll have different ways of conducting business and different lifestyles. It's all taste, choices, and judgment combined.
I think that's what defines us as human beings. We are not robots. If our future world is going to be ruled by 1 standard, with a taste dictated by 1 company, we turn ourselves into an army of robots. That's very depressing to me.
I think what separates Homo sapiens from other species is creativity—the deep desire to pursue new things and discover new ways of living. That defines us as human beings, and that part cannot be copied.
That's why, in Silicon Valley, there's so much creativity. Across the world, there's so much creativity in building new businesses. What is a new business?
I had an interesting conversation with Jensen after his GTC keynote. We actually recorded it, and it was interesting.
I watched it. It was great.
Recording with Jensen is not really recording. He just started having a conversation with me. I didn't know his crew had already started recording, and we just kept talking. It's so easy.
We talked about specialized intelligence. He said 1 thing to me: “Lin, you're right. There's no such thing as a general company. Every company is built on a special belief about doing things; otherwise, there's no reason it should exist.”
It feels logical, but then I started to think back about what he said. It's profound, because every single company is doing something unique that justifies its existence. This something unique is deeply baked into its product design, its software design, and its system building. It's deeply baked into the data and its interaction with its users, and into its deep understanding of its users' intent, their interaction with the product, their engagement, and so on.
All of that is the fundamental basis of why a company should exist. That is not learnable or assured by another company sitting outside.
3. Is AGI Really the End Goal?
Can you help me understand, then? As a podcaster, I specialize in asking basic questions, so forgive me. Why, then, do people like Dario, Sam, Larry, and Sergey talk about AGI in the way that they do—as inevitable?
I think what they build is fantastic, because they are basically building a power line to distribute a really great source of intelligence that everyone else can build on top of. That's how I view their contribution.
4. The AI Infrastructure Race Is Just Getting Started
If we don't have this fundamental infrastructure, then we will not have all kinds of appliances living in our homes. I love my coffee machine, and it's specially branded, right? But without that power, we don't get to do the things that are fun, that are unique, and that are special—things that ingrain and encode our taste.
5. Why One Company Should Never Control Intelligence
So I do think that's very, very important. But the question is: is this power line going to replace everything we do? I don't think so.
The question for me as an investor is: are power lines good businesses? You said something about PyTorch and the open ecosystem. Open source, in the last, I would say, 3 months, we've all realized, is accelerating so fast, and the capabilities have increased to such an extent that it's not quite comparable, but it's getting 90% as efficient and, as Chamath stated, 15 times more cost-effective.
Are power lines good businesses in a world of open source?
6. Open Source vs Frontier Models: Who Wins?
Here's how I view open source. Early on, when we founded the company, we had a pretty deep debate among the co-founders: what do we do? Do we build our own models, or do we build on top of open models?
At that time, open models were almost in their infancy. If we were going to take that direction, it was a huge bet that they were going to do well. But with our PyTorch experience, we believed in the open community. We believe in openness. That's a fundamental principle we operate with.
Openness gives control. Openness gives control to the user. Think about open models: once a model is released, you have full control of the weights. You can change it however you want. It's yours.
And then you can build on top of it. That is a fundamentally different operating principle that we believe in because of our roots in open source. We took that bet, and it did pay off in the sense that the quality of both open and closed models has significantly improved over the past 2 years.
Both of these streams crossed a quality threshold to the point where they can solve so many problems. Within Fireworks, we do our own product development. We use open models to drive our recruiting process, candidate sourcing, and feedback collection. We use open models to drive some internal finance processes and, obviously, for coding and reasoning models to help us debug.
We have a ton of agents within Fireworks ourselves, and we are cost-conscious. Both model categories have crossed a threshold and can solve a wide variety of problems. Second, once open models cross a threshold, they are much easier to tune. The ability to steer a model is part of its intelligence, and models are much easier to steer, especially with a small amount of data.
A particular company might have a small amount of unique data, and then we can hill-climb toward its eval. Oftentimes, the end result of hill-climbing is that, to solve your unique problem with your data, you are better than a general-purpose model.
7. Are AI Giants Massively Overvalued?
When 90% of enterprise workflows can be done, as you said, through the incredible array of functions that you now use open source and open models for, the usage for frontier models will not be as large as it would be if they were needed for everything. Are these companies actually dramatically overvalued and overestimated if the majority can just go through open models?
I think people are starting to realize it. I remember, 2 years ago, I went to different places and talked about an interesting phenomenon that did not exist in the SaaS era. During the SaaS era, product-market fit and a durable business were almost equivalent to each other. The hardest thing was finding product-market fit, and then, once you found it, scaling as fast as you could, because CPU is a commodity. The infrastructure you build on top is almost a commodity, so you do not even have to worry about it.
Now, product-market fit and a durable business are 2 separate concepts. For startups, we have great companies that have product-market fit. Customers want to pay them, and they really value their products, but they cannot scale because, once they scale, they could scale into bankruptcy. Have you heard about “scaling to bankruptcy”? That is a real problem.
It is an even bigger problem for incumbents, the big digital-native companies, because they have the traffic. They got a winner a decade ago when they were startups, and they have a huge amount of traffic. Once they roll out those AI features, they are going to reach their entire customer base, and they cannot afford to do it because their CFO looks at the cost proposal and cost forecasting and says, “There is no way you can justify this.”
It becomes a real problem for all those innovators. They really want to plug into this new, disruptive technology, but they cannot afford it. They need to find an alternative to be able to afford it, and the alternative is to have control over an open-weights model and roll out their own model.
Or you see what Sam Altman released in the last few days, which is dramatically lower-cost models. I cannot remember the amount, but I think it is half as expensive, or maybe 3 times cheaper. Is the next step that we just see a massive reduction in price from the frontier models?
It could be, but at the same time, it is just a very different operating principle. For open-weights models, model acquisition has no cost. Obviously, some companies train those models and are willing to open them up. I know that, within the US, there are multiple companies doing that, including NVIDIA, which is training Nemotron. We are working very closely with them.
8. Why Open Models Could Beat Closed AI
Once the model is there, whoever uses that model has literally no cost. There is a fundamental cost for the frontier labs to invest in those models and recoup their R&D costs. Second, you just cannot customize those general-purpose models. You use them as is, on top of an API that you have no control over.
With an open model, you have full control. You can tune it however you want, and you can use it however you want. Fireworks is a specialized intelligence platform. We offer all sorts of tools for you to easily customize the model for one specific use case. After that model is tuned with high quality, we further optimize it for inference deployment.
Think about Fireworks this way: We think about every single model deployment as “one size fits one.” It is unique for your workload only, and it is optimized for your workload only from a quality, speed, and cost point of view. We believe that is absolutely needed, because once you think about production scale—reaching millions of users, tens of millions, or billions of users—even a 5% cost reduction means a lot. It is a massive amount, let alone the 5–10 times cost reduction that we have seen in the past.
9. Should We Trust Chinese AI Models?
The 1 question that I do have to ask is about the concern enterprises have around national security. When you look at OpenRouter, I think the top 6 models today are Chinese models, and they are incredible quality. The speed of development is incredible, but they are Chinese models. Do we have serious national security concerns when analyzing the power of Chinese open source?
I think there is a huge debate happening right now across the industry. Once the model is open, you can put all kinds of guardrails specialized to your business around it. I would say that, for all models, it does not matter if they are open or closed: You should put your own guardrails around them.
The fundamental reason is that a model provider will infuse its own judgment and taste into the model training process. You cannot guarantee that it matches yours. Remember, it goes back to Jensen’s comment: There is no such thing as a general company. Every company is special. Every company will have a special design principle, a special taste, and a special target audience to serve.
10. How Cheap Will AI Become?
Because of that specialty, it is guaranteed that the judgment, taste, and design principles from one company will mismatch or misalign with your company, which is a special problem. That is the reason you need to tune those models to match yours.
I really believe the future will not be a small number of AGI models dominating the world. I really believe the future may be scary, but I think that is true: It will be millions of specialized models, 1 per application and per use case.
We saw reports in the last week that China was looking at restricting access to its open models because it was seeing the development happening so quickly and at such a high quality. What would happen in a world where China actually started restricting access to its open models, given the lack of open models we have in the US?
I think it would have a big impact in the short term. But the beauty of the open ecosystem is that it is not 1 provider. That is why it is open. It usually attracts many interested parties to participate.
In terms of talent density and resources, I do believe the US will be able to build that open system by itself, and we should. I have seen this happening again in many open systems: There are a thousand flowers blossoming, and that is the beauty of it.
When we talk about the specialization of intelligence within enterprises, as you have described, let us take a prime example—which I do not particularly want to take because I am an investor in Legora, and I think I know which side you are going to fall on here. You have 2 companies that compete in the legal space: Harvey and Legora. Harvey has committed to building its own model, and Legora has not.
A year ago, it looked like companies that did not commit to their own model were right, because frontier models were increasing so quickly in capability. Now it looks like they are wrong. Should companies like Harvey and Legora be building their own model? And, actually, if you do not, what happens?
Here is 1 observation I have, and many people have as well: Software development, and especially the SaaS space, has been significantly disrupted because of the general intelligence of coding. The application development life cycle has significantly collapsed in terms of the timeline and resources needed.
In the past, it required tens of very strong product engineers and product managers to convert an idea into an implementation and then into a production-scale product. That required multiple quarters or years of investment. That was a deep moat. Today, 1 person in a few weeks can possibly launch their idea into a product and scale quickly.
This is unprecedented, and it also creates interesting dynamics and redefines where the competition is, because it is really hard just to compete on the idea of an application. The application by itself is no longer enough, because many people have similar ideas. Implementation is no longer such a big barrier.
Is that actually true, though, when you are looking at enterprise deployment and enterprise rollout, if you are working with some of the biggest law firms in the world? The enterprise sales cycle is at least multiyear, with relationship-building, which is very tough, and then you have deployment that is very customized. It is not like ElevenLabs, where you pick it up and go. It is different.
Also, I think the legal space is particularly challenging because lawyers are usually more conservative.
Legal is also not tolerant at all of errors, right? Because that’s why lawyers get paid, right? You need to build a very rock-solid case. If something hallucinates and generates the wrong judgment, then you’re in trouble. So I do think the legal space is a very interesting space to penetrate, and these companies are both doing a great job.
But on the flip side, I do think both companies own proprietary knowledge and information on how to build those assistants, to do case studies, to go deep in driving legal research and all this, right? My understanding of legal is so shallow, but there are so many different versions or flavors of cases. So I do think they are in a unique position to convert that deep understanding, and they all have data.
It’s not just about how defensive their business is. It’s about, hey, oftentimes when they build those assistants, there’s a harness integrating, deciding, and orchestrating which AI tool to use, which tools to call, and this is bespoke, this is customized. The accuracy of calling those tools, and calling what kind of tools, is important. Even that harness needs to be co-trained with the model powering it, right?
So there are just ample examples of driving that business to excellence by owning their own intelligence of how to do that in the workflow layer. So maybe it’s a timing issue. Coding, for example, I think in the coding space, Cursor is probably one of the pioneers.
11. Will AI Model Breakthroughs Ever Slow Down?
Starting to tune their model, and now almost all coding companies tune their own models. Does that pace of model development slow down? Because every single day, it seems like we have a new model with a new capability, and it’s like, “Oh my gosh, Cursor’s newest model is amazing.” Next, we have someone else’s—Mistral’s—newest model is amazing. Gemini’s newest model is amazing.
In 3 years’ time, will the pace of model development still be so fast, and model superiority be so transient, where one day it’s one and the next day it’s another?
There are a few layers of model advancement. There’s base general IQ advancement, so those will take step functions. That’s why when they release, there are always major releases or minor releases, right? The major releases are step functions.
As you remember, at the beginning of last year, this whole thinking process was new, right? The model just doesn’t spit out an answer immediately. The model will think by itself and spit out an answer; it’s much better that way. So that’s one step function, and there are many step functions we have seen through.
But I see those as every year or every 3 quarters, there’s a major leap. At the same time, building on top of those base models, I can see specialization start to accelerate because, as I said, it’s really like a tree, right? There are so many branches and leaves that can possibly hang on the trunk. As the base-model quality starts to have step-function leaps, there’s so much more we can do to specialize.
So I do see specialization in the world accelerating much faster than the general-intelligence part.
When we think about the general-intelligence part, just before we move further into the stack of multi-model, you saw Sam proffer the 5% kind of gifting of OpenAI and others to the administration. Do you think we’ve reached a stage where model development is so advanced and so important to society that they will in part be government- or administration-owned?
That’s a very interesting question. I think there were precedents of that. If we think about the foundation tier of those general-intelligence models as fundamentally a base infrastructure for the big economy to operate around, there have been precedents of, like, PG&E owning electricity and gas, and so on, right?
I actually don’t know, but what I don’t want to see is that there’s only 1 company that owns intelligence. I think that doesn’t make sense to me because there are different flavors of intelligence. There’s this general, common intelligence that benefits everyone, and then there’s specialized intelligence that actually helps us advance in history, to think differently, to create new paradigms of living or new paradigms of doing business and shaping industry.
I don’t want that to die because there’s only 1 company that can do that. I don’t think that makes sense.
With the many models blooming, theoretically, there’s the idea that you will route a task to different models depending on what they specialize in.
I think so.
With that in mind, will you not build your own OpenRouter for the world to cater to that?
Yes, you can argue they’re the best builders because they deeply understand their use case and they have the evals. So, again, my thinking about what the frontier is is not just this 1 model. The frontier could be your special routing mechanism for your business.
You decompose that based on, hey, in order to fulfill this task, usually you need a highly intelligent layer, maybe the most expensive open or closed models, to judge at the highest complexity. Usually, people will also build sub-agents to solve smaller problems. Then those can go to smaller open models, and those can also be further customized to fit into your special design.
I’ve seen a lot of people already doing that today, and we also think there’s a space to build an automatic routing system that can learn by itself. Compounded with an automatic tuning system, eventually we think it should all be automated. Then you can see a self-evolving system based on what flows through your product.
Your product keeps evolving. Your product is alive, right? So you keep deploying and launching new features to interact with your users, and it will be a totally self-evolving, automated system.
Do you think, then, that routing layer of the stack is valuable if it can be automated or it can be built on its own? Is that a valuable layer to have?
I definitely think so.
You do think so?
I do think so.
If it can be automated or companies can build it themselves, why would you need a Requesty or an OpenRouter?
You probably don’t.
Yeah, we’re not there yet. But I do think this could be an area of innovation.
You said Cursor was one of the front-runners in terms of how innovative they’ve been. I completely agree with you, but I heard—and I really stalk you before shows—that your CTO, Dima, was embedded at Cursor for months building the RL infrastructure. Is that how it has to be done, and is that scalable?
What’s happening is usually, in the early adoption curve of new technology, the early adopters are all hackers. Hacker is not meant in a bad way; it doesn’t have a negative connotation. They have deep expertise in a certain area, and they want to control a lot of things.
Versus in the late stage of a new tech adoption curve, it starts to get more accessible to a much bigger cohort of users that doesn’t have deep expertise, and they need less control. So it always goes into deep control first, usually, and a little control later.
We definitely are aiming toward the later stage as the ultimate end state we want to target, but it’s also extremely valuable to understand what is required to get there. So that’s why we partner deeply with Cursor. They are the pioneer trying those ideas. They do have researchers from frontier labs, and they want to control every single thing.
At the same time, we’re also pushing to the boundary. We’re doing things that never existed before. We’re doing systems that never existed before because we push the boundary; that is unique to this particular setting.
Okay. What is unique here?
Typically, if you think about training, training happens; training is very capital-intensive. It usually happens in big companies. They have a lot of money, and they put that money toward buying very expensive training clusters interconnected with each other. Super expensive.
Once you have those expensive, large fleets, usually you don’t need to think too deeply about how to be efficient; you just focus on doing your work. Cursor is like us; they’re a startup. Both of us are very capital-conscious, and we want to be efficient while we don’t want to slow down research innovation.
So together, we figure out a very smart way to drive their training process. They do massive post-training, which is reinforcement-learning-based, and we break the reinforcement learning into 2 pieces. One is the trainer that is tweaking the weights of the model, and it basically generates new model versions constantly.
That new model will be deployed to what we call the RL rollout. It basically deploys that new version to interact with a synthetic environment—a synthetic coding environment—or a real coding environment, and then gets the reward back to judge if that model version is good or bad, right? So that’s a rough process.
We decouple these 2. In the past, in large hyperscalers, they ran that all together. If you think about it, you get 10,000 or 100,000 chips all interconnected through Infinity Band. It’s extremely expensive and really hard to find, but they need to go really quickly.
We designed a fully distributed system, and we run it across 5 or 6 data center regions globally, tapping into scattered GPUs, and they’re able to run massive jobs—our jobs. But the challenge there is that we need to sync model weights across all these different regions.
How hard can that be? It matters because the latency, the delay of sending these weights over, is going to dictate how fresh the rewards are. If it’s too stale, then you’re too far off. So it’s a balance, but we innovated a way that we can distribute fresh model weights quickly.
12. Is AI Coding Already Yesterday's Biggest Trend?
It’s not too off. Numerically, it’s still sound, while we’re not limited by a very expensive deployment of a GPU fleet. Those are the innovations we worked on together with Cursor to push the boundary, leading to their recent model launches. We’re very proud of them.
Can I ask a blunt question? Cursor is an incredible customer to have, given the amazing progress they’ve had with you, and it’s wonderful to see that partnership. It’s a very large customer for you. How do you think about the concern of Cursor churn in the wake of a SpaceX acquisition?
Yeah, everyone’s concerned. The whole industry, in terms of application innovation, is driven by models, in the sense that there are a few companies that are very, very successful. They achieve escape velocity, but there are only a few of them. So that’s the shape of the whole industry, and last year, Cursor was one of the few.
I would say all model companies are concentrated on Cursor. We concentrate on the same group of app companies, and since then, it does change, right? So we do have a very healthy, diversified customer base.
Especially, I think last year was the year of coding. I think all major coding companies are on us, and this year is the year of co-work. Co-work is much more diversified by itself than coding because there’s general-purpose co-work. For example, general-purpose co-work to help you do all kinds of research.
You want to ask, “Hey, what will the NVIDIA GPU price be 2 years later?” or “What will be Anthropic’s stock price after its IPO?” Those are deep-research, general-purpose applications. There are also so many different categories of special-purpose co-work: legal—we just talked about 2 great legal companies—finance, customer support, recruiting, sales, marketing, and healthcare. So there’s a very broad set of co-work applications and innovative app companies. They’re doing really well, and we have them as our customer base.
More interestingly, we’re starting to see an uptick in consumer-facing companies looking into GenAI technology. They’re changing how they think about their traditional businesses of doing recommendations, for example. That’s very interesting to me because we’ve obviously worked at Meta, which has one of the biggest recommendation systems in the world, and we’re very eager to see how that transforms into a new economy for us.
I’m sorry for being naive here. Do people work with just one provider in the inference space, like you, or do they work with you and together—or anyone else in the space?
I think people are more attuned to a multivendor strategy in this space because they don’t know what’s happening. It feels safe to have multiple providers to balance things out.
But we don’t view ourselves as an inference provider. Again, we view ourselves as delivering this specialized intelligence, where we help companies tune their models. To give you some numbers, today we process more than 40 trillion tokens a day. The majority of those tokens are coming from customized models, not off-the-shelf models.
What will that token count be at the end of next year?
Anywhere ranging from 20x to 100x could be possible.
20x to 100x.
Yeah, we’re at a very early stage of S-curve explosion right now.
20x to 100x. If it’s 20x to 100x, the idea that we’re in a capex bubble is ridiculous, and we desperately need far more capex than we’re ever suggesting for compute. Is that right?
That is right. At the same time, I think Jensen has a 5-layered cake—a 5-layered AI cake. From top down: application, model, infrastructure, chips, and energy. We’re bottlenecked by the lower part of the AI cake in terms of the supply chain.
Being energy—
Being energy, being chips. I think, in the physical world, how fast we can manufacture is the question because, historically, all these industries were not designed for massive scaling. Speaking about 100x scaling, no one was designed for that.
I talk with many manufacturers, and we’re bottlenecked by small parts—transistors, the smallest, tiniest parts that hold up the whole manufacturing line of servers that can be deployed to data centers and used to generate tokens.
Do you have to be full-stack, according to Jensen’s 5-layered AI cake? Do you have to then be full-stack to win or to reduce dependencies? We’ve seen OpenAI come out with Jalapeño—terrible name. Anthropic is talking to Samsung about building its own chips. DeepSeek is building its own chips. Zuck came out with Meta building its own chips. Do you have to be all of it?
It really depends on the company’s philosophy. To us, agility is everything, and we need to earn the right to build anything. Focus is everything for us, and we want to focus on where we add the biggest amount of value based on our strengths. We’d like to leverage other people’s strengths to build on top of them.
In particular, we want to run everywhere on all possible AI chips in the world, and we don’t want to be limited by how many chips we can bring into our data center, whether we construct it or rent it.
Over time, when the business grows very big, I still remember when Meta was young. They didn’t build everything, and when they’re big, it makes sense to build. You earn the right to build your own giant infrastructure, and if it saves 5x more cost, then you’re sure to go do it, right?
At the early stage, that’s why I’ll tell you an interesting story. In the coding space, I would say Cursor was the first company that decided to work with us early on. I remember when they worked with us, they were a single-digit-million-dollar company.
Wow.
They were very small. This was only 2 years ago. They grew by 100x to 1,000x over 2 years, something like that. But they decided to work with us early on because they recognized they only wanted to focus on product innovation and, later on, research.
They did not want to focus on platform innovation. They knew we were putting all our R&D in there, and they wanted to find the best partner to win big. So I do think that’s the right mentality: to specialize. We want to specialize. We do not want to own the whole stack. That’s not our goal as a company.
I’m sorry to be harping on. Why does Jensen skip your layer of the cake? Because he’s doing NeMo with models.
Well, Jensen isn’t building a cloud either, right? You can say, “Hey, Jensen, you probably have all the rights to build an NVIDIA cloud.” He’s not building cloud infrastructure. I think he mentioned that as well, and many people ask him that question. He also mentioned that he wants to specialize in what they have the rights to do.
Why models? I think it’s purely a supply-chain question. If the US doesn’t have a US-native open model, it’s a problem. It’s a supply-chain problem. So he’s solely there to solve the supply-chain problem.
But if there’s no supply-chain problem because companies like us are providing this specialized intelligence platform layer, then he doesn’t need to worry about it. He just wants to make sure the whole 5 layers of the AI cake are flowing. There’s no blockage, and if there’s a blockage, he’s interested in solving those problems.
Marc Benioff, one of your investors, I think in the new round—which obviously will come out after the round is announced—said that he spends about 3.8% of developer salaries at Salesforce on Anthropic and Claude Code.
I think it’s a useful analogy because, if you assume that that’s what’s spent on Claude Code and coding tools, that says one side of the market. But if it’s 20%, wow, we’re underestimating how big these companies can be.
When you think forward 1 or 2 years, what percentage of developer salaries do you think we’ll spend? Is it less because these tools will get cheaper, or is it more because they’ll get better and better?
I do think the cost of tokens will go down drastically because, again—
It hasn’t so far.
It hasn’t so far because of supply-chain constraints, but we’re living in a free economy. Think about it: whenever there’s a shortage, the price is high. High prices will invite a lot of people to come solve the problem, and they’ll invite competition. Competition will bring down the cost, and eventually it will lead to a very economical solution.
Actually, that’s good for everyone because much more affordable infrastructure will invite more usage. My prediction is that infrastructure costs will go down and usage will explode because of that. The moment you don’t think about that as a problem and, if it’s a utility, you just use it.
How much will token costs come down? Help me understand. Is it like a halving? Is it like, “Oh, it’ll be a hundredth of the cost”?
There are different ways to think about this. It’s not that all tokens are equal. I think we should establish best practices to evaluate the token economy per task because different models have different ways of spitting out tokens. Some are much more verbose than others.
You can imagine one model is 2x cheaper than the other, but it takes 2x more tokens to solve the same task, and then they’re at the same cost.
Yeah.
Right. But overall, as model quality improves, I think being precise is going to be part of the optimization. That’s one level of optimization: to solve one task, we should need fewer tokens.
The second is how to do that for one token. You need to customize the model to solve your problem especially better and more precisely. That goes into model tuning.
Second, for each token spit out from those models and processed by those models, we also specialize in making the unit economics much better through our platform. Third is the underlying infrastructure, like the GPUs, the surrounding memory, and all these things. Today, it’s under stark supply chain constraints, but the situation will get much better. I don’t think it will probably change in the next 1 year or 1.5 years, but in the long term, 2 to 3 years, it should change, and that cost will compress. So overall, I can imagine a 10x cost reduction in the next 3 years, and this 10x cost reduction will drive 100x usage.
You said there about token efficiency and how you enable your customers to be much more efficient. With that efficiency, you do charge more. When I did the research and compared you to competitors, I got that Together’s the price king. I don’t mean this disparagingly, but they’re cheaper. If you want cheap, you go there, and respectfully, if you want a better-quality product, you go to Fireworks. But it is more expensive. Do you think that’s a fair assessment and a fair analogy?
I think we’re probably not comparing apples to apples, in the sense that it again goes back to our business. The majority of our traffic is customized models, and we optimize for quality—number 1, always quality. Quality as in model quality toward your applications, your specific business, your use case, and so on.
The second is, when we deliver those models in inference, it’s also quality. We care about quality so much that we do extreme things. For example, during training, there’s a very hard thing to achieve called zero KLD. It’s a little bit technical; the idea here is—
Zero KLD?
KLD is a measure of quality. What it means is that between the training system and the inference system, when the model moves over, we have bit equivalence. The numerics are fully the same; we do not lose a bit of accuracy. That’s really hard to achieve.
The reason we push that and deliver that is because we know our primary business is in model customization and inference of customized models. We want our customers to maximize every single dollar invested in training. If, across the training-inference boundary, it’s not bitwise equivalent, they just drop the quality down, and then it’s like you pay your training investment with discounted quality. Why do you do that?
So, quality first. Quality does bring additional value, and that’s why we are not interested in commoditized, one-size-fits-all. This off-the-shelf model deployed in the same way for everyone—that kind of business will always customize the model and deploy it in a unique way for your particular workload.
Two questions. Do you have to have an FDE model to make the customized model efficient?
As a matter of fact, we do have an FDE team. It’s called the Applied Machine Learning Engineering team. Their primary job is to accelerate this customized deployment and, as a matter of fact, also build the agent to automate a lot of deployments.
13. Hypergrowth vs Profit: Why Margins Can Wait
Given where we are in the stack and the amount of complexity that we have, we have a margin structure that’s a little bit different from traditional SaaS being 80%. I don’t know the margins precisely here, but they’re traditionally in the 30% to 40% range for where we are. Is that the new normal for where we are?
I don’t think that’s the new normal. I think that, at least for us—I don’t know about other companies—it’s a reflection of the fact that we are in a hypergrowth phase.
During a hypergrowth phase, you have the choice. To me, margin optimization is a constraint problem: We want to go to 70% margin, we want to go to 80% margin, and then we’re going to go backward and impose those constraints to guarantee those margins. Usually, constraints slow down innovation.
To give you an example, during system development and in a high-velocity, system-expanding phase, we don’t want to overbuild because we’re in a high-experimentation phase. We’re testing what will stay and what will not stay. Optimization doesn’t make any sense. Once we know this is a system that we want to build 100%, and we’re going to scale this 1,000 times bigger, then we go optimize the heck out of it. I think you think about business the same way.
In hypergrowth, if our focus is only on optimizing gross margin, we absolutely can do that, but we’re sacrificing the speed of growth as well. We want to go everywhere. We want to go into different geographies, tackle different use cases, constantly create different product lines, and those are not the times for optimization. That’s my opinion.
So we will be able to increase margin without moving into different layers of the stack?
We absolutely are not going to move into the application layer. That’s very clear to us. Whether we will move down into, for example, building data centers and so on, that could always be on the table, but the question is timing.
Isn’t the statement, “You either die, or you live long enough to build your own data centers”? Elon and Zuck are now spending, I think, $10 billion on the latest data center in Canada. Would you like to build data centers?
I have built data centers that matter, and there’s also lots of innovation possible there. There’s no one-size-fits-all as well, and building a GPU-native data center is also interesting. I think there is a potential direction of building one.
It’s a trade-off, right? From an operations point of view, it’s much better to build a homogeneous deployment. It’s all the same chips, all the same SKU, as big as possible, and you run multiple workloads so it’s fungible. It’s very easy to manage. You build 1 principle, 1 process to do maintenance operations.
But again, it goes to optimization. Once it’s so big, then any optimization is going to drive a lot of economic return. For example, we’re talking about NVIDIA recently acquiring a company also called Groq, with a Q. It’s a large SRAM-based ASIC accelerator.
I spoke to Jonathan before this show. Jonathan is excellent.
He said, “What a fan he is of yours.”
Oh, also a fan of his.
It’s a great combination between a FLOPs-intensive GPU and an SRAM-intensive ASIC. FLOPs-intensive is really good for the first half of LLM processing—it’s called prefill processing, processing the prompts and so on—and SRAM-intensive is really good for generation. That’s just the nature of the model architecture.
It’s great to combine these 2 instead of running homogeneously on the same chip, right? But that requires a very unique system design and deployment into a data center. It is heterogeneous, actually. I really mean homogeneous design is much better for operation. This is heterogeneous, and operating this heterogeneous design requires unique innovation in data center deployment and so on.
So data centers aren’t commoditized. You can specialize in data center deployment, and one data center is better than another. Data center deployment can be done well and badly. Data centers are so complicated, right? If you think about it, from the beginning all the way through construction and power deployment, you need to have the right power come in, the fiber channels, the right cooling—especially with new chips, which require liquid cooling—to get all this right, and the parts can fall apart and you need to know how to replace them. It is all very deep expertise. It’s no joke. It’s not that tomorrow I can be a data center operator. I cannot.
14. Can the West Keep Up With China's Infrastructure Speed?
Is that not where you would bet long on China, with the greatest of respect? Especially in the US, one of the biggest barriers to data center deployment is policy and local legal infrastructure that prevents it. In China, you don’t have any of that, and data center deployment is much, much faster.
I think, in general, infrastructure—the physical infrastructure construction in China—is going really fast. I literally see some kind of crossover bridge being built within a week. The velocity is very, very high there.
There’s a highway close to my home that, after 1 year, is not done yet. So this is also a crossover. I do think there’s a unique strength, probably because of the population density, and they are specializing in those kinds of construction-related work.
But I do think here we also have those specialty people. It’s just, I heard even electricians are under severe shortage.
Yeah.
We are under global supply chain constraints here.
What change would moving into the data center layer cause to margins? Would that take it from 30% to 50%? Would it be not that meaningful? What would that change do to margins?
15. Why AI Hardware Depreciates Faster Than Ever
How we calculate gross margin is interesting these days because how long hardware depreciates has changed significantly.
Yeah.
In the past, it was 6 years.
Yeah.
A solid 6 years. Hardware release is usually 3 years, and that’s fast. Now, within 1 year, from one vendor alone, we have 3 SKUs. The newer model usually runs best on the newest hardware model, and depreciation is also very fast.
Every week, we’re launching a new model, and then the model is at its peak in value before the next model comes out. Imagine this cadence after 2 years. Which model runs on 2-year-old hardware? It’ll be a 2-year-old model. Are those models still valuable?
I think that’s the real dynamic we’re facing right now. The hardware will last for 6 years still, but—
But what you’re saying is that the speed of model development far outstrips the speed of chip and hardware depreciation.
The speed of model development is definitely the fastest, but even the hardware innovation itself is the fastest.
So after 3 years, if every year there are 3 hardware SKUs, after 3 years there are 9 hardware SKUs in between. Do you still want to go back to 9-generation-older hardware running a 3-year-old model? That's questionable.
Maybe there's a world where it's still valuable, but with this pace of innovation, it's questionable. Now, with a different depreciation cycle, it changes the dynamics of build versus own, build versus buy. Again, it goes back to my original thesis: do you optimize for growth, or do you optimize for gross margin? It's all about timing.
How do you think about that question for yourself when you're sitting there in an armchair on a Sunday afternoon thinking, "We're optimizing for growth. Now, when is that time to optimize for gross margin?"
I would say we want to optimize for both. [Laughter] Here's how I think about it. Optimizing for growth requires a lot of business planning, assuming there's product-market fit. Optimizing for gross margin is optimizing for differentiation.
I think I want to avoid overoptimizing for gross margin, but we should optimize for gross margin continuously. In other words, we should optimize for product differentiation continuously. There's no question about it.
We want to continuously optimize toward a healthy gross margin that allows us to grow really fast. It's a trade-off, and we don't want to make compromises. The compromise would be overoptimizing for gross margin, resulting in very slow growth.
One possible way to optimize gross margin is not to grow at all. We just optimize the heck out of it. I know we can climb to a high number, but that's an absolute disaster outcome.
16. Why AI Will Create More Jobs, Not Fewer
Okay, interesting. If we just said, "Hey, for gross margin, we're going to take it from 30% to 10%," is it a winner-take-all market where we could eat up everyone else's lunch and then optimize gross margin later?
I think a winner is probably not a snapshot in time. It's going to be a long-term situation. We do see a particular industry oscillate and start to settle with a few good ones.
Take legal, for example. I was at a dinner table, and interestingly, it seemed like there were a lot of those companies around 2 years ago, but now it's pretty much 2.
I think it's a long, long game.
How do you see the more mature state of your market? Is it like a cloud market, where you have Azure, AWS, and GCP, or is it an Uber and Lyft situation, where one takes 90% and the others fight for scraps?
We're in the adoption curve where a lot more companies in the AI space are starting to seriously think about moving to specialized intelligence and starting to seriously think about whether owning their intelligence is better than renting.
Going back to this optimization question—when is the good timing?—it's the same question we're answering for ourselves in build versus buy. Our customers are also thinking about build versus buy, or build versus rent, or own versus rent.
I think the AI journey, or AI adoption journey, has gone further along. A lot of companies have meaningful traffic. A lot of companies are deploying AI into production, and a lot of companies are at the phase of scaling.
That's where optimization kicks in. When optimization kicks in, you need to have control to optimize. If you don't have control, you just don't have the range to optimize.
For you to have control, you have to build on top of some open model. You have to turn your data into your intelligence. That's pretty much the path we've seen so many companies across industries take. They reach the same conclusion: they're moving in this direction.
Speaking of owning your own intelligence versus renting it, that does apply to a national layer. We've seen Fable be banned in some cases by the administration, briefly for 19 days.
Especially in Europe, we suddenly went, "Oh my gosh, we cannot be at the hands of OpenAI and Anthropic, where we can just be banned. Our health services could sit on the infrastructure of something that an administration can turn off." Do we see a future of sovereign models, where large nations or nation blocs own sovereign models?
I definitely see that possibility. I also see that, if we think about the general-intelligence model as the electricity layer—as a power line—every country should own its own power line.
I think that is a very scary moment: my power line is going to be cut off, and all my fundamental day-to-day needs are not going to work. I feel so frustrated whenever there's a power outage in my home. I feel so frustrated when I cannot access my Wi-Fi. I feel so anxious. [Laughter]
17. The Biggest Mistakes AI Founders Are Making
Obviously, operating a country is extremely important when it's built on top of this fundamental baseline. For every single company, it's the same thing. It's not just about whether a country should have its unique sovereign independence; every single company should have its independence.
You don't want any single person to cut you off. That's an extremely scary moment.
Why would you move into the data center space, but you wouldn't move into the chip space?
Because I know building a chip is extremely hard.
I thought so, too. Okay, again, I admit to being an outsider, which is why I think the show is a little bit successful. I thought so, too. But then how come everyone is seemingly doing it as if it's just another product? As I said, OpenAI, Anthropic, DeepSeek, and Meta are building their own chips now.
I think Meta has been building its chips for more than 5 years—way more than 5 years. MTIA has been a project since 2018, maybe earlier.
Meta has been investing in AI for a long time, pre-GenAI, and it has a huge AI workload focused on ranking and recommendation. Meta has been building other hardware as well in the past.
Whenever the usage has passed a certain threshold, it makes economic sense for you to build the underlying supply. You can specialize toward your workload, and that's another form of specialization: specializing by baking your logic into hardware.
This hardware is purpose-built for your particular workload, and you better make sure this workload doesn't change, because it's really hard once the hardware is taped out. It's possible to change it, but it's very costly.
Once your workload has stabilized and your business has stabilized and doesn't change too often, then that's the time to consider building a chip. I still see the whole AI world, especially model customization, as very dynamic—very, very dynamic. Workload patterns are very dynamic.
Think about how much energy there is in the application space. People are experimenting with all kinds of things. You don't know which one is going to take off, and they will just take off quickly. Once they take off, which one is going to sustain? A few will sustain, and then that's the time when we say, "Now we know this is the pattern, and now we should probably encode this pattern into hardware," bring this hardware into a data center, and so on.
It's all cascading, and then it's going to cascade down to me. It's a fundamental question: where are we in terms of workload maturity? We're still in the early stage of workload maturity to warrant a chip that will be durable.
Now you go back to the fact that we have so many accelerators. They are successful; some are really successful. But remember, those companies started before GenAI. They started with something to optimize some workload, and then they pivoted to AI and tried to fit the AI workload.
It's almost like you bet before this AI workload emerged, and now it becomes a serendipity question. Are you lucky enough that this just works? Some really worked.
Some fundamental designs, like putting a lot of SRAM on the chip, are great for AI models because they are memory-hungry. This really accelerates the execution of inference, and so on. Those worked, and some didn't work.
18. Do AI Startups Need to Build Their Own Models?
What do you see as the greatest bottleneck today? I think it was when I had Jonathan from Groq on the show, who said HBM was the greatest bottleneck, and that's why you've seen a 5x increase in price. What do you see as the greatest bottleneck that people don't talk about enough?
I still think we don't have a great system for very large models. I really believe the fundamental lower-level infrastructure cost will go down. For solving tasks, we should need fewer tokens. That will increase. Collectively, the cost will significantly reduce.
Therefore, we can run the highest-intelligence models much more ubiquitously in the future, but we don't have a system designed for that. For example, we don't have a great system designed for 10-trillion-parameter models today.
That will require very smart engineering and code design, from the model to the customization and serving-platform layer, all the way to the chip layer. The chip is not an individual chip, but a system—a collection of chips in a system—and all of it as a total package.
I think there's still a lot of innovation we can do.
I think recently you announced that you were at $800 million in AR. Incredible feat, and you've scaled so fast. What is that by the end of this year?
We think we can at least double it.
By the end of the year. Wow. You know what's so interesting for me as a venture investor? I've been investing for 10 years. We used to be in the days when Slack was the golden child, where going from $1 million to $10 million in revenue in 18 months was amazing.
Now we have companies like Fireworks, where you scale to $800 million in revenue in a matter of years. You mentioned Cursor scaling to billions in revenue in a matter of years.
The speed of company revenue growth is just unparalleled.
I think it's because there's a fundamental disruption in this technology that is all-empowering. And all-empowering in the sense that it reaches out to every individual one of us to be creative, and it unleashes a lot of creativity that we just don't have access to. That's why we're seeing this phenomenon of extremely fast growth: because of the demand.
Final one before we do a quick fire. You hired George Hu, who was president of Salesforce. He's exceptional. He's one of the most direct, no-BS operators I've ever met. But you met him a couple of years before, or a year before, and you were like, “Oh, we're not ready for you yet.” Why did you say that, and why did you decide now was the time?
Right. So, a year ago, I think we were probably just 50 people. Today, we're at 200 people. We're still not that big.
Wow. You're 4x ahead.
Yeah. So, at 50 people, I was more thinking about scaling the product first, then getting to massively scale the business. I have huge respect for him. I know he's a legend. He's legendary. He's a legendary operator in Silicon Valley.
I just felt like we were too small for him. I told him, “Hey, we're probably too small for you, but I would like to work with you at some capacity.” So, he helped me actually build up the team and interview a lot of executives. His feedback is always well-balanced and very thoughtful, and we started to work together in that capacity.
I think by the end of last year, we were growing really fast, and we started talking seriously. That early relationship paid off. He's really cool. He's really cool in the sense that—
He's so cool.
19. Why Great Leaders Stay Close to the Work
He did a lot of things, with great accomplishments in the past, but I find a unique characteristic about him: he's extremely experienced, has high aptitude and business vision, but he's also very curious. He doesn't make assumptions—“I know it all. I've seen all the movies; it's the same movie, so let me just direct this as I did in the past.” He didn't come with that attitude.
He knows AI goes at an insanely fast pace, and he's learning along the way, but he also fully embraces AI. His team—our GTM team—is using all kinds of AI agents. They're sharing skills, so they maximize their productivity.
He knows we have a superlinear demand curve, and there's just a certain pace at which we can build our GTM team. In order for us to catch this curve, we need to build a team, but the team also needs to have increasing productivity to match it. That's a problem he's solving, and I feel very fortunate to work with him.
In general, I feel that in the AI space, the unique part is that people need to have very special traits, almost like contradictory characteristics. For example, they need to be very experienced but super curious, with a fast learning curve. Or Dima, who we talked about a little bit earlier: he is brilliant, with high intellectual horsepower, but extremely humble. It's a weird combination.
He's amazing too.
He's almost cynical in an Eastern European way, but also, at the same time, very humble.
Can I do a quick-fire round?
Okay, let's do it.
Okay, what have you changed your mind on most in the last 12 months?
I think it's how fast we grow. I changed my mind because I had been quite worried about having too big a team too early. That's why, when I met George, I told him we were too small for him, because I didn't intend to grow very fast in terms of people. I worried about getting slowed down and losing our agility and velocity very deeply.
Since then, we've been very aggressively using our tools. We've developed our own unique way of hiring certain types of people who we know will be charging forward with high velocity, an extreme sense of ownership, strong communication, and who never take no for an answer. We also learned how to get those people. Now I feel much more comfortable scaling really fast.
What's your type of people? I know that sounds weird, but our type of people is actually really specific. We pretty much only hire immigrants. British people don't work very hard—sorry. They're very scientific and rigorous, use data for most things. I actually think creativity often comes from data and is informed by data, and they're unwaveringly accountable and ownership-oriented. Nothing is anyone else's fault; it's all my fault, even if it's someone else's fault. That's a 20VC person. What would you say yours is?
It's not, in a weird way, competence. It's weird: we want people with high confidence. But more importantly, the strong indicator of whether they will do well in this wave, especially at Fireworks, is whether they are really built for taking extreme ownership.
Extreme ownership means we're not putting anybody in any boxes; we're just stacking the boxes together into a tower. We want people to automatically claim, “Hey, this is an end-to-end problem. I'm going to see through the whole thing and work with a bunch of people to make it happen, and I'm going to deliver it no matter what.”
Those kinds of people have the highest, longest mileage, and their growth curve is amazing.
What's your biggest lesson from working with Jensen Huang on what makes him so special?
He's everywhere. I seriously think he has a clone—like hundreds of Jensens somehow plugged in. [Gasps and laughter.] For example, I send him an email, and he'll reply in 1 minute. I just don't understand how he's constantly in the details.
But now I've operated a company for 4 years, and I understand why he's doing that. That defines velocity, because what is leadership? Leadership is just judgment. It's not privilege; it's judgment. You basically have the context. You need to have the right context to make the right judgment for the team.
Especially in a high-velocity space, if you don't know what's happening, what works, what doesn't work, or what the gaps are, you make the wrong call. In a slow-moving space, you can wait for cascading information up and down and make those calls. But in a fast-moving space, you just cannot wait, because there's guaranteed to be information loss in transition layers. People are people; it always happens. Not knowing exactly what is happening and having the position to make judgments makes for bad leadership.
He's demonstrated that through his own example. Even before this crazy AI thing, he's been operating that way. Before, I admired him for his sheer amount of capability in doing that. Now I understand the wisdom behind it, because I also operate that way. I need to know what's happening on the ground to make the judgment for the company.
What did you wait on in the Fireworks journey that you wish you hadn't waited on?
Marketing. [Laughter.] We talk about it. At the very beginning of our journey, we didn't discuss it, but we felt the product spoke for itself. At the end, the product stands, and we wanted to devote all our effort and focus to building the product, working with customers, validating product-market fit, and going from there.
We didn't spend much time marketing at all. We didn't prioritize educating our customers about the right direction to think about the trend and the value. But now I do think it's important. Marketing is not about fluff; it's more about education, more about clarity, and we're working on that.
What area of AI is underinvested in today, in your mind? You mentioned cooling or servers. What else is underinvested in?
I think AI has a sexy part because it's such an innovative, creative technology, and building something on top of it is the focus. But monitoring the ROI—I think the industry is starting to pay attention to it. Eventually, that's what matters: it's not how much you spend; it's what the return is, what the cost is, and what the attribution is.
I think in the next couple of years, as AI gets more and more into production, there will be a lot of focus on getting that clarity and getting that discipline out. Token maxing is just a thing in time, I think, but we'll quickly move into ROI maxing, which is all about running a business.
What large customer do you not have that you would most like to have?
We haven't spent too much time in the traditional enterprise segment. I think that's just because we were very small. Now, as we build out our company, I do think that, even without us investing, we have customers like GEICO, Capital One, Mercury Insurance, and RBI. All these companies, even without us pursuing traditional enterprise, come to us and are customers. I do think that's a very big market.
What has to happen before the end of the year that hasn't happened for you to consider it a good year?
I'm confident in our capability to drive the business. To me, this is a year when I want to prove we can scale quickly while keeping the same velocity, and that's very important to me. If we reach that proof point, next year I'll have a lot more confidence to continue to scale extremely aggressively. I want to make sure we do it right this year.
Final one for you. What does no one see about the next 3 years that you see very clearly happening or not happening?
I really see that every single company will own its own intelligence as a must-have. It's not optional. That's a trend I'm seeing, because there's an analogy to software: there's a reason why every company builds its own software stack.
There’s no standardized software you can just use off the shelf to solve their problem, because every single company is solving a unique problem, and they want to build software because they want to have full control.
Obviously, they will pick and choose which part of the stack they want to build themselves and which part of the stack is common knowledge—there’s no point in building that. But every single company owns their own software stack. Obviously, we’re talking about this in the SaaS era, right? So, same, I think, in the AI era, every single company should own their own intelligence.
Lin, you know, it was Matt who introduced us first. I’ve had the joy of getting to know you and, obviously, George. I can’t thank you enough for joining me and for coming in person. It is so wonderful to do it in person, and you’ve been fantastic.
That’s an amazing studio. You asked a lot of interesting questions. I had a lot of fun talking with you.
We do a lot of research beforehand, huh?
Yes, you did.
Thank you so much for that, and you’re fantastic.