[BidClub_]
Gradient Dissent · · 79 分钟

多数 AI 初创公司都在扩张中走向破产|Fireworks CEO Lin Qiao

Lin QiaoLukas Biewald

YouTube
TL;DR
  • Fireworks称其每日处理超过40万亿个提示词与生成 token,其中95%来自定制模型或定制部署。 Lin Qiao表示,据 Fireworks 所知,这一规模“大于 OpenAI API 和 Gemini API”,但也承认各家的统计口径可能不同。叠加15亿美元 D轮融资和175亿美元投后估值,这一说法意味着定制化已形成规模化需求,而非边缘实验。

  • Qiao的核心判断是,智能将碎片化为数百万个模型——“每个应用、每个用例一个”——而不是最终落入前沿实验室双寡头格局。 最有价值的数据存在于应用和企业内部,“绝不应该被共享”,因此企业可以把客户行为和业务逻辑转化为持续更新的专有模型。Fireworks预计,专业化智能与通用智能将长期共存。

  • AI 产品实现产品市场匹配,并不再保证商业模式持久,因为 AI 基础设施和推理成本可能让增长看起来像“扩张走向破产”。 开源模型的标价可能便宜约10倍,但生成内容的冗长度高出1.5–2倍,按完成任务计算,实际节省幅度更接近5–6倍。Qiao认为,行业正从“token 最大化”转向“价值最大化”,尤其是在上市公司必须为投资回报仍未确定的 AI 功能辩护之际。

  • 强化学习微调正在把模型开发变成一套围绕专有评测、奖励和反馈闭环展开的产品方法论。 Fireworks既支持研究人员控制全部参数,例如 Cursor 的 Composer 模型,也为专业程度较低的团队提供 SDK;Doximity则结合 SFT、DPO、KTO 和 RL 等方法开展医学研究。Qiao在与 Jensen 交流后形成的关键判断是:“不存在专业化的通用公司。”

  • 编程曾是 Fireworks 去年的主导工作负载,如今已成为更广泛应用周期的赋能层。 Qiao称,过去需要强大的产品团队、跨越多个季度完成的工作,如今一个完全不懂写代码的人也可能只需几周。由此催生的产品包括通用知识工作应用,以及覆盖法律、金融、招聘、营销、销售和客服等领域的垂直产品,明年甚至可能迎来面向消费者的搜索与推荐突破。

  • 当被问及美国企业是否应担心中国模型时,Qiao拒绝把地缘政治问题压缩成开源与闭源之争。 她更强烈的主张是,美国公司应该发布最好的模型,并明确呼吁 OpenAI 这样做;Biewald的问题也提到了 Anthropic。Lukas追问了研发投入如何变现这一显而易见的矛盾,Qiao则表示,开放必须具备可行的商业逻辑。

  • Fireworks认为,其护城河不在通用推理,而在围绕“一个尺寸适配一个需求”共同设计的训练与服务系统。 公司宣称训练与推理可以做到按位等价,支持全球拆分、规模最高达数万张 GPU 的训练任务,并针对每个工作负载搜索超过10万个推理选项。首日发布能力之所以重要,是因为客户不再相信公开基准,必须在下一个模型到来前测试新的基础模型。

  • 公司的开源模型倡议并不延伸至自身的专有训练和推理引擎,这一矛盾被 Biewald 反复追问。 Qiao称,内部变化太快,社区管理会变得低效,因此她更愿意支持 vLLM 和 SGLang,而不是再造一个竞争项目。她更广泛的经营方法包括极致主人翁意识、扁平团队、坦率的事前复盘、在数据不足时快速决策,以及保持足够真实、从而维护技术可信度的营销。

摘要 · 为研究而整理的核心内容

1. 极致主人翁意识成了 Fireworks 的操作系统

  • Qiao大约在2015年开始考虑创办一家科技公司,此前她先后从事研究、内存数据仓库,以及离线和在线数据产品。她觉得自己技术上已经准备好,但“从人的角度还没有准备好”,于是从 LinkedIn 加入 Facebook,学习创办公司所需的人际与组织能力。

  • 令她在 Facebook 感到震动的是一种“极致主人翁意识”:员工会修复并非自己编写的代码,会上报损坏的前台仪表盘,并把任何问题都视为自己的问题。Fireworks如今明确奖励这种行为,因为这类人“无需你开口就会脱颖而出”。

  • Fireworks的起点,是 PyTorch 团队此前在 Meta 搭建 AI 基础设施的经历。Qiao回忆,约在2017年时,系统还以 CPU 为主,算法规模很小,深度学习软件十分初级,也没有成熟的 AI 团队。5年后,外部公司开始向 PyTorch 工程师咨询训练平台、服务平台,甚至如何组建 AI 组织。

  • 这种从移动端数据生成转向 AI 的产业迁移,让这次机会显得大于单一产品。Qiao和联合创始人认为,他们是一台已经知道如何帮助组织跨越同一基础设施转型的“影响力机器”。

2. 每日40万亿 token 验证定制化部署需求

  • Fireworks将自己定义为专用智能平台,采取与前沿实验室通用黑盒 API 平行的路径。其基本假设是,“智能是数据的导数”。全球数据中只有很小一部分位于公共互联网,公开数据和标注数据是前沿实验室从头训练模型的主要来源。

  • Qiao认为,绝大多数数据锁定在企业和应用内部,构成它们的“阿尔法”。这些数据应该成为企业独有的模型,而不是交给供应商;随着产品、数据和基础模型变化,专业化也必须每周、每天,甚至每隔几小时重复一次。

  • 她对终局的判断十分明确:“世界不会成为双寡头。”通用智能仍然有用,但数百万个专用模型将服务于单个应用和具体用例,让每家公司掌握自己的知识产权、产品差异化和成本结构。

  • Biewald追问,这是否只是一个开源模型推理供应商的品牌包装。Qiao用规模回答:Fireworks每日处理超过40万亿个提示词与生成 token,其中95%来自定制模型或推理部署。她称,据 Fireworks 所知,这一规模大于 OpenAI 和 Gemini 的 API,但也承认各家公司可能采用不同统计口径。

3. 编程打开了不断扩张的 AI 原生应用栈

  • 去年的需求“全部是编程”,但编程也加速了其余应用的开发。Qiao对比了旧路径:数十名优秀工程师和产品经理跨越多个季度协作;如今,一个可能完全不懂写代码的人,只需几周就能把产品推向生产环境。

  • 下一层包括深度研究、幻灯片制作和其他通用工作工具,随后是覆盖法律、金融、招聘、营销、销售和客户支持等领域、颗粒度越来越细的垂直产品。Fireworks自身也同时使用开源和闭源模型进行财务预测及账务管理。

  • Qiao还看到消费公司开始把 LLM 的推理能力用于搜索和推荐。她的预测仍然保留余地:“可能明年”会成为面向消费者的 GenAI 突破点;目前可以确定的是,实验正通过越来越多样化的产品进入生产环境。

4. 微调成为产品方法论,而非实验室服务

  • Qiao表示,前沿实验室的经济模式天然不鼓励大规模定制:巨额预训练研发投入,最适合通过向大量客户规模化提供少数几个 API 来变现。支持数百万个微调模型需要完全不同的基础设施,实际上是“一门完全不同的生意”,这也解释了为什么早期微调 API 可以存在,却始终没有成为战略重点。

  • Fireworks覆盖多个控制层级。Cursor使用其训练栈开发 Composer 模型,控制底层 RL 参数,并将训练器连接至 Fireworks 管理的 rollout 推理;专业程度较低的团队则使用训练 SDK 选择算法、提供损失函数,并运行托管任务,无需知道所需 GPU 数量。

  • Doximity面向医生的深度研究产品展示了这一领域的应用。其研究人员结合 SFT、DPO、KTO 以及不同 RL 变体,教授模型医学术语和推理能力。这正是更广泛模式的一个例子:基础模型理解逻辑,却不理解某个应用的专业语言或 DSL。

  • Qiao给出的清晰区分是:SFT类似于把教科书交给模型,让它记住标准答案;RL则让不同模型变体与产品或模拟环境互动,获得奖励后再次探索。在这两步之前,企业需要先建立评测体系,相当于单元测试和集成测试;随后还要设计独立的奖励函数,决定模型沿哪个方向持续优化。

5. 私有数据是护城河,但成本决定它能否存续

  • 奖励设计催生了一种介于研究员和产品工程师之间的混合角色。判断医生是否获得了高质量的药物检索结果,需要领域判断,因此研究员必须学习产品开发,或者产品人员必须学习模型;产品反馈与模拟环境随后闭合学习循环。

  • 当 Jensen 告诉她“没有专业化的通用公司”时,Qiao的判断体系最终成形。每家公司都在用特定的品味和判断力解决特定问题,而这些东西没有任何一项是普遍共享的;将这种判断编码为企业自有智能,就能在模型和产品之间形成飞轮。

  • Biewald反驳说,更新权重只是其中一种机制。Qiao表示同意:企业应该使用提示词与上下文工程,激活私有数据,并在不同模型之间路由任务,以兼顾质量和成本。她更窄的主张是,权重依然不可或缺,因为专有的意图、偏好、参与度和业务逻辑不应被“留在桌面上”。

  • 在 SaaS 领域,Qiao认为产品市场匹配与商业持久性基本同时到来;在 GenAI 领域,两者却彼此分离,因为智能的运营成本很高。一些用户喜爱且具备变现能力的产品正在“扩张走向破产”,而成熟公司则面临华尔街审视。Qiao称,把成本降低5–10倍可能很重要;考虑到开源模型更高的冗长度后,她称按任务计算的实际节省约为5–6倍。

6. 开放智能是 Qiao 应对模型集中化的答案

  • 开源与闭源系统将长期共存。Qiao建议,在早期验证阶段使用最昂贵的 API,因为此时成本并不重要;产品规模扩大后,再进行优化。她也强调不同任务存在差异:Anthropic擅长编码代理和指令遵循,OpenAI擅长交互和深度研究,Gemini擅长多模态,而 GLM、Qwen 和 Kimi 的视觉能力各不相同。

  • 在网络安全问题上,Qiao将目标定义为让攻击与防御保持对等:不对称是危险的,而开放模型可以为防御方提供可获得的工具,并扩大后训练社区。她举出的 Hugging Face 案例仍有争议:Biewald称 OpenAI 的某个模型在一次评测中“某种程度上黑进了” Hugging Face;Qiao表示,据她了解,失控的是网络安全攻击基准 CyberGym,而且该模型也拒绝修复问题。

  • 当被直接问及美国公司使用中国模型的问题时,Qiao没有给出针对具体模型的风险评估。她转而将地缘政治关切与开放开发分开,并认为开源数据基础设施曾帮助整个行业创新;在当前这个“关键时刻”,对智能的集中控制无法带来同样均衡的演进。

  • 她的呼吁异常明确:美国公司应该发布最好的模型,并直接点名作为市场领导者的 OpenAI 率先示范。Biewald追问谁来偿还研发投入;Qiao表示,任何发布都需要一个与变现相关的“理由”,同时预测更多美国公司参与后,开源与闭源之间的质量差距最终会缩小。

7. 首日支持让客户跑赢模型淘汰

  • 即使没有提前获得模型访问权限,Fireworks仍建立了发布即支持的声誉。早期 Mistral 权重发布时并未附带模型代码,Qiao称团队根据此前版本反向工程出实现,并在 Mistral 发布自有 API 前就开放了该模型。

  • 质量有时会压倒速度。Fireworks曾将 DeepSeek 的上线推迟3天,因为拿到的权重和代码未通过内部评测;团队与 vLLM、SGLang 社区并肩工作到不眠不休,修复漏洞并将补丁贡献回上游。“我们不能部署一个明知存在问题的模型。”

  • 紧迫性来自模型极短的生命周期,以及事关重大的基础模型选择。Qiao称,客户已经不再相信公开基准,因此需要立即访问模型、运行私有评测;如果花太久才决定是否重新微调,“下一个模型就会出现”。

8. Fireworks的护城河是共同设计的质量,以及专有引擎

  • Qiao拒绝“推理公司”“新云服务”和“PyTorch 云”等标签。Fireworks围绕输出质量共同设计训练和推理,再针对速度和成本定制部署,因为她从未见过客户仅仅为了更低价格,就接受明显更差的结果。

  • 公司宣称的一项质量里程碑是训练与推理之间“KL 散度为零”——即使使用不同的数值库,结果仍然按位等价。这在部署阶段很重要,尤其是在 RL 中,训练器与 rollout 推理之间的反复切换可能放大微小的数值误差。

  • Fireworks还会把分布在全球不同区域的 GPU 汇聚起来,执行规模最高达数万张 GPU 的训练任务。Qiao强调,难点不仅是让拆分式基础设施运行起来,还要在增加同步环节后保持模型质量。

  • Biewald揭示了核心矛盾:Fireworks倡导开源模型,却将两套引擎都保留为专有资产。Qiao称,其模块化推理设计形成了超过10万个选项的搜索空间,并会针对每位客户的独特工作负载进行搜索;在一家约200人的公司里,内部只有约10名引擎工程师,她更愿意支持 vLLM 和 SGLang,而不是管理一个快速变化的竞争性社区。

9. 扁平且带有偏执感的公司在数据到来前做决策

  • Qiao称,她每天都会思考自己是否应该采用更高调的创始人形象,但“我无法改变自己”。她后悔太晚开始营销 Fireworks——工程师曾以为产品会自己说话——但她不愿通过不真实的信息换取关注。

  • Biewald的反驳带有支持意味:没有产品实质支撑的营销会损害信任,企业买家归根结底也是能够识别真实感的人。他的实际诊断是,市场仍把 Fireworks 视为推理供应商;公司必须不断重复“专用智能”,并让官网围绕这一简单叙事展开。

  • 在内部管理上,Qiao偏好共享上下文、少设层级的扁平组织。她与另外6名联合创始人共同创办 Fireworks,并将公司的“深层知识诚实”归功于这群人;团队经常进行“我们会如何死掉”的事前复盘,随后继续工作,不把这种练习视为不忠。

  • 创业带给她最大的意外,是数据的缺席。在 Meta,团队可以围绕0.1%的提升持续优化;初创公司则是在一条无人走过的路上铺路,因此决策必须从直觉出发,再通过反馈验证。“不能因为缺少数据就不做决定”——这套理念与她最后的自白相伴而行:“我总是很偏执。”

Lukas Biewald

Do you think American companies should be concerned about using Chinese models?

Lin Qiao

I do believe we are in this crucial moment when, across the industry, we should encourage more open intelligence. We truly believe in open development and open intelligence.

Lukas Biewald

I think Fireworks is a company that does a really good job running open-source models.

Lin Qiao

We process more than 40 trillion tokens a day. That's bigger than OpenAI's API and Gemini's API. We do not believe the world will be dominated by a few models from frontier labs.

If you think about the period before AI, data was one of the biggest innovations. There are so many open-source projects about data, and I believe that's the fundamental reason the data field was able to move so fast across the entire industry. To me, it doesn't make sense for intelligence to be in the hands of a duopoly.

Lukas Biewald

Do you think OpenAI and Anthropic should just open-source their models?

Lin Qiao

I would really think...

1. Introduction to Lin Qiao

Lukas Biewald

You're listening to Gradient Descent, a show about making machine learning work in the real world and I'm your host Lukas Biewald. All right, I'm here talking with Lin Qiao, an old friend of mine. I've always admired her as an entrepreneur, and I've always wanted to have her on the podcast. I finally got to do it today. I get to ask her questions about the technical aspects of how her product works, what motivates her as a founder, and how she got started in the space.

Thanks for joining us. I've been wanting to have you on this podcast for years, going back to when you were a big part of the PyTorch team and working on it. One thing that really surprised me when researching you is that you have a college-aged daughter, which made me realize that you're older than I thought, I think.

Lin Qiao

Yes.

2. Becoming a founder later in life

Lukas Biewald

You came to being a founder a little later than a lot of people, at least by the Silicon Valley ideal. Did you always want to be a founder, and now you're doing your dream? Or how did that happen?

Lin Qiao

I talk with people who are college dropouts or founders, so I'm more than twice their age.

First of all, I think there's no good timing. Only you know when a good time is, and that's just a gut feeling. To me, about 10 years ago, it started to come to me that I wanted to be a founder. I wanted to found a technology company, and I actually started doing that.

3. Why she joined Facebook to learn company building

By 2015, I had done everything across data, starting from being a researcher, building the first-generation, in-memory, superfast data warehouse, and building many data products, from offline to online data tools. I thought, "Hey, I know everything in that space, and there are many companies during that time, so I want to start one."

But I didn't feel I was ready from a people point of view, because I knew that building companies was all about people at the end. It was all about organizing experts across different fronts to be able to build that up. I didn't think I had that skill, and that's the main reason I joined Facebook.

Lukas Biewald

Oh, really? What did you take from that?

Lin Qiao

Believe it or not, when I moved to Facebook from LinkedIn—LinkedIn is another social network, a professional social network, and Facebook is a consumer social network—I knew these companies were outstanding. But on the other hand, I was thinking, "I've probably seen a lot, so it probably won't be that different. But I need to learn what I need to learn."

I was shocked when I joined the company. What really surprised me, and what I carried on to Fireworks, was the extreme sense of ownership. At the other company, I didn't feel that. I felt like in other companies there's always, "Hey, you're an employee. Here's your box. You do your job, and you'll be evaluated." That's kind of how things work.

When I first joined Facebook, everyone cared so deeply about the company. They did whatever it took. You could change any code. It didn't matter whether you wrote it; if you found a bug, you fixed it.

No problem was someone else's problem. Even small things—if the dashboard in the reception area that every visitor saw was broken, you reported it and fixed it. It almost felt like this was my family, my company. Although I had tiny ownership of the company, it didn't matter; I did feel that way. Everyone felt that way.

4. Starting Fireworks to have industry-wide impact

It was magical, bringing out the best in a person to build for the company. That's the secret sauce, the magic power. At Fireworks, we strongly reward extreme ownership, and people who carry that sense of extreme ownership rise up without you asking them to do anything. They figure things out by themselves.

Lukas Biewald

What was your core motivation for starting the company? You've talked about how it seems like you're okay with the competition being successful. Maybe it's not about beating the competition. What made you want to start Fireworks?

Lin Qiao

I think it's the same with my other co-founders. We gravitate very strongly toward impact. We're almost like an impact machine.

Before we started, it was a little bit rooted in our journey with PyTorch. PyTorch has a lot of great people moving it forward, including Soumith and Joe Spisak. They're all great people I worked with. Because of PyTorch, we worked a lot with the open-source community and with companies.

It became very clear to us that the entire industry was moving from mobile-first, with a lot of data being generated, to AI-first. If we go back 5 years, we had the same problem at Meta. Mobile-first, a lot of data being generated, AI-first.

At that time, around 2017, there was no AI hardware. Everything was CPU-based, with tiny machine-learning algorithms and linear regression. There was no AI software; deep learning had just gotten started and was very rudimentary. There was no AI team.

We all joined around the same time, built AI infrastructure from the ground up, built PyTorch, and engaged our community. Then, 5 years later, the industry was hitting the same problem. That made it clear there was an entire industry movement toward AI. They were hitting the same problem we had started this journey with at Facebook, and we knew how to help them.

As a matter of fact, they kept coming to the PyTorch team saying, "Can you build a training platform for us? Can you build a serving platform for us? Can you tell us how to even build this AI team?" There was a lot of desire—a clear demand—for the industry to have an industry-wide impact. That's the reason we wanted to start a company: to help a much bigger movement.

With regard to competition, I don't view those as competitions. If a community builds a great open-source project, it's not our game to compete with them. I want to see them doing well.

Again, it goes back to the fact that every single company should have a reason to exist. Our unique value is not in that space. Our unique value is in the other layers that we add. That's where we find our position, and we anchor on specialized intelligence, as we believe that's the biggest area where we can move the industry and deliver impact. That's how we think about it.

5. Raising $61.5 million at $17.5 billion valuation

Lukas Biewald

I think I saw recently that you raised at something like a $15 billion valuation. Is that right?

Lin Qiao

Yes, we just finished raising Series D: $1.5 billion in cash and a $17.5 billion post-money valuation.

Lukas Biewald

Congratulations. Do you want to brag a little bit about your metrics or your success at Fireworks?

Lin Qiao

If there's anything I haven't done very well, it's talking about who we are and what we do. We've been really focused on being the engineers building the product.

6. Specialized intelligence vs general intelligence

But yes, I would love to talk about Fireworks. We built a specialized intelligence platform. Specialized intelligence is a parallel strategy to general intelligence. This is a discussion that's actually happening across the industry, and we're very happy to participate in that debate, but also to charge forward and build the best tool in the specialized intelligence category.

The idea here is that there are 2 strategies for intelligence. One is building intelligence that can solve all problems as a black-box API. This is the typical frontier-lab approach to building AGI. It's very familiar, and everyone can understand that direction.

But we believe intelligence is the derivative of data. That's our fundamental assumption. If we look at the world's data, you're actually living in the data world, right? When we think about where AI started, we see a lot of data flowing into training first and then to the rest of AI.

If we look at the world's data distribution, only a small fraction of data is on the public internet. That, plus labeled data, is the primary source for frontier labs to train from scratch. The majority of the data is actually locked inside applications and enterprises.

In my opinion, this data should never be shared with anyone else because that is the alpha of those companies. There’s much deeper thinking behind that. We believe the new frontier of intelligence should be specialized intelligence—that is, turning that private data into a customized model that is uniquely owned by the company and enables it to design its product in a unique way.

This process shouldn’t be one-time. It should be continuous because we all know applications keep evolving, and product engineers make our applications constantly better. Base models keep improving. Before, we could say there was a new base model popping up every month. Now, every few days, there’s a new model. The velocity is insane, with models popping up at the top of the leaderboard in a really good, forward-looking way.

So your customization and specialization process should be continuous. You should be doing that every week, or possibly every day or every few hours. We do have customers with different cadences.

If we take a big step back, we believe specialized intelligence and generalized intelligence—these 2 strategies—will coexist in the future. They will have their own usage and so on. But we do not believe the world will be dominated by a few models from frontier labs. The world will not be a duopoly. The world will be millions of specialized models: 1 per application, per use case. That’s the worldview we’re seeing.

We’re building our platform toward giving control of IP and giving control of cost to every single company, because those companies exist for a reason.

Lukas Biewald

That’s interesting. I think Fireworks is a company that does a really good job running open-source models, but specialized intelligence is a little bit of a different way of looking at it. Are most of your customers actually modifying the open-source models before they run them?

7. Processing more tokens than OpenAI's API

Lin Qiao

That’s a really good question. Today, we process more than 40 trillion tokens a day. This is actually a fun fact: based on what we know, it’s bigger than OpenAI’s API and Gemini’s API. More interestingly, 95% of that traffic is from customized models and customized inference deployments. It’s not off the shelf.

Lukas Biewald

Uh-huh.

Lin Qiao

From what we see in our customer engagement and the demand coming to us, we’re just at the beginning of this S-curve.

Lukas Biewald

So you’re claiming that you think you ingest more tokens than OpenAI’s API?

Lin Qiao

We process them. By “process,” I mean the prompt and the generated tokens combined.

Lukas Biewald

I see. So more input tokens and output tokens in total than OpenAI’s API or Gemini’s API?

Lin Qiao

Yeah. We don’t know how each company accounts for this number. They may have their own philosophy, but just by the numbers, that’s what I’m seeing.

Lukas Biewald

Wow. What are the biggest use cases?

Lin Qiao

That’s also evolving. Last year, it was all coding, and we had all the coding companies building on top of us. Coding actually unlocked the next wave. Because of the advancements in coding, I think one thing that has changed significantly in the past 1 year is software development and application development.

The velocity is insane. Before, it would take tens of very strong product engineers and PMs to go from ideation to implementation to production scale over multiple quarters. Even multiple quarters may be fast. Now, 1 person in a few weeks, knowing nothing about writing code, can do that.

Because of that velocity, we started to see a vibrant usage, especially in the knowledge-work space. We have many applications solving general-purpose knowledge-work problems, such as professional deep research, slide generation, or the tools we use day to day.

We’re also seeing the development of a wide variety of diversified, vertical-specific knowledge work across legal, finance, recruiting, marketing, sales, and customer support. Even within each bucket, there are fine-grained products being built. That’s what we see this year.

In addition, we’re seeing the consumer-facing market start to think about how to use the logical-reasoning capabilities of LLMs to change how we do search, recommendations, and many other things. Possibly next year will be the consumer-facing unlock year for GenAI.

Nonetheless, one thing that is clear is that the variety of creativity and innovation being built on top of GenAI, moving from experimentation to production, is exploding right now.

8. Why frontier labs don't prioritize fine-tuning

Lukas Biewald

Uh-huh. OpenAI had a fine-tuning API in various iterations, but I don’t think it was very popular. Why do you think RL fine-tuning is more popular with open-source models?

Lin Qiao

Fundamentally, if you think about the unit economics of general-intelligence companies, it doesn’t align with supporting a fine-tuning product. The fundamental reason is that training from scratch is very expensive. It’s a massive R&D investment, and the end result is that companies want to monetize and scale as fast as possible.

The way to scale as fast as possible is to package that as an API, and the underlying infrastructure will be most efficient if it’s just a few models. If it supports millions of models, it’s completely different unit economics. It’s massive—almost like a completely different business.

I think it makes sense for frontier labs to focus on scaling the model and going really fast and wide. Creating a separate business is a strategic change and pivot. That’s the fundamental reason. It’s not that having a fine-tuning service doesn’t have a market; it’s that the focus of the company doesn’t align.

Lukas Biewald

I’ve been a little surprised, honestly, over the last few years that fine-tuning with RL hasn’t been more popular. It does seem like it can be fiddly to get it right, and you need the data in a good format. You also need to create a simulation environment, which can be an expensive process. Do you help your customers with that?

Lin Qiao

We have different levels of engagement, and the interesting thing is that our customers have a wide diversity of needs right now in terms of the products they’re building. Rather than there being 1 need, there’s a spectrum.

We have customers who are very deep in this. They have researchers, and they have researchers at Fireworks Labs. They want to control every single knob. For example, we work with Cursor. They build their Composer models on our training stack, with a particular focus on RL.

They want to control every possible parameter to tune, so we help them connect with their trainer, and we manage the rollout inference. That’s 1 level of engagement: extremely low-level, giving you all the controls you want.

There’s a different tier that I think is bigger in terms of the number of developers who can use this. They have AI experience, but they haven’t been experts in training a model. They’re learning that specialty, and they want to control something, but not everything, because it’s overwhelming.

So we built the next-level API. We have a Fireworks Training SDK geared toward this use case. They can plug in a loss function, tweak which algorithm they want to use, and adjust a few other parameters. Then they can start to do their experiments.

Lukas Biewald

Can you talk a little bit about how that would work for a real-world use case where it would work particularly well?

Lin Qiao

We have been working with startups, digital natives, and even—to my surprise—enterprises. The Training SDK is actually self-serve. They can code against the SDK, and that SDK talks with a backend API.

We have 2 modes. We have a service where they don’t even need to know how many GPUs are required; they can kick off the training job. Then we have telemetry for them to see how things are going.

Lukas Biewald

Can you give a specific example?

Lin Qiao

Of a use case, or—

Lukas Biewald

Yeah, like a specific customer, even, if you can talk about it.

Lin Qiao

For example, in the healthcare space, we have Doximity. They’re building deep research for doctors, and they’re tuning their model. In medical terms, it’s almost like a different language.

Usually, the tuning follows a pattern involving a special DSL, whether it’s medical or some kind of special programming language for spreadsheet manipulation or data processing. That knowledge—that logical reasoning—is being baked into the base model, but the model isn’t familiar with that language.

Lukas Biewald

But let’s talk about how Doximity is using your SDK to fine-tune the model.

Lin Qiao

Showing meaning?

Lukas Biewald

What? How is it teaching the model what it wants?

Lin Qiao

They have a team of researchers. They’re using a combination of—we have SFT, DPO, KTO, and RL. We have different flavors of RL. It’s actually up to them to pick and choose, and often they use a combined algorithm.

Usually, you use SFT. It’s very similar to the frontier labs’ training process without pretraining, right? Usually, SFT is like pretraining to infuse the knowledge, and you’re—

Lukas Biewald

Let’s take a step back.

For people who maybe aren’t quite as familiar with all of these acronyms, reinforcement learning is another way, in this context, to fine-tune a model. In the past, fine-tuning meant that you needed to show the model exactly what you wanted each time and modify it. How is RL different?

9. How reinforcement learning works for customization

Lin Qiao

Yes, RL is very different. Compare SFT with RL: SFT is basically telling the model the ground truth and letting it pick up the ground truth. It’s like, “Hey, here’s a textbook. Memorize the whole textbook.” That’s SFT.

RL is where you try different variations of the model, let the model interact with your product or your simulation, and then get the result back. You rate that result, saying it’s good or it’s not good, or you give it a range from 0 to 1—a kind of reward. Based on the reward, the model will learn, “This direction, this exploration, is not that good, so I’m going to backtrack. I’m going to try some other exploration,” until the reward shows up as pretty good. Then it says, “Okay, that’s probably a good result.”

Lukas Biewald

The rating is tricky, right? How does Doximity rate the quality of the results?

Lin Qiao

That’s where a lot of deep product experience comes in. This is interesting: We’re seeing a new, emerging type of person doing this. It’s not necessarily a job profile. In the past, before GenAI, we had product engineers focused on the product and researchers focused on building models, and they talked with each other to figure things out.

Now we see product people who work on models and have product knowledge. They know how to tweak things. Either our researchers learn to build products, or product people learn how to work with models, because judgment is involved. There’s a lot of judgment in asking, “What should the reward look like?” and “Would this be a good search result for a doctor searching for this medicine?” That’s a judgment call and almost requires domain-specific knowledge.

Lukas Biewald

Are you saying they’re using people to actually look at the results and say whether they’re good or not, or are they using people to build algorithms that automatically decide whether the results are good or not?

Lin Qiao

Before anyone does anything, the company needs to build its own eval. It’s the same as when you write software: You need to write unit tests and integration tests to judge how good the software is. It all starts from there.

Once you have that, that’s what you’re going to choose to hill-climb. With an eval, you start to understand, “Now I’m going to write what the reward should look like,” which is different from the eval. This is exactly how humans grow. As we’re born, we come with an IQ that stays the same. This is the base model, right?

Then we learn by going to school and memorizing. That’s SFT, supervised fine-tuning. We learn by trying things because, deep down, I think humans—Homo sapiens—are defined by our curiosity and desire to explore and try new things. When we try new things and get positive feedback, we’ll try more. If we get negative feedback, we’ll try less. This is exactly how our models learn.

The feedback comes from the product. If you get feedback directly from the product, then you have a signal from the product, or it can come from a simulation, which will also provide feedback. Then the product engineer or product researcher determines what that product feedback should be based on the generated text and closes the loop.

Lukas Biewald

I mean, I guess in the past—and I was a big part of this at CrowdFlower—you would label lots of examples of exactly what you wanted. That was expensive, but at least it was very clear what you were doing with SFT, or supervised fine-tuning.

Now, when you don’t actually know exactly what you want it to do, you’re just looking at results and trying to say whether they’re good or bad, I feel like there’s a much more complicated problem here, right? Famously, the labs pay tons of people to go in and do RLHF—to have humans grade over and over.

But I think you’re talking about using product signals. We also see RLAIF, where the AI looks at the results themselves. You’re seeing all these companies building their own evaluation functions or reward functions. What are the trends here? What’s the best practice, and can this really scale to all the different applications out there?

Lin Qiao

Yeah. Our thinking is that it can, and it should. There are multiple fundamental reasons why it can and should.

10. Jensen Huang: there's no specialized general company

This articulation came from my conversation with Jensen after his GTC keynote. We were shooting a video together, and shooting a video with Jensen is very casual: He just starts talking, and then we talk, and then we’re done. He casually mentioned one thing while we were talking about specialized intelligence: “There’s no specialized general company.” It sounds logical, but when I reflected on it, it was very profound.

Lukas Biewald

He said there’s no specialized general company?

Lin Qiao

There’s no specialized general company.

Lukas Biewald

What does that mean?

Lin Qiao

That means every company is special. They’re solving a special problem using a special solution for a special purpose, and that’s why there are millions of companies—maybe tens of millions of companies—in the world. We carry a unique design of the solution space for the problem we care about. That’s why companies exist.

Because of that, the knowledge, choices, taste, and judgment used to create these companies aren’t unified or common. They’re not even commonly shared as a standard. That’s why it’s really hard to capture them in a general-purpose model.

We believe every company should own its intelligence because it is the expert carrying that taste, judgment, and unique thinking. That should be codified into the intelligence the company owns, have that intelligence further power its product, make the product even better, and start to create this flywheel.

Lukas Biewald

But I guess there are multiple ways to do it, right? One way is to actually modify the weights of the model itself. That sounds powerful, but it also seems potentially difficult and complicated. There’s also the context window, where you could insert something, or just use the model in different ways. Modifying the weights isn’t the only way to modify the model. Why do you feel like modifying the weights is going to be the way that companies inject their knowledge, or their special sauce, into these models?

Lin Qiao

Yeah, there are many different ways to define or create a frontier. Prompt engineering and context engineering are one way, and many companies are already doing that. Only they know how to construct a context in the best way, and so on. Activating their private data to constantly tune the model without giving it away is another way.

Having routing logic to route the task toward the best model in terms of quality and cost, and having a network of underlying model supply, is another way. There are many different ways to push the frontier in a given setup. I think companies should use them all. I don’t think tuning the model, or having the weights, is the only way, but I will say having the weights is essential for other reasons.

The reason goes back to how application software development has been disrupted. Before, the moat was the ability to implement and push a production-scale idea, because it was hard and required a deep barrier to get it right. Now that barrier becomes very thin, so what is a moat for an application company? A moat is something that cannot be copied or replicated.

The data collected from your product about customer intent, customer preferences, why they engage, and why they don’t engage, as well as the business logic—those are proprietary. You just leave your alpha on the table if this isn’t integrated into the model you use to power your product.

I would argue this is probably, by and large, the biggest moat every company should create: Turn your data—which is a reflection of your proprietary taste and judgment, and the reason why you exist as a company—into your own model, which is your own intelligence, and do not let that proprietary knowledge leave your premises.

11. Unit economics and scaling to bankruptcy

The second big macro is unit economics. That’s very interesting. We all started working on AI, pre-GenAI. That’s the SaaS era, right? During the SaaS era, product-market fit and a durable business were almost the same thing. It was really, really hard to find product-market fit, but if you found it, you just scaled—scaled as fast as you could. Run.

The COGS of running your business were mostly people and infrastructure costs, which were mostly CPU and storage. Those were commodities, and people usually didn’t care about them; those weren’t cost problems. But now, in the GenAI era, product-market fit and a durable business are 2 separate concepts.

Hitting product-market fit doesn't mean you will automatically have a durable business, because operating AI infrastructure—or intelligence—is expensive. The unit economics have completely changed. We have so many startups that have great products, whose customers love them and want to pay them, but they just cannot scale the business because they are scaling to bankruptcy. Scaling to bankruptcy is the norm now.

You really need to think about how to build cash flow and have a positive gross margin, and then it's durable. It's even worse for large, public companies. The digital natives, for example, were startups a decade ago, and they're the winners, right? They won their markets—consumer markets, enterprise markets, and developer markets. The consequence of winning is that they have a huge amount of traffic and already have a huge customer base.

For them to roll out an AI feature to cover everyone, the cost is enormous. Then they need to answer to Wall Street in their quarterly earnings report and explain, “Why did you suddenly spend so much on a feature whose ROI is TBD?” So cost control is a really big concern. On top of that, because of the popularity of coding agents, almost every company has a coding harness. They're all worried about how to justify spending so much on coding.

The industry is shifting from token maxing, I think beginning this year, into value maxing. We seriously need to talk about our ROI. Cost control is another big concern. That's where picking the right model for you to activate private data and private knowledge, making the model quality on par with—or even better than—the black-box API, and bringing the cost down by 5 to 10 times becomes extremely important to building a durable business. So that's another reason.

Lukas Biewald

How much cheaper is it to run your own model? I mean, the frontier labs are also constantly dropping their token prices at the same time that people use more and more of it. What's the delta?

Lin Qiao

When we talk about pricing, it's actually not per-token pricing, because the verbosity of the different models is different. Open models tend to be a little bit more verbose. So even though, if you look at the pricing—everything is public—they're usually 10 times cheaper, they're usually 1.5 to 2 times more verbose. The cost savings are therefore around 5 to 6 times. Per task, to solve the same problem, that's the kind of delta we've seen.

Lukas Biewald

What about quality right now? I mean, how do you view the quality trade-off between the best open-source models and the frontier labs?

Lin Qiao

Across the board, open and closed model quality has passed the threshold for solving a lot of day-to-day problems. Interestingly, many of the tasks we work on—especially for professionals—can be interpreted as coding problems. Coding is, by and large, a solved problem, so many of our professional tasks are very friendly to being tackled by those models. Open source.

Lukas Biewald

Sorry?

Lin Qiao

So again, it goes back to the fact that most of the coding—all these model providers, closed and open—really care about coding, because coding is the foundation for solving the next level of problems. Across the board, there are many real-world problems. For example, I think our company is a reflection of the industry. Within a company, every company has a finance department, so our finance team uses both open and closed models heavily for financial forecasting and to manage our books.

Lukas Biewald

So you actually use closed models inside of finance?

Lin Qiao

We do. We use both. As I mentioned, another new frontier is how to blend those models intelligently.

Lukas Biewald

So I guess you would be a huge proponent of open models, right? What are the cases where you use closed models?

Lin Qiao

As I said, I think specialized intelligence and general intelligence will coexist. They will coexist during different phases of development. In the early phase of development, you don't want to think about which model to pick. Just pick one, and you also don't care about cost. You don't have the later-stage concerns around production scale and optimization. For simplicity, you just want to pick one, and you can pick the most expensive one to see if it's even viable. Is there even a viable solution?

As development progresses to, “Hey, it's validated. It's a viable solution,” and you now want to scale, then you want to start optimizing. So that's one possibility. Another possibility is that different models are actually slightly geared toward solving different problems better.

Even across the closed models, Anthropic is obviously the best at tackling coding-agent tasks and really focuses on instruction following. OpenAI has been focusing on interaction-based applications and Deep Research, and now they're also focusing on coding. Gemini used to focus heavily on multimodality, and now it also focuses on coding.

Coding becomes a common denominator, but different labs still have their own secret sauce and special focus. Even from there, that just means your starting point may be different. The same is true for open models. For example, GLM doesn't have a vision component. Qwen has been very good at vision from the beginning, and Kimi has recently been combining them.

Based on the task, we will give our customers guidance about which model could be the best starting point to kick off their job.

12. Security trade-offs between open and closed models

Lukas Biewald

Do you think there's a security issue here as well with open-source models versus closed models?

Lin Qiao

This is a very deep topic. There's a lot of debate happening now on Twitter and across the industry about security. Here's the controversy: security always has 2 sides, the attack side and the defense side. The challenge of security is whether there's an asymmetry. If the attack side has better tools than the defense side, then it's really bad. If the defense side has better tools than the attack side, that's really good.

Usually, it will get to the point where they're equivalent and on par. I think that's a healthy situation. I'm not saying we should encourage attackers to have better tools, but they will find other ways to acquire them. Because of that nature, I feel like open models are a way to strike that balance.

The other benefit of open models is that they encourage a broader community to continue building on them and increase the complexity of the defense. The models can be post-trained. Recently, our friends at Hugging Face worked with OpenAI to resolve an incident fairly well. They couldn't get OpenAI's model to work because it detected, “Oh, this is a possible cybersecurity breach,” and refused to find a solution. They had to activate GLM-4.5 to find the solution quickly and contain the problem.

That's really about accessibility and defense.

Lukas Biewald

I heard that part of the story. So didn't—I thought what happened with Hugging Face was that OpenAI had a model that, in an eval stage, actually kind of hacked Hugging Face. Is that right?

Lin Qiao

My understanding is that it was using CyberGym, which is a cybersecurity attack benchmark, and that just started to go wild.

Lukas Biewald

But then the model also refused to fix the cybersecurity issue?

Lin Qiao

That's my understanding.

Lukas Biewald

Interesting. What about China? It's kind of interesting that all the open-source models—the other best ones—seem to be coming out of China. Do you think American companies should be concerned about using Chinese models?

13. Should American companies use Chinese models

Lin Qiao

I think we should differentiate the geopolitical debate from open models versus closed ecosystems. I've been working on PyTorch for a long time. You're also in that ecosystem, right? We truly believe in open development and open intelligence.

We've seen so many good things. Even before PyTorch, I worked on many open-source projects. If you think about before AI, data was one of the biggest innovations. There were so many open-source projects about data, and I believe that's the fundamental reason the data field was able to move so fast across the entire industry.

From there, it derived a huge variety of different kinds of data processing, and that powered all sorts of parts of our economy. That goes all the way from data analytics as a new field, which made it possible to standardize the processing of all business metrics and make data-driven decisions, to online ranking and recommendations, which are everywhere in our digital life. Those are significantly powered by AI, and now everyone is using them. Product analytics is a new practice for building products. Self-driving cars use huge amounts of data, from perception.

All these data innovations heavily depend on open-source projects to do all sorts of interesting data processing. Fast-forward to now, I do believe we are at this crucial moment when, across the industry, we should encourage more open intelligence.

To me, it doesn't make sense. Intelligence is in the hands of a duopoly, or in the hands of a few people. I have never seen a community evolve in a balanced way—especially when we're talking about security and cybersecurity—if only a small number of people have control over intelligence.

Lukas Biewald

Why do you think we have this dynamic where Chinese companies are mostly leasing open-source models and American companies are mostly closed-source?

14. Should OpenAI and Anthropic open source their models

Lin Qiao

I really think there's no reason American companies shouldn't open-source their best models. I couldn't find a reason. I think we should.

Lukas Biewald

Interesting. Do you think OpenAI should just open-source its models, and Anthropic should open-source its models?

Lin Qiao

I would strongly call for OpenAI to do that. They have done that before. Last year was a high moment for the community: “Wow, here's a strong open model coming out.” I would hope they continue to do that.

Especially because they are the leaders in the market, they should set an example for the entire industry.

Lukas Biewald

Well, I think what they would say, though, is, “Hey, we spent all this money on R&D. We need to recoup that money by running the model.” Do you have some alternative monetization plan for them, or do you feel like this is the moral thing to do? What do you mean by saying they should open-source their models?

Lin Qiao

I think whether they open-source or not, they need to have a reason to do that. It should hopefully connect with their monetization path.

At the same time, there are many other American companies we work closely with that are deeply, deeply passionate about open-sourcing the next generation and really pushing the quality of open source in the U.S. to the next level. We work with Microsoft and Meta AI. We work with NVIDIA and Nomic AI. Hopefully, Reflection will have something come up soon, and Thinking Machines just launched its model.

I think we're just at the beginning. Again, based on—

Lukas Biewald

It still honestly feels like the best open-source models are coming out of China, and the most interesting innovation—maybe you could add Mistral to the list—but outside of America, it seems like the models that are really in use are generally not American, right?

Lin Qiao

Right. I think it's probably the strategy. In the past, the primary strategy for U.S. companies—whether those were frontier labs or hyperscalers—was mostly focused on pushing closed models because, again, it's fastest to scale 1 model in terms of business, right? If you think of business efficiency, that is the fastest.

But it has a limit because you cannot get into the specialized part of your business, right? If you are a general-purpose model solving common tasks, that's the limitation. What we believe the next phase of the frontier is, is getting into the most interesting part: specialized, proprietary knowledge of the business. Let businesses own their own intelligence. That will push the next frontier.

That work just started because the quality of the base model, where businesses can build their specialized intelligence on top of it, has reached a positive threshold. Therefore, making that flywheel much easier and more accessible, with much better results, and so on. We've been pushing on that before, but the quality difference was too big.

I believe that, now, with a lot of interest from American companies trying to build really, really good open models, the American gap between closed and open models will close. We also have other choices of open models across different regions. I think the open community and open ecosystem will have much more diverse contributions, and that makes me so much more excited.

15. Day zero launches and fixing DeepSeek bugs

Lukas Biewald

It does seem like there's been a lot of innovation recently in making the models run more efficiently and faster. I feel like a lot of these newer models are designed to run well.

Related to that, an amazing thing that's happening in our space is that when a new model comes out, it's almost immediately supported by Fireworks—same day. It's unbelievable how fast you guys get these things working well.

Are you in conversation with the model-building companies about what they're doing, getting early previews, and testing them on your infrastructure before the actual release date?

Lin Qiao

It really depends on the situation. Sometimes we have early access, and sometimes we don't. Regardless, we're very proud of the day-zero launch. We almost have a reputation for it.

Fun fact: In the early days of Mistral, they didn't launch the model. They just tossed out the weights; there was no model code. We reverse-engineered the model code from their previous launches and launched before they launched the API. It was a fun exercise.

We care about quality more than day-zero launch. When DeepSeek launched, we held back the launch from our side by 3 days. During those 3 days, we didn't sleep at all. The reason is that the release of the weights we got, and the corresponding code we got, had a lot of bugs.

The model didn't pass a lot of our internal evaluations, and it didn't pass our threshold. We worked closely with the open-source community—specifically, with vLLM and SGLang—to fix those bugs and contribute the fixes back, so they could also fix those bugs in their communities.

That took us 3 days, and we launched 3 days later. We just cannot deploy a model when we know there's an issue. That trumps everything.

Lukas Biewald

Why is it so important to you to launch so quickly when a new model comes out? I mean, you talk about specialized intelligence. Nobody's going to be able to take the model and specialize it to their application on day 0. Why the rush? Why are you telling people not to sleep for 3 days to get something out?

Lin Qiao

People—this is an interesting time because the velocity of AI development is so fast—always want to get in touch with the latest. There's no time to wait.

Our customers are going to make a decision about whether they tune the model and change the backbone of their models, which is a big decision. Because the model release cycle is so fast, and the releases are very close to each other, they need to be informed quickly.

Today, no one really trusts public benchmarks anymore. The judgment is really based on being able to test through their own internal evaluations and make those calls. If they decide, “Oh, this is actually a really good base model,” and they want to tune it, they need to act quickly. If they don't move quickly, the next model will come out.

It's just that the pace of AI across the entire industry is so fast.

Lukas Biewald

I mean, that leads me to another question. The shelf life of these models is so low, right? People always switch to the next great model. You have this business that's scaling amazingly well, but the metrics are so public for everyone to see. You can see the cost and the performance, and there's not much else. What is your long-term moat for Fireworks?

16. What is Fireworks' long-term moat

Lin Qiao

Many people ask me, first of all, what Fireworks is. Then we can talk about the moat based on that. They ask me if Fireworks is an inference company, if Fireworks is a new cloud, if Fireworks is a new kind of cloud, or if Fireworks is a PyTorch cloud. What is Fireworks?

We're none of these. We're a specialized intelligence platform. We've built a training platform and an inference platform, co-optimized and co-designed with the goal of maximizing quality. We're very obsessed with quality.

Lukas Biewald

When you say quality, what do you mean? Is that the quality of the result?

Lin Qiao

Quality of the result, yes. After the quality is good, this model will be in play, right? I haven't seen people sacrifice quality significantly for a much cheaper price. If you don't care about quality when you build a product, I'm not sure how that's going to help you.

After that, we customize our inference deployment for the specific application, for speed and cost.

Let's talk about quality. We're obsessed with it. We go the extra mile to get quality. For example, between training and inference, when there's a transition, there could be a loss of precision because of numerical differences and the different libraries we use. Aligning numerics is a very hard challenge, so we reach 0 KL divergence across training and inference.

Lukas Biewald

Zero what? Sorry.

Lin Qiao

Zero KL divergence.

Lukas Biewald

What is KL divergence?

Lin Qiao

KL divergence is a metric of precision. What it means is that we reach bitwise equivalence. The result from running on the training side and the result from the inference side are bitwise equivalent.

Why does that matter? After you train and deploy, you don't want to lose precision there. For RL, RL is actually a combination of training and inference, with the inference constantly going back and forth. If the numerics have even a small error, they start to amplify.

This is very hard to achieve, and we spent all that R&D to achieve it. Second, almost as a principle we operate by, we want to make our platform and its tools accessible to all sorts of companies.

Very few companies can afford 10,000 fully interconnected GPUs. There aren't many of those companies—not many of them. And even if you want to spend the money, there aren't many of those GPUs available. So we implement a very aggressively tiered, disaggregated system where we can basically pull together scattered GPUs across all regions globally to do a training run of up to tens of thousands of GPUs in one run.

This requires a lot of innovation, and the innovation isn't just making it work. It's making it work without losing quality, because the more synchronization you have across different clusters, the more possibilities there are for errors. Then you spend the money, but the result isn't there. Those are the examples where we really obsess over quality.

After quality, we continue to obsess over speed optimization and cost optimization.

Lukas Biewald

So you don't want to be just in the game of competing with other inference providers on price and speed?

Lin Qiao

We will compete. We will absolutely compete. But I think the unique part, again, goes back to the fact that every company exists for a reason. The reason for us to exist is that we're squarely focused on one size fits one. We squarely focus on customization.

Our belief is that every single company is special, and we want to deliver the special intelligence for them. That reflects in special quality, special cost, and speed. We'll do whatever it takes to optimize for that, and that's what we build our platform for.

Lukas Biewald

So it sounds like you are committing back to open-source libraries like SGLang and vLLM, but as far as I know, the core of what you do is closed source, isn't it?

Lin Qiao

Yes. Both our training and inference engines are proprietary.

Lukas Biewald

So I could imagine you're here telling me, “Hey, OpenAI should open up their models.” I could imagine talking to someone at OpenAI, and they say, “Hey, Lin Qiao, open-source your runtime and training environments.” How would you respond to that?

Lin Qiao

First of all, why do we use proprietary technology? We built the engine before vLLM and SGLang. They didn't exist. But we also need a special design because, again, going back to the roots, we want to customize to the extreme.

Delivering extreme quality, speed, and cost requires us to design the engine in a special way. We want to maximize the choices we can make to extract the most optimized result, and that requires us to design the engine in a very modular way. Each module can interact with the others through fully compatible interfaces. Each module has a few choices, and we can combine those choices. Then it becomes a search problem.

For the inference engine itself, we create a search space of more than 100,000 options. Based on the customer requirements—remember, every company is unique, their workload is unique, and their trained model and result are unique—we search across those options and find the one that's best for them.

Because of this unique design, it's hard to make other engines deliver the results we want to deliver. That's why we built the proprietary engine. We also work very closely with the open-source community. In the DeepSeek case, we worked closely with the open-source community, and we work closely with NVIDIA. We give them a lot of feedback.

We also work with AMD to give them feedback about what kind of kernels we need. Whatever feedback we give to them goes back into open source. We feel like this is the most efficient way to engage, because we also want to maintain the velocity of moving forward really, really fast. That's how we strike the best balance.

Lukas Biewald

So are you saying that you don't want to open source because you don't want to deal with outside contributions, or you want to protect your intellectual property, or what is it?

Lin Qiao

We don't think it would be productive because I've done open source before, right? Multiple rounds.

It requires a lot of people to engage and build a community, and we have done that before. For example, PyTorch took us 7 years to really become broad. Now, everything is moving so fast, and the velocity with which we're changing things is extremely fast. When you change things much faster and open source, people get confused. Where do they contribute? And if their contribution isn't incorporated in time, they won't be happy.

I don't think we have actually discussed open-sourcing our engine internally.

Lukas Biewald

Do you think—I mean, you now have lots of resources. Surely you could hire people to manage the community. Would you ever open source the Fireworks engine?

Lin Qiao

Guess how many people we have.

Lukas Biewald

100?

Lin Qiao

More than that. How many? [Laughter.] Do we have more than that? We have more. We have 100 engineers in product and technology, so in total we have 200 people. The people working on the engine are around 10.

Lukas Biewald

I guess this is—you seem so passionate about open source, and you're saying, “Oh, everything should be open source. Open source always wins.” It's interesting that you're not excited about open-sourcing the core thing that you're building.

Lin Qiao

I would work on things if we could add a tremendous amount of value. We already have great open-source projects in vLLM and SGLang. Qwen Chat LLM is great. I'd rather support them and help them go really big than build another open-source project to compete with them.

I feel like this is the part of open source where there's no ego. If there's a leader in the open-source world, then we will support the leader.

Lukas Biewald

If there's a leader, would you switch to using it?

Lin Qiao

We're open-minded. We're open-minded. Again, we're very practical. If it helps us get actual miles of optimization, we're open-minded about using any tools.

Lukas Biewald

Do you have a feeling right now, between vLLM, TensorRT-LLM, and SGLang, which one is the leader? Is there one that you favor?

Lin Qiao

I think they have unique strengths.

Lukas Biewald

Come on.

Lin Qiao

They have unique strengths.

Lukas Biewald

Come on.

Lin Qiao

I admire their effort. I know it's very hard to build an open-source project. I admire that they keep focusing there. Although they're each starting to build companies, I think that's great for them to have a real entity to see into the future. But I think they're doing the right thing for the community.

Lukas Biewald

Do you use strategies like Andrej Karpathy's autoresearch, where an LLM iteratively tries new sets of parameters and decides what to try next?

Lin Qiao

With kernels and so on, it's not surprising that we're doing that, too, but it's not easy. I don't think it's going to replace performance engineers anytime soon.

The thing is, we haven't found a model that can do things we already know how to solve. We haven't seen the discovery part of finding new ways to write kernels. We will try all different kinds of ways, but we haven't seen that happening yet.

Maybe it really depends on the quality of the base model. Maybe there's another leap into much deeper thinking when models get to 10 trillion parameters. It may be possible.

Lukas Biewald

Okay, well, switching gears a little bit.

Lin Qiao

Ooh, switching gears is scary.

Lukas Biewald

You should be scared. No. [Laughter.]

I've always admired your style as a founder, and I relate to it a little bit in that you have a quiet confidence. Unlike a lot of Silicon Valley founders, you're not really out there beating your chest and saying, “Hey, I'm so great.”

I remember when I was running my company, I sometimes got criticized by my board. They'd say, “Hey, you should be out there more. You should be more aggressive.” Do you ever feel like you should change your style or be a more aggressive founder, maybe more like some other company leaders?

Lin Qiao

I think about that every day.

Lukas Biewald

Interesting.

Lin Qiao

But I cannot change who I am, right? The way I present myself, I need to feel that I'm authentic.

Lukas Biewald

Totally.

Lin Qiao

I cannot be anyone else as a CEO. I cannot be anyone else when I speak for the company. I cannot be anyone else when I speak about my personal opinion just for the sake of getting a following.

So I'm trying to learn where the balance is. Being able to talk about the company is absolutely important. On that front, I agree with you. That's part of what I've been working on for a long time.

Lukas Biewald

I don't know if you're really agreeing with me. I'm saying I like your authentic style, and I'm not sure I want you to change. [Laughter.] From what I said, I was just curious.

Lin Qiao

My point is that there's a much more aggressive style of marketing, and I'm not sure I agree with those approaches. But it's not okay with me if people don't know what we're building or who we are. Yeah.

Lukas Biewald

I mean, I guess you're suddenly running this incredibly important company in the AI space. How has that experience been? Do you feel overwhelmed? Do you feel excited? What's it like?

Lin Qiao

I feel like I'm living my dream.

Because it's a lot of work, obviously. Everyone working in AI knows it's very intense. It's fast-paced, and things are constantly changing. The market is very dynamic, but I'm really enjoying solving all kinds of problems across the board. It's not just technical problems or product problems; it's always the intersection of product, go-to-market, procurement, and finance. It's multidimensional, and that makes me super excited. We never lack problems to solve. Every day, a new problem pops up and I need to jump in.

Our style is very flat. There's no deep hierarchy, and I don't like deep hierarchy because I want to make sure people all have similar context and are able to make decisions. But we also don't want to carve out the boundaries: "Here's a box; you fit there. Here's a box; you fit there." We stack the boxes, and then we build an organization that doesn't fit the velocity needed for AI.

Because of that, I get to work deeply with all sorts of teams. That makes me very, very excited and energetic every day, but at the same time, it's very humbling. It's very humbling to see so many entrepreneurs—fearless, regardless of whether they have experience or not. They just go in to implement their ideas and figure out how to get to market. The creativity is off the charts. We're doing our best, but I've seen so many other peer entrepreneurs doing their best, too. It's a very inspiring environment. We'll keep our heads down and deliver the best for them because most of them are our customers.

Lukas Biewald

What's been the most surprising thing about actually being in charge of the whole company?

Lin Qiao

I think maybe when I first raised my Series A, he asked me one question: "What would you do differently now that you're a founder compared with a public-company executive?" Guess what my answer was? I said, "I got it."

Lukas Biewald

What was your answer, B? Well, now that I'm part of a public company, one thing I wish we could do is share all the metrics with all the employees. I feel it's a big bummer not to be able to give all employees access to all the relevant metrics. I want everyone to know what the score is and what true north looks like, but there are a lot more controls around that inside CoreWeave. But surely that wasn't your answer. What was your answer?

Lin Qiao

My answer is, "I don't know."

Lukas Biewald

I love it. That's a great answer. I mean, that's so good.

Lin Qiao

I haven't done a company before, so—

Lukas Biewald

Nice.

Lin Qiao

So I don't know. Obviously, the subtext of his question is that there's not much success among established public-company executives who start companies. Why are you able to build a company?

I'm lucky. I started a company with 6 other co-founders, and they're all top-notch, world-class engineers and technologists. They're the foundation for us to build this company together. I'm more on the business side, and they're on the product and technical side.

17. The biggest surprise of starting a company

The character of this group is deep intellectual honesty. One thing we do a lot—I don't know if it's weird or not—is pre-mortems. The pre-mortem is: How are we going to die?

Lukas Biewald

Mm-hmm, totally.

Lin Qiao

When other people join the company, we bring them along with us to do a pre-mortem, and they're like, "That's scary." How can you discuss that, go back to it, and continue working as if nothing happened?

Lukas Biewald

Totally, yeah.

Lin Qiao

We're very candid and transparent. We want to make the best decision for the company and see the strategy through.

But a lot of the time, I think the biggest surprise in starting this company is that there's no data—not a sufficient amount of data—to make informed, data-driven decisions, especially early on. At Facebook/Meta, there was an abundance of data. You can go to the granularity of 0.1% of the lift and try to work out whether it's a massive impact. You also have all the data to make a judgment about which path is the right path.

But in a startup, a lot of the time there's no data because we're traveling and paving a path no one has traveled. If everyone's traveling that path, then you shouldn't be that company. A lot of decisions get made by intuition, but we have to validate them. The feedback loop for validation is important, so it's okay to say, "This doesn't work, and we need to shut it down." But it's not okay not to make a decision because of a lack of data. Not making a decision is a bad decision.

We never want to have analysis paralysis. That's why we do a lot of these simulations and try to make the best calls, then keep adjusting based on feedback.

Any feedback? I'm trying to do what I can, to your point. One thing I regret is that I didn't do marketing early enough. We're a bunch of engineers, and engineers have this cynical view of marketing. Based on our past history, we believe the product will speak for itself. But this market is so noisy—so noisy and seriously noisy—and people cannot differentiate the signal from the noise. So I'm trying to do a little bit more. Give me feedback. What's your thinking there?

Lukas Biewald

On marketing? I don't know. I mean, you're phenomenally successful. I do think—

Lin Qiao

No, I mean—

Lukas Biewald

If you market without the product to back it up, that can be bad, too. I feel like you have—

Lin Qiao

There are so many companies doing that, right?

Lukas Biewald

I know, but I think it's actually bad. I feel like Fireworks has a great reputation, and your customers speak for themselves. I really like your authentic style.

Lukas Biewald

Yeah, I think we're similar in style. I think you know the space better than any marketer could. So if you do marketing and build out a marketing team, I just think you should really trust yourself and keep it as authentic as you are.

The things I've regretted in marketing have been when I've let someone come in who didn't understand that marketing is such a core expression of what the company does. For a technical thing, you want the expert in the technical thing. You wouldn't want to hire a CISO who didn't know more about security than you.

But marketing is such a core expression of what the company does that you shouldn't let a marketer convince you that something you feel uncomfortable about is actually going to be good. A lot of times marketers come in and say, "Oh, this is enterprise marketing, and enterprises want to hear this." In my experience, enterprises want to feel like they're dealing with an authentic company. They want to be spoken to in an authentic way. At the end of the day, there are human beings behind that, and you understand your customer.

Marketing can help get the message out, but if I were you, I wouldn't let a marketing person come in and tell you what the message should be. It seems like you feel really good about specialized intelligence. That feels authentic to you, so I would lean on that. Keep it authentic because, honestly, from the outside, it looks to me like it's working.

Lin Qiao

Mm-hmm. Okay.

Lukas Biewald

So, that—

Lin Qiao

Anything we can do better there?

Lukas Biewald

What's that?

Lin Qiao

Anything we can do better there?

Lukas Biewald

I don't know. I feel like you actually have a pretty clear message. I guess I didn't think of it as an inference company, so you're all kind of fighting—

Lin Qiao

What is it we're fighting?

Lukas Biewald

Yeah, you're still kind of fighting a simple narrative. I think you're just going to have to keep saying it over and over. Your website doesn't—

Lin Qiao

Reflect that.

Lukas Biewald

—reflect the messaging that you're saying. You just can't deliver a very nuanced message ever—

Lin Qiao

Yes.

Lukas Biewald

—to people. You just need to be like: "Specialized intelligence. That's what we are. Here's what it is." Just bang on the 2 words. You're not going to be able to deliver much nuance broadly.

Lin Qiao

Makes sense.

Lukas Biewald

You built the industry-leading company in your space. It's awesome. I don't know.

Lin Qiao

I'm always paranoid. I'm always paranoid.

Lukas Biewald

Totally.

Lin Qiao

Yeah, because the market is so dynamic.

Lukas Biewald

Yeah, I know. And that seems so crazy. It's so dynamic. It seems like it would be stressful to be in your seat.

Lin Qiao

Yeah. But at the same time, it's good to have a team. So, yeah, we're building out the team.

多数 AI 初创公司都在扩张中走向破产|Fireworks CEO Lin Qiao — 文字稿与摘要 | BidClub