[BidClub_]
20VC · · 77 min

The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao

Harry StebbingsLin Qiao

YouTube
TL;DR
  • Lin Qiao's founding thesis is a direct challenge to AGI maximalism: "if you think intelligence is a derivative of data," the majority of the world's data is private, "locked inside enterprise," and will never be shared — so the frontier is specialized, private intelligence, and the endgame is "millions of specialized models — one per application, per use case," not one model that rules everything. Fireworks already processes 40+ trillion tokens a day, the majority from customized models, not off-the-shelf ones.
  • On Harry's are-OpenAI-and-Anthropic-overvalued question, Lin reframes frontier labs as "power lines" — essential infrastructure that won't replace what's built on top — and leaves the investor question standing: are power lines good businesses when open and closed models have both "crossed the quality threshold" and open weights can be tuned on small proprietary data to beat general models on your eval? "What I don't want to see is there's only one company owns intelligence."
  • The headline economics call: token costs haven't fallen yet "because of supply chain constraint," but competition will deliver 10x cost reduction in the next three years, driving 100x usage — and Fireworks' own token count could go "anywhere ranging from 20 to 100x" by end of next year. "We're at a very early stage of S-curve explosion." Harry's inference — a capex bubble thesis is "ridiculous" — she endorses, with the caveat that the real bottleneck is the bottom of Jensen's five-layer cake: energy, chips, "we're bottlenecked by small parts — transistors."
  • "Scaling to bankruptcy" is the mechanism pushing enterprises to open weights: PMF and durable business have decoupled because inference isn't a commodity like CPU was in SaaS — incumbents with huge traffic can't get AI features past the CFO, so they must own and tune their own models. Her three-year non-consensus: "every single company will own their own intelligence as a must-have. It's not optional."
  • On app companies training their own models (Harvey vs Legora, where Harry disclosed his Legora position): Cursor pioneered tuning and "now almost all coding companies tune their own models" — the tool-calling harness "needs to be co-trained with the model powering it." Harry said Fireworks' CTO Dima was embedded at Cursor for months; Lin described decoupled RL infrastructure running across five-six data center regions on scattered GPUs instead of an InfiniBand supercluster.
  • The business: $800M in AR, expecting to "at least double" by year-end with just 200 people. The 30-40% gross margin (vs SaaS's 80%) is framed as a hyper-growth choice, not the new normal — "constraints slow down innovation." Hard line: "we absolutely are not going to move into application layer"; data centers are "always on the table — the question is timing"; chips are out because workloads are too dynamic and hardware depreciation math has broken ("within a year from one vendor alone we have three SKUs").
  • Chinese open models (the top six on OpenRouter) get a pragmatic answer: guardrail every model, open or closed, because every provider "will infuse their own judgment, their own taste" into training. If China restricts access, it's "a big impact in the short term" — but the US "will be able to build that open system by ourselves, and we should" (Nvidia is training a model likely called Nemotron amid that supply-chain gap).
Digest · the substance, structured for research

1. The founding thesis: most of the world's data will never touch a frontier model

  • Lin's core argument, in response to Harry's question: "If you think intelligence is a derivative of data, then the majority of the data is actually not used for training a general intelligence model." Training corpora are public internet plus labels — "a very small corpus" against the world's data, most of which is "private, locked inside applications, locked inside enterprise — it will never get shared with anyone else because this is the company's proprietary IP." Fireworks exists to activate it: "the frontier of the intelligence is actually private intelligence, specialized intelligence."
  • The backstory explains the conviction: PhD in distributed systems, LinkedIn, and by 2015 a business proposal and a co-founder list — but she paused because "I don't think I have the skill set on people." She joined Facebook "secretly planning to learn for one year or two and leave"; she stayed seven, then founded Fireworks at 48. Eric Vishria broke his no-big-tech-directors rule after his advisor warned "how many big tech executives have you seen being successful? Very few."
  • Harry's own marker on the company: a $10 million check after a 15-minute meeting — "one of the easiest investment decisions that I've made in a 10-year investing career."

2. AGI maximalism vs "an army of robots"

  • Harry's sharpest early push: isn't activating private enterprise data exactly Anthropic's thesis with Claude co-work? Lin's reframe — Anthropic "fully believes in AGI," defined as one model that solves all problems, which by definition means never specializing. Her rebuttal is civilizational, not technical: regions differ in values, policy, taste, and "if our future world is going to be ruled by one standard, a taste dictated by one company, we turn ourselves into an army of robots."
  • The line she keeps returning to came from Jensen Huang after his GTC keynote: "There's no specialized general company" — every company exists on a unique belief, baked into product design, data, and understanding of user intent, "not learnable or assured by another company sitting outside."
  • So why do Dario, Sam, Larry and Sergey talk AGI as inevitable? Lin's answer: they're "building power lines to distribute a really great source of intelligence" — vital, like electricity enabling her beloved coffee machine — "but is this power line going to replace everything we do? I don't think so." Harry's investor translation, left deliberately open: "are power lines good businesses?" On government stakes (Sam's 5% offer to the administration), she cites the PG&E precedent but lands on: "what I don't want to see is there's only one company owns intelligence."

3. Open source crossed the threshold — and "scaling to bankruptcy" pushes everyone there

  • The founding bet — build on open models when they were "almost at infancy" — came from PyTorch roots: "openness gave control." It paid off twice: both open and closed models "crossed a quality threshold" solving real problems, and open models became easy to steer — with a small amount of unique company data "you hill-climb towards your eval" and "often the end result is: to solve your unique problem with your data, you are better than a general purpose model." Fireworks itself runs open models for recruiting, finance, and internal coding agents.
  • Her signature framing of the economic forcing function: "Have you heard about scaling to bankruptcy?" In SaaS, product-market fit and durable business were equivalent because CPU was a commodity; now they're separate concepts. Startups with real PMF "could scale into bankruptcy" — and it's worse for digital-native incumbents whose CFOs can't justify rolling AI features out to a decade of accumulated traffic. The alternative: "have control over your open weights model."
  • Harry's counter: doesn't Sam Altman's newly released, dramatically cheaper models solve this? "It could be," she concedes, but open weights have no acquisition cost while frontier labs must recoup R&D — and "you just cannot customize those general purpose models... with open model you have full control." At billions of users, "even 5% of cost reduction means a lot — let alone what we have seen in the past, five times to ten times."

4. Chinese models, sovereignty, and the millions-of-models future

  • On the national-security question — "the top six models on OpenRouter today are Chinese" — Lin refuses the framing: guardrail every model, open or closed, because any provider "will infuse their own judgment, their own taste into the model training process; you cannot guarantee it matches yours." Then the episode's biggest call: "It may be scary, but I think that's true. It will be millions of specialized models — one per application, per use case."
  • If China restricts open-model access (reports surfaced last week): "a big impact in the short term," but open ecosystems attract many parties — "in terms of talent density and resources, I do believe US will be able to build that open system by ourselves, and we should." Nvidia is training a model likely called Nemotron amid this supply-chain gap: a missing US-native open model "is a supply chain problem."
  • On sovereign AI, Harry cites Fable being briefly banned by the administration for 19 days as Europe's wake-up call. Lin extends the power-line metaphor: "every country should own their own power line... you don't want any single person to cut you off. That's an extremely scary moment" — and the same holds for every company.

5. Should app companies train their own models? Harvey, Legora, and the Cursor precedent

  • Harry disclosed his Legora position before asking: Harvey committed to its own model, Legora didn't — a year ago the non-builders looked right, "now it looks like they're wrong." Lin's context: the application lifecycle has collapsed from "tens of very strong product engineers and PMs, multiple quarters" to "one person, a few weeks" — so implementation is no longer the moat, and competition has moved elsewhere.
  • Harry's pushback, worth keeping: enterprise legal is multi-year relationship sales with bespoke deployment — "it's not like 11 Labs where you pick it up and go." Lin concedes legal is uniquely unforgiving ("lawyers are usually more conservative... legal is not tolerant at all on errors — that's why lawyers get paid") but insists both firms own proprietary workflow knowledge, and the orchestration harness deciding which tools to call "needs to be co-trained with the model powering it."
  • The proof point: "in the coding space, Cursor probably is one of the pioneers" tuning their own model — "now almost all coding companies tune their own models." Maybe Harvey-vs-Legora is just a timing question.

6. Inside the Cursor build: distributed RL on scattered GPUs

  • Harry said Fireworks' CTO Dima was embedded at Cursor for months building RL infrastructure. Both companies being "capital conscious," they broke reinforcement learning into trainer and rollout — instead of hyperscaler-style 10,000-100,000 chips on InfiniBand ("extremely expensive and really hard to find"), they run "fully distributed across five, six data center regions globally, tapping into scattered GPUs." The hard part is syncing fresh weights across regions so rewards don't go stale: "if it's too stale, then you are too off."
  • Her adoption-curve frame for why this partnership matters: early adopters are hackers — "they have researchers from frontier labs and they want to control every single thing" — while the late-stage mass market needs little control. Fireworks targets the latter, but partnering with the pioneer teaches "what is required to get there."
  • On Cursor concentration risk after the SpaceX acquisition, an honest answer: "everyone's concerned — the whole entire industry" is shaped by a handful of escape-velocity apps, and "all model companies were concentrated on Cursor." Since then: "last year is the year of coding, and this year is the year of co-work" — general-purpose deep research plus legal, finance, customer support, recruiting, healthcare — and an uptick of consumer companies rethinking recommendation systems with genAI.

7. 40 trillion tokens a day — and why the capex-bubble thesis fails

  • The numbers: "we today process more than 40 trillion tokens a day," the majority from customized models. End of next year? "Anywhere ranging from 20 to 100x could be possible... we're at a very early stage of S-curve explosion." Harry draws the conclusion — then the capex bubble idea is ridiculous — and she agrees, redirecting to Jensen's five-layer cake: "we are bottlenecked by the lower part" — energy, chips, manufacturing lines never designed for 100x scaling. "We're bottlenecked by small parts — transistors."
  • Why haven't token costs fallen yet, per Harry's challenge? "Supply chain constraint — but we are living in a free economy": shortage invites competition, competition compresses cost. Her prediction: "10x cost reduction in the next three years, and this 10x cost reduction will drive a 100x usage." (The demand anchor Harry offers: an investor says Salesforce spends ~3.8% of developer salaries on Anthropic and Claude Code.)
  • The nuance she insists on: "not all tokens are equal" — evaluate token economy per task, since a model 2x cheaper but 2x more verbose solves the same task at the same cost. As quality improves, "being precise is going to be part of the optimization."
  • The under-discussed bottleneck: "we don't have a great system designed for 10-trillion-parameter models today" — solving it needs co-design from model through serving platform down to systems of chips, and that's where she sees the remaining infrastructure innovation.

8. The business: premium on quality, 30-40% margins, and where the stack ends

  • On Together being cheaper: "we're probably not comparing apples to apples" — most Fireworks traffic is customized models, optimized quality-first to the point of "zero KLD": bit-equivalence between training and inference systems, "we do not lose a bit of accuracy." Otherwise "you pay your training investment by discounted quality — why do you do that?" One-size-fits-one deployment, backed by an applied-ML (FDE-style) team now building agents to automate deployments.
  • On 30-40% gross margins vs SaaS's 80%: "I don't think that's the new normal" — it reflects hyper-growth. "Margin optimization is a constraint problem... and constraints slow down innovation"; you optimize the heck out of a system only once you know you'll scale it a thousand-x. The disaster case: hill-climb margin to a high number and stop growing. Hard boundaries: "we absolutely are not going to move into application layer"; data centers "could always be on the table, but the question is timing."
  • Chips are categorically different: "I know building a chip is extremely hard." Meta's MTIA dates to ~2018 and serves ranking/recommendation workloads; you tape out only "once your workload stabilized," and today's AI workloads are "very, very dynamic." Pre-genAI accelerator startups were serendipity bets — the SRAM-heavy designs happened to fit memory-hungry models (she notes Nvidia's recent acquisition of Groq, whose SRAM-intense chips pair naturally with flops-intense GPUs: prefill on one, generation on the other — a heterogeneous data-center design she finds genuinely interesting).
  • Depreciation math is breaking build-vs-buy: hardware used to depreciate over six years against three-year release cycles; "now within a year from one vendor alone we have three SKUs," models peak weekly, and "the newer model usually runs the best on the newest hardware." After three years, "do you still want nine-generations-older hardware running a three-year-old model? That's questionable."

9. $800M in AR doubling, George Hu, and the Jensen operating system

  • The trajectory: $800M in AR, "we think we can at least double" by year-end, at 200 people (50 a year ago). The formative customer story: Cursor signed when they were "single-digit million dollars... only two years ago — they grew by 100 to a thousand-x over two years," choosing a partner for platform R&D and focusing on product. Harry's context for the era: Slack's 1-to-10M in 18 months was once venture's golden child.
  • On hiring ex-Salesforce president George Hu: a year ago she told him "we're probably too small for you" — he spent that year helping interview executives before joining as growth made it serious. Her people filter isn't competence: "whether they are really built for extreme ownership... we are not putting people in boxes and stacking the boxes into a tower."
  • The Jensen lesson, from his one-minute email replies: "Leadership is just judgment. It's not privilege." In a high-velocity space you can't wait for information to cascade through layers — "not knowing what exactly is happening and having the position of making judgment makes bad leadership." Her own change of mind this year: shedding the fear that fast headcount growth kills agility. Her admitted mistake: waiting on marketing — "marketing is not about flows, it's about education... clarity."
  • The closing three-year call: "every single company will own their own intelligence as a must-have. It's not optional" — the same logic by which every company owns its software stack. And the industry's next phase: "token maxing is just a thing in time — we'll quickly move into ROI maxing, which is all about running a business."
Lin Qiao

What I don't want to see is only 1 company owning intelligence. That doesn't make sense to me. I think last year was the year of coding, and this year is the year of co-work.

I do think the cost of tokens will go down drastically. There will be a 10x cost reduction in the next 3 years, and this 10x cost reduction will drive 100x usage. We absolutely are not going to move into the application layer. It's very unclear to us whether we will move down into data centers and so on. That could always be on the table, but the question is—

Harry Stebbings

Ready to go, Lin? I am so excited for this. I heard so many great things. I just got off the phone with your co-founder, Dima. I spoke to Alfred Lin, Sonia, Matt Miller, and many more, so thank you for joining me.

Lin Qiao

Thanks for having me.

Harry Stebbings

Now, I heard that Eric Vishria has a rule: don't invest in big tech executives. But he broke that rule with you, which is very special.

Lin Qiao

I think so, too. A funny story: after we decided to shake hands, he called me and said he had talked with one of his advisers. His adviser questioned him: “Hey, how many big tech executives have you seen being successful in starting a company?” Very few.

He told me that, and I was surprised. I said, “Are we breaking our handshake now?” He said, “No.” Since then, we have worked very closely with each other.

Harry Stebbings

Eric is one of the best. You also started the company when you were 48.

Lin Qiao

Yeah.

1. Why Starting a Company at 48 Was an Advantage

Harry Stebbings

That's quite late. Can I ask you how you reflect on being a 48-year-old founder when we glorify starting a company when you're pretty much 15 these days?

Lin Qiao

I didn't think deeply about that. I always wanted to have a tech business myself. I actually wanted to start a business in 2015. I'm a first-generation immigrant and came to the US in 2000. I did my PhD in distributed systems and computer science, especially focused on databases.

Databases are a very complex system to build, with a lot of different objectives to optimize for. I pretty much touched every single aspect of processing data. Then I moved to LinkedIn to further drill down and build systems and products to be used to drive real impact.

At that time, I felt I was ready to start a company. I knew all the tech, I knew what product to build, I had a business proposal, and I had a list of people I wanted to start a company with. I spent time thinking about it, and I paused because I didn't think I had the skill set around people to build a company.

It's not just about the product. It's not just about the tech. It's actually about people. I decided I wanted to go to a place where I could learn the most about people, and the best company at that time was Facebook. It was a rising star in Silicon Valley.

Secretly, I was planning to learn for 1 or 2 years, leave, and go back to do my own business. I stayed there for 7 years.

2. The AI Layer Everyone Is Overlooking

Harry Stebbings

With Fireworks, you saw something in inference that the world was not focused on. The world was focused on training. I think it's helpful for people to understand the stack, because beneath you, there are obviously chip providers—the NVIDIAs of the world—and above you, you've got the model providers, while you sit in between. Why is that a valuable part of the stack and not a commodity?

Lin Qiao

That's a really good question. Why bother with specialized intelligence? Why not just use generalized intelligence and worry about fewer things? You could just build on top of an API provided by frontier labs. Wouldn't that be much easier?

The argument is the following: if you think intelligence is a derivative of data, then the majority of the world's data is actually not used for training a general-intelligence model. The training data is coming from the public internet and labeled data. The public internet is a very small corpus of data compared with the world's data.

The majority of the world's data is actually private, locked inside applications and locked inside enterprises. It will never get shared with anyone else because it's a company's proprietary IP. So then the space becomes very interesting, because the majority of data is not being activated to derive any intelligence.

That's where we believe our role is: to activate that data. We believe the future frontier of intelligence is actually private intelligence and specialized intelligence. That's where Fireworks came from. From the beginning, we have been focusing on driving the value.

Harry Stebbings

I have so many questions to ask you. I totally understand you in terms of the value in private data within some of these largest companies. Is that not the premise of Anthropic's enterprise business, though, with Claude for Work and with a lot of the adjacencies that they're building? Would Dario not say that that's exactly what we're going after?

Lin Qiao

That's interesting, because I view Anthropic as a company that fully believes in AGI. The definition of AGI is that there's 1 model that can solve all problems in the best way. To me, that's the definition of AGI.

That means you do not need to specialize. That 1 model should be able to solve all the problems. If it's so intelligent and has so much knowledge of every part of the businesses and every part of the jobs, and it can fulfill all of them, then why do you need to bother specializing?

That itself is a validation that we're living in a world that's not ruled by 1 principle. We are living in a fully diversified world. Give you 1 example: different regions will have different value systems and different policies. They'll have different ways of conducting business and different lifestyles. It's all taste, choices, and judgment combined.

I think that's what defines us as human beings. We are not robots. If our future world is going to be ruled by 1 standard, with a taste dictated by 1 company, we turn ourselves into an army of robots. That's very depressing to me.

I think what separates Homo sapiens from other species is creativity—the deep desire to pursue new things and discover new ways of living. That defines us as human beings, and that part cannot be copied.

That's why, in Silicon Valley, there's so much creativity. Across the world, there's so much creativity in building new businesses. What is a new business?

I had an interesting conversation with Jensen after his GTC keynote. We actually recorded it, and it was interesting.

Harry Stebbings

I watched it. It was great.

Lin Qiao

Recording with Jensen is not really recording. He just started having a conversation with me. I didn't know his crew had already started recording, and we just kept talking. It's so easy.

We talked about specialized intelligence. He said 1 thing to me: “Lin, you're right. There's no such thing as a general company. Every company is built on a special belief about doing things; otherwise, there's no reason it should exist.”

It feels logical, but then I started to think back about what he said. It's profound, because every single company is doing something unique that justifies its existence. This something unique is deeply baked into its product design, its software design, and its system building. It's deeply baked into the data and its interaction with its users, and into its deep understanding of its users' intent, their interaction with the product, their engagement, and so on.

All of that is the fundamental basis of why a company should exist. That is not learnable or assured by another company sitting outside.

3. Is AGI Really the End Goal?

Harry Stebbings

Can you help me understand, then? As a podcaster, I specialize in asking basic questions, so forgive me. Why, then, do people like Dario, Sam, Larry, and Sergey talk about AGI in the way that they do—as inevitable?

Lin Qiao

I think what they build is fantastic, because they are basically building a power line to distribute a really great source of intelligence that everyone else can build on top of. That's how I view their contribution.

4. The AI Infrastructure Race Is Just Getting Started

If we don't have this fundamental infrastructure, then we will not have all kinds of appliances living in our homes. I love my coffee machine, and it's specially branded, right? But without that power, we don't get to do the things that are fun, that are unique, and that are special—things that ingrain and encode our taste.

5. Why One Company Should Never Control Intelligence

So I do think that's very, very important. But the question is: is this power line going to replace everything we do? I don't think so.

Harry Stebbings

The question for me as an investor is: are power lines good businesses? You said something about PyTorch and the open ecosystem. Open source, in the last, I would say, 3 months, we've all realized, is accelerating so fast, and the capabilities have increased to such an extent that it's not quite comparable, but it's getting 90% as efficient and, as Chamath stated, 15 times more cost-effective.

Are power lines good businesses in a world of open source?

6. Open Source vs Frontier Models: Who Wins?

Lin Qiao

Here's how I view open source. Early on, when we founded the company, we had a pretty deep debate among the co-founders: what do we do? Do we build our own models, or do we build on top of open models?

At that time, open models were almost in their infancy. If we were going to take that direction, it was a huge bet that they were going to do well. But with our PyTorch experience, we believed in the open community. We believe in openness. That's a fundamental principle we operate with.

Openness gives control. Openness gives control to the user. Think about open models: once a model is released, you have full control of the weights. You can change it however you want. It's yours.

And then you can build on top of it. That is a fundamentally different operating principle that we believe in because of our roots in open source. We took that bet, and it did pay off in the sense that the quality of both open and closed models has significantly improved over the past 2 years.

Both of these streams crossed a quality threshold to the point where they can solve so many problems. Within Fireworks, we do our own product development. We use open models to drive our recruiting process, candidate sourcing, and feedback collection. We use open models to drive some internal finance processes and, obviously, for coding and reasoning models to help us debug.

We have a ton of agents within Fireworks ourselves, and we are cost-conscious. Both model categories have crossed a threshold and can solve a wide variety of problems. Second, once open models cross a threshold, they are much easier to tune. The ability to steer a model is part of its intelligence, and models are much easier to steer, especially with a small amount of data.

A particular company might have a small amount of unique data, and then we can hill-climb toward its eval. Oftentimes, the end result of hill-climbing is that, to solve your unique problem with your data, you are better than a general-purpose model.

7. Are AI Giants Massively Overvalued?

Harry Stebbings

When 90% of enterprise workflows can be done, as you said, through the incredible array of functions that you now use open source and open models for, the usage for frontier models will not be as large as it would be if they were needed for everything. Are these companies actually dramatically overvalued and overestimated if the majority can just go through open models?

Lin Qiao

I think people are starting to realize it. I remember, 2 years ago, I went to different places and talked about an interesting phenomenon that did not exist in the SaaS era. During the SaaS era, product-market fit and a durable business were almost equivalent to each other. The hardest thing was finding product-market fit, and then, once you found it, scaling as fast as you could, because CPU is a commodity. The infrastructure you build on top is almost a commodity, so you do not even have to worry about it.

Now, product-market fit and a durable business are 2 separate concepts. For startups, we have great companies that have product-market fit. Customers want to pay them, and they really value their products, but they cannot scale because, once they scale, they could scale into bankruptcy. Have you heard about “scaling to bankruptcy”? That is a real problem.

It is an even bigger problem for incumbents, the big digital-native companies, because they have the traffic. They got a winner a decade ago when they were startups, and they have a huge amount of traffic. Once they roll out those AI features, they are going to reach their entire customer base, and they cannot afford to do it because their CFO looks at the cost proposal and cost forecasting and says, “There is no way you can justify this.”

It becomes a real problem for all those innovators. They really want to plug into this new, disruptive technology, but they cannot afford it. They need to find an alternative to be able to afford it, and the alternative is to have control over an open-weights model and roll out their own model.

Harry Stebbings

Or you see what Sam Altman released in the last few days, which is dramatically lower-cost models. I cannot remember the amount, but I think it is half as expensive, or maybe 3 times cheaper. Is the next step that we just see a massive reduction in price from the frontier models?

Lin Qiao

It could be, but at the same time, it is just a very different operating principle. For open-weights models, model acquisition has no cost. Obviously, some companies train those models and are willing to open them up. I know that, within the US, there are multiple companies doing that, including NVIDIA, which is training Nemotron. We are working very closely with them.

8. Why Open Models Could Beat Closed AI

Once the model is there, whoever uses that model has literally no cost. There is a fundamental cost for the frontier labs to invest in those models and recoup their R&D costs. Second, you just cannot customize those general-purpose models. You use them as is, on top of an API that you have no control over.

With an open model, you have full control. You can tune it however you want, and you can use it however you want. Fireworks is a specialized intelligence platform. We offer all sorts of tools for you to easily customize the model for one specific use case. After that model is tuned with high quality, we further optimize it for inference deployment.

Think about Fireworks this way: We think about every single model deployment as “one size fits one.” It is unique for your workload only, and it is optimized for your workload only from a quality, speed, and cost point of view. We believe that is absolutely needed, because once you think about production scale—reaching millions of users, tens of millions, or billions of users—even a 5% cost reduction means a lot. It is a massive amount, let alone the 5–10 times cost reduction that we have seen in the past.

9. Should We Trust Chinese AI Models?

Harry Stebbings

The 1 question that I do have to ask is about the concern enterprises have around national security. When you look at OpenRouter, I think the top 6 models today are Chinese models, and they are incredible quality. The speed of development is incredible, but they are Chinese models. Do we have serious national security concerns when analyzing the power of Chinese open source?

Lin Qiao

I think there is a huge debate happening right now across the industry. Once the model is open, you can put all kinds of guardrails specialized to your business around it. I would say that, for all models, it does not matter if they are open or closed: You should put your own guardrails around them.

The fundamental reason is that a model provider will infuse its own judgment and taste into the model training process. You cannot guarantee that it matches yours. Remember, it goes back to Jensen’s comment: There is no such thing as a general company. Every company is special. Every company will have a special design principle, a special taste, and a special target audience to serve.

10. How Cheap Will AI Become?

Because of that specialty, it is guaranteed that the judgment, taste, and design principles from one company will mismatch or misalign with your company, which is a special problem. That is the reason you need to tune those models to match yours.

I really believe the future will not be a small number of AGI models dominating the world. I really believe the future may be scary, but I think that is true: It will be millions of specialized models, 1 per application and per use case.

Harry Stebbings

We saw reports in the last week that China was looking at restricting access to its open models because it was seeing the development happening so quickly and at such a high quality. What would happen in a world where China actually started restricting access to its open models, given the lack of open models we have in the US?

Lin Qiao

I think it would have a big impact in the short term. But the beauty of the open ecosystem is that it is not 1 provider. That is why it is open. It usually attracts many interested parties to participate.

In terms of talent density and resources, I do believe the US will be able to build that open system by itself, and we should. I have seen this happening again in many open systems: There are a thousand flowers blossoming, and that is the beauty of it.

Harry Stebbings

When we talk about the specialization of intelligence within enterprises, as you have described, let us take a prime example—which I do not particularly want to take because I am an investor in Legora, and I think I know which side you are going to fall on here. You have 2 companies that compete in the legal space: Harvey and Legora. Harvey has committed to building its own model, and Legora has not.

A year ago, it looked like companies that did not commit to their own model were right, because frontier models were increasing so quickly in capability. Now it looks like they are wrong. Should companies like Harvey and Legora be building their own model? And, actually, if you do not, what happens?

Lin Qiao

Here is 1 observation I have, and many people have as well: Software development, and especially the SaaS space, has been significantly disrupted because of the general intelligence of coding. The application development life cycle has significantly collapsed in terms of the timeline and resources needed.

In the past, it required tens of very strong product engineers and product managers to convert an idea into an implementation and then into a production-scale product. That required multiple quarters or years of investment. That was a deep moat. Today, 1 person in a few weeks can possibly launch their idea into a product and scale quickly.

This is unprecedented, and it also creates interesting dynamics and redefines where the competition is, because it is really hard just to compete on the idea of an application. The application by itself is no longer enough, because many people have similar ideas. Implementation is no longer such a big barrier.

Harry Stebbings

Is that actually true, though, when you are looking at enterprise deployment and enterprise rollout, if you are working with some of the biggest law firms in the world? The enterprise sales cycle is at least multiyear, with relationship-building, which is very tough, and then you have deployment that is very customized. It is not like ElevenLabs, where you pick it up and go. It is different.

Lin Qiao

Also, I think the legal space is particularly challenging because lawyers are usually more conservative.

Legal is also not tolerant at all of errors, right? Because that’s why lawyers get paid, right? You need to build a very rock-solid case. If something hallucinates and generates the wrong judgment, then you’re in trouble. So I do think the legal space is a very interesting space to penetrate, and these companies are both doing a great job.

But on the flip side, I do think both companies own proprietary knowledge and information on how to build those assistants, to do case studies, to go deep in driving legal research and all this, right? My understanding of legal is so shallow, but there are so many different versions or flavors of cases. So I do think they are in a unique position to convert that deep understanding, and they all have data.

It’s not just about how defensive their business is. It’s about, hey, oftentimes when they build those assistants, there’s a harness integrating, deciding, and orchestrating which AI tool to use, which tools to call, and this is bespoke, this is customized. The accuracy of calling those tools, and calling what kind of tools, is important. Even that harness needs to be co-trained with the model powering it, right?

So there are just ample examples of driving that business to excellence by owning their own intelligence of how to do that in the workflow layer. So maybe it’s a timing issue. Coding, for example, I think in the coding space, Cursor is probably one of the pioneers.

11. Will AI Model Breakthroughs Ever Slow Down?

Harry Stebbings

Starting to tune their model, and now almost all coding companies tune their own models. Does that pace of model development slow down? Because every single day, it seems like we have a new model with a new capability, and it’s like, “Oh my gosh, Cursor’s newest model is amazing.” Next, we have someone else’s—Mistral’s—newest model is amazing. Gemini’s newest model is amazing.

In 3 years’ time, will the pace of model development still be so fast, and model superiority be so transient, where one day it’s one and the next day it’s another?

Lin Qiao

There are a few layers of model advancement. There’s base general IQ advancement, so those will take step functions. That’s why when they release, there are always major releases or minor releases, right? The major releases are step functions.

As you remember, at the beginning of last year, this whole thinking process was new, right? The model just doesn’t spit out an answer immediately. The model will think by itself and spit out an answer; it’s much better that way. So that’s one step function, and there are many step functions we have seen through.

But I see those as every year or every 3 quarters, there’s a major leap. At the same time, building on top of those base models, I can see specialization start to accelerate because, as I said, it’s really like a tree, right? There are so many branches and leaves that can possibly hang on the trunk. As the base-model quality starts to have step-function leaps, there’s so much more we can do to specialize.

So I do see specialization in the world accelerating much faster than the general-intelligence part.

Harry Stebbings

When we think about the general-intelligence part, just before we move further into the stack of multi-model, you saw Sam proffer the 5% kind of gifting of OpenAI and others to the administration. Do you think we’ve reached a stage where model development is so advanced and so important to society that they will in part be government- or administration-owned?

Lin Qiao

That’s a very interesting question. I think there were precedents of that. If we think about the foundation tier of those general-intelligence models as fundamentally a base infrastructure for the big economy to operate around, there have been precedents of, like, PG&E owning electricity and gas, and so on, right?

I actually don’t know, but what I don’t want to see is that there’s only 1 company that owns intelligence. I think that doesn’t make sense to me because there are different flavors of intelligence. There’s this general, common intelligence that benefits everyone, and then there’s specialized intelligence that actually helps us advance in history, to think differently, to create new paradigms of living or new paradigms of doing business and shaping industry.

I don’t want that to die because there’s only 1 company that can do that. I don’t think that makes sense.

Harry Stebbings

With the many models blooming, theoretically, there’s the idea that you will route a task to different models depending on what they specialize in.

Lin Qiao

I think so.

Harry Stebbings

With that in mind, will you not build your own OpenRouter for the world to cater to that?

Lin Qiao

Yes, you can argue they’re the best builders because they deeply understand their use case and they have the evals. So, again, my thinking about what the frontier is is not just this 1 model. The frontier could be your special routing mechanism for your business.

You decompose that based on, hey, in order to fulfill this task, usually you need a highly intelligent layer, maybe the most expensive open or closed models, to judge at the highest complexity. Usually, people will also build sub-agents to solve smaller problems. Then those can go to smaller open models, and those can also be further customized to fit into your special design.

I’ve seen a lot of people already doing that today, and we also think there’s a space to build an automatic routing system that can learn by itself. Compounded with an automatic tuning system, eventually we think it should all be automated. Then you can see a self-evolving system based on what flows through your product.

Your product keeps evolving. Your product is alive, right? So you keep deploying and launching new features to interact with your users, and it will be a totally self-evolving, automated system.

Harry Stebbings

Do you think, then, that routing layer of the stack is valuable if it can be automated or it can be built on its own? Is that a valuable layer to have?

Lin Qiao

I definitely think so.

Harry Stebbings

You do think so?

Lin Qiao

I do think so.

Harry Stebbings

If it can be automated or companies can build it themselves, why would you need a Requesty or an OpenRouter?

Lin Qiao

You probably don’t.

Harry Stebbings

Yeah, we’re not there yet. But I do think this could be an area of innovation.

You said Cursor was one of the front-runners in terms of how innovative they’ve been. I completely agree with you, but I heard—and I really stalk you before shows—that your CTO, Dima, was embedded at Cursor for months building the RL infrastructure. Is that how it has to be done, and is that scalable?

Lin Qiao

What’s happening is usually, in the early adoption curve of new technology, the early adopters are all hackers. Hacker is not meant in a bad way; it doesn’t have a negative connotation. They have deep expertise in a certain area, and they want to control a lot of things.

Versus in the late stage of a new tech adoption curve, it starts to get more accessible to a much bigger cohort of users that doesn’t have deep expertise, and they need less control. So it always goes into deep control first, usually, and a little control later.

We definitely are aiming toward the later stage as the ultimate end state we want to target, but it’s also extremely valuable to understand what is required to get there. So that’s why we partner deeply with Cursor. They are the pioneer trying those ideas. They do have researchers from frontier labs, and they want to control every single thing.

At the same time, we’re also pushing to the boundary. We’re doing things that never existed before. We’re doing systems that never existed before because we push the boundary; that is unique to this particular setting.

Harry Stebbings

Okay. What is unique here?

Lin Qiao

Typically, if you think about training, training happens; training is very capital-intensive. It usually happens in big companies. They have a lot of money, and they put that money toward buying very expensive training clusters interconnected with each other. Super expensive.

Once you have those expensive, large fleets, usually you don’t need to think too deeply about how to be efficient; you just focus on doing your work. Cursor is like us; they’re a startup. Both of us are very capital-conscious, and we want to be efficient while we don’t want to slow down research innovation.

So together, we figure out a very smart way to drive their training process. They do massive post-training, which is reinforcement-learning-based, and we break the reinforcement learning into 2 pieces. One is the trainer that is tweaking the weights of the model, and it basically generates new model versions constantly.

That new model will be deployed to what we call the RL rollout. It basically deploys that new version to interact with a synthetic environment—a synthetic coding environment—or a real coding environment, and then gets the reward back to judge if that model version is good or bad, right? So that’s a rough process.

We decouple these 2. In the past, in large hyperscalers, they ran that all together. If you think about it, you get 10,000 or 100,000 chips all interconnected through Infinity Band. It’s extremely expensive and really hard to find, but they need to go really quickly.

We designed a fully distributed system, and we run it across 5 or 6 data center regions globally, tapping into scattered GPUs, and they’re able to run massive jobs—our jobs. But the challenge there is that we need to sync model weights across all these different regions.

How hard can that be? It matters because the latency, the delay of sending these weights over, is going to dictate how fresh the rewards are. If it’s too stale, then you’re too far off. So it’s a balance, but we innovated a way that we can distribute fresh model weights quickly.

12. Is AI Coding Already Yesterday's Biggest Trend?

It’s not too off. Numerically, it’s still sound, while we’re not limited by a very expensive deployment of a GPU fleet. Those are the innovations we worked on together with Cursor to push the boundary, leading to their recent model launches. We’re very proud of them.

Harry Stebbings

Can I ask a blunt question? Cursor is an incredible customer to have, given the amazing progress they’ve had with you, and it’s wonderful to see that partnership. It’s a very large customer for you. How do you think about the concern of Cursor churn in the wake of a SpaceX acquisition?

Lin Qiao

Yeah, everyone’s concerned. The whole industry, in terms of application innovation, is driven by models, in the sense that there are a few companies that are very, very successful. They achieve escape velocity, but there are only a few of them. So that’s the shape of the whole industry, and last year, Cursor was one of the few.

I would say all model companies are concentrated on Cursor. We concentrate on the same group of app companies, and since then, it does change, right? So we do have a very healthy, diversified customer base.

Especially, I think last year was the year of coding. I think all major coding companies are on us, and this year is the year of co-work. Co-work is much more diversified by itself than coding because there’s general-purpose co-work. For example, general-purpose co-work to help you do all kinds of research.

You want to ask, “Hey, what will the NVIDIA GPU price be 2 years later?” or “What will be Anthropic’s stock price after its IPO?” Those are deep-research, general-purpose applications. There are also so many different categories of special-purpose co-work: legal—we just talked about 2 great legal companies—finance, customer support, recruiting, sales, marketing, and healthcare. So there’s a very broad set of co-work applications and innovative app companies. They’re doing really well, and we have them as our customer base.

More interestingly, we’re starting to see an uptick in consumer-facing companies looking into GenAI technology. They’re changing how they think about their traditional businesses of doing recommendations, for example. That’s very interesting to me because we’ve obviously worked at Meta, which has one of the biggest recommendation systems in the world, and we’re very eager to see how that transforms into a new economy for us.

Harry Stebbings

I’m sorry for being naive here. Do people work with just one provider in the inference space, like you, or do they work with you and together—or anyone else in the space?

I think people are more attuned to a multivendor strategy in this space because they don’t know what’s happening. It feels safe to have multiple providers to balance things out.

Lin Qiao

But we don’t view ourselves as an inference provider. Again, we view ourselves as delivering this specialized intelligence, where we help companies tune their models. To give you some numbers, today we process more than 40 trillion tokens a day. The majority of those tokens are coming from customized models, not off-the-shelf models.

Harry Stebbings

What will that token count be at the end of next year?

Lin Qiao

Anywhere ranging from 20x to 100x could be possible.

Harry Stebbings

20x to 100x.

Lin Qiao

Yeah, we’re at a very early stage of S-curve explosion right now.

Harry Stebbings

20x to 100x. If it’s 20x to 100x, the idea that we’re in a capex bubble is ridiculous, and we desperately need far more capex than we’re ever suggesting for compute. Is that right?

Lin Qiao

That is right. At the same time, I think Jensen has a 5-layered cake—a 5-layered AI cake. From top down: application, model, infrastructure, chips, and energy. We’re bottlenecked by the lower part of the AI cake in terms of the supply chain.

Harry Stebbings

Being energy—

Lin Qiao

Being energy, being chips. I think, in the physical world, how fast we can manufacture is the question because, historically, all these industries were not designed for massive scaling. Speaking about 100x scaling, no one was designed for that.

I talk with many manufacturers, and we’re bottlenecked by small parts—transistors, the smallest, tiniest parts that hold up the whole manufacturing line of servers that can be deployed to data centers and used to generate tokens.

Harry Stebbings

Do you have to be full-stack, according to Jensen’s 5-layered AI cake? Do you have to then be full-stack to win or to reduce dependencies? We’ve seen OpenAI come out with Jalapeño—terrible name. Anthropic is talking to Samsung about building its own chips. DeepSeek is building its own chips. Zuck came out with Meta building its own chips. Do you have to be all of it?

Lin Qiao

It really depends on the company’s philosophy. To us, agility is everything, and we need to earn the right to build anything. Focus is everything for us, and we want to focus on where we add the biggest amount of value based on our strengths. We’d like to leverage other people’s strengths to build on top of them.

In particular, we want to run everywhere on all possible AI chips in the world, and we don’t want to be limited by how many chips we can bring into our data center, whether we construct it or rent it.

Over time, when the business grows very big, I still remember when Meta was young. They didn’t build everything, and when they’re big, it makes sense to build. You earn the right to build your own giant infrastructure, and if it saves 5x more cost, then you’re sure to go do it, right?

At the early stage, that’s why I’ll tell you an interesting story. In the coding space, I would say Cursor was the first company that decided to work with us early on. I remember when they worked with us, they were a single-digit-million-dollar company.

Harry Stebbings

Wow.

Lin Qiao

They were very small. This was only 2 years ago. They grew by 100x to 1,000x over 2 years, something like that. But they decided to work with us early on because they recognized they only wanted to focus on product innovation and, later on, research.

They did not want to focus on platform innovation. They knew we were putting all our R&D in there, and they wanted to find the best partner to win big. So I do think that’s the right mentality: to specialize. We want to specialize. We do not want to own the whole stack. That’s not our goal as a company.

Harry Stebbings

I’m sorry to be harping on. Why does Jensen skip your layer of the cake? Because he’s doing NeMo with models.

Lin Qiao

Well, Jensen isn’t building a cloud either, right? You can say, “Hey, Jensen, you probably have all the rights to build an NVIDIA cloud.” He’s not building cloud infrastructure. I think he mentioned that as well, and many people ask him that question. He also mentioned that he wants to specialize in what they have the rights to do.

Why models? I think it’s purely a supply-chain question. If the US doesn’t have a US-native open model, it’s a problem. It’s a supply-chain problem. So he’s solely there to solve the supply-chain problem.

But if there’s no supply-chain problem because companies like us are providing this specialized intelligence platform layer, then he doesn’t need to worry about it. He just wants to make sure the whole 5 layers of the AI cake are flowing. There’s no blockage, and if there’s a blockage, he’s interested in solving those problems.

Harry Stebbings

Marc Benioff, one of your investors, I think in the new round—which obviously will come out after the round is announced—said that he spends about 3.8% of developer salaries at Salesforce on Anthropic and Claude Code.

I think it’s a useful analogy because, if you assume that that’s what’s spent on Claude Code and coding tools, that says one side of the market. But if it’s 20%, wow, we’re underestimating how big these companies can be.

When you think forward 1 or 2 years, what percentage of developer salaries do you think we’ll spend? Is it less because these tools will get cheaper, or is it more because they’ll get better and better?

Lin Qiao

I do think the cost of tokens will go down drastically because, again—

Harry Stebbings

It hasn’t so far.

Lin Qiao

It hasn’t so far because of supply-chain constraints, but we’re living in a free economy. Think about it: whenever there’s a shortage, the price is high. High prices will invite a lot of people to come solve the problem, and they’ll invite competition. Competition will bring down the cost, and eventually it will lead to a very economical solution.

Actually, that’s good for everyone because much more affordable infrastructure will invite more usage. My prediction is that infrastructure costs will go down and usage will explode because of that. The moment you don’t think about that as a problem and, if it’s a utility, you just use it.

Harry Stebbings

How much will token costs come down? Help me understand. Is it like a halving? Is it like, “Oh, it’ll be a hundredth of the cost”?

Lin Qiao

There are different ways to think about this. It’s not that all tokens are equal. I think we should establish best practices to evaluate the token economy per task because different models have different ways of spitting out tokens. Some are much more verbose than others.

You can imagine one model is 2x cheaper than the other, but it takes 2x more tokens to solve the same task, and then they’re at the same cost.

Harry Stebbings

Yeah.

Lin Qiao

Right. But overall, as model quality improves, I think being precise is going to be part of the optimization. That’s one level of optimization: to solve one task, we should need fewer tokens.

The second is how to do that for one token. You need to customize the model to solve your problem especially better and more precisely. That goes into model tuning.

Second, for each token spit out from those models and processed by those models, we also specialize in making the unit economics much better through our platform. Third is the underlying infrastructure, like the GPUs, the surrounding memory, and all these things. Today, it’s under stark supply chain constraints, but the situation will get much better. I don’t think it will probably change in the next 1 year or 1.5 years, but in the long term, 2 to 3 years, it should change, and that cost will compress. So overall, I can imagine a 10x cost reduction in the next 3 years, and this 10x cost reduction will drive 100x usage.

Harry Stebbings

You said there about token efficiency and how you enable your customers to be much more efficient. With that efficiency, you do charge more. When I did the research and compared you to competitors, I got that Together’s the price king. I don’t mean this disparagingly, but they’re cheaper. If you want cheap, you go there, and respectfully, if you want a better-quality product, you go to Fireworks. But it is more expensive. Do you think that’s a fair assessment and a fair analogy?

Lin Qiao

I think we’re probably not comparing apples to apples, in the sense that it again goes back to our business. The majority of our traffic is customized models, and we optimize for quality—number 1, always quality. Quality as in model quality toward your applications, your specific business, your use case, and so on.

The second is, when we deliver those models in inference, it’s also quality. We care about quality so much that we do extreme things. For example, during training, there’s a very hard thing to achieve called zero KLD. It’s a little bit technical; the idea here is—

Harry Stebbings

Zero KLD?

Lin Qiao

KLD is a measure of quality. What it means is that between the training system and the inference system, when the model moves over, we have bit equivalence. The numerics are fully the same; we do not lose a bit of accuracy. That’s really hard to achieve.

The reason we push that and deliver that is because we know our primary business is in model customization and inference of customized models. We want our customers to maximize every single dollar invested in training. If, across the training-inference boundary, it’s not bitwise equivalent, they just drop the quality down, and then it’s like you pay your training investment with discounted quality. Why do you do that?

So, quality first. Quality does bring additional value, and that’s why we are not interested in commoditized, one-size-fits-all. This off-the-shelf model deployed in the same way for everyone—that kind of business will always customize the model and deploy it in a unique way for your particular workload.

Harry Stebbings

Two questions. Do you have to have an FDE model to make the customized model efficient?

Lin Qiao

As a matter of fact, we do have an FDE team. It’s called the Applied Machine Learning Engineering team. Their primary job is to accelerate this customized deployment and, as a matter of fact, also build the agent to automate a lot of deployments.

13. Hypergrowth vs Profit: Why Margins Can Wait

Harry Stebbings

Given where we are in the stack and the amount of complexity that we have, we have a margin structure that’s a little bit different from traditional SaaS being 80%. I don’t know the margins precisely here, but they’re traditionally in the 30% to 40% range for where we are. Is that the new normal for where we are?

Lin Qiao

I don’t think that’s the new normal. I think that, at least for us—I don’t know about other companies—it’s a reflection of the fact that we are in a hypergrowth phase.

During a hypergrowth phase, you have the choice. To me, margin optimization is a constraint problem: We want to go to 70% margin, we want to go to 80% margin, and then we’re going to go backward and impose those constraints to guarantee those margins. Usually, constraints slow down innovation.

To give you an example, during system development and in a high-velocity, system-expanding phase, we don’t want to overbuild because we’re in a high-experimentation phase. We’re testing what will stay and what will not stay. Optimization doesn’t make any sense. Once we know this is a system that we want to build 100%, and we’re going to scale this 1,000 times bigger, then we go optimize the heck out of it. I think you think about business the same way.

In hypergrowth, if our focus is only on optimizing gross margin, we absolutely can do that, but we’re sacrificing the speed of growth as well. We want to go everywhere. We want to go into different geographies, tackle different use cases, constantly create different product lines, and those are not the times for optimization. That’s my opinion.

Harry Stebbings

So we will be able to increase margin without moving into different layers of the stack?

Lin Qiao

We absolutely are not going to move into the application layer. That’s very clear to us. Whether we will move down into, for example, building data centers and so on, that could always be on the table, but the question is timing.

Harry Stebbings

Isn’t the statement, “You either die, or you live long enough to build your own data centers”? Elon and Zuck are now spending, I think, $10 billion on the latest data center in Canada. Would you like to build data centers?

Lin Qiao

I have built data centers that matter, and there’s also lots of innovation possible there. There’s no one-size-fits-all as well, and building a GPU-native data center is also interesting. I think there is a potential direction of building one.

It’s a trade-off, right? From an operations point of view, it’s much better to build a homogeneous deployment. It’s all the same chips, all the same SKU, as big as possible, and you run multiple workloads so it’s fungible. It’s very easy to manage. You build 1 principle, 1 process to do maintenance operations.

But again, it goes to optimization. Once it’s so big, then any optimization is going to drive a lot of economic return. For example, we’re talking about NVIDIA recently acquiring a company also called Groq, with a Q. It’s a large SRAM-based ASIC accelerator.

Harry Stebbings

I spoke to Jonathan before this show. Jonathan is excellent.

Lin Qiao

He said, “What a fan he is of yours.”

Harry Stebbings

Oh, also a fan of his.

Lin Qiao

It’s a great combination between a FLOPs-intensive GPU and an SRAM-intensive ASIC. FLOPs-intensive is really good for the first half of LLM processing—it’s called prefill processing, processing the prompts and so on—and SRAM-intensive is really good for generation. That’s just the nature of the model architecture.

It’s great to combine these 2 instead of running homogeneously on the same chip, right? But that requires a very unique system design and deployment into a data center. It is heterogeneous, actually. I really mean homogeneous design is much better for operation. This is heterogeneous, and operating this heterogeneous design requires unique innovation in data center deployment and so on.

So data centers aren’t commoditized. You can specialize in data center deployment, and one data center is better than another. Data center deployment can be done well and badly. Data centers are so complicated, right? If you think about it, from the beginning all the way through construction and power deployment, you need to have the right power come in, the fiber channels, the right cooling—especially with new chips, which require liquid cooling—to get all this right, and the parts can fall apart and you need to know how to replace them. It is all very deep expertise. It’s no joke. It’s not that tomorrow I can be a data center operator. I cannot.

14. Can the West Keep Up With China's Infrastructure Speed?

Harry Stebbings

Is that not where you would bet long on China, with the greatest of respect? Especially in the US, one of the biggest barriers to data center deployment is policy and local legal infrastructure that prevents it. In China, you don’t have any of that, and data center deployment is much, much faster.

Lin Qiao

I think, in general, infrastructure—the physical infrastructure construction in China—is going really fast. I literally see some kind of crossover bridge being built within a week. The velocity is very, very high there.

There’s a highway close to my home that, after 1 year, is not done yet. So this is also a crossover. I do think there’s a unique strength, probably because of the population density, and they are specializing in those kinds of construction-related work.

But I do think here we also have those specialty people. It’s just, I heard even electricians are under severe shortage.

Harry Stebbings

Yeah.

Lin Qiao

We are under global supply chain constraints here.

Harry Stebbings

What change would moving into the data center layer cause to margins? Would that take it from 30% to 50%? Would it be not that meaningful? What would that change do to margins?

15. Why AI Hardware Depreciates Faster Than Ever

Lin Qiao

How we calculate gross margin is interesting these days because how long hardware depreciates has changed significantly.

Harry Stebbings

Yeah.

Lin Qiao

In the past, it was 6 years.

Harry Stebbings

Yeah.

Lin Qiao

A solid 6 years. Hardware release is usually 3 years, and that’s fast. Now, within 1 year, from one vendor alone, we have 3 SKUs. The newer model usually runs best on the newest hardware model, and depreciation is also very fast.

Every week, we’re launching a new model, and then the model is at its peak in value before the next model comes out. Imagine this cadence after 2 years. Which model runs on 2-year-old hardware? It’ll be a 2-year-old model. Are those models still valuable?

I think that’s the real dynamic we’re facing right now. The hardware will last for 6 years still, but—

Harry Stebbings

But what you’re saying is that the speed of model development far outstrips the speed of chip and hardware depreciation.

Lin Qiao

The speed of model development is definitely the fastest, but even the hardware innovation itself is the fastest.

So after 3 years, if every year there are 3 hardware SKUs, after 3 years there are 9 hardware SKUs in between. Do you still want to go back to 9-generation-older hardware running a 3-year-old model? That's questionable.

Maybe there's a world where it's still valuable, but with this pace of innovation, it's questionable. Now, with a different depreciation cycle, it changes the dynamics of build versus own, build versus buy. Again, it goes back to my original thesis: do you optimize for growth, or do you optimize for gross margin? It's all about timing.

Harry Stebbings

How do you think about that question for yourself when you're sitting there in an armchair on a Sunday afternoon thinking, "We're optimizing for growth. Now, when is that time to optimize for gross margin?"

Lin Qiao

I would say we want to optimize for both. [Laughter] Here's how I think about it. Optimizing for growth requires a lot of business planning, assuming there's product-market fit. Optimizing for gross margin is optimizing for differentiation.

I think I want to avoid overoptimizing for gross margin, but we should optimize for gross margin continuously. In other words, we should optimize for product differentiation continuously. There's no question about it.

We want to continuously optimize toward a healthy gross margin that allows us to grow really fast. It's a trade-off, and we don't want to make compromises. The compromise would be overoptimizing for gross margin, resulting in very slow growth.

One possible way to optimize gross margin is not to grow at all. We just optimize the heck out of it. I know we can climb to a high number, but that's an absolute disaster outcome.

16. Why AI Will Create More Jobs, Not Fewer

Harry Stebbings

Okay, interesting. If we just said, "Hey, for gross margin, we're going to take it from 30% to 10%," is it a winner-take-all market where we could eat up everyone else's lunch and then optimize gross margin later?

Lin Qiao

I think a winner is probably not a snapshot in time. It's going to be a long-term situation. We do see a particular industry oscillate and start to settle with a few good ones.

Take legal, for example. I was at a dinner table, and interestingly, it seemed like there were a lot of those companies around 2 years ago, but now it's pretty much 2.

I think it's a long, long game.

Harry Stebbings

How do you see the more mature state of your market? Is it like a cloud market, where you have Azure, AWS, and GCP, or is it an Uber and Lyft situation, where one takes 90% and the others fight for scraps?

Lin Qiao

We're in the adoption curve where a lot more companies in the AI space are starting to seriously think about moving to specialized intelligence and starting to seriously think about whether owning their intelligence is better than renting.

Going back to this optimization question—when is the good timing?—it's the same question we're answering for ourselves in build versus buy. Our customers are also thinking about build versus buy, or build versus rent, or own versus rent.

I think the AI journey, or AI adoption journey, has gone further along. A lot of companies have meaningful traffic. A lot of companies are deploying AI into production, and a lot of companies are at the phase of scaling.

That's where optimization kicks in. When optimization kicks in, you need to have control to optimize. If you don't have control, you just don't have the range to optimize.

For you to have control, you have to build on top of some open model. You have to turn your data into your intelligence. That's pretty much the path we've seen so many companies across industries take. They reach the same conclusion: they're moving in this direction.

Harry Stebbings

Speaking of owning your own intelligence versus renting it, that does apply to a national layer. We've seen Fable be banned in some cases by the administration, briefly for 19 days.

Especially in Europe, we suddenly went, "Oh my gosh, we cannot be at the hands of OpenAI and Anthropic, where we can just be banned. Our health services could sit on the infrastructure of something that an administration can turn off." Do we see a future of sovereign models, where large nations or nation blocs own sovereign models?

Lin Qiao

I definitely see that possibility. I also see that, if we think about the general-intelligence model as the electricity layer—as a power line—every country should own its own power line.

I think that is a very scary moment: my power line is going to be cut off, and all my fundamental day-to-day needs are not going to work. I feel so frustrated whenever there's a power outage in my home. I feel so frustrated when I cannot access my Wi-Fi. I feel so anxious. [Laughter]

17. The Biggest Mistakes AI Founders Are Making

Obviously, operating a country is extremely important when it's built on top of this fundamental baseline. For every single company, it's the same thing. It's not just about whether a country should have its unique sovereign independence; every single company should have its independence.

You don't want any single person to cut you off. That's an extremely scary moment.

Harry Stebbings

Why would you move into the data center space, but you wouldn't move into the chip space?

Lin Qiao

Because I know building a chip is extremely hard.

Harry Stebbings

I thought so, too. Okay, again, I admit to being an outsider, which is why I think the show is a little bit successful. I thought so, too. But then how come everyone is seemingly doing it as if it's just another product? As I said, OpenAI, Anthropic, DeepSeek, and Meta are building their own chips now.

Lin Qiao

I think Meta has been building its chips for more than 5 years—way more than 5 years. MTIA has been a project since 2018, maybe earlier.

Meta has been investing in AI for a long time, pre-GenAI, and it has a huge AI workload focused on ranking and recommendation. Meta has been building other hardware as well in the past.

Whenever the usage has passed a certain threshold, it makes economic sense for you to build the underlying supply. You can specialize toward your workload, and that's another form of specialization: specializing by baking your logic into hardware.

This hardware is purpose-built for your particular workload, and you better make sure this workload doesn't change, because it's really hard once the hardware is taped out. It's possible to change it, but it's very costly.

Once your workload has stabilized and your business has stabilized and doesn't change too often, then that's the time to consider building a chip. I still see the whole AI world, especially model customization, as very dynamic—very, very dynamic. Workload patterns are very dynamic.

Think about how much energy there is in the application space. People are experimenting with all kinds of things. You don't know which one is going to take off, and they will just take off quickly. Once they take off, which one is going to sustain? A few will sustain, and then that's the time when we say, "Now we know this is the pattern, and now we should probably encode this pattern into hardware," bring this hardware into a data center, and so on.

It's all cascading, and then it's going to cascade down to me. It's a fundamental question: where are we in terms of workload maturity? We're still in the early stage of workload maturity to warrant a chip that will be durable.

Now you go back to the fact that we have so many accelerators. They are successful; some are really successful. But remember, those companies started before GenAI. They started with something to optimize some workload, and then they pivoted to AI and tried to fit the AI workload.

It's almost like you bet before this AI workload emerged, and now it becomes a serendipity question. Are you lucky enough that this just works? Some really worked.

Some fundamental designs, like putting a lot of SRAM on the chip, are great for AI models because they are memory-hungry. This really accelerates the execution of inference, and so on. Those worked, and some didn't work.

18. Do AI Startups Need to Build Their Own Models?

Harry Stebbings

What do you see as the greatest bottleneck today? I think it was when I had Jonathan from Groq on the show, who said HBM was the greatest bottleneck, and that's why you've seen a 5x increase in price. What do you see as the greatest bottleneck that people don't talk about enough?

Lin Qiao

I still think we don't have a great system for very large models. I really believe the fundamental lower-level infrastructure cost will go down. For solving tasks, we should need fewer tokens. That will increase. Collectively, the cost will significantly reduce.

Therefore, we can run the highest-intelligence models much more ubiquitously in the future, but we don't have a system designed for that. For example, we don't have a great system designed for 10-trillion-parameter models today.

That will require very smart engineering and code design, from the model to the customization and serving-platform layer, all the way to the chip layer. The chip is not an individual chip, but a system—a collection of chips in a system—and all of it as a total package.

I think there's still a lot of innovation we can do.

Harry Stebbings

I think recently you announced that you were at $800 million in AR. Incredible feat, and you've scaled so fast. What is that by the end of this year?

Lin Qiao

We think we can at least double it.

Harry Stebbings

By the end of the year. Wow. You know what's so interesting for me as a venture investor? I've been investing for 10 years. We used to be in the days when Slack was the golden child, where going from $1 million to $10 million in revenue in 18 months was amazing.

Now we have companies like Fireworks, where you scale to $800 million in revenue in a matter of years. You mentioned Cursor scaling to billions in revenue in a matter of years.

The speed of company revenue growth is just unparalleled.

Lin Qiao

I think it's because there's a fundamental disruption in this technology that is all-empowering. And all-empowering in the sense that it reaches out to every individual one of us to be creative, and it unleashes a lot of creativity that we just don't have access to. That's why we're seeing this phenomenon of extremely fast growth: because of the demand.

Harry Stebbings

Final one before we do a quick fire. You hired George Hu, who was president of Salesforce. He's exceptional. He's one of the most direct, no-BS operators I've ever met. But you met him a couple of years before, or a year before, and you were like, “Oh, we're not ready for you yet.” Why did you say that, and why did you decide now was the time?

Lin Qiao

Right. So, a year ago, I think we were probably just 50 people. Today, we're at 200 people. We're still not that big.

Harry Stebbings

Wow. You're 4x ahead.

Lin Qiao

Yeah. So, at 50 people, I was more thinking about scaling the product first, then getting to massively scale the business. I have huge respect for him. I know he's a legend. He's legendary. He's a legendary operator in Silicon Valley.

I just felt like we were too small for him. I told him, “Hey, we're probably too small for you, but I would like to work with you at some capacity.” So, he helped me actually build up the team and interview a lot of executives. His feedback is always well-balanced and very thoughtful, and we started to work together in that capacity.

I think by the end of last year, we were growing really fast, and we started talking seriously. That early relationship paid off. He's really cool. He's really cool in the sense that—

Harry Stebbings

He's so cool.

19. Why Great Leaders Stay Close to the Work

Lin Qiao

He did a lot of things, with great accomplishments in the past, but I find a unique characteristic about him: he's extremely experienced, has high aptitude and business vision, but he's also very curious. He doesn't make assumptions—“I know it all. I've seen all the movies; it's the same movie, so let me just direct this as I did in the past.” He didn't come with that attitude.

He knows AI goes at an insanely fast pace, and he's learning along the way, but he also fully embraces AI. His team—our GTM team—is using all kinds of AI agents. They're sharing skills, so they maximize their productivity.

He knows we have a superlinear demand curve, and there's just a certain pace at which we can build our GTM team. In order for us to catch this curve, we need to build a team, but the team also needs to have increasing productivity to match it. That's a problem he's solving, and I feel very fortunate to work with him.

In general, I feel that in the AI space, the unique part is that people need to have very special traits, almost like contradictory characteristics. For example, they need to be very experienced but super curious, with a fast learning curve. Or Dima, who we talked about a little bit earlier: he is brilliant, with high intellectual horsepower, but extremely humble. It's a weird combination.

Harry Stebbings

He's amazing too.

Lin Qiao

He's almost cynical in an Eastern European way, but also, at the same time, very humble.

Harry Stebbings

Can I do a quick-fire round?

Lin Qiao

Okay, let's do it.

Harry Stebbings

Okay, what have you changed your mind on most in the last 12 months?

Lin Qiao

I think it's how fast we grow. I changed my mind because I had been quite worried about having too big a team too early. That's why, when I met George, I told him we were too small for him, because I didn't intend to grow very fast in terms of people. I worried about getting slowed down and losing our agility and velocity very deeply.

Since then, we've been very aggressively using our tools. We've developed our own unique way of hiring certain types of people who we know will be charging forward with high velocity, an extreme sense of ownership, strong communication, and who never take no for an answer. We also learned how to get those people. Now I feel much more comfortable scaling really fast.

Harry Stebbings

What's your type of people? I know that sounds weird, but our type of people is actually really specific. We pretty much only hire immigrants. British people don't work very hard—sorry. They're very scientific and rigorous, use data for most things. I actually think creativity often comes from data and is informed by data, and they're unwaveringly accountable and ownership-oriented. Nothing is anyone else's fault; it's all my fault, even if it's someone else's fault. That's a 20VC person. What would you say yours is?

Lin Qiao

It's not, in a weird way, competence. It's weird: we want people with high confidence. But more importantly, the strong indicator of whether they will do well in this wave, especially at Fireworks, is whether they are really built for taking extreme ownership.

Extreme ownership means we're not putting anybody in any boxes; we're just stacking the boxes together into a tower. We want people to automatically claim, “Hey, this is an end-to-end problem. I'm going to see through the whole thing and work with a bunch of people to make it happen, and I'm going to deliver it no matter what.”

Those kinds of people have the highest, longest mileage, and their growth curve is amazing.

Harry Stebbings

What's your biggest lesson from working with Jensen Huang on what makes him so special?

Lin Qiao

He's everywhere. I seriously think he has a clone—like hundreds of Jensens somehow plugged in. [Gasps and laughter.] For example, I send him an email, and he'll reply in 1 minute. I just don't understand how he's constantly in the details.

But now I've operated a company for 4 years, and I understand why he's doing that. That defines velocity, because what is leadership? Leadership is just judgment. It's not privilege; it's judgment. You basically have the context. You need to have the right context to make the right judgment for the team.

Especially in a high-velocity space, if you don't know what's happening, what works, what doesn't work, or what the gaps are, you make the wrong call. In a slow-moving space, you can wait for cascading information up and down and make those calls. But in a fast-moving space, you just cannot wait, because there's guaranteed to be information loss in transition layers. People are people; it always happens. Not knowing exactly what is happening and having the position to make judgments makes for bad leadership.

He's demonstrated that through his own example. Even before this crazy AI thing, he's been operating that way. Before, I admired him for his sheer amount of capability in doing that. Now I understand the wisdom behind it, because I also operate that way. I need to know what's happening on the ground to make the judgment for the company.

Harry Stebbings

What did you wait on in the Fireworks journey that you wish you hadn't waited on?

Lin Qiao

Marketing. [Laughter.] We talk about it. At the very beginning of our journey, we didn't discuss it, but we felt the product spoke for itself. At the end, the product stands, and we wanted to devote all our effort and focus to building the product, working with customers, validating product-market fit, and going from there.

We didn't spend much time marketing at all. We didn't prioritize educating our customers about the right direction to think about the trend and the value. But now I do think it's important. Marketing is not about fluff; it's more about education, more about clarity, and we're working on that.

Harry Stebbings

What area of AI is underinvested in today, in your mind? You mentioned cooling or servers. What else is underinvested in?

Lin Qiao

I think AI has a sexy part because it's such an innovative, creative technology, and building something on top of it is the focus. But monitoring the ROI—I think the industry is starting to pay attention to it. Eventually, that's what matters: it's not how much you spend; it's what the return is, what the cost is, and what the attribution is.

I think in the next couple of years, as AI gets more and more into production, there will be a lot of focus on getting that clarity and getting that discipline out. Token maxing is just a thing in time, I think, but we'll quickly move into ROI maxing, which is all about running a business.

Harry Stebbings

What large customer do you not have that you would most like to have?

Lin Qiao

We haven't spent too much time in the traditional enterprise segment. I think that's just because we were very small. Now, as we build out our company, I do think that, even without us investing, we have customers like GEICO, Capital One, Mercury Insurance, and RBI. All these companies, even without us pursuing traditional enterprise, come to us and are customers. I do think that's a very big market.

Harry Stebbings

What has to happen before the end of the year that hasn't happened for you to consider it a good year?

Lin Qiao

I'm confident in our capability to drive the business. To me, this is a year when I want to prove we can scale quickly while keeping the same velocity, and that's very important to me. If we reach that proof point, next year I'll have a lot more confidence to continue to scale extremely aggressively. I want to make sure we do it right this year.

Harry Stebbings

Final one for you. What does no one see about the next 3 years that you see very clearly happening or not happening?

Lin Qiao

I really see that every single company will own its own intelligence as a must-have. It's not optional. That's a trend I'm seeing, because there's an analogy to software: there's a reason why every company builds its own software stack.

There’s no standardized software you can just use off the shelf to solve their problem, because every single company is solving a unique problem, and they want to build software because they want to have full control.

Obviously, they will pick and choose which part of the stack they want to build themselves and which part of the stack is common knowledge—there’s no point in building that. But every single company owns their own software stack. Obviously, we’re talking about this in the SaaS era, right? So, same, I think, in the AI era, every single company should own their own intelligence.

Harry Stebbings

Lin, you know, it was Matt who introduced us first. I’ve had the joy of getting to know you and, obviously, George. I can’t thank you enough for joining me and for coming in person. It is so wonderful to do it in person, and you’ve been fantastic.

Lin Qiao

That’s an amazing studio. You asked a lot of interesting questions. I had a lot of fun talking with you.

Harry Stebbings

We do a lot of research beforehand, huh?

Lin Qiao

Yes, you did.

Harry Stebbings

Thank you so much for that, and you’re fantastic.

The Open-Source AI Reality | How Token Costs Will Fall 10X & Usage Will Explode 100X | Lin Qiao | BidClub