Baseten CEO Tuhin Srivastava谈定制模型与构建推理云
Baseten 30倍的增长反映了AI原生应用的爆发,但更大的企业推理市场仍基本缺席。 Sarah Guo表示,Baseten预计今年营收将超过10亿美元,Tuhin Srivastava对此予以确认;他估算,应用公司仍占推理调用量的99%,因此“市场的大多数还没有上线”。
应用层真正持久的护城河是专有工作流反馈,而不只是模型权重的获取权限。 Abridge的临床医生修改记录和后续EMR操作形成了前沿模型公司可能无法获得的奖励信号;基于这一信号进行后训练,有望打造专用的长时程智能体。“只有他们自己能收集到的用户信号”,才是保护应用层的关键。
生产推理如今已高度定制化,模型部署也因此与日益持续的后训练闭环相连。 Baseten超过95%的token通过专用推理运行,几乎每个客户都会为质量、性能或两者同时改造模型:“没人只是直接运行原版开源权重。” 顺序很重要——“在产品市场匹配之前不要做后训练”——因为公司应先用最好的模型验证价值,再将其专门化,做到“更好、更快、更便宜”。
中国开源模型具有经济战略意义,但Guo指出,定义绝对能力前沿的仍是美国闭源实验室。 Srivastava称自己“可能会判断错”,但表示如果这些模型被网络边界隔离,就不会“神奇地”跨越网络边界;他认为DeepSeek的生产成本可能只有Anthropic的20%,延迟相当甚至更低,可靠性可能还更好。他的立场是:美国开源模型既有必要,也不可避免,但忽视今天已经可用的智能,就会“只见树木、不见森林”。
算力稀缺如今已直接传导至合同期限、营运资金以及可能的上市时点。 Baseten在18朵云上运行90个集群,利用率高得令人不适——接近但通常还没到95%区间——并每天下午4点开会分配供给。要从一家可信云厂商获得1,024块B200,可能需要签下3至5年承诺,并预付TCV的约20%;当被问及这是否意味着公司应更早上市时,Srivastava回答:“早点上市。”
Baseten的护城河逻辑是软件叠加稀缺算力,而不是单纯出租GPU。 Srivastava称GPU即服务是商品化业务,而推理软件让Baseten前30大客户实现零流失,年度NDR约为400%:“如果算力全在我们手里,祝你推理跑得起来。” 他预计专用芯片会出现,但NVIDIA的供应链、CUDA和生态意味着基础设施运营商“今天用NVIDIA能跑得最快”。
推理效率呈现出类似杰文斯悖论的特征:单位成本下降带来更长的智能体和更多总认知,而不是需求饱和。 成本下降后,开发者会加入“多得多的智能”,因为更好的答案能改善体验并带来更多收入。因此,Srivastava称推理是“最后一个市场”——即使实现AGI,推理仍然存在;他认为消费者将迎来“万物皆管家”,而那些不加入智能的工作流公司则会遭遇“灭绝时刻”。
1. 工作流数据让应用层保持生命力
Guo开场提到Baseten在12个月内增长了30倍,并表示公司预计今年营收将超过10亿美元;Srivastava确认了这一说法。他对整个市场的解释是:开放权重跨过了能力“鸿沟”,RL和后训练走向主流,应用开始越来越多地将智能内置于自身。Baseten“在一定程度上押注于”这轮扩张。
Srivastava对应用层的判断是,专有模型权重本身不是护城河,真正稀缺的是嵌入工作流中的用户信号。Abridge是一款供医生使用的环境式医疗记录工具,按他的描述已进入美国几乎所有医院;它捕捉医生修改记录以及EMR中的后续操作,形成前沿模型公司可能无法获得的反馈。这些奖励信号最终可能用于后训练专用的长时程智能体。
客户支持领域也遵循同一逻辑:一张进来的工单可能触发“1个、2个、10个、20个动作”,而不是只生成一次性答案。真正有差异化的资产,是这一连串动作和决策,它让应用公司拥有明确的优化对象。
按推理调用量计算,Srivastava认为今天的市场中99%来自AI原生应用,而不是企业。市场已经从AI工具走向闭源模型API,下一步是定制模型普及。“Stripe式演进”让Baseten能够跟随前沿客户扩张,而Abridge、OpenEvidence等公司则把客户对数据留存、模型、部署、GPU、延迟和透明度的要求转化为产品需求。
2. 能力优先于来源,生产环境中的权重都是定制的
Guo描述了市场从Mistral、Llama转向中国来源模型的变化;Srivastava表示,客户会先从能力而非成本出发,因为能力能够释放经济价值,之后才进行优化。Baseten的客户覆盖GPT-4o、Moonshot、DeepSeek以及Canopy的Orpheus文本转语音模型:他们使用“任何处于前沿的模型”。
Baseten将业务分为专用推理、共享推理和训练,但超过95%的token都在专用推理中运行。几乎所有客户都会用自己的数据改造模型,通常还会通过编译或其他方式同时优化质量和性能。“没人只是直接运行原版开源权重。”
谈到中国模型的风险时,Srivastava表示“我可能会判断错”,但称如果对这些模型实施网络边界隔离,它们不会“神奇地”跨越边界;除了少数早期案例被迅速发现外,他没有看到其中嵌入明确议程或偏见的真实证据。不过,他仍称美国开源模型既有必要,也不可避免。Guo指出,中国补贴实际上把剩余价值转移给了美国企业;Srivastava则表示,DeepSeek的运行成本可能只有Anthropic的20%,延迟相当甚至更低,可靠性可能还更好。Guo另行强调,绝对能力前沿仍掌握在闭源模型手中。
3. 推理与后训练构成不断复利的闭环
被收购的研究团队来自Parea。Parea既是Baseten的客户,也一直在进行模型后训练并通过Baseten部署模型。Parea意识到,自己最终需要成为一家推理公司;Baseten则意识到,研究能力可以帮助它更早触达客户,并加速满足市场对后训练软件和实操型专业能力的双重需求。
Srivastava给出的生命周期建议首先是克制:在优化任何东西之前,先用现有最好的模型证明价值——“在产品市场匹配之前不要做后训练”。一旦公司拥有独特的用户信号,专门化就能让工作负载“更好、更快、更便宜”;例如,客服模型不需要保留前沿编程能力。
训练选择会影响模型在推理阶段应如何量化,而部署后的推理又会为下一轮后训练提供数据、评测结果和奖励函数。Baseten希望通过训练API,让持续学习“在某种程度上成为已经解决的问题”,并由Braintrust负责评测、合作伙伴负责沙箱。它的产品逻辑是“向下打通以释放供给、创造利润”,同时“向上延伸以释放价值”。
4. 算力紧张把推理变成融资业务
Baseten在18朵云上运营90个集群,利用率处于“令人不适的高位”——大多数时候还没到95%区间,但已经接近。公司最初搭建统一运行时层,是为了抽象不同云之间的可靠性、延迟和故障切换;如今可以在半天或更短时间内让另一个国家的供应商上线。公司仍每天举行下午4点的算力分配会议。当被问及什么让他夜不能寐时,Srivastava立即回答:“算力。”
名义上的GPU可用量高估了真正有用的供给,因为许多供应商从未运营过数据中心,也不了解推理SLA。Srivastava称市场中有一部分“有点像骗钱”:可能只有约12朵云真正合格,其中只有3至4朵达到他所谓的“黄金层级”,导致市场既受供给约束,又在供应商和运营两端同时吃紧。
长期供给可以买到,但市场变化太快,意味着每家公司都在下注:已经服役4年半的H100价格仍在上涨,可能拥有9年的有效寿命。如今,要从一家优质云厂商获得1,024块B200,不可能签不到3至5年合同,而且预付比例可能达到TCV的20%。能否匹配供需并以低成本融资至关重要;谈到IPO时,Srivastava的回答是:“早点上市。”
GPU租赁本身“没有粘性”,客户把它当作商品。捆绑软件的推理业务则不同:Srivastava表示,Baseten前30大客户无一流失,年度NDR约为400%。他还强调,买家需要有足够需求来消化买下的算力,同时拥有较低的资本成本。
5. 近期开局仍由NVIDIA领先,运行时则持续碎片化
Srivastava预计会出现专用于推理和解码的芯片,但认为NVIDIA的供应链、CUDA和开发者生态将在未来数年构成强大优势:基础设施公司“今天用NVIDIA能跑得最快”。竞争者需要建立生态;如果某家实验室锁定了一款芯片90%的供给,它就有动力把锁定比例提高到95%,确保所有东西都围绕自己构建,其他人无法使用。
工作负载的变化决定运行时路线图:扩散Transformer、编程智能体沙箱、推测式推理、感知KV cache的路由,以及将prefill和decode作为两个独立问题处理。Baseten称,拆分这两类任务带来了“巨大的收益”;将异步批量推理作为一等能力,则能提高Baseten及其客户的利用率。
规模扩大后,最先暴露的往往是普通系统边缘问题,而非模型层面的奇异故障:Baseten第一次遇到kernel panic,是因为2个Fluent Bit worker同时向同一节点产生了过多日志。Srivastava认为,LLM运行时和当前对KV cache的利用仍不成熟,但目前观察到的大多数边缘案例仍属于系统和内核层面;下一代基础原语需要同时提升规模、安全性和性能。
6. 更便宜的推理扩大市场,也抬高运营门槛
Baseten在大约8至18个月前仍基本停滞,直到Srivastava放弃了工程师常有的判断:管理者只是额外开销。他现在的检验标准是,一名高管能否独立负责一个完整问题;创始人长期事无巨细地介入,往往意味着“你可能只是没有找到合适的人”。
招聘标准强调第一性原理思考,而不是过去是否做过同一份工作;要求把工作放在高优先级,同时坚持善意、协作、低自我,以及“拒绝英雄文化”。Srivastava认为,这种明确性让匹配者和不匹配者都更容易被识别,尽管公司快速扩张,仍将不必要的人员流失控制在较低水平。
运营文化则没有那么多讨价还价的空间:一次45分钟的会议中,多名AWS高管的寻呼机反复响起。在Baseten,所有人都要值班,寻呼机被戏称为办公室警报器。Srivastava的联合创始人意识到这套规范已经传到家里,是因为他7岁的孩子听到寻呼机后问:“这是P0吗?”
谈到杰文斯悖论时,Srivastava认为不存在“这个答案已经够好了”的上限:推理成本下降后,开发者会加入“多得多的智能”,智能体运行时间更长,而更好的答案会带来更好的体验和更多收入。他设想的终局是“万物皆管家”——提供个性化照护、教育和工具,软件会更多而非工程师更少;对于那些无法加入智能的工作流公司,则会迎来“灭绝时刻”。
Hi, listeners. Today, Elad and I are here with Tuhin Srivastava, the founder and CEO of Baseten, the AI inference cloud. We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open-source and perhaps multichip future, and what 30× scale in a year looks like. Tuhin, welcome back.
Hi.
Good to see you.
Thanks for having me.
All right, you are in one of the craziest markets: AI inference. It's very important, and there's a lot going on. You guys have grown 30× over the last year, and I think I can say you're expecting to do more than $1 billion in revenue this year.
Mm-hmm.
What's going on? Tell us about scale.
Yeah, it's been nuts. I think what's happened over the last 24 months—and this keeps getting bigger and bigger—is that everyone is realizing that you can put AI everywhere. You have all these great options available, from closed-source to open-source models. The open-source models have crossed some sort of chasm in terms of their baseline capability, and then RL techniques and post-training for specialized models have become mainstream enough. There are enough examples of it working.
Customers are realizing they can own their inference more and more. What that's meant for us is more of the long-tail models coming through, customers in-housing a lot of that intelligence themselves, and the application layer just getting bigger and bigger. As that grows, we are just somewhat indexed on that, and we've been around to be able to collect the demand.
There's an existential question in here that I think everybody is continually asking: Does the independent application layer get to exist at all, versus the labs? You have to believe this. Why do you believe it?
Yeah, look, I think it would be a sad thing if it didn't exist in general, and that's my—sadness is fine.
Sad all the time.
Yeah, sadness is fine. But that's not the reason why I think the application layer will exist. I think the application layer will exist for a number of reasons.
One is this idea that what is valuable to a company is the user signal that they can gather, which only they can gather. To the extent that signal is encoded in a model, I think a lot of their business will be at risk. But to the extent that it is encoded in workflows, that is where they will be able to develop a moat.
A good example of that is a company like Abridge, where the edits clinicians make to the notes, and what they do with the notes after the fact in the thing that happens inside the EMR three steps down, becomes a workflow that only—
Can you explain what Abridge does?
Sorry. Abridge is an ambient scribe that is used by physicians in almost all hospitals in the U.S. I think they're a lot to invest in. Shiv Rao is amazing. Great company, great team, great product. They've basically got this very deep integration into hospitals and clinician workflows.
My argument here is that it's very hard for a frontier model company to eat that because they just don't have access to that user signal. What will happen over time is that folks who have access to that user signal can start to post-train models on that reward signal and start to get long-horizon agentic models running.
To the extent that is possible, and that signal is differentiated, unique, and somewhat rare to get access to, there will be an application layer. Support companies are another example of that. A support task isn't one-shotted. Usually, at a company like Basecamp, when a ticket comes in, there are one, two, ten, or twenty actions that get taken. That is where someone can develop a specialized model.
So there are almost two versions of this. There's the new companies, like Abridge, Decagon, and some of these other things that you mentioned, that are doing these new types of applications using AI and selling them to customers. The other is enterprises building things in-house or building their own models.
What proportion of the market today do you think is these new application companies—AI natives, the fast-growing companies, some of which are at considerable scale now, like Abridge, Cursor—
OpenEvidence.
OpenEvidence—those types of companies? What do they teach you? What does that push the company to do? How do you think about serving them versus evolving for the enterprise?
I think you asked me the same question two years ago. It's crazy that the answer is still the same. If you look by inference count, it would be 99% the former. That represents the scope of the opportunity here: the majority of the market hasn't come online and added AI into this market.
Yeah, this is just so much still to come, and people are underestimating that, I think.
100%, and what's cool is that we're seeing the transition happen. Before, it was like, "Hey, are they using AI tools?" I don't think that was immediately obvious two years ago. I think that's obvious now: yes, they are. Are they using closed-source model APIs? I think they're starting to get there. And then once you do that and see what is possible, then comes the whole custom model adoption. I think that is all that is ahead of us today.
Yeah. I think, firstly, you just learn a lot by building with the companies at the greatest scale and doing the most interesting things. We think about it in two ways.
The most obvious way is to build for the highest scale. The customers that push you the most technologically will push everything else into place. I think the evolution of Stripe as a company showed that. Stripe now serves so many enterprises, but 12 years ago that wasn't the case. They just built for the frontier and went with them.
The second way we think about this is to build for companies that are serving enterprises. We don't serve the enterprise, but our customers serve enterprises. Abridge serves enterprises, OpenEvidence, Decagon, Writer, Gamma—all these companies serve enterprises at scale.
What we actually get is a translation of the requirements from them. They're saying, "We need this sort of data retention. We need these types of models deployed. These are the types of GPUs or latencies we're okay with. These are the model requirements, from a transparency perspective, that we care about."
I think that is actually the more nuanced answer. If you listen to what their needs are, we get a full translation of what the enterprise would require. I would say that by serving companies like Abridge and OpenEvidence, we're probably pretty well suited to go serve the healthcare system, given that they are selling into healthcare and selling to those customers.
How much of a shift are you seeing in terms of the types of open-source models that are being used? Two or three years ago, I think the main thing was Mistral and then a few other things. Then Meta came along with Llama, and it really shifted in terms of the best-performing models. Now, the best-performing models are of Chinese origin in different ways. Do you see that mix reflected in what's being used by our customers?
Yeah. The customers we are serving—these are the fastest-growing AI companies in the world—are very forward-thinking. They want to use the best models, and they are optimizing.
There is a subset of tasks, which I think is small today, where people really start with cost. But everyone comes for capability first, because that's really where economic growth is being unlocked and where value is being delivered. Then they optimize.
We've seen customers use everything from GPT-4o to Moonshot models to DeepSeek to Canopy's Orpheus, which is a really good text-to-speech model. Customers generally want to use whatever is at the frontier.
The difference is that we have a lot more visibility into how to run these models and how to run them really well. Secondly, they're good now.
There have been a number of concerns raised about the use of Chinese models, in particular security concerns—whether there's something embedded in the models, Trojan horses, or other things.
First, do you think there's any real concern there? Second, people often talk about how there should be U.S. counterweights to this. From a geopolitical perspective, do you think that's legitimate—something we should be worried about? How do you think about the origins of these models versus their uses?
Yeah, look, I think these models, firstly, are fantastic. They're amazing. We work with these teams. They're truly awesome.
I'd say, look, I don't know. It is hard for me to see, and I could be wrong, but if I network-bound these models, they're not magically going to be able to cross the network boundaries. Data is data, and I've never seen any real evidence—except from some very early models that I think people picked up on very quickly—that there is some agenda or bias built into them.
I do think that, to some extent, it is important for the U.S. that we develop our own models. I think it would be a massive loss if there are 5 companies—5 different labs in China—creating open-source models and we're struggling to get 1 set up. It's necessary. I also think it's inevitable.
You know, the DeepSeek moment a year ago—I remember someone saying to me, and I thought it was very well said, and the world has changed a lot, but they said, “Hey, we should just forget
Mm-hmm.
that this is a Chinese model. We should just act like this came from
Mm-hmm.
Meta and build with that in mind.”
Mm-hmm.
It's like, I think you're missing the forest for the trees. There are 2 scenarios, right? Either America does not ever come up with good open-source models and there's probably a fundamental problem there, or we will get there and we need to be ready for that world.
Yeah, that makes sense. It's interesting because I think it's very important for the U.S. to have a strong open-source footprint here. At least for now, it looks like the Chinese government is effectively subsidizing at least a large subset of these models. That subsidy, or surplus, is effectively just being passed on to U.S. enterprises that are adopting these models.
In other words, it's a way for the Chinese government to effectively subsidize U.S. enterprise in an indirect manner, and I think that's a little bit lost right now. But it's always interesting to weigh that against some of the other concerns that are raised. I appreciate your comments on this.
Well, yeah, and I think the concern also becomes: What happens if we aren't able to? I think if you think about the economics here, DeepSeek, by most measures, is a very good model. You can argue whether it's at the absolute frontier or not, but let's go back 3 months. It was doing a whole lot of things 3 months ago.
You could run DeepSeek at probably 20% of the cost of running Anthropic models in production, with comparable or better latency and probably better reliability. If we don't have access to that intelligence in that form, I think it's just a massive loss.
As a country, we won't be able to innovate as fast, because the cost of intelligence going down and control of intelligence—what we have seen—just means more intelligence. Intelligence is being embedded in more places.
Yeah, an important note here that we didn't mention explicitly is that the state-of-the-art models—the ones that are furthest ahead on the frontier—are actually still the closed-source Anthropic, OpenAI, Google, et cetera.
Yeah.
Can you characterize the workload a little bit? What percentage of tokens are being served on Baseten, and how many of them are from custom models of some kind versus vanilla open source today?
It's all custom.
Okay.
95% plus.
95%. I think that's really cool, to be honest.
We have 3 businesses right now.
Do we have 3 now?
No, no. Dedicated inference, which is basically customer-owned inference—your SLA is your SLA. We have shared inference, which is shared inference with shared SLAs, and we have a training business. I'd say 95% of the tokens today are on the first business, and almost all of them involve the customer making some modifications to the model with their own data, specialized for the use case. What's even more important is that they might be compiling in different ways. No one is just running the vanilla open-source weights. You might be customizing it for quality, but you must also be customizing it for performance.
You made an acquisition of a research team a few months ago. You've mentioned post-training customization. What was the rationale behind the acquisition, and what is that team doing today?
Yeah. The rationale for the acquisition was that we're infrastructure and product people, and now we're really good infrastructure people, but we didn't have much of a research capability ourselves. What we saw was the market moving heavily toward post-training, and that we could accelerate the market itself with post-training resources, either productized or even just as resources for that market.
Parea was a company that was a Baseten customer. They were post-training models and running them on Baseten. I think what they realized was that they would eventually need to become an inference company. What we realized was, “Hey, we really needed that expertise, too, because it represents a way for us to get closer to the customer earlier and be able to support them more.”
It just made sense as a fit—pairing them together. As I said in the opening statement, as more and more post-training models have come up, we've realized that the demand for software tools to do post-training or for post-training expertise is very high, and we're really investing in that. There's also a bunch of Australians. I like to think that we had a bit of alpha there.
They're working with all sorts of customers, and it's also very interesting. We were doing a lot of research on the performance side and less so on the post-training side. As we've started to do a lot more research on the post-training side, you start to see how linked inference and post-training are. Even when you think about stuff like quantization—when you should do that, and how training the model affects how you need to quantize for inference—it has become very apparent how paired these problems are.
Mm-hmm.
More and more, we realize that post-training and inference are both sides of the same problem. Inference will ideally beget more post-training: inference creates data, you do evals, and you can now post-train on the reward function that you found with those evals. Hopefully, this will play out entirely.
Plenty of folks from Anthropic and OpenAI—Sam, Greg, et cetera—have said in recent months that inference is super strategic, inference talent is strategic, and capacity is strategic. So, between that and post-training, these are very difficult-to-gather capabilities.
Yeah.
I imagine that lots of your customers go to you for advice on how to make this progression toward custom models. What do you tell people about the lifecycle and when they should invest in that?
Yeah, I think it's: go prove to yourself with the best-in-class model that you have something worth optimizing. A lot of customers, if they come to us—there was that meme from 2 years ago: “It feels like there's no GPUs pre-product-market fit.” It's like, no post-training pre-product-market fit as well.
Yeah, yeah, yeah. Some people that you're working with are very, very at scale first.
Yeah, they have a user signal that they know how to optimize. They've shown that they can serve customer value and that they have something special around that value. Once you have that value, it's like, “Okay, now how can I do that better, faster, and cheaper?”
The idea is that if you need to be very good at customer support, maybe you don't need to be that good at coding, and a specialized model might be a better fit for that problem. You can do it better, faster, and cheaper.
What about the capacity side? You started with unifying capacity across all the clouds and new clouds. How do you think about this when you keep talking about a supply crunch and a multiyear supply crunch?
I think there's so much narrative around the supply crunch.
No matter how much we hear about the supply crunch, I don't think people realize how bad it really is. There is very, very little slack compute available. We run pretty large clusters ourselves, and we run them at uncomfortably high utilization. We're not saying we're in the mid-90s most of the time, but it's close.
We now sit in 18 different clouds. We have 90 clusters around the world across 18 different clouds. Initially, we built this technology to create one runtime fabric that spans all these different clouds and abstracts that away from our customers as a way to think about reliability, latency, failover, and all these things that we think are going to be very important for mission-critical use cases.
That same technology—our ability to get compute wherever humanly possible—has been really, really helpful in our ability to get supply. What I mean by that is, we can be introduced to a new provider in a different country and have it up and running with the whole Baseten inference stack—
—as part of the fabric.
Part of the fabric in half a day? Half a day, maybe less. Even for us, it is hard to grow. We have a 4:00 p.m. standing meeting for the company where we basically ask, "How do we manage capacity for the demand right now?"
I think the second part that people don't really understand is that there are also a lot of suppliers right now that are kind of grifty. They haven't run data centers before. They don't understand SLAs, especially for inference.
Even when there is capacity available, there's a lot of doubt. We run a lot more of this than we've ever done in-house, so it's fine. But there's probably a dozen good clouds, and I would put three or four of them in the gold tier.
Mhm.
I think that just means that we're not only supply-crunched; we're supplier- and operationally crunched onto people who can run these data centers as well.
How far ahead can you actually buy capacity right now? In other words, is there any slack in the market if you buy 2 years ahead or 5 years ahead?
You mean the actual contract length, or actually saying, "I want this in January 2028"?
Yeah, either one. It's more, "I want this in January 2028," or at least, "I have some visibility into my future supply."
Yeah. You could buy that, but you also have to remember how quickly the market is moving. That gets balanced somewhat by the fact that the H100 is such a great chip.
Yeah.
It's crazy—it's 4 and a half years old, and the price is still going up.
Yeah.
Maybe it has a useful life of 9 years.
Yeah.
That's good, but at the same time, yes, you can do that, but you're making a lot of bets as part of that. In terms of what's changed over the last 6 months, the term length that people want has just gone up.
If you wanted 1,024 B200s from a good cloud right now, you're not getting that for less than a 3- to 5-year contract. Right now, it probably comes with 20% of the TCV prepaid. What becomes important when acquiring capacity is that you need to have enough demand to serve it, and you also need a low cost of capital, which is actually changing the dynamic pretty significantly.
Does that impact how you think about going public as a company? Because arguably—
Yeah. I think you'd go sooner.
Yeah, exactly.
Yeah, I think you need—I think there was demand for that. But I think the pool will also—one of the realizations that we had recently, especially as software people, and so we don't think about this all the time, is that our business has very interesting working-capital requirements.
As a result of that, it has very interesting financing requirements, and we're not, at least right now, even going down to the—
Yeah, there's also things you could do in terms of debt or other structures.
Yeah. I've learned a lot about debt recently.
Given the supply crunch, and inference being one of the top couple of markets you're going after, you have plenty of people who understand this problem and therefore some competition. How do you think about what factors create a dominant player here, or a winning player? Is it, as you mentioned, cost of capital? Is it access to supply? Is it software? Is it demand?
Yeah.
Just being excellent at everything?
Yeah, I think so.
Is it operations? Like a special—
Yeah, yeah, yeah. GPUs as a service are not sticky. I think that's been seen. Customers generally just see that as a commodity.
Inference with the software layer included is incredibly sticky. None of our top 30 customers have ever churned. We're talking about 400% annual NDR around our business. It's very, very sticky, so I think that software layer is very important.
The optimist in me is like, "Oh, there's so much value in the software." I think we will build the best software layer for inference that exists.
As it's becoming clear now, access to inference compute is—
Yeah.
—is a strategic advantage. I think that is the strategy that even the labs are going after, which is, "If we have all the compute, good luck running inference."
Yeah, yeah. In a world of constrained compute, the number one thing to own is compute.
Yeah.
Just owning it in and of itself is an asset, and I think people underappreciate that.
Yeah. You can't make good hot chocolate without milk.
[Laughter.] Unless you're vegan. No one wants the vegan inference.
[Laughter.] Well, I've got to ask you. People might want alternative milk, right? The H100 is a great chip. People want a B200. They want a GB200. They want, of course, tons and tons of NVIDIA.
When you think about making a bet several years in the future, do you believe that there's a multichip world? What do you think happens from a compute perspective on the chip side?
Yeah. I think diversification everywhere is a good thing. In the same way I want a world of many models, I think we want a world of many, many things.
It'd be sad if it didn't happen.
Yeah, and I think everyone would be sad. I will say that, to some extent, I think there will be inference-specific chips. I think you'll have decode-specific chips. We're looking at it—I mean, NVIDIA said this.
Yeah, yeah. I mean, that was a whole—
It's like, you know, I think that is very straightforward and makes sense. I think people really, really, really underestimate NVIDIA's supply-chain capabilities, how good CUDA is, and the developer ecosystem around it.
To me, one of the most important things as an infrastructure company in this moment is how fast you can move, and you can move fastest with NVIDIA today. I think that is the reality. Given the scale at which they operate, it's hard to see, in the short term—in the next couple of years—how anyone will be able to compete with that.
I'm not saying it won't happen, but especially with so many of the other players, what you need to be able to compete here is for the ecosystem to form around you. If you tie up all your supply with one buyer, which a bunch of the other chip providers have done, it's actually hard for that ecosystem to form.
If you're a big lab and you have a proprietary deal with one chip type where you get 90% of the supply, it's actually in your best interest to make sure you get 95% of the supply, everything gets built for you, and no one else can ever use it.
When you think about reacting to the market, what do you think is happening with the actual workloads that you have to invest in? Obviously, code agents and long-horizon agents have become a big deal over time. People talk a lot more about CPU compute. Video inference is different. I don't know if it's that—
Sandbox is like—what’s important for you guys to invest in now?
Yeah, look, I think for us, all the runtime stuff is obviously very important. That means what chips we run on, how we run, and what kinds of workloads we support. Do we get very good at diffusion transformers? Yes. Coding agents need sandboxes, which could be called sandboxes.
There are all sorts of new speculative techniques to get faster inference. We need to do that. Even things like KV-cache-aware routing—that stuff is a bit old now, but we need to continue to be very good at it. We’re also somewhat disentangling prefill and decode and starting to treat them as separate problems. I think that’s something we’re very focused on, and we’re seeing massive gains there.
That’s at the runtime level. Beyond that, everything we think about is how to create more of that loop between inference and post-training, because we think that just begets more inference. We will build or partner in almost everything. We’re going to work with the best eval companies in the world to make sure that’s very well integrated, like Braintrust, into and around Baseten.
On the sandbox side, we will partner to build the best sandbox experience that will exist. Then we’ll create the best training APIs to make it so continual learning becomes somewhat of a solved problem, not just a discrete thing. I think that’s the core Baseten product thesis: How do we build that loop?
Everything else around that becomes: How do we make sure we can do everything we can to ensure that it gets as big as possible? That’s access to compute. That’s our own infrastructure. We need to make sure we can get compute anywhere and that we have access to our own compute.
Then I think it’s all the primitives that come after that, which become incredibly margin-accretive, both for us and our customers. That’s stuff like sandboxes and async batch inference. How do we drive utilization by having a first-class batch-inference experience?
To me, this is what an inference cloud looks like: You are very good at inference, and then you start to do all the things that are tangential to, or loop into, inference. You partner when necessary and build when necessary. But we really do want to own that core inference story, then go down to unlock supply and create margin, and go up the stack to unlock value.
What would surprise people about some of the issues you discover only at scale? I was surprised when you guys ran into scale limitations—fundamental limitations—with some of the hyperscaler products that you were consuming. I kind of think of the AWS and GCPs of the world as supporting infinite scale.
Yeah. I mean, I think very large companies that run services at big scale probably see the same stuff.
Mhm.
Yeah, all the edge cases just become—you actually experience them. You experience them.
You start seeing things like—yesterday, for the first time ever, we saw a kernel panic. That only happened because a Fluent Bit worker was creating too many logs, the scale was too big, and everything was going into one node. It was happening 2 times at the same time with 2 different workers, so you see all these systems-level and kernel-level problems.
But I think the craziest thing is that you start to see, with LLMs, that these runtimes are pretty immature. Even how we use KV cache is probably a little less sophisticated than most people realize. We’re starting to see the limitations of the current and next set of primitives that need to be built from a scale, security, or performance perspective.
I think it’s really at the runtime level and the systems level. The edge cases are a lot more systems-level than they are LLM-specific.
What are the things that keep you up at night?
Capacity.
Quick answer. Yeah.
Yeah, I think capacity. Probably just that this market is so big, and it represents a moment when you should be as aggressive as possible. We’ve grown a ton this year, in the last 12 months, in the last few months, but the answer is always just: Go bigger, go faster. I think that’s really, really fun.
It’s also a little exhausting, and we are all in somewhat uncharted territory in terms of how fast and how big you can go and how things can get. But I think we need to do that to get the amount of value that we want to get out of LLMs in the next 5 to 10 years.
Or we have to invent a lot of new stuff.
Yeah.
Maybe if we just talk a little bit about what you’re learning from scaling, 30× is an aggressive thing to go through as a company. You’ve brought in a lot of really amazing talent, like Danny, Samir, and Stephen Day—folks on both the technical and go-to-market sides. What do you think is working about how you are recruiting and scaling, or what’s your philosophy on that?
We were very, very flat until 8 to 18 months ago. I remember when I worked with Elad a lot, actually. A lot of it is that you just need leaders. It’s actually so contrary to everything engineers think. They’re like, “Oh—”
It’s all overhead.
It’s all—everything is overhead.
You once told me, I think, that you didn’t—you were like, “Hey, Sarah, what about we just have engineers instead of salespeople?”
Yeah. Bad.
Everybody learns that.
Everyone—we all know about it. I remember, you said it so clearly at the time, Elad, and I think that’s what we noticed: Actually having a leadership team that you can trust is so important.
I think the 2 or 3 things that I’ll say are: You want people where you can give them whole problems. If you feel like you’re micromanaging, if you feel like you need to be involved in everything, I think that’s a bit of a cop-out as a founder, because you’re just like, “I just need to be involved in everything.” No, you probably just don’t have the right people.
I think the second thing is: Be very, very clear about what you’re optimizing for. When you’re very, very clear about what you’re optimizing for, the people who fit become apparent, and the people who don’t also become apparent.
If it’s something generic like, “We want the smartest, hard-working people,” you can’t do much with that. What we cared about was, “Hey, actually, we don’t care about a lot of people who’ve done this before. We care about first-principles people who think from first principles.” Work has to be a high priority, but they also have to be very kind and nice and care about the collaborative environment. We don’t have a hero culture. We’re very low-ego. If you need a manager, it’s probably not the right place to be.
Once you have that clear rubric, the people who fit into it become very apparent, and the people who don’t fit into it also become very apparent. We’ve hired amazing people, like you mentioned, but I think what’s a lot more interesting is that we haven’t had a ton of unnecessary turnover. People tend to stay because we’re very clear on what we want. It took us a while to get there, though.
What about the idea of an operations culture? We were talking to Alyssa and Henry about this, and she was like, “Well, the hard thing about cloud is actually just operations. I slept with a pager under my pillow for a decade.” I don’t think I’ve seen you detached from your Slack channel.
Yeah, my phone is buzzing right now. I’m getting anxious.
And you’ve been concerned before: Do people get it? What is distinctive about that?
I think, one, if you have worked at an infrastructure company—for example, we were once in a meeting with a bunch of AWS executives, and these were very senior AWS folks. All their pagers went off multiple times during our 45-minute meeting. It’s very much a cultural thing.
Our infrastructure can’t go down. I think my co-founder, when his pager goes off, has a 7-year-old who says, “Is that a P0?”
You just have to get used to it. That’s the culture you live in, and it changes the speed, but it also becomes a cultural thing.
I think it rejects people who don't fit into it very quickly.
Like engineers who avoid pagers.
Yeah. When we have pagers, everyone's on call. There's been a joke that they may as well be a siren that goes off in the office.
People have been talking ad nauseam in the AI community about Jevons paradox.
Yeah.
It's really a question around price elasticity and availability: if you decrease the cost of a good—say, intelligence as a good—people actually consume more of it.
Yeah.
The personal or business ROI of it goes up, and the demand for it goes up, not down. Do you see this, and are you working against yourself trying to make these models more efficient? People just use them more or less?
Yeah, I think you have to think about this from a developer perspective and a consumer perspective. I think consumers just want the best answers and the best experience, which is somewhat governed by more intelligence to some extent.
From the developers' perspective, they would insert more intelligence if you made it cheaper. They would insert more intelligence anyway, but if you make it cheaper, they'll insert a hell of a lot more intelligence. You see this with agents and stuff. Agents are just longer-running now, and I think that's what we have seen with the cost of inference going down: folks are just like, okay, we can run this for longer, we can make it do a bit more work, and we'll get to a larger end result.
I think as compute scales from an inference perspective as well, and we're seeing that with almost all our customers, they either start with, “This is the quality of answer I need to get to, and this is the amount of inference I need to do to get that,” or, “This is the base-level model that I can start with or work with to get there.” I think the more we drive down the cost, what they realize is that more intelligence just means a better user experience.
I just want a better answer.
Better answers, better experiences, more dollars, more dollars, even more revenue. So, yeah, I think inference going down just begets more. But I truly think we're kind of in a world that is the last market, right? Even if there's AGI, all that's left is inference.
So, you do not see in your customers a, like, “This answer is enough and this action is enough” dynamic?
No. No.
Yeah, it's going to keep going for a long time, I think. How do you view all this evolving toward the future? This seems like it's going to be one of the biggest markets of all time. We have this massive shift where we're moving from software and seats and digitalization into actual intelligence—selling units of cognition, selling agentic workflows. What does this all look like in a couple of years? What is your view of this future world?
I think for consumers, it's the best possible thing, right? Everything is somewhat smarter. You get better care because your doctors have access to better tools. There's all this stuff about there being fewer software engineers. I think we just build more software.
I think we just build a ton more software, and we're not slowing down hiring software engineers; we're just building more things. For consumers, that just means better tools, more software, all those good things.
Like everybody has their own team for everything, right? You have an agent that helps with your doctor. You have an agent that helps you learn stuff. You have an agent that helps you organize your life.
It's concierge. It's concierge everything.
Yeah, concierge everything for everyone.
Yeah, and I think what that means is amazing. I think that's great. And education, same thing: you have concierge education. You get personalized access to everything.
I think when you go one step back to how it affects developers and companies, if you don't embrace this, I think it's the extinction moment for a bunch of folks. Everything needs to change, and I don't think that means that product design needs to figure it out. I don't think that's a thing.
I think what's more interesting is that all these workflow and software companies need to figure out what the intelligent, or intelligence-inserted, versions are that drive all that user value for those consumers we talked about.
Yeah, very exciting. Thank you so much for joining us today.
Yeah, thanks, guys.