[BidClub_]
No Priors · · 43 min

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud

Tuhin SrivastavaSarah GuoElad Gil

YouTube
TL;DR
  • Baseten’s 30x growth reflects an AI-native application boom, while the larger enterprise inference market is still mostly absent. Sarah Guo said Baseten expects more than $1 billion in revenue this year, which Tuhin Srivastava acknowledged; he estimates application companies still represent 99% of inference count, so “the majority of the market hasn’t come online.”

  • The durable application moat is proprietary workflow feedback, not merely access to model weights. Abridge’s clinician edits and downstream EMR actions create a reward signal that frontier-model companies may not have access to; post-training on that signal could produce specialized, long-horizon agents. “The user signal that they can gather that only they can gather” is what protects the application layer.

  • Production inference is already overwhelmingly custom, linking deployment to an increasingly continuous post-training loop. More than 95% of Baseten’s tokens run through dedicated inference, and almost every customer modifies models for quality, performance or both: “No one is just running the vanilla open-source weights.” The sequencing matters—“no post-training pre-product-market fit”—because companies should first prove value with the best model, then specialize it to become “better, faster, and cheaper.”

  • Chinese open models are economically strategic, while Guo noted that closed-source U.S. labs still define the absolute frontier. Srivastava hedged that he “could be wrong,” but said that if these models are network-bounded, they will not “magically” cross network boundaries; he said DeepSeek could run at probably 20% of Anthropic’s production cost, with comparable or better latency and probably better reliability. His stance: U.S. open models are both necessary and inevitable, but ignoring today’s available intelligence would be “missing the forest from the trees.”

  • Capacity scarcity now reaches directly into contract duration, working capital and likely public-market timing. Baseten runs 90 clusters across 18 clouds at uncomfortably high utilization—close to, but not generally in, the mid-90s—and holds a daily 4 p.m. meeting to allocate supply. A 1,024-B200 block from a credible cloud can require a three-to-five-year commitment and probably 20% of TCV prepaid; asked whether that argues for going public earlier, Srivastava answered, “go sooner.”

  • Baseten’s moat thesis is software plus scarce compute, not bare GPU rental. Srivastava calls GPU-as-a-service a commodity, while inference software has delivered no churn among Baseten’s top 30 customers and roughly 400% annual NDR: “If we have all the compute, good luck running inference.” He expects specialized chips, but NVIDIA’s supply chain, CUDA and ecosystem mean infrastructure operators “can move fastest with NVIDIA today.”

  • Inference efficiency is behaving like Jevons paradox: lower unit costs produce longer agents and more total cognition, not demand saturation. Developers insert “a hell of a lot more intelligence” when it gets cheaper because better answers improve experiences and revenue. Srivastava therefore calls inference “the last market”—even with AGI, inference remains—and sees “concierge everything” for consumers but an “extinction moment” for workflow companies that fail to add intelligence.

Digest · the substance, structured for research

1. Workflow data keeps the application layer alive

  • Guo opened with Baseten’s 30x growth in 12 months and said the company expects more than $1 billion of revenue this year; Srivastava acknowledged the statement. His market-wide explanation: open weights crossed a capability “chasm,” RL and post-training became mainstream, and applications increasingly in-house intelligence. Baseten is “somewhat indexed on that” expansion.

  • Srivastava’s application-layer thesis is that proprietary model weights alone are not the moat; scarce user signal embedded in workflows is. Abridge, an ambient scribe used by physicians in almost all U.S. hospitals as he described it, captures clinician edits and downstream actions inside the EMR, creating feedback a frontier-model company may not have access to. That reward could eventually post-train specialized, long-horizon agents.

  • Customer support carries the same argument: one incoming ticket may trigger “one, two, 10, 20 actions,” not a one-shot answer. The differentiated asset is the sequence of actions and decisions, which gives an application company something specific to optimize.

  • By inference count, Srivastava puts today’s market at 99% AI-native applications rather than enterprises. Adoption has progressed from AI tools toward closed-source model APIs; custom-model adoption comes next. A “Stripe evolution” lets Baseten follow frontier customers while Abridge, OpenEvidence and others translate customers’ data-retention, model, deployment, GPU, latency and transparency requirements.

2. Capability beats provenance, and production weights are custom

  • Guo described a shift from Mistral, then Llama, toward Chinese-origin models; Srivastava said customers begin with capability rather than cost because capability unlocks economic value, then optimize. Baseten customers range across GPT-4o, Moonshot, DeepSeek and Canopy’s Orpheus text-to-speech model: they use “whatever [is] at the frontier.”

  • Baseten describes three businesses—dedicated inference, shared inference and training—but more than 95% of tokens sit in dedicated inference. Almost every customer changes the model with its own data, often compiling or otherwise customizing it for both quality and performance. “No one is just running the vanilla open-source weights.”

  • On Chinese-model risk, Srivastava hedged—“I could be wrong”—but said that if he network-bounded these models, they would not “magically” cross network boundaries, and he had seen no real evidence of an embedded agenda or bias beyond some early cases that people detected quickly. He nevertheless called U.S. open models necessary and inevitable. Guo noted that Chinese subsidies effectively pass surplus to U.S. enterprises; Srivastava said DeepSeek could run at probably 20% of Anthropic’s production cost, with comparable or better latency and probably better reliability. Guo separately stressed that the absolute frontier remains with closed-source models.

3. Inference and post-training form one compounding loop

  • The acquired research team came from Parea, a Baseten customer that had been post-training models and deploying them through Baseten. Parea realized it would eventually need to become an inference company; Baseten realized that research expertise would let it reach customers earlier and accelerate a market demanding both post-training software and hands-on expertise.

  • Srivastava’s lifecycle advice starts with restraint: prove value using the best available model before optimizing anything—“no post-training pre-product-market fit.” Once a company has a distinctive user signal, specialization can make the workload “better, faster, and cheaper”; a support model, for example, need not retain frontier coding ability.

  • Training choices affect how a model should be quantized for inference, while deployed inference yields data, evals and reward functions for the next post-training pass. Baseten wants training APIs that make continual learning “somewhat of a solved problem,” with Braintrust around evals and partners around sandboxes. Its product thesis is to “go down to unlock supply and create margin” and “go up the stack to unlock value.”

4. The capacity crunch turns inference into a financing business

  • Baseten operates 90 clusters across 18 clouds at “uncomfortably high” utilization—not mid-90s most of the time, but close. Its common runtime fabric was initially built to abstract reliability, latency and failover across clouds, and it can bring a provider in another country online in half a day or less. The company still holds a daily 4 p.m. capacity-allocation meeting. Asked what keeps him awake, Srivastava answered immediately: “Capacity.”

  • Nominal GPU availability overstates useful supply because many providers have never operated data centers or understood inference SLAs. Srivastava called part of the market “kind of grifty”: perhaps a dozen clouds are good, and only three or four qualify for his “gold tier,” leaving the market both supply-constrained and “supplier- and operationally crunched.”

  • Long-dated supply is buyable, but the rapidly moving market means making a number of bets: the four-and-a-half-year-old H100 still has a rising price and might have a useful life of nine years. A 1,024-B200 block now cannot be obtained from a good cloud for less than a three-to-five-year contract and probably comes with 20% of TCV prepaid. Matching supply to demand and funding it cheaply matters; on an IPO, Srivastava said, “go sooner.”

  • GPU rental itself “is not sticky”; customers treat it as a commodity. Inference bundled with software is different: Srivastava said none of Baseten’s top 30 customers has churned and annual NDR is around 400%. He also emphasized that buyers need enough demand to serve acquired capacity and a low cost of capital.

5. NVIDIA leads the near term while the runtime keeps fragmenting

  • Srivastava expects inference-specific and decode-specific chips, but sees NVIDIA’s supply chain, CUDA and developer ecosystem as a formidable next-few-years advantage: infrastructure companies “can move fastest with NVIDIA today.” Rivals need an ecosystem; if one lab locks up 90% of a chip’s supply, its incentive is to take 95%, making sure everything gets built for it and no one else can use it.

  • Workload change dictates the runtime roadmap: diffusion transformers, coding-agent sandboxes, speculative inference, KV-cache-aware routing and treating prefill and decode as separate problems. Baseten reports “massive gains” from that disaggregation, while first-class asynchronous batch inference can raise utilization for both Baseten and its customers.

  • Scale exposes mundane systems edges before exotic model failures: a first-ever kernel panic came from two Fluent Bit workers simultaneously producing too many logs into one node. Srivastava says LLM runtimes and current KV-cache use remain immature, but most observed edge cases are still systems- and kernel-level; the next primitives must improve scale, security and performance.

6. Cheaper inference expands the market—and raises the operating bar

  • Baseten remained very flat until roughly 8–18 months ago, when Srivastava abandoned the engineering instinct that leaders were merely overhead. His test now is whether an executive can own a whole problem; persistent founder micromanagement is often a sign that “you probably just don’t have the right people.”

  • Its hiring rubric prizes first-principles thinking over having done the job before, makes work a high priority, and insists on kindness, collaboration, low ego and “no hero culture.” Srivastava argues that this specificity makes both fits and misfits obvious—and has limited unnecessary turnover despite rapid scaling.

  • Operations culture is less negotiable: during one 45-minute meeting, multiple senior AWS executives’ pagers repeatedly fired. At Baseten, everyone is on call, and pagers are jokingly likened to an office siren. Srivastava’s co-founder knew the norm had reached home when his seven-year-old heard the pager and asked, “Is that a P0?”

  • On Jevons paradox, Srivastava sees no “this answer is enough” ceiling: cheaper inference lets developers insert “a hell of a lot more intelligence,” agents run longer, and better answers produce better experiences and revenue. His end-state is “concierge everything”—personalized care, education and tools, with more software rather than fewer engineers—and an “extinction moment” for workflow companies that fail to insert intelligence.

Sarah Guo

Hi, listeners. Today, Elad and I are here with Tuhin Srivastava, the founder and CEO of Baseten, the AI inference cloud. We're here to talk about capacity constraints for AI compute, why inference is the last market, how the workload is changing, the open-source and perhaps multichip future, and what 30× scale in a year looks like. Tuhin, welcome back.

Tuhin Srivastava

Hi.

Sarah Guo

Good to see you.

Tuhin Srivastava

Thanks for having me.

Sarah Guo

All right, you are in one of the craziest markets: AI inference. It's very important, and there's a lot going on. You guys have grown 30× over the last year, and I think I can say you're expecting to do more than $1 billion in revenue this year.

Tuhin Srivastava

Mm-hmm.

Sarah Guo

What's going on? Tell us about scale.

Tuhin Srivastava

Yeah, it's been nuts. I think what's happened over the last 24 months—and this keeps getting bigger and bigger—is that everyone is realizing that you can put AI everywhere. You have all these great options available, from closed-source to open-source models. The open-source models have crossed some sort of chasm in terms of their baseline capability, and then RL techniques and post-training for specialized models have become mainstream enough. There are enough examples of it working.

Customers are realizing they can own their inference more and more. What that's meant for us is more of the long-tail models coming through, customers in-housing a lot of that intelligence themselves, and the application layer just getting bigger and bigger. As that grows, we are just somewhat indexed on that, and we've been around to be able to collect the demand.

Sarah Guo

There's an existential question in here that I think everybody is continually asking: Does the independent application layer get to exist at all, versus the labs? You have to believe this. Why do you believe it?

Tuhin Srivastava

Yeah, look, I think it would be a sad thing if it didn't exist in general, and that's my—sadness is fine.

Sarah Guo

Sad all the time.

Tuhin Srivastava

Yeah, sadness is fine. But that's not the reason why I think the application layer will exist. I think the application layer will exist for a number of reasons.

One is this idea that what is valuable to a company is the user signal that they can gather, which only they can gather. To the extent that signal is encoded in a model, I think a lot of their business will be at risk. But to the extent that it is encoded in workflows, that is where they will be able to develop a moat.

A good example of that is a company like Abridge, where the edits clinicians make to the notes, and what they do with the notes after the fact in the thing that happens inside the EMR three steps down, becomes a workflow that only—

Sarah Guo

Can you explain what Abridge does?

Tuhin Srivastava

Sorry. Abridge is an ambient scribe that is used by physicians in almost all hospitals in the U.S. I think they're a lot to invest in. Shiv Rao is amazing. Great company, great team, great product. They've basically got this very deep integration into hospitals and clinician workflows.

My argument here is that it's very hard for a frontier model company to eat that because they just don't have access to that user signal. What will happen over time is that folks who have access to that user signal can start to post-train models on that reward signal and start to get long-horizon agentic models running.

To the extent that is possible, and that signal is differentiated, unique, and somewhat rare to get access to, there will be an application layer. Support companies are another example of that. A support task isn't one-shotted. Usually, at a company like Basecamp, when a ticket comes in, there are one, two, ten, or twenty actions that get taken. That is where someone can develop a specialized model.

Sarah Guo

So there are almost two versions of this. There's the new companies, like Abridge, Decagon, and some of these other things that you mentioned, that are doing these new types of applications using AI and selling them to customers. The other is enterprises building things in-house or building their own models.

What proportion of the market today do you think is these new application companies—AI natives, the fast-growing companies, some of which are at considerable scale now, like Abridge, Cursor—

Tuhin Srivastava

OpenEvidence.

Sarah Guo

OpenEvidence—those types of companies? What do they teach you? What does that push the company to do? How do you think about serving them versus evolving for the enterprise?

Tuhin Srivastava

I think you asked me the same question two years ago. It's crazy that the answer is still the same. If you look by inference count, it would be 99% the former. That represents the scope of the opportunity here: the majority of the market hasn't come online and added AI into this market.

Sarah Guo

Yeah, this is just so much still to come, and people are underestimating that, I think.

Tuhin Srivastava

100%, and what's cool is that we're seeing the transition happen. Before, it was like, "Hey, are they using AI tools?" I don't think that was immediately obvious two years ago. I think that's obvious now: yes, they are. Are they using closed-source model APIs? I think they're starting to get there. And then once you do that and see what is possible, then comes the whole custom model adoption. I think that is all that is ahead of us today.

Yeah. I think, firstly, you just learn a lot by building with the companies at the greatest scale and doing the most interesting things. We think about it in two ways.

The most obvious way is to build for the highest scale. The customers that push you the most technologically will push everything else into place. I think the evolution of Stripe as a company showed that. Stripe now serves so many enterprises, but 12 years ago that wasn't the case. They just built for the frontier and went with them.

The second way we think about this is to build for companies that are serving enterprises. We don't serve the enterprise, but our customers serve enterprises. Abridge serves enterprises, OpenEvidence, Decagon, Writer, Gamma—all these companies serve enterprises at scale.

What we actually get is a translation of the requirements from them. They're saying, "We need this sort of data retention. We need these types of models deployed. These are the types of GPUs or latencies we're okay with. These are the model requirements, from a transparency perspective, that we care about."

I think that is actually the more nuanced answer. If you listen to what their needs are, we get a full translation of what the enterprise would require. I would say that by serving companies like Abridge and OpenEvidence, we're probably pretty well suited to go serve the healthcare system, given that they are selling into healthcare and selling to those customers.

Sarah Guo

How much of a shift are you seeing in terms of the types of open-source models that are being used? Two or three years ago, I think the main thing was Mistral and then a few other things. Then Meta came along with Llama, and it really shifted in terms of the best-performing models. Now, the best-performing models are of Chinese origin in different ways. Do you see that mix reflected in what's being used by our customers?

Tuhin Srivastava

Yeah. The customers we are serving—these are the fastest-growing AI companies in the world—are very forward-thinking. They want to use the best models, and they are optimizing.

There is a subset of tasks, which I think is small today, where people really start with cost. But everyone comes for capability first, because that's really where economic growth is being unlocked and where value is being delivered. Then they optimize.

We've seen customers use everything from GPT-4o to Moonshot models to DeepSeek to Canopy's Orpheus, which is a really good text-to-speech model. Customers generally want to use whatever is at the frontier.

The difference is that we have a lot more visibility into how to run these models and how to run them really well. Secondly, they're good now.

Sarah Guo

There have been a number of concerns raised about the use of Chinese models, in particular security concerns—whether there's something embedded in the models, Trojan horses, or other things.

First, do you think there's any real concern there? Second, people often talk about how there should be U.S. counterweights to this. From a geopolitical perspective, do you think that's legitimate—something we should be worried about? How do you think about the origins of these models versus their uses?

Tuhin Srivastava

Yeah, look, I think these models, firstly, are fantastic. They're amazing. We work with these teams. They're truly awesome.

I'd say, look, I don't know. It is hard for me to see, and I could be wrong, but if I network-bound these models, they're not magically going to be able to cross the network boundaries. Data is data, and I've never seen any real evidence—except from some very early models that I think people picked up on very quickly—that there is some agenda or bias built into them.

I do think that, to some extent, it is important for the U.S. that we develop our own models. I think it would be a massive loss if there are 5 companies—5 different labs in China—creating open-source models and we're struggling to get 1 set up. It's necessary. I also think it's inevitable.

You know, the DeepSeek moment a year ago—I remember someone saying to me, and I thought it was very well said, and the world has changed a lot, but they said, “Hey, we should just forget

Sarah Guo

Mm-hmm.

Tuhin Srivastava

that this is a Chinese model. We should just act like this came from

Sarah Guo

Mm-hmm.

Tuhin Srivastava

Meta and build with that in mind.”

Sarah Guo

Mm-hmm.

Tuhin Srivastava

It's like, I think you're missing the forest for the trees. There are 2 scenarios, right? Either America does not ever come up with good open-source models and there's probably a fundamental problem there, or we will get there and we need to be ready for that world.

Sarah Guo

Yeah, that makes sense. It's interesting because I think it's very important for the U.S. to have a strong open-source footprint here. At least for now, it looks like the Chinese government is effectively subsidizing at least a large subset of these models. That subsidy, or surplus, is effectively just being passed on to U.S. enterprises that are adopting these models.

In other words, it's a way for the Chinese government to effectively subsidize U.S. enterprise in an indirect manner, and I think that's a little bit lost right now. But it's always interesting to weigh that against some of the other concerns that are raised. I appreciate your comments on this.

Tuhin Srivastava

Well, yeah, and I think the concern also becomes: What happens if we aren't able to? I think if you think about the economics here, DeepSeek, by most measures, is a very good model. You can argue whether it's at the absolute frontier or not, but let's go back 3 months. It was doing a whole lot of things 3 months ago.

You could run DeepSeek at probably 20% of the cost of running Anthropic models in production, with comparable or better latency and probably better reliability. If we don't have access to that intelligence in that form, I think it's just a massive loss.

As a country, we won't be able to innovate as fast, because the cost of intelligence going down and control of intelligence—what we have seen—just means more intelligence. Intelligence is being embedded in more places.

Sarah Guo

Yeah, an important note here that we didn't mention explicitly is that the state-of-the-art models—the ones that are furthest ahead on the frontier—are actually still the closed-source Anthropic, OpenAI, Google, et cetera.

Tuhin Srivastava

Yeah.

Sarah Guo

Can you characterize the workload a little bit? What percentage of tokens are being served on Baseten, and how many of them are from custom models of some kind versus vanilla open source today?

Tuhin Srivastava

It's all custom.

Sarah Guo

Okay.

Tuhin Srivastava

95% plus.

Sarah Guo

95%. I think that's really cool, to be honest.

Tuhin Srivastava

We have 3 businesses right now.

Sarah Guo

Do we have 3 now?

Tuhin Srivastava

No, no. Dedicated inference, which is basically customer-owned inference—your SLA is your SLA. We have shared inference, which is shared inference with shared SLAs, and we have a training business. I'd say 95% of the tokens today are on the first business, and almost all of them involve the customer making some modifications to the model with their own data, specialized for the use case. What's even more important is that they might be compiling in different ways. No one is just running the vanilla open-source weights. You might be customizing it for quality, but you must also be customizing it for performance.

Sarah Guo

You made an acquisition of a research team a few months ago. You've mentioned post-training customization. What was the rationale behind the acquisition, and what is that team doing today?

Tuhin Srivastava

Yeah. The rationale for the acquisition was that we're infrastructure and product people, and now we're really good infrastructure people, but we didn't have much of a research capability ourselves. What we saw was the market moving heavily toward post-training, and that we could accelerate the market itself with post-training resources, either productized or even just as resources for that market.

Parea was a company that was a Baseten customer. They were post-training models and running them on Baseten. I think what they realized was that they would eventually need to become an inference company. What we realized was, “Hey, we really needed that expertise, too, because it represents a way for us to get closer to the customer earlier and be able to support them more.”

It just made sense as a fit—pairing them together. As I said in the opening statement, as more and more post-training models have come up, we've realized that the demand for software tools to do post-training or for post-training expertise is very high, and we're really investing in that. There's also a bunch of Australians. I like to think that we had a bit of alpha there.

They're working with all sorts of customers, and it's also very interesting. We were doing a lot of research on the performance side and less so on the post-training side. As we've started to do a lot more research on the post-training side, you start to see how linked inference and post-training are. Even when you think about stuff like quantization—when you should do that, and how training the model affects how you need to quantize for inference—it has become very apparent how paired these problems are.

Sarah Guo

Mm-hmm.

Tuhin Srivastava

More and more, we realize that post-training and inference are both sides of the same problem. Inference will ideally beget more post-training: inference creates data, you do evals, and you can now post-train on the reward function that you found with those evals. Hopefully, this will play out entirely.

Sarah Guo

Plenty of folks from Anthropic and OpenAI—Sam, Greg, et cetera—have said in recent months that inference is super strategic, inference talent is strategic, and capacity is strategic. So, between that and post-training, these are very difficult-to-gather capabilities.

Tuhin Srivastava

Yeah.

Sarah Guo

I imagine that lots of your customers go to you for advice on how to make this progression toward custom models. What do you tell people about the lifecycle and when they should invest in that?

Tuhin Srivastava

Yeah, I think it's: go prove to yourself with the best-in-class model that you have something worth optimizing. A lot of customers, if they come to us—there was that meme from 2 years ago: “It feels like there's no GPUs pre-product-market fit.” It's like, no post-training pre-product-market fit as well.

Sarah Guo

Yeah, yeah, yeah. Some people that you're working with are very, very at scale first.

Tuhin Srivastava

Yeah, they have a user signal that they know how to optimize. They've shown that they can serve customer value and that they have something special around that value. Once you have that value, it's like, “Okay, now how can I do that better, faster, and cheaper?”

The idea is that if you need to be very good at customer support, maybe you don't need to be that good at coding, and a specialized model might be a better fit for that problem. You can do it better, faster, and cheaper.

Sarah Guo

What about the capacity side? You started with unifying capacity across all the clouds and new clouds. How do you think about this when you keep talking about a supply crunch and a multiyear supply crunch?

Tuhin Srivastava

I think there's so much narrative around the supply crunch.

No matter how much we hear about the supply crunch, I don't think people realize how bad it really is. There is very, very little slack compute available. We run pretty large clusters ourselves, and we run them at uncomfortably high utilization. We're not saying we're in the mid-90s most of the time, but it's close.

We now sit in 18 different clouds. We have 90 clusters around the world across 18 different clouds. Initially, we built this technology to create one runtime fabric that spans all these different clouds and abstracts that away from our customers as a way to think about reliability, latency, failover, and all these things that we think are going to be very important for mission-critical use cases.

That same technology—our ability to get compute wherever humanly possible—has been really, really helpful in our ability to get supply. What I mean by that is, we can be introduced to a new provider in a different country and have it up and running with the whole Baseten inference stack—

Sarah Guo

—as part of the fabric.

Tuhin Srivastava

Part of the fabric in half a day? Half a day, maybe less. Even for us, it is hard to grow. We have a 4:00 p.m. standing meeting for the company where we basically ask, "How do we manage capacity for the demand right now?"

I think the second part that people don't really understand is that there are also a lot of suppliers right now that are kind of grifty. They haven't run data centers before. They don't understand SLAs, especially for inference.

Even when there is capacity available, there's a lot of doubt. We run a lot more of this than we've ever done in-house, so it's fine. But there's probably a dozen good clouds, and I would put three or four of them in the gold tier.

Sarah Guo

Mhm.

Tuhin Srivastava

I think that just means that we're not only supply-crunched; we're supplier- and operationally crunched onto people who can run these data centers as well.

Sarah Guo

How far ahead can you actually buy capacity right now? In other words, is there any slack in the market if you buy 2 years ahead or 5 years ahead?

Tuhin Srivastava

You mean the actual contract length, or actually saying, "I want this in January 2028"?

Sarah Guo

Yeah, either one. It's more, "I want this in January 2028," or at least, "I have some visibility into my future supply."

Tuhin Srivastava

Yeah. You could buy that, but you also have to remember how quickly the market is moving. That gets balanced somewhat by the fact that the H100 is such a great chip.

Sarah Guo

Yeah.

Tuhin Srivastava

It's crazy—it's 4 and a half years old, and the price is still going up.

Sarah Guo

Yeah.

Tuhin Srivastava

Maybe it has a useful life of 9 years.

Sarah Guo

Yeah.

Tuhin Srivastava

That's good, but at the same time, yes, you can do that, but you're making a lot of bets as part of that. In terms of what's changed over the last 6 months, the term length that people want has just gone up.

If you wanted 1,024 B200s from a good cloud right now, you're not getting that for less than a 3- to 5-year contract. Right now, it probably comes with 20% of the TCV prepaid. What becomes important when acquiring capacity is that you need to have enough demand to serve it, and you also need a low cost of capital, which is actually changing the dynamic pretty significantly.

Sarah Guo

Does that impact how you think about going public as a company? Because arguably—

Tuhin Srivastava

Yeah. I think you'd go sooner.

Sarah Guo

Yeah, exactly.

Tuhin Srivastava

Yeah, I think you need—I think there was demand for that. But I think the pool will also—one of the realizations that we had recently, especially as software people, and so we don't think about this all the time, is that our business has very interesting working-capital requirements.

As a result of that, it has very interesting financing requirements, and we're not, at least right now, even going down to the—

Sarah Guo

Yeah, there's also things you could do in terms of debt or other structures.

Tuhin Srivastava

Yeah. I've learned a lot about debt recently.

Sarah Guo

Given the supply crunch, and inference being one of the top couple of markets you're going after, you have plenty of people who understand this problem and therefore some competition. How do you think about what factors create a dominant player here, or a winning player? Is it, as you mentioned, cost of capital? Is it access to supply? Is it software? Is it demand?

Tuhin Srivastava

Yeah.

Sarah Guo

Just being excellent at everything?

Tuhin Srivastava

Yeah, I think so.

Sarah Guo

Is it operations? Like a special—

Tuhin Srivastava

Yeah, yeah, yeah. GPUs as a service are not sticky. I think that's been seen. Customers generally just see that as a commodity.

Inference with the software layer included is incredibly sticky. None of our top 30 customers have ever churned. We're talking about 400% annual NDR around our business. It's very, very sticky, so I think that software layer is very important.

The optimist in me is like, "Oh, there's so much value in the software." I think we will build the best software layer for inference that exists.

As it's becoming clear now, access to inference compute is—

Sarah Guo

Yeah.

Tuhin Srivastava

—is a strategic advantage. I think that is the strategy that even the labs are going after, which is, "If we have all the compute, good luck running inference."

Sarah Guo

Yeah, yeah. In a world of constrained compute, the number one thing to own is compute.

Tuhin Srivastava

Yeah.

Sarah Guo

Just owning it in and of itself is an asset, and I think people underappreciate that.

Tuhin Srivastava

Yeah. You can't make good hot chocolate without milk.

Sarah Guo

[Laughter.] Unless you're vegan. No one wants the vegan inference.

[Laughter.] Well, I've got to ask you. People might want alternative milk, right? The H100 is a great chip. People want a B200. They want a GB200. They want, of course, tons and tons of NVIDIA.

When you think about making a bet several years in the future, do you believe that there's a multichip world? What do you think happens from a compute perspective on the chip side?

Tuhin Srivastava

Yeah. I think diversification everywhere is a good thing. In the same way I want a world of many models, I think we want a world of many, many things.

Sarah Guo

It'd be sad if it didn't happen.

Tuhin Srivastava

Yeah, and I think everyone would be sad. I will say that, to some extent, I think there will be inference-specific chips. I think you'll have decode-specific chips. We're looking at it—I mean, NVIDIA said this.

Sarah Guo

Yeah, yeah. I mean, that was a whole—

Tuhin Srivastava

It's like, you know, I think that is very straightforward and makes sense. I think people really, really, really underestimate NVIDIA's supply-chain capabilities, how good CUDA is, and the developer ecosystem around it.

To me, one of the most important things as an infrastructure company in this moment is how fast you can move, and you can move fastest with NVIDIA today. I think that is the reality. Given the scale at which they operate, it's hard to see, in the short term—in the next couple of years—how anyone will be able to compete with that.

I'm not saying it won't happen, but especially with so many of the other players, what you need to be able to compete here is for the ecosystem to form around you. If you tie up all your supply with one buyer, which a bunch of the other chip providers have done, it's actually hard for that ecosystem to form.

If you're a big lab and you have a proprietary deal with one chip type where you get 90% of the supply, it's actually in your best interest to make sure you get 95% of the supply, everything gets built for you, and no one else can ever use it.

Sarah Guo

When you think about reacting to the market, what do you think is happening with the actual workloads that you have to invest in? Obviously, code agents and long-horizon agents have become a big deal over time. People talk a lot more about CPU compute. Video inference is different. I don't know if it's that—

Sandbox is like—what’s important for you guys to invest in now?

Tuhin Srivastava

Yeah, look, I think for us, all the runtime stuff is obviously very important. That means what chips we run on, how we run, and what kinds of workloads we support. Do we get very good at diffusion transformers? Yes. Coding agents need sandboxes, which could be called sandboxes.

There are all sorts of new speculative techniques to get faster inference. We need to do that. Even things like KV-cache-aware routing—that stuff is a bit old now, but we need to continue to be very good at it. We’re also somewhat disentangling prefill and decode and starting to treat them as separate problems. I think that’s something we’re very focused on, and we’re seeing massive gains there.

That’s at the runtime level. Beyond that, everything we think about is how to create more of that loop between inference and post-training, because we think that just begets more inference. We will build or partner in almost everything. We’re going to work with the best eval companies in the world to make sure that’s very well integrated, like Braintrust, into and around Baseten.

On the sandbox side, we will partner to build the best sandbox experience that will exist. Then we’ll create the best training APIs to make it so continual learning becomes somewhat of a solved problem, not just a discrete thing. I think that’s the core Baseten product thesis: How do we build that loop?

Everything else around that becomes: How do we make sure we can do everything we can to ensure that it gets as big as possible? That’s access to compute. That’s our own infrastructure. We need to make sure we can get compute anywhere and that we have access to our own compute.

Then I think it’s all the primitives that come after that, which become incredibly margin-accretive, both for us and our customers. That’s stuff like sandboxes and async batch inference. How do we drive utilization by having a first-class batch-inference experience?

To me, this is what an inference cloud looks like: You are very good at inference, and then you start to do all the things that are tangential to, or loop into, inference. You partner when necessary and build when necessary. But we really do want to own that core inference story, then go down to unlock supply and create margin, and go up the stack to unlock value.

Sarah Guo

What would surprise people about some of the issues you discover only at scale? I was surprised when you guys ran into scale limitations—fundamental limitations—with some of the hyperscaler products that you were consuming. I kind of think of the AWS and GCPs of the world as supporting infinite scale.

Tuhin Srivastava

Yeah. I mean, I think very large companies that run services at big scale probably see the same stuff.

Sarah Guo

Mhm.

Tuhin Srivastava

Yeah, all the edge cases just become—you actually experience them. You experience them.

You start seeing things like—yesterday, for the first time ever, we saw a kernel panic. That only happened because a Fluent Bit worker was creating too many logs, the scale was too big, and everything was going into one node. It was happening 2 times at the same time with 2 different workers, so you see all these systems-level and kernel-level problems.

But I think the craziest thing is that you start to see, with LLMs, that these runtimes are pretty immature. Even how we use KV cache is probably a little less sophisticated than most people realize. We’re starting to see the limitations of the current and next set of primitives that need to be built from a scale, security, or performance perspective.

I think it’s really at the runtime level and the systems level. The edge cases are a lot more systems-level than they are LLM-specific.

Sarah Guo

What are the things that keep you up at night?

Tuhin Srivastava

Capacity.

Sarah Guo

Quick answer. Yeah.

Tuhin Srivastava

Yeah, I think capacity. Probably just that this market is so big, and it represents a moment when you should be as aggressive as possible. We’ve grown a ton this year, in the last 12 months, in the last few months, but the answer is always just: Go bigger, go faster. I think that’s really, really fun.

It’s also a little exhausting, and we are all in somewhat uncharted territory in terms of how fast and how big you can go and how things can get. But I think we need to do that to get the amount of value that we want to get out of LLMs in the next 5 to 10 years.

Sarah Guo

Or we have to invent a lot of new stuff.

Tuhin Srivastava

Yeah.

Sarah Guo

Maybe if we just talk a little bit about what you’re learning from scaling, 30× is an aggressive thing to go through as a company. You’ve brought in a lot of really amazing talent, like Danny, Samir, and Stephen Day—folks on both the technical and go-to-market sides. What do you think is working about how you are recruiting and scaling, or what’s your philosophy on that?

Tuhin Srivastava

We were very, very flat until 8 to 18 months ago. I remember when I worked with Elad a lot, actually. A lot of it is that you just need leaders. It’s actually so contrary to everything engineers think. They’re like, “Oh—”

Sarah Guo

It’s all overhead.

Tuhin Srivastava

It’s all—everything is overhead.

Sarah Guo

You once told me, I think, that you didn’t—you were like, “Hey, Sarah, what about we just have engineers instead of salespeople?”

Tuhin Srivastava

Yeah. Bad.

Sarah Guo

Everybody learns that.

Tuhin Srivastava

Everyone—we all know about it. I remember, you said it so clearly at the time, Elad, and I think that’s what we noticed: Actually having a leadership team that you can trust is so important.

I think the 2 or 3 things that I’ll say are: You want people where you can give them whole problems. If you feel like you’re micromanaging, if you feel like you need to be involved in everything, I think that’s a bit of a cop-out as a founder, because you’re just like, “I just need to be involved in everything.” No, you probably just don’t have the right people.

I think the second thing is: Be very, very clear about what you’re optimizing for. When you’re very, very clear about what you’re optimizing for, the people who fit become apparent, and the people who don’t also become apparent.

If it’s something generic like, “We want the smartest, hard-working people,” you can’t do much with that. What we cared about was, “Hey, actually, we don’t care about a lot of people who’ve done this before. We care about first-principles people who think from first principles.” Work has to be a high priority, but they also have to be very kind and nice and care about the collaborative environment. We don’t have a hero culture. We’re very low-ego. If you need a manager, it’s probably not the right place to be.

Once you have that clear rubric, the people who fit into it become very apparent, and the people who don’t fit into it also become very apparent. We’ve hired amazing people, like you mentioned, but I think what’s a lot more interesting is that we haven’t had a ton of unnecessary turnover. People tend to stay because we’re very clear on what we want. It took us a while to get there, though.

Sarah Guo

What about the idea of an operations culture? We were talking to Alyssa and Henry about this, and she was like, “Well, the hard thing about cloud is actually just operations. I slept with a pager under my pillow for a decade.” I don’t think I’ve seen you detached from your Slack channel.

Tuhin Srivastava

Yeah, my phone is buzzing right now. I’m getting anxious.

Sarah Guo

And you’ve been concerned before: Do people get it? What is distinctive about that?

Tuhin Srivastava

I think, one, if you have worked at an infrastructure company—for example, we were once in a meeting with a bunch of AWS executives, and these were very senior AWS folks. All their pagers went off multiple times during our 45-minute meeting. It’s very much a cultural thing.

Our infrastructure can’t go down. I think my co-founder, when his pager goes off, has a 7-year-old who says, “Is that a P0?”

You just have to get used to it. That’s the culture you live in, and it changes the speed, but it also becomes a cultural thing.

Tuhin Srivastava

I think it rejects people who don't fit into it very quickly.

Sarah Guo

Like engineers who avoid pagers.

Tuhin Srivastava

Yeah. When we have pagers, everyone's on call. There's been a joke that they may as well be a siren that goes off in the office.

Sarah Guo

People have been talking ad nauseam in the AI community about Jevons paradox.

Tuhin Srivastava

Yeah.

Sarah Guo

It's really a question around price elasticity and availability: if you decrease the cost of a good—say, intelligence as a good—people actually consume more of it.

Tuhin Srivastava

Yeah.

Sarah Guo

The personal or business ROI of it goes up, and the demand for it goes up, not down. Do you see this, and are you working against yourself trying to make these models more efficient? People just use them more or less?

Tuhin Srivastava

Yeah, I think you have to think about this from a developer perspective and a consumer perspective. I think consumers just want the best answers and the best experience, which is somewhat governed by more intelligence to some extent.

From the developers' perspective, they would insert more intelligence if you made it cheaper. They would insert more intelligence anyway, but if you make it cheaper, they'll insert a hell of a lot more intelligence. You see this with agents and stuff. Agents are just longer-running now, and I think that's what we have seen with the cost of inference going down: folks are just like, okay, we can run this for longer, we can make it do a bit more work, and we'll get to a larger end result.

I think as compute scales from an inference perspective as well, and we're seeing that with almost all our customers, they either start with, “This is the quality of answer I need to get to, and this is the amount of inference I need to do to get that,” or, “This is the base-level model that I can start with or work with to get there.” I think the more we drive down the cost, what they realize is that more intelligence just means a better user experience.

Sarah Guo

I just want a better answer.

Tuhin Srivastava

Better answers, better experiences, more dollars, more dollars, even more revenue. So, yeah, I think inference going down just begets more. But I truly think we're kind of in a world that is the last market, right? Even if there's AGI, all that's left is inference.

Sarah Guo

So, you do not see in your customers a, like, “This answer is enough and this action is enough” dynamic?

Tuhin Srivastava

No. No.

Sarah Guo

Yeah, it's going to keep going for a long time, I think. How do you view all this evolving toward the future? This seems like it's going to be one of the biggest markets of all time. We have this massive shift where we're moving from software and seats and digitalization into actual intelligence—selling units of cognition, selling agentic workflows. What does this all look like in a couple of years? What is your view of this future world?

Tuhin Srivastava

I think for consumers, it's the best possible thing, right? Everything is somewhat smarter. You get better care because your doctors have access to better tools. There's all this stuff about there being fewer software engineers. I think we just build more software.

I think we just build a ton more software, and we're not slowing down hiring software engineers; we're just building more things. For consumers, that just means better tools, more software, all those good things.

Sarah Guo

Like everybody has their own team for everything, right? You have an agent that helps with your doctor. You have an agent that helps you learn stuff. You have an agent that helps you organize your life.

Tuhin Srivastava

It's concierge. It's concierge everything.

Sarah Guo

Yeah, concierge everything for everyone.

Tuhin Srivastava

Yeah, and I think what that means is amazing. I think that's great. And education, same thing: you have concierge education. You get personalized access to everything.

I think when you go one step back to how it affects developers and companies, if you don't embrace this, I think it's the extinction moment for a bunch of folks. Everything needs to change, and I don't think that means that product design needs to figure it out. I don't think that's a thing.

I think what's more interesting is that all these workflow and software companies need to figure out what the intelligent, or intelligence-inserted, versions are that drive all that user value for those consumers we talked about.

Sarah Guo

Yeah, very exciting. Thank you so much for joining us today.

Tuhin Srivastava

Yeah, thanks, guys.

Baseten CEO Tuhin Srivastava on Custom Models, and Building the Inference Cloud | BidClub