Lukas Biewald
In a world where we're shifting toward models doing everything, with all the value being at the application layer, if there is AGI, the only market that will exist to some extent is what the model needs to do. And the model can only run on infrastructure.
1. Infrastructure and Runtime Challenges
There are 2 different parts of infrastructure. On the infrastructure level, I have a workload running across 5, 10, or 100,000 GPUs. How's this thing going to scale? The second is the runtime-level problem: how fast do these models actually run on a given GPU? Who's using open source and who's using closed source in your experience, and how do people think about that trade-off?
Tuhin Srivastava
I think everyone's on this curve of maturity. A lot of people start with custom models and open-source models, where they're doing their own things. I think, at the lower end, people start by going with Anthropic or OpenAI, and they have these great models that can do a lot, but they're either too expensive or have a lot of reliability issues themselves.
The third piece is that our customers care that we aren't just piping all this data off to someone who's going to train models on it. That matters to us.
2. The Journey of Baseten
Lukas Biewald
You're listening to Gradient Descent, a show about making machine learning work in the real world. And I'm your host, Lukas Biewald. Today I'm talking with Tuhin Srivastava, who is the CEO and founder of Baseten. Baseten is currently one of the fastest-growing companies in the inference space, which itself is a fast-growing space. It's phenomenally successful and just raised a giant venture round, but I was especially interested to talk to Tuhin because he's been at it for a lot longer than people think. I think you see those graphs where you go from 0 to 100 million ARR in a couple months, but you kind of don't see the years that I've known him, when the growth wasn't there and they were looking for what to do. As someone who has pivoted myself and gone through a lot of struggles, I was excited to talk to him about how he worked through that and how he eventually got the company growing. I was also really excited to talk to him about inference and how it works. I feel like that changes every couple months. You know, everyone calls the space a commodity space, but it certainly doesn't look like a commodity space with the growth you see in certain companies like Baseten. So this turned out to be a really interesting interview. It went in a lot of directions I didn't expect, and I hope you enjoy it.
I was thinking about this interview. I have a lot of technical questions that I want to get into, which I think will be really interesting, but I thought it might be fun to start with a little inspiration for people who are feeling stuck in their companies.
I feel like you have this amazing story of suddenly taking off after a couple of years of not being the hot company, and I thought that would be a great story to share with everyone who might be feeling at an impasse. Could you tell me about your journey?
3. Differentiating in the Inference Market
Tuhin Srivastava
Yeah, it's actually quite interesting. Lukas, you could probably empathize with this. I've been working in early-stage companies since 2011, actually. That was a series of small companies where I was a very early employee. Most of them didn't work. They really didn't work.
Then I had another company from 2015 to 2018. I did make friends along the way, and that was mostly the upside of it.
Even at Baseten, which was actually founded in 2019—honestly, not that long after Weights & Biases was probably founded—the world was just very different in terms of the state of machine learning and what people thought about machine learning. Even who we were targeting was different, to some extent.
We were trying to work with data scientists working with small models for internal use cases, and that's what we thought the important market was. Obviously, in 2022, everything had changed. I'm sure we'll talk about the last 3 years in this interview.
The reason I bring up the fact that we've been doing this for 15 years is that I don't think the inputs have changed that much since 2011, at least—definitely not since 2015. My motivation toward work, how I think about work, and even what I do haven't changed very much. The velocity of everything around it has changed.
Every day, we still come in, I check the support channel, and I go dig into some product problem or user feedback that we've gotten. I still go and try to sell Baseten to a bunch of people.
I think the only change was that a market arrived and we didn't give up. There was plenty of capital to keep us going. My inspiration, honestly, is that the desperation is even greater now than it was when things weren't working.
I don't think it should be that inspirational, but my inspirational thing would just be: don't give up. Work with people you like and follow markets.
I think one of the big things that we did in the early stages of this company was not scaling too quickly. When markets shift, fundamental truths change underneath you, and the weight of the company is probably what gets in the way of you being dynamic with those things.
From 2019 to 2023, when we raised our Series B at the end of 2023, I think that's when I thought, “Oh, we have something real here.” We were still 18 people, and I think we were lucky in that way.
Lukas Biewald
Totally. Do you know my story with CrowdFlower? I also had a long slog before the market took off. It takes real skill to keep a company cohesive. Even at 18 people, I think it takes real skill to keep people going.
I think there's an incredibly underrated skill—maybe the core entrepreneurial skill—to jump onto an opportunity that's in front of you. Could you talk about what you were seeing, how you did that, and how much of a pivot it was versus how ready you were for it?
Tuhin Srivastava
A lot of people think it was a pivot. It kind of wasn't. We already had a lot of the infrastructure built for it.
From 2019 to 2022, obviously no one was calling it inference back then. We called it model deployment and model serving. We were going around telling everyone that model serving was a commodity. It was easy. We'd give that away for free. But we had built all of it.
What really changed was that, if you think about those 3 things I mentioned initially—data scientists, smaller models, and internal use cases—you went from small models to big models, which was one massive market expansion. That made serving actually quite hard.
Then we went from internal use cases to production, which meant that SLAs and infrastructure mattered a lot. The third thing was that we went from data scientists to engineers, and all of a sudden you had people who could actually exercise agency to change the things in front of them.
Around those 3 shifts, it was more of a refocusing as opposed to a pivot. We had other parts of the product that we just killed and said, “All right, we're going all in on this,” but the fundamental product was still there.
Lukas Biewald
That makes sense. What happened in late 2023? Was it 1 customer that kind of pulled it out of you, or what was going on?
4. The Impact of ChatGPT and Stable Diffusion
Tuhin Srivastava
I do give us credit here. I think we did a good job. There were a couple of things. It was the end of 2022, when ChatGPT happened.
That didn't actually have that big an impact on us, except that it made people pay attention to AI. At the end of 2022, the narrative really shifted around AI.
I think ChatGPT set up 2 things that made it very interesting. First, it set the standards for consumers of what they expected out of AI products. People building things who are now our customers had this standard that they had to beat.
It's pretty crazy that the first mainstream AI product was ChatGPT because it was so good, even from a user-experience perspective. If Stripe was the first company to ever do payments, that would be kind of unreal.
On the other hand, ChatGPT also started creating these developer APIs, which again were high-quality APIs.
I think the standards for developers were set very, very high, and that was one moment around ChatGPT. The bigger moment was probably Stable Diffusion, to be honest, because all of a sudden you had an open-source model that was approximately as good. It still wasn't as good as DALL-E 2 at the time, or whichever DALL-E we were on, but it was approximately as good. All of a sudden, people started paying a lot of attention to the ecosystem that formed around it. I think that injected a lot of excitement.
I'd say what's really cool is that there were 2 customers who came to us and said, “This is going to be really cool.” The first one was Patreon, which was experimenting with Whisper for generating subtitles. We were like, “Oh, that's awesome. That's very cool.”
The second was this guy who was a friend of ours called Seth. He had taken a year off between companies, and I think he called it his “year of yes.” As part of that, he wasn't supertechnical, but he figured out how to fine-tune Stable Diffusion to generate music. It was called Riffusion, and he put it on Hacker News, and it kind of took off.
Lukas Biewald
Oh, I remember Riffusion totally. Yeah.
Tuhin Srivastava
Yeah. At the time, I remember he came to us the day before his launch and said, “I might need a lot of GPUs.” We were like, “Yeah, sure you will. You'll probably need 5 A10Gs.” I think he ended up scaling up to 100 or 150 A10Gs, and we thought that was mind-blowing—how many GPUs that was at the time.
I think that was a real spur, like, “Oh, there's going to be a lot more of this. Let's really think about that market as a core market.” I think we really refocused around that moment. It was just like, “All right, let's change everything.” So I think over the course of 6 weeks—
Lukas Biewald
You completely changed the surface area of the product, which is pretty cool. Wow, that's awesome. I mean, good for you for making that.
Tuhin Srivastava
Yeah, you don't have that many customers, but still—
Lukas Biewald
Still, a lot of people don't. I mean, I guess before we get into the technical stuff, you kind of immediately jumped to something I always think about. I'm the same person. I remember taking a lot of criticism when stuff wasn't working and sort of getting praised for almost exactly the same stuff when things were working. The inputs are the same, and I really relate to that.
Is there anything more you want to say about that? Do you feel like you've learned something from watching that experience?
Tuhin Srivastava
Yeah. I think everyone needs to just figure out what drives them, to some extent. At my core, what I'm driven by is not really having a successful company as much as working with people you care about on problems you care about, and everything else will kind of change.
I think it's really easy to focus on the input when you're asking, “What do I need to do?” I was like, “Well, I need to go hire more people I like, and let me go find hard problems to solve.” I'd say where I get stuck—and I don't know if you relate to this at all—the darkest moments are almost when you're trying to do company-building as the core focus, if that makes sense. It's like when you're—
Lukas Biewald
No, what do you mean by company-building? Like, when you start chasing quotas and revenue targets? Because with this artificial Gregorian calendar that we've come up with, or like—this actually happens a lot with go-to-market teams in startups—is when you start applying generic playbooks to things.
You're not really coming at it from—you’re almost trying to treat everything like science, where I think a lot of it is still the art of it, which is, yeah, hire good people, work on stuff you care about, and really optimize around that. But that can be a bit idealistic at times as well.
I love that attitude. I think I relate to that when things are small, like, “Oh, this is just my project.” It's kind of hard to hold on to—or, in my experience, it's hard to hold on to—that feeling as a company grows. How big are you today, and how are you still feeling like it's a fun project with people you care about?
Tuhin Srivastava
Yeah. Look, we're 110 people now. We have our company off-site next week, and we do it every 6 months. We were less than half that at our last company off-site, so it's growing pretty fast.
I think I'm lucky in that I am still a bit like that. One of the really great things about Baseten is that my 2 co-founders—I’ve known them between 15 and 25 years. These are very close friends of mine. A lot of the people who were early in the company are still around. Even our investors, like Sarah Guo, who is on our board—I’ve talked to her every day for 6 years.
Lukas Biewald
That's amazing. Really? Every day?
Tuhin Srivastava
And so, you know, I think that keeps that stuff alive. I've also learned this from another founder: I kind of refuse to do things that I don't want to do.
Being a founder is actually quite challenging, just because all the problems, to some extent, fall upon you at the end. But I still take a lot of agency around it. I will only work on stuff that is business-critical for the company or stuff I want to do. It's okay if I skip meetings. That's the only thing I ask of myself and for myself: that agency. I think that allows it to feel a bit more like that.
I'll caveat all this by saying that everything is really fun when you're growing a lot, and the last 2 years have felt like that.
Lukas Biewald
Yeah. Okay, so we're taking this in a much different direction than I thought I would ask. What's the thing that you don't do that you think someone like me would feel most guilty about not doing?
Tuhin Srivastava
I skip 1-on-1s all the time.
Lukas Biewald
Nice.
Tuhin Srivastava
Usually, they're of the form, “Hey, if there's anything pressing, come catch me at my desk.”
Lukas Biewald
It's funny. I feel like every founder CEO that I've gotten to know well enough actually says the same thing, all the way up to Jensen. He was actually on this podcast saying that, and that kind of gave me freedom to feel like I could do that.
But it's funny because I feel like it was Ben Horowitz, years ago, who said, “It's just outrageous to miss 1-on-1s,” and everyone just started feeling guilty about skipping 1-on-1s. But if it drains energy and there's nothing to talk about, it certainly doesn't make sense to do a standard meeting.
Tuhin Srivastava
And maybe there are things that are rituals, which are important, and there's stuff that's just performative at times. The idea that you need to meet with someone you work with every day for 30 minutes to reflect or problem-solve—I think it's a bit much.
We have our own version of that with me and my co-founders. We have this weekly meeting for 45 minutes on Friday mornings. It would seem like a 1-on-1, but it's purely—we just shoot the shit. We kind of just do whatever we want in that time. Sometimes it's work, sometimes it's not.
That part's a lot more important to me from a ritual perspective than the performative part of, “What are the 3 things that we need to talk about this week?”
Lukas Biewald
And yet you have a daily 1-on-1 with your investor Sarah. What are you guys talking about there?
Tuhin Srivastava
[Laughter.] I like Sarah so much as a friend at this point. We end up just blurring the boundaries between everything as founders. I lean 100% into that. There are very few boundaries in my life, for better or for worse.
5. Operational Discipline and Company Culture
Lukas Biewald
Interesting. Well, I feel like someone might listen to this and think, “Wow, Baseten is really lacking some kind of operational discipline.” Do you feel like that's a fair criticism?
Tuhin Srivastava
No, I don't think so at all. I think we have a lot of operational discipline where it matters. For example, in our sales force and the way we do sales, it's very, very customer-centric.
I'd say if you went and talked to our customers, they'd be like, “Hey, no one pays as much attention to us as Baseteners.” I hope they'd say that.
I think a lot of the things that we are trying to solve for are: How do we do the most impactful thing possible without it draining energy? For us, at least for engineers, a lot of the time those things are very much about structure—for example, too many meetings, too many one-on-ones, too many referrals. On the flip side, there’s a healthy tension between what that needs to look like in every company, which may be different, as opposed to just the industrialization of the product, if that makes sense.
Lukas Biewald
Totally. Totally. Yeah. All right. Switching gears to what you guys do, I was looking at your website recently, preparing for this, and you do a lot more than I knew. I think you’re expanding the product surface area lately, but I sort of think of you as an inference company. Is that fair?
Tuhin Srivastava
That’s fair. I’d say our north star is inference. That’s probably the best way to put it.
I know a lot of the founders in the space, and I’ve had a lot of friends come in and out of that category. I think people feel like it’s commoditizing, and I think you even said you thought of it as a commodity. Is inference a commodity service? What are the differentiators?
You need to step back and think about inference of what. If you think of generic inference of an open-source model—if all we did was serve Llama behind an endpoint—I think that’s 100% a commodity. Given some quality benchmarks and performance benchmarks, people will go for the lowest price. I think there’s still a lot to differentiate on from a performance perspective, but in the fullness of time, I believe that generic inference, or vanilla inference, of vanilla models is probably commoditized.
The truth for us is that dedicated deployments of custom and fine-tuned models are not commoditized. Every workload is slightly different, and every model is slightly different. Serving those models, and the tooling around that, is quite differentiated.
We differentiate on 3 different things today. One is infrastructure: What does it actually take to run this reliably at scale? We take a lot of pride in the fact that we have many nines of reliability. We don’t go down when the clouds go down. We’ve designed everything to be fault-tolerant and pretty elastic, and we align our infrastructure around the customer’s needs as opposed to a cloud’s needs.
That’s the first piece, and there are a lot of infrastructure problems we can talk about here. The second piece is performance. I think a lot of inference providers differentiate around performance. We are very performant; in any head-to-head, we don’t generally lose on that.
Lukas Biewald
Now, when you say performance, do you mean—
Tuhin Srivastava
Tokens per second?
Lukas Biewald
Tokens per second, latency. Not the quality of the model.
Tuhin Srivastava
Not the quality of the model. We mean, how fast does this move?
There’s a lot of talk around how fast these things move. If you normalize over time, what you’ll see is that, if you look at the history of software, open-source runtimes seem to win. The best-quality, best-performing runtimes will eventually just be the open-source ones, because that’s how developers want to adopt things and where they want to push them forward.
You still need to meet a minimum viable performance level there. We think we’re A-plus, but we don’t think that’s going to win over the market in the fullness of time. You need to be very good in that class.
The third one is developer experience and platform. We think of this as a software problem. Most of our customers have between 3 and 20 models deployed on Baseten. Most of them are serving their customers with different hardware, different scaling needs, and different traffic patterns, and we provide a lot of software to manage all that.
We think that is very differentiated, the same way that a lot of other CI/CD tools or software tools add value. So it’s infrastructure, performance, and developer experience. When we look at the intersection of those 3, that’s where we think a massive inference company exists in the fullness of time.
We do other things, too. Those things are very important, but they happen in the service of doing inference. We want to be the production-grade inference company. We did this marketing campaign recently with these obnoxious buses all over the city, which I thought was great, but the message was, “Inference is everything. Inference is Baseten.” That’s how we think every day when we come to work.
We have a training product where we create workloads for training, but we train those things so they can be served on Baseten, to some extent. We provide fine-tuning scripts so those models can be fine-tuned and eventually run on Baseten.
Lukas Biewald
Okay. But there’s infrastructure, performance, and developer experience. Why do you make this distinction between an open-source model and somebody’s custom or fine-tuned model? I would think that with your newest open-source model, all those things could differentiate: reliability, performance could potentially be better, and developer experience probably means you’re managing a bunch of these things anyway.
Even with a fine-tuned model or something, it seems like those could also differentiate, right? I mean, I could imagine that if everybody’s using vLLM or something to run their models, then maybe it doesn’t matter.
Tuhin Srivastava
Yeah, I think you’re right. Today, that is the case. You can differentiate on those things today, and we do. But in the fullness of time, I do think a lot of those things will become standardized. We’ll know, for a given model type, how to run it very fast.
Reliability will have to be there. The unreliable companies won’t exist, and what will be left are reliable endpoints that are pretty fast, with switching costs between them that are pretty low.
Lukas Biewald
Wait, so are you telling me that in the fullness of time, you don’t think that you will differentiate?
Tuhin Srivastava
No, no. Sorry, maybe I should separate this differently. I’m talking about dedicated and shared. When I think about shared endpoints, I don’t think there’s much differentiation in the shared market. I mean, there’s a lot of differentiation in the dedicated market.
Lukas Biewald
I see.
Tuhin Srivastava
Yeah, sorry, maybe that wasn’t clear. If you want to use a Llama endpoint that a bunch of other people can hit, and it’s providing the same tokens per second at the same quality, to me, they’re very much the same. That’s the shared-endpoint business.
Ninety-nine percent of our business is dedicated capacity, where people have single-tenant, not multitenant, endpoints. Only those customers are using those endpoints.
Lukas Biewald
That’s actually a different axis: dedicated versus shared, although I guess you wouldn’t have a shared endpoint with a custom model.
Tuhin Srivastava
Yeah, exactly.
Lukas Biewald
But, okay, in a dedicated situation, why is it more possible to differentiate?
Tuhin Srivastava
Because all those workloads look very different. Firstly, not everything is a language model, but let’s assume everything is a language model for now. I’m setting up—I’m deploying—a model for use case X. Another customer comes along and says, “I want to deploy it with this,” and their inputs are slightly different, their outputs are slightly different, their SLA is different, they can only run capacity in a certain cloud, or they need things to be HIPAA-compliant in a particular way.
All those things form different attributes of the workload, and that gives you ways to differentiate in serving those customers.
Lukas Biewald
Interesting. So when I come in and I’m talking to you, you’re having a conversation about my requirements and then setting something up custom for me?
Tuhin Srivastava
Yeah. Or we’re giving you the tools through our software to configure those things yourself. At Baseten, when you come in, you have full control over your runtime.
We can give you the tools to make your runtime better, or we can give you some default runtimes, or we can look at vLLM or TensorRT-LLM and tell you, “Here are some good configurations to run those.” But you can do whatever you want there, and you might change those configurations based on your runtime—
Lukas Biewald
In a way that would make sense for another customer?
Totally. Totally. Yeah. Okay. So could we dive a little deeper into how inference works on a modern LLM? First of all, I guess let's talk about how it works and maybe what the possible optimizations are.
Tuhin Srivastava
Totally. I think there are 2 different parts of inference. You need to think about infrastructure-level problems and runtime-level problems.
At the infrastructure level, I have a workload running across 5, 10, or 100,000 GPUs. How is this thing going to scale? You need to set up your infrastructure to allow inference to be done. What that might mean is setting up your infrastructure in a way that allows the same user to go back to the same GPU to reuse the KV cache for their workload as much as possible.
There is an infrastructure problem we see with some of our customers: they can only use GPUs in a certain region because they want to minimize the number of hops in the world. You have these Cloudflare-esque or Vercel problems around running inference workloads. In that way, it's very much an infrastructure problem, and we've done a lot of work there.
You need to acquire capacity. What happens when you need 2,000 B200s and one cloud can only give you 500? What do you do? How do we give you all the tools to solve that?
I think the second set, which is a bit more research-y, is runtime-level problems: how fast do these models actually run on a given GPU? That's where things like vLLM, TensorRT-LLM, and SGLang come in. That's where some of the proprietary runtimes from other companies come in. That's even where the chip-level companies start to come in, saying, “Actually, at the hardware level, we've changed how inference will be done.”
6. Optimizing Inference Performance
The question then becomes: what makes this challenging? There are a couple of things that matter here. I think people care about utilization, and they care about speed, and there's somewhat of a trade-off there.
Lukas Biewald
What do you mean by speed?
Tuhin Srivastava
The general metrics that people care about are time to first token, which is how quickly the first response comes back. The second piece that people care about is time per output token. After that first token, how long does each additional token take to come back?
The third one is throughput, which is the memory question: how much capacity can we put through this without degrading performance? The last one is cost per token, which is the cost of the underlying hardware and the KV cache, and how well you're using that.
There are lots of different things to optimize here. With time to first token, that's really dominated by the prefill, which is the first forward pass through a neural net. Time per output token is dominated by the decode, which is the repeated single-token steps that are happening. Throughput is largely about memory bandwidth and quantization: how much memory can you use, and how many FLOPs can you process?
7. The Role of Hardware in AI Inference
The cost per token comes back to the cost of the hardware, how big the weights are, how much KV cache you're using, and how much work you can reuse. There's so much optimization that can go into it. When you come to some of these open-source frameworks that I'm sure you're familiar with, and your listeners are familiar with, like vLLM, SGLang, and TensorRT-LLM, they're all coming up with their own frameworks to optimize across these things.
Whether that's quantization, using FlashAttention-fused kernels, continuous batching, or speculative decoding, they all have the same research-y mechanisms to do that. These problems at the runtime level today are really, really challenging, and they're still pretty research-intensive. Oftentimes, we're pushing things that are less than a week old into production, which is pretty terrifying.
What's probably even harder right now, and I'm sure you've heard of this, is that the talent that knows how to do this is pretty limited. We're competing with some pretty crazy people for that talent.
Lukas Biewald
I've heard of that. Yeah, I've heard that's an issue.
Tuhin Srivastava
There are a bunch of different ways to make LLMs run fast, and that's what I'd call the runtime problems. Then there are the infrastructure problems. We think of both of these things as related to some degree, but they're not interchangeable. Both are necessary.
Lukas Biewald
I guess it's kind of interesting that there are multiple competing open-source projects to do this runtime thing. Are there fundamental differences in opinion that they have, or situations where one works better than the other?
Tuhin Srivastava
Yeah. We love them all. I think a lot of it comes back to that classic difference between usability and control. Usability needs speed and control.
From what we've seen at scale, who's pretty good at running inference on NVIDIA chips? It's NVIDIA. TensorRT-LLM is the lowest level, especially with all the new Dynamo changes that they're pushing through. At scale, once some time has elapsed from when a model has dropped to when you need it—maybe not day 1, but day 90—TensorRT-LLM is by far the fastest, and you can do the most with it. That makes sense because it's the lowest level of what NVIDIA provides.
On the flip side, I'd say vLLM and SGLang both—but vLLM especially—really lean toward usability. You can use vLLM and, honestly, anyone can do inference with it, but you're going to take a bit of a performance hit. You can get around a lot of that by configuring it, but then you really have to become a master. I'd say SGLang sits somewhere in between.
The question might become, over time, why do all 3 exist to some extent? I think it's somewhat of a nuanced market, with slightly different requirements for the people who are using them. A healthy ecosystem matters. How long did it take for front-end frameworks to converge? I think we could argue that Next.js has dominated today. It definitely dominates today, but we're 30 years into that journey—or 20 years into that journey. It will just take a bit for the market to converge.
Given how fast the ground is shifting underneath us, it's really important that all these frameworks continue to exist.
Lukas Biewald
Do you work with hardware besides NVIDIA GPUs?
Tuhin Srivastava
Yeah. We are enthusiasts. We've done experimental work with everything, and we've worked with a lot of the chips coming out of the cloud. We've done some work with AMD, and we've even done some work with some of the new set of chips.
Today, what we keep coming back to is that we love them all. Again, we want a thriving ecosystem. The reliability and versatility of CUDA, along with the developer ecosystem around CUDA, are just very, very hard to overlook.
When you're pushing something out into production tomorrow, the last thing you want to be doing is working with anything besides CUDA. That being said, some of the performance and cost trade-offs we've seen with some of the other providers are really, really promising. I think it will just take a minute for them to catch up here in terms of versatility in production.
Lukas Biewald
And why is that? Because, I guess, from where I sit, I've had 5 different companies on this podcast pitching me on better inference. I probably have 15 other pitches that I've heard in my life. I go out to dinner and get pitched on, “Hey, I've got faster inference.”
It seems like a very testable thing, and the market's there. I look at that and I'm like, “Okay, there's no market risk if you have faster inference, as far as I can tell. Faster, cheaper.” So what do you actually practically run into?
Tuhin Srivastava
Yeah. It's a really good question. I was thinking about that just yesterday, actually. I think it comes down to 3 or 4 things.
First, you need the chips. Then you need the manufacturing capability to scale up. And then you need the software layer for versatility. With the chips, you need to know what an inference chip looks like. Second, with the manufacturing capabilities, we need tens of thousands, if not hundreds of thousands, of these. How do you do that at scale? And third, how do you build the API so that people can mess around with these things themselves?
I think all 3 of those things contribute to this being a very challenging problem for all the folks that you've probably had on the podcast. But I think the biggest thing is—
It’s just that everything is moving so fast. Chip turnaround times are, what, 6 to 9 months if you’re lucky. With that in mind, in the market we’re in, I can’t project more than 60 days forward in this business. Who knows what model will come out or what people will want to use?
I think the speed of the market probably just makes things harder. You see that with NVIDIA right now as well: NVIDIA has 3 or 4 different chips that people are talking about. You have the Hopper series, then the B200s and B300s, and the GB series.
I think customers are also a bit stuck. Things are moving so fast that customers are confused about what they want as well. Projecting 9 months forward is very hard, but I agree: in the fullness of time, as soon as things slow down just a bit, we’ll get there. I hope.
Lukas Biewald
I feel like “in the fullness of time” is your catchphrase. I like it.
Tuhin Srivastava
Yeah, I like it.
Lukas Biewald
Do you know, it’s funny: I was talking to a CoreWeave customer recently, and they were talking about how they don’t like to change chips. Another thing they told me was interesting: they were able to get a lot more performance out of any particular chip over time.
I was surprised. They seemed like they really did not want to be on the latest chip, so they didn’t want to move their workloads to new chips. But it seems like, if somebody’s working with you, shouldn’t you abstract away the hardware they’re running on?
Tuhin Srivastava
Totally. I think there’s something we’ll move toward over time, which is big, small, extra-small—T-shirt sizing for compute. I think that’s the case.
Again, everything is moving so fast that you can’t just take a workload that’s running on an H100, apply it to a B200, and say, “Oh, look, now I have 3× the speedup.” Unfortunately, it takes a bit of finagling. The FlashAttention versions need to change to support FP4 or this new form of hardware.
I think we will abstract that away. Again, everything changes so fast underneath us that it’s kind of hard to do that today. What customers want changes from a week-to-week perspective as well.
Lukas Biewald
Do most customers show up saying, “Hey, I really want a GB200”?
Tuhin Srivastava
A lot of people do. A lot of engineers just want the latest and greatest hardware all the time.
But we actually see 2 sets of customers. We have customers who show up saying, “Give me the best hardware I could possibly buy. Give me the Ferrari. I want the Ferrari.”
The majority of our customers, especially the ones we want to sell to—companies either in the enterprise or selling to the enterprise—are a little more pragmatic. They say, “I need to solve a problem for my customer. This is my workload, this is the size of my model, and this is the SLA I need to meet. Could you help me get there?”
That’s probably the better way to think about the problem, and it fits back into what you were saying about there being some lock-in to old hardware, especially if you’re meeting the SLAs you care about.
Lukas Biewald
I’m really amazed to hear that in your case because I feel like, in my limited experience with this stuff, there’s nothing fun about being on the latest chip with the least support and the hardest setup. I’m surprised people want that. Do they think there’s going to be a better performance trade-off, or is it more of an emotional connection?
Tuhin Srivastava
We have a customer we work with that’s doing hundreds of millions of tokens a minute, which is a lot of throughput. For those really, really high-volume use cases, the performance trade-offs are huge.
Even with a lot of video workloads, B200s provide a pretty significant 40% to 50% speedup right off the bat if you can get them running. Those trade-offs can drive a fundamentally better user experience for their end customers with better hardware.
Lukas Biewald
I see.
Tuhin Srivastava
Yeah. Again, the game comes back to a lot of that Anthropic, OpenAI, and Google stuff. From a speed perspective, they’re setting the bar for what custom and open-source models need to be able to hit, and they’re all pretty good at that piece—at least the speed part.
8. Market Dynamics and Future Predictions
Lukas Biewald
That’s a good segue into my next question, which I’m sure people ask you all the time, and people ask me as well: who’s using open source, and who’s using closed source in your experience? How do people think about that trade-off?
If they’re coming to you, they’ve kind of decided on open source, but some people ask me why anyone would ever not use Anthropic or GPT. Maybe we’ll start there.
Tuhin Srivastava
Everyone’s on this curve of maturity. A lot of people start with custom models or open-source models, where they’re doing their own things. They’re training their own reranking models or their own synthesis models that are purpose-built for what they’re doing.
Other people start with Anthropic or OpenAI. You have these great models that can do a lot, but they’re either too expensive, or the speed is fine but OpenAI and Anthropic have a lot of reliability issues themselves because they’re serving massive workloads at enormous scale.
People come to it because they want more control over costs and more control over their own reliability. They want dedicated infrastructure at a decent price with some transparency, so they know what’s happening.
The third piece is that our customers care that we aren’t just piping all this data off to someone who’s going to train models on it. That matters to us a lot. I think all of those reasons are why people shift from closed to open: something custom, something cheaper, something more reliable, data privacy, and SLAs for the enterprise.
The reality is that everyone’s going to use a bit of both.
Lukas Biewald
Do you have a prediction, in the fullness of time, where the balance ends up?
Tuhin Srivastava
I don’t know—40/60. I don’t know which way.
Lukas Biewald
[Laughter]
Tuhin Srivastava
42. I think there are 2 worlds that we live in. Either AGI comes, and all we have left to do is go on podcasts because everything else is done. AGI would make a better podcast of me than me coming on here anyway.
I think that’s a pretty low probability. Models are already very, very good, and the next unlock is probably further away than we think.
Lukas Biewald
Why do you think that? That’s an interesting point of view.
Tuhin Srivastava
I think we’re out of data. I don’t know where the data is going to come from. We kind of went through this period over the last 5 years as we ingested more and more data, so I think there needs to be an architectural unlock. Then we’re really going back to being a research problem again.
That’s fine, because I still think we have 1,000× of economic value to unlock from what we’ve already unlocked. We could stop today and still have more productivity gains than the Industrial Revolution, which is fine.
Going back to your original question, the more reasonable outcome is that this is going to be a long tail of models, and people want that. You don’t need the most powerful model in the world to do everything for you all the time.
Lukas Biewald
Are you seeing reinforcement learning start to change your business?
Tuhin Srivastava
Yeah, I think there’s a lot of demand there. Again, everything comes down to a data-collection problem now. I think RL, in a lot of ways, is going to turn into an inference problem.
We’re not there today in terms of first-class support for those things, but all those things are possible using Baseten as a primitive. I think there’s a lot of demand there, especially with folks who are able to collect real preference data.
What’s really interesting about working with so many customers at the application layer is that most of them are collecting real, valuable preference data, which they want to use to make their models better.
9. The Importance of Inference in AI
Lukas Biewald
All right. I sent over my questions, and unlike most guests, you actually engaged with my pre-read. You added some questions yourself that I was intrigued by. I don't know if this was you or your team, but you added a question: Why is inference the fastest-growing part of the AI market? It seems incredibly obvious that it would be, but I'm wondering if there's something deeper there that you wanted to talk about.
Tuhin Srivastava
No, no. I think my team probably came up with that one.
Lukas Biewald
Good job, team. But I actually think that's the wrong question. I think the better question really is why it's the most important market.
Tuhin Srivastava
In a world where we're shifting toward models doing everything and all the value being at the application layer, there just needs to be a hell of a lot more inference. We're very lucky that we kind of locked into this market—or this market showed up.
It's interesting as a thought experiment: It might be the final market, right? If there is AGI, we're just going to need even more inference, and the only market that will exist, to some extent, is what the model needs to do. The model can only run on inference.
I think that is why it's the biggest market. Why it's the fastest-growing market, I think, is pretty obvious: We're just rushing as a society to inject models into everything we do. Every day, I now use a dozen products that are powered by Baseten in some shape or form, and I think that's somewhat of an indication of how big inference is getting.
Lukas Biewald
Yeah, that's incredibly cool. I was thinking, as you were saying that, tool use also might be part of it. Do you also run the tools for your customers?
Tuhin Srivastava
Yeah, yeah. I think the other one is sandboxes and actual code execution. All these problems are, the way I would say it, inference problems. I was saying this to someone on our team today: most startups have this problem where they go into a market and saturate it, and then start thinking about what else they can build alongside it. For me, honestly, it feels like we're just swimming into the ocean and it just keeps getting deeper and deeper and deeper and deeper, which is really cool, but it's also somewhat frightening. Tool use, sandboxes, RL—all these things are inference problems. Once they're all built out, I think that's why this market just continues to get bigger and compound over time.
Lukas Biewald
I do want to talk about—you had this quote, “You've got to burn the boats,” which I thought was evocative. I think a lot of people talk about that, but it's actually really hard to do in practice. Can you talk about burning the boats, how that works, and how you get people bought in on that?
Tuhin Srivastava
Yeah, we're just really odd, emotional people. No, I'm joking. I'd say we're just so future-looking, honestly. We forward-look all the time. Everything we've built until now is in the past, and what I owe the people who work for us, our customers, and even our investors is chasing what we think is the largest opportunity.
Lukas Biewald
Okay, great, but what's the most beautiful boat that you've burned down? Where you were like—
Tuhin Srivastava
In 2022, we killed 3 products out of 4 in 1 year. We basically spent 3 years, going back to what you said, building this Retool-type application builder alongside the serving engine. It was kind of bizarre that we built all this, and 2 dozen-odd people spent 2 and a half years building it.
I think we just killed it. Honestly, within 6 weeks, we were like, “Let's offboard customers, let's get them new places to do this work. We're going all in on this.”
Another one was that we launched a fine-tuning product called Blueprint in 2022. It was early for fine-tuning back then. We had 17 people in the company, and 6 people—about a third of the company at the time—spent 6 months building the thing. We launched it, and it didn't really go anywhere. We realized we had the abstraction wrong, and 3 months later we were like, “All right, we're not working on that anymore.”
I think we constantly do this as much as possible. Rewrites are part of the job. Throwing away stuff is okay. We're actually pretty okay with, “Let's sprint this out. It's okay if we have to throw it away.”
The trade-off there is that it allows you to be a little distracted and go on side quests, and be okay with that side quest not being super valuable, as long as you can reenter and go back to the thing you tried to do.
Lukas Biewald
Oh, that's kind of interesting, because earlier, when we were talking, you talked about “don't give up” and “keep going.” I can see how the 2 are related, but it's a tricky question to know when to—
Tuhin Srivastava
Stop. I say don't give up, and don't be emotionally attached. Those are probably the 2 things I'd say. Just be forward-looking to some extent. I think that's where I've always gotten stuck when I've been building things: The worst thing that could happen is you get stuck at a local maximum.
I think acknowledging that is important. The second piece, just to go alongside this, is that it's so interesting—I don't know how you think about this, because you've been around venture for how long? Like 10 years, 15 years? Around venture-funded companies for a while.
Lukas Biewald
A fair amount of time, yeah.
Tuhin Srivastava
I feel like, as a venture-backed founder, you're really committing to a type of company you're trying to build. I meet founders often who are like, “Oh, it's not really working, so we're going to start to become cash-flow positive.” I'm like, “Oh, I don't think I'm in the cash-flow-positive business for a while here.”
I signed up for something that's going to be very, very big, and I'm just swinging big. I think that's something that everyone on our team will agree with. We try not to take the conservative tack as much as possible.
Lukas Biewald
Yeah. It's also interesting that you say that because earlier, at the very beginning of this conversation, you were talking about—how did you put it? You're sort of like, “I like to think of it more as this project that I'm doing for this purpose,” rather than optimizing for growth or something. But I think probably your venture investors are thinking, “Hey, we're going to optimize for growth here.” Don't you think?
Tuhin Srivastava
I think so, but I don't think they're as detached as you think. With respect to the former thing—I'm not thinking about an IPO, but I need to be happy if I'm going to get there.
I think the thing that drives happiness for me is working on big things. It's not necessarily about how I squeeze the most amount of sales efficiency out of the machine that we have, and, you know, we need a new leader there because we mis-hired. It's more just: Are we chasing the biggest thing? Are we chasing important problems, to some extent?
10. Conclusion and Final Thoughts
Lukas Biewald
I think that makes sense. Yeah. All right, I think that's a great place to stop. That was an awesome interview. I really appreciate it.
Tuhin Srivastava
Yeah, no worries. Thanks for having me.