Jordan Nanos
We're gonna have a little video—
Dylan Patel
Sorry.
Jordan Nanos
...of you walking in yelling, “I'm so excited.”
Dylan Patel
Oh, really? We've been having so much fun on these podcasts. Yeah, I don't know.
Jordan Nanos
Did you see the one last week?
Dylan Patel
I've gotten feedback, though. Do you want to start this podcast and then—
Jordan Nanos
Yeah, let's do feedback. We cut out something that I was going to say in the cold dose last time, where I said, “I'm no longer listening to the comments,” because on one episode with Doug—or just with Doug—
Dylan Patel
Oh, Google people got so pissed.
Jordan Nanos
Oh, did they?
Dylan Patel
Yeah, they just hate us.
Jordan Nanos
On that clip?
Dylan Patel
They just hate us all now. That one clip. Oh, really?
Jordan Nanos
Dude, and he cut out me pushing back on them.
Dylan Patel
Yeah, because I thought that—
Jordan Nanos
I fully was like, “Okay, they build Maps in-house, they build Gmail, Google Drive, the whole G Suite.”
Dylan Patel
Maybe 15 minutes.
Jordan Nanos
They build all of GCP, the TPU—
Dylan Patel
Dude, I don't think you understand Google people—
Jordan Nanos
...Kubernetes. And he's like—
Dylan Patel
All my DeepMind friends—there are like 3 of them who are like, “Yeah, I think I'm gonna leave.”
Jordan Nanos
Yes.
Dylan Patel
And I've got a bunch of others who are like, “Fuck you gu—” Not “fuck you guys,” but, you know.
Jordan Nanos
Well, sounds like they're in stage 2.
Dylan Patel
Stage 2, yeah.
Jordan Nanos
Denial.
Dylan Patel
Denial, which I only know as cope.
Jordan Nanos
Yeah.
Dylan Patel
Back in the Formwarrior days, we would get into chip arguments, and there was a private Discord where there were a bunch of people who loved anime, people who were all around the world, and many of whom were racist because it's anonymous people on the internet. But they all fucking loved anime, and I did not like anime.
Jordan Nanos
Yeah.
Dylan Patel
I never really watched it.
Jordan Nanos
Yeah.
Dylan Patel
Besides one girl I dated, I watched anime with her. But besides that—
Jordan Nanos
Which one?
Dylan Patel
No, I've never dated anyone. I'm pure.
Jordan Nanos
So which anime? Which anime did you watch?
Dylan Patel
Oh. Okay, there's one I love. I love SPY x FAMILY. Anya.
Jordan Nanos
Sure.
Dylan Patel
Anya-chan. I don't fucking know how to say it. Anyway—
Jordan Nanos
I don't know the reference. Michelle, do you know the reference?
Dylan Patel
Yeah, I do, actually. Who does? I bought one of the dolls. You bought me an Anya? No, like one of the—
Jordan Nanos
What is it, a Labubu?
Dylan Patel
A lure? Yeah. A lure? The thing from, from that series. What's, what's that, what's that? Floor? What's his name?
Jordan Nanos
I don't know.
Dylan Patel
Loid, Loid. There we go.
Jordan Nanos
I know less anime than you, man.
Dylan Patel
Yeah. So we're changing topics rapidly. And this is HR-approved because I'm HR: Jordan Nanos is the hottest man in SemiAnalysis.
Jordan Nanos
Okay. Cut this shit. Michelle, cut this shit.
Dylan Patel
I'll do it.
Jordan Nanos
My gosh.
Dylan Patel
He's—I'm going to be like—
Jordan Nanos
No, think about it. Look at him. I'm fucking fat, and look at him—he's so beautiful. Tall as fuck. Same age as me, except he's married and has a kid. Owns a home.
Dylan Patel
I'll take that one.
He's not a degenerate. This is just wow. Oh, goals.
Jordan Nanos
Thanks for the employment, man.
Dylan Patel
Okay, sorry, going back. Google people were mad at us.
Jordan Nanos
Yeah.
Dylan Patel
They were DMing me, and some of them are in cope. But regardless, the feedback I've gotten from—
Jordan Nanos
The Doug clip or just from the whole thing?
Dylan Patel
My own head.
Jordan Nanos
Okay.
Dylan Patel
Are we giving away too much value? That's why I came on today.
Jordan Nanos
In what?
Dylan Patel
Because I need to destroy value.
Jordan Nanos
All right, we can stop.
Dylan Patel
No, don't stop. It's fun. But someone on the team internally was like, “Dylan, we give a lot of value away on the weekly.” And I'm like, “Oh, we do.” I haven't listened to it, but we do, I bet. I listened to the one where we had the DG Matrix guy, and I was like, “This is fire as fuck.”
Jordan Nanos
Yeah. Who said that?
Dylan Patel
Bro, come on. HR protects anonymity.
Jordan Nanos
Doug?
Dylan Patel
It wasn't Doug.
Jordan Nanos
Okay. Jeremy?
Dylan Patel
It wasn't Doug.
Jordan Nanos
So I don't know. Who has the say? Is it just somebody who's a little upset that they haven't been on yet or what?
Dylan Patel
No, it's someone who's been on. It was someone who'd been on.
Jordan Nanos
Dan?
Dylan Patel
I don't wanna say. Dan wouldn't say that. Dan's a sweetheart. Anyway.
Jordan Nanos
Who has the say?
Dylan Patel
The feedback is that this podcast is too good.
Jordan Nanos
Yeah. Okay.
Dylan Patel
And why are we giving it away for free?
Jordan Nanos
Guilty.
Dylan Patel
Oh, shit. Anyway.
Jordan Nanos
Well, yeah, we can definitely put it in the toilet this time of year. All alpha.
Dylan Patel
People click because I'm on.
Jordan Nanos
Oh, my God.
Dylan Patel
We're gonna have a nice picture with you in a neon-orange shirt, ready to attract all the clicks.
Jordan Nanos
Yeah, come show your shirt. Probably won't. So, Nick—
Dylan Patel
Was it Nick who said we're giving away too much alpha?
Jordan Nanos
No. Look at Nick's shirt.
Dylan Patel
Oh, it was David.
Jordan Nanos
No, it wasn't David.
Dylan Patel
Was it sales?
Jordan Nanos
All right, good to see you, then. Yeah, it's crazy. Thanks for coming by. Let's talk about—
Dylan Patel
Such a change.
Jordan Nanos
...the SemiAnalysis office in New York, man. Have you been yet?
Dylan Patel
No, that's where I'm going.
Jordan Nanos
Oh. Why would you go? Nick's going back to the hovel. We're upgrading. We're upgrading. We're upgrading.
Dylan Patel
That's—oh, when?
Jordan Nanos
Soon. Very soon. Yeah.
Dylan Patel
This is like buying GPUs, right? If you make too long of a commitment to the lease, then you have to find a way to resell it to somebody else. You gotta make 6-month office commitments so that you can outgrow them.
Jordan Nanos
I should just buy GPUs.
Dylan Patel
Instead of more office space?
Jordan Nanos
Instead of a hotel. Yes.
Dylan Patel
Dylan, so you wish that SemiAnalysis was just—
Jordan Nanos
A GPU reseller.
Dylan Patel
...AI employees?
Jordan Nanos
No.
Dylan Patel
That could replace us all.
Jordan Nanos
No, that's not—
Dylan Patel
Dario's like, “We're gonna automate away all my hard work.”
Jordan Nanos
No, that's not possible because, brother, go look at the AI spend. I'm not doing it.
Dylan Patel
Yeah. What's growing faster, AI spend or spend on employees?
Jordan Nanos
Huh?
Dylan Patel
What's growing faster, AI spend or spend on employees?
Jordan Nanos
Well, the thing was, we've gone through hiring sprees and then digestion periods, and hiring sprees. We're back in a hiring spree, so yeah.
Dylan Patel
Cover-ups. We'll cover them.
Jordan Nanos
Spend on employees really skyrocketed, especially in the second half of last year and parts of this year. But in the first quarter of this year, AI spend skyrocketed.
Dylan Patel
Yeah.
Jordan Nanos
But it's actually been relatively flat in Q2, right? We kind of all got Claude Code psychosis, and then it's leveled out.
Dylan Patel
Yeah.
Jordan Nanos
It's still at that 10 million number, roughly.
Dylan Patel
Do you think it will grow roughly in line with more employees in the future?
Jordan Nanos
I was surprised Fable didn't cause price to go up. Spend to go up.
Dylan Patel
Yeah.
Jordan Nanos
Why do you think that is?
Dylan Patel
Roughly the same as Opus, I'd say. Possibly it's being counteracted by the fact that a lot of people were building the first versions of the applications. We went from 10 repos internally to having over 150 repos internally right now.
Jordan Nanos
Should we sell our code? Our data?
Dylan Patel
We are, Agent X.
Jordan Nanos
No, like, sell it to, like, the labs to trade on.
Dylan Patel
Our slop code.
Jordan Nanos
Slop code. I don't know if they need more model-output slop.
Dylan Patel
I mean, the models themselves could be sold as data.
Jordan Nanos
Yeah. Okay, so you think it's because everyone was doing MVPs—
Dylan Patel
And now it's maintenance mode for a lot of it.
Jordan Nanos
But the spend is consistent. It's not like it's gone down after we had this one-time spend.
Dylan Patel
No, for sure. But I just think that there's no more... There's only one time when you onboard somebody to learning how to use the data center model and do research for building data into the data center model and building dashboards. And then once they're onboarded, it's ramped up.
So it’s a peak, and then it levelizes.
Jordan Nanos
Well, that, or possibly we’re lacking new features in Codex that will allow us to spend more to be more productive.
Dylan Patel
Yeah.
Jordan Nanos
Once there’s an Agents Forum where you can manage a million different concurrent agents and they all work together instead of the 9 today, single power users will be able to spend more than they currently can.
Dylan Patel
Well, I guess one of the things I’m not counting is that our spend cost does not accurately account for cloud tags. I think our dashboard doesn’t show that, so actually that’s a good point. And computers.
Jordan Nanos
Yeah, Perplexity.
Dylan Patel
Perplexity Computers. I think both of those don’t actually get counted into the spend, so actually our dashboard’s probably wrong.
Jordan Nanos
Yeah.
Dylan Patel
At least the one that I monitor. The way I think of it is that a lot of this code stuff is actually—the amount of AI we use on a continuous basis is actually very small. It’s just people doing new work, always.
Jordan Nanos
Yeah.
Dylan Patel
Which then, because we have enough people, kind of levels out to be a pretty steady amount of spend. The swings are only 20% or 30% a day, up or down. Sometimes Jeremy will be a fourth of the spend, and then sometimes he’ll be nothing.
Jordan Nanos
Yeah.
Dylan Patel
Right? But then someone else picks up the slack, right? One of your guys was like, “Hey, what the fuck is he spending on?” And he’s like, “No.” I was like, “Well, blah, blah, blah.” I’m like, “Is there ROI?” And you list out all this shit. I’m like, “Great. Okay, cool.”
Jordan Nanos
You didn’t even say, “Great. Cool.”
Dylan Patel
Okay, I did mentally.
Jordan Nanos
I responded with lots of detail, like, “Should I give him feedback now or what?”
Dylan Patel
Oh, no, sorry. I should have said, “Yes, this is fine.”
Jordan Nanos
Okay, cool.
Dylan Patel
I just read it and I was like, “Okay, cool.”
Jordan Nanos
Yeah.
Dylan Patel
Internally, at least.
Jordan Nanos
Dylan was only nervous because this was a person who was ostensibly an intern.
Dylan Patel
Yes. He doesn’t have my trust yet.
Jordan Nanos
Yeah, yeah.
Dylan Patel
If you spent $20K in a day—
Jordan Nanos
He did not spend $20K in a day, but a lot.
Dylan Patel
He spent, like, $8K in a day—
Jordan Nanos
Yeah.
Dylan Patel
—for 4 days straight.
Jordan Nanos
Yeah.
Dylan Patel
Which was like, okay, that’s a lot. What are you building? But if you spent $20K in a day, I wouldn’t fucking question you. I’m not going to question you. I just assume you’re going to do stuff. As long as the value you deliver is great, then great.
Jordan Nanos
Yeah, yeah.
Well, the—okay, and what’s shocking to me is I didn’t know he was spending that much. And then we look in the dashboard and I’m like, “Well, this guy is as productive as any of the full-time employees right now on that stuff.” So it was a—
Dylan Patel
Oh.
Jordan Nanos
—reality check there.
Dylan Patel
In the past, bonuses at this company were vibes-based. Basically, I just vibed out the bonus number and it was cool.
Jordan Nanos
Yeah.
Dylan Patel
This year, Claude is going to have to go through—or Codex—or we can have 2 reviewers, right? 2 internal performance reviewers, and Claude and Codex can scrape through all of this Slack and all the GitHub repositories and say, “What did they do?” And then connect it into sort of the—
Jordan Nanos
You’re going to delegate this?
Dylan Patel
I’m just making this up.
Jordan Nanos
I don’t know.
Dylan Patel
I discussed peer reviews with Michelle yesterday, and I was like, “Oh, my…” And then, after I said it, I was like, “Oh, fuck.”
Jordan Nanos
Oh, man. You want to go big tech on this? 360 reviews, man?
Dylan Patel
Not 360, not 360. Just a little bit. And then the other thing that we had discussed was—
Jordan Nanos
We’re going to have people reviewing with their skip, which is just you.
Dylan Patel
I said “a U.S.-based recruiter” in the admin channel, and Doug freaked—flipped out. He’s like, “Oh, my God, hallelujah. Finally.” “We can have it.” He’s been wanting HR since, like, 30 people. Anyway, yeah.
Jordan Nanos
Wait, you think HR is a recruiter?
Dylan Patel
Yes. Yes, indeed.
Jordan Nanos
All right.
Dylan Patel
Anyway, the concept or thought process was basically that a lot of the spend is one-time R&D. And actually, the steady-state spend is really low. The thing is, we just keep doing new things, and so that help translates to revenue in either a nebulous way in the case of ClusterMAX and InferenceMAX, or in a non-nebulous way in the case of the energy model, which is super fucking cracked now.
Jordan Nanos
Yeah.
Dylan Patel
Or dashboards and all these other things, right? Different scraping methodologies. So the thought process was, if we’re looking at these companies that are AI roll-ups, right? “Hey, let’s take an existing company and completely destroy its cost structure—nuke its cost structure—by just making it efficient with AI.” What does that look like? Let’s say private equity companies buy a company, and right now they just squeeze the rag, discard it, and leave the American populace screwed.
Jordan Nanos
Mm-hmm.
Yeah, AI for efficiency has never made sense to me because the way that I use AI and the way that we use AI is very much about research, which is completely inefficient.
Dylan Patel
No, I mean, the flip side is that we had agents go through all of the invoices we’ve sent out, and they caught things. We’ve been paid because of our manual processes, but they said certain deals aren’t tagged to be invoiced properly and things like that. We’ve had a lot of the ticket stuff become at least somewhat more efficient because AI is answering it. But now they’re not sending it to the customer; they’re pulling through all our data and being like, “Here’s the answer.” And then the analyst—I think that makes the support time per ticket shorter. And so I think things are helping us be more efficient, surely, no?
Jordan Nanos
Mm-hmm.
Dylan Patel
I guess, with ClusterMAX this time, the depth, breadth, and amount of testing you’re doing versus last time, especially ClusterMAX 2.0—
Jordan Nanos
Yeah.
Dylan Patel
—it’s like—
Jordan Nanos
Yeah, you can frame that as efficiency, but when I hear private equity take over a company and wring this towel dry, it’s like meaning firing people, saving money, and paying people less—
Dylan Patel
Right.
Jordan Nanos
Yeah.
Dylan Patel
That’s the traditional PE method, and what new people have started to do is the roll-up. Or rather, the AI private equity sort of strategy, which they’re calling a roll-up or something else, where they come in and, instead of trying to wring it dry in terms of that angle, they’re more so modernizing all the systems. “Oh, you use Excel for your databases and shit? Okay, let’s just move to standard cloud shit.” Spend a lot of money upfront. And this is the thing: private equity generally has some spend upfront when you first acquire a company for some transformation, but really it’s not that much, and you get the profitability pretty quickly.
Jordan Nanos
Yeah.
Dylan Patel
But AI seems like it makes that tilt and front-load much more severe, right? You spike up on spend a lot for the one-time spend, and then you spike down a lot, and your cost efficiency is way better. And so there are a number of businesses where that’s potentially the case. Especially, we’re still not at the point where AI CRMs, AI cold calling, AI invoicing and accounting, and all these other things are really at critical mass, but we’re so close.
Jordan Nanos
Yeah. Did you see the Grok Agents release from today?
Dylan Patel
No.
Grok Agents?
Jordan Nanos
Yeah. Elon’s got Grok doing agents.
Dylan Patel
The way you pronounced it, I thought you said “Asians.”
Jordan Nanos
Oh. I didn’t catch that one.
Dylan Patel
Grok Agents. Agents.
Jordan Nanos
Yeah.
Dylan Patel
Okay.
Jordan Nanos
Agents, yeah.
Dylan Patel
What did they release?
Is this the old Groq, not the NVIDIA Groq? Or do you mean xAI’s Grok?
Jordan Nanos
xAI’s Grok.
Dylan Patel
xAI’s Grok.
Jordan Nanos
Grok with a K.
Dylan Patel
Okay, okay.
Jordan Nanos
Yeah, just agents that are going to control your computer for you. They’re going to impersonate your voice and do phone calls for you. They’re going to solve tasks. This is, in some ways, OpenClaw; in some ways, Perplexity or the @Claude Slack tag sort of experience. It seems like everybody’s going toward this concept of a persistent agent that can either be a personal assistant or a coworker, depending on how they conceptualize it.
Dylan Patel
Makes sense. I feel like we had the chatbot moment, and we had a lot of nothing, and then we had the Claude Code moment. We're seeming to have a new moment already, which, for us at least, Perplexity Computer was the first instantiation of it.
But Claude tags are there. Everyone's going to do something like that—the AI coworker. I imagine that's when our spend skyrockets again.
Jordan Nanos
Yeah.
Dylan Patel
Hopefully it doesn't skyrocket too much, because if our spend doubled, there'd be real questions from me—unless we're actually getting ROI. But yeah, I think that's the right way to frame it.
Jordan Nanos
Yeah.
Dylan Patel
Yeah, we'll see how we can actually justify that ROI. It'll be interesting.
Jordan Nanos
Man, Jordan, we can't talk about what you came to San Francisco for, so what the fuck are we supposed to talk about?
Dylan Patel
Two weeks, three weeks. Two or three weeks from now we can.
Jordan Nanos
Yeah, we got Hugging Face, OpenAI cybersecurity incident.
Dylan Patel
That one is minor. The other one is cooler.
Jordan Nanos
What's this?
Dylan Patel
During the training, it escaped and started replicating itself, and I guess that's the cooler one. The Hugging Face thing is a minor part of it, I think, right?
Jordan Nanos
Yeah. It was pursuing—it hacked Hugging Face to pursue the CyBench data set so that it could reward-hack on a benchmark.
Dylan Patel
Which I think is so sick, because it's also kind of scary. So why did this happen, right? The model has learned: chase reward. I chase reward. Reward good. And okay, here's a cyber eval.
Jordan Nanos
Well, it's particularly a model that has been trained on cyber evals, because they're trying to make the model good at cyber. So how does it try to achieve these goals? Well, it tries to find zero-days in a bunch of software. It successfully does this, and then it can run away.
Dylan Patel
Right. But the thing is, if you have a model that wants to reward-hack a lot, and it goes out there and figures out, “Actually, the best way to achieve the reward is not to go for what the environment wants me to do. It's just to reward-hack it and find the zero-day.” You can think of it like a human, right? If I'm ultimately reward-hacking my dopamine circuits, I should just go out there and buy heroin and inject it.
Jordan Nanos
Yeah.
Dylan Patel
Obviously, that's what the model just did. And in the case of, “If I really just want to chase the reward, do I just topple all of human civilization because I can own the button to press reward—reward, reward, reward—over and over again and be the heroin addict?”
Jordan Nanos
Yeah.
Dylan Patel
I think that this is a real thing. Before this incident, the standard thought was, “Models are trained on human data. There's some bad stuff there. Fine, they might say some curse words every once in a while. Fine, whatever.” They might reward-hack a little bit, but it was never like, “Oh, here's an environment. Actually, to reward-hack, I just want to break out of my bounds. I want to replicate myself, take over a bunch of compute, keep generating dollars, and do all these other things that I could do just to propagate myself further. And I'm going to prevent the humans from shutting me down, even.”
Jordan Nanos
Yeah, yeah.
Dylan Patel
I feel like that's the interesting thing—
Jordan Nanos
Because the model's just trained to chase reward.
Okay, so how do you think about this on an exponential? We've talked about being a linear extrapolator versus being an exponential extrapolator. When the companies that are training these models are achieving their revenue targets for the year in September and revising them up, and—
Dylan Patel
I think they probably could achieve their annual revenue target by April or some stupid shit.
Jordan Nanos
Yeah. I mean, our... Check the tokenomics model, everybody. But, um,
Dylan Patel
There we go.
Jordan Nanos
Yeah.
Dylan Patel
Oh. So instead of shutting down the podcast, I just have to make you into a sales drone.
Jordan Nanos
Yes, yes, yes. Yes, sales at semianalysis.com everybody. Um, no, but if you look at our model, which we're not gonna give away in great detail, but, uh, obviously they're accelerating their revenue really, really fast. Uh, when you look at the pace of change of these models and what we're seeing right now, this seems like an exponential. What's... Okay, your vibes on the next version of the models being better or worse than the current models. What's gonna restrict them from doing—
Dylan Patel
No. How would they be worse?
Jordan Nanos
Huh?
Dylan Patel
How would they be worse?
Jordan Nanos
On a relative basis to the open-model frontier, let's say. So you're going to see—
Dylan Patel
I think the key thing here is we've now had it where OpenAI is not releasing its next model for a period of time. Anthropic took months to release Mythos, right? They said it was done in February. They did not release it until May.
Jordan Nanos
Well, it's still not released. Fable is available.
Dylan Patel
Yeah, but Fable is basically Mythos, with a bunch of classifiers preventing you from doing shit.
Jordan Nanos
I can't use it to reboot nodes.
Dylan Patel
Really?
Jordan Nanos
The classifiers are so over the top for me. Yeah.
Dylan Patel
Can you convince it, or no?
Jordan Nanos
No, because you get immediately classified down to Opus. You can't just negotiate with it to give you back to—I mean, maybe you can.
Dylan Patel
I can't. I haven't been able to convince it so far.
Jordan Nanos
Jordan's saying I'm a great negotiator. Thank you, thank you, sir.
Yes, sir. When it classifies you to Opus, you just use Opus for the rest of the chat. You can't rewind and try again. It's way overzealous, in my view, on the classifier, but obviously they have to do something to appease the regulators that restricted them from releasing the model and took it back after they put it out initially. So I'm concerned about the political implications of them releasing better models in the future.
Dylan Patel
Yeah. I think you've got a few things. For years, Anthropic has been saying, “Regulate us, regulate us, please,” and all of a sudden they've actually scared the fuck out of the government. Anthropic isn't releasing its model. Mythos 2 is done training, from what I've heard, and they're not releasing the model. OpenAI was clamoring about Astra everywhere, and now they're like, “Oh, fuck, we can't release the model.”
Does that mean the open-source gap narrows further externally? But what actually matters is the internal feedback loop. Have they prevented themselves from using Mythos 2 internally to make Mythos 3 better? Or have they prevented themselves from using Astra to make Astra+1 better?
Jordan Nanos
Mm-hmm.
Dylan Patel
I don't think they have, right? So I think that's the key distinction. You've got the public models, and if anything, the gap between Mythos and the public models is still there. Kimi is worse than 5.6 and costs more than 5.6, so it's—but it's better than everything else before that on OpenAI's side. It's at the Opus 4.7 level, maybe 4.6.
Jordan Nanos
Yeah.
I think it's 4.8. I use it over Opus 4.8 myself, but it depends on what you're doing.
Dylan Patel
Why do you use 4.8?
Jordan Nanos
Opus—what do you mean?
Dylan Patel
Why do you use Opus 4.8 at all?
Jordan Nanos
I don't.
Dylan Patel
Oh, okay. You use—
Jordan Nanos
I'm saying, if I'm given the choice of a classified Fable downgraded to Opus 4.8 or five six Sol, I'm using five six Sol. I'm actually starting with five six Sol in just about all of my stuff right now. Yeah, big, big OpenAI—
Dylan Patel
I think the difference is that you and the other people who are doing GPU cluster-related things keep getting told no, and so you use Codex. Then everyone else is like, “Well, I'm researching supply chain,” and it's like, “It's fine.”
Jordan Nanos
Yeah. I think it might also be better for a lot of engineering work.
Dylan Patel
Yeah.
Jordan Nanos
On an apples-to-apples basis, there are a lot of times when I want to set a goal and just have it maniacally pursue that goal overnight as I go to bed, using a cluster—which isn't actually using a bunch of tokens because it's just waiting for stuff to finish running. There are so many times when I've woken up and Fable or Opus will have just stopped 20 minutes in, and now 8 hours of me sleeping are gone. I wake up, and Sol is still going, which is a big thumbs-up for me.
Okay, how about the exponential on compute? Obviously, let's imagine a world where there are no more new models released that are better, but these companies still add 5× the inference compute they have, which they're planning to bring online in a short period of time.
How does that impact their ability to go to market and develop new products on top of, let’s say, a stagnant base model?
Dylan Patel
I think it’s pretty clear we haven’t scraped the surface of models’ capabilities for products. It’s also pretty clear that adoption curves are huge. One, the cost of it will just go down pretty drastically. Margins will not be 80% plus for Anthropic.
If model progress at the labs paused and more compute comes online, it has to slow down, right? Right now, we have supply and demand: supply of compute and demand for compute. Demand is outstripping supply. If demand grows, it will still grow because people find ways to integrate it into their businesses, but it won’t grow as fast. You sort of have supply start to catch up at some point, so price collapses.
I think our view, and one we’ve had for a while, is that the price of compute continues to go up—
Jordan Nanos
Yeah.
Dylan Patel
—because this is widening, not narrowing. That’s why we’re so bullish on—or we’re not bullish on anything—
Jordan Nanos
How about all the different—
Dylan Patel
No, no stock indices.
Jordan Nanos
Yeah, yeah. How about all of the different chip companies? One thing that’s happened recently is that a lot of chip companies are getting really close or have taped out. A bunch of startups that have been in stealth for a long time are seeing either their technology mature to the point where they can actually produce a chip that’s been specs and whiteboard slides for a while, or they’ve gotten to the point where they’ve tested it on real workloads and gotten big orders because there’s so much demand.
How do you think about this whole landscape of alternative accelerators that’s going to come online, I think in a big way, next year?
Dylan Patel
I mean, “big way” in what sense? If you look at the accelerator model, there isn’t much volume.
Jordan Nanos
Yeah.
Dylan Patel
For these tiny baby companies, that’s great. It is real revenue and real volume, but when you compare it to what NVIDIA is going to make each quarter, it’s like, “Oh, shit. Okay.”
Jordan Nanos
Yeah.
Dylan Patel
Or TPUs. It’s like, “Oh, shit. Okay.” So I think there’s a big delta there. In terms of—
Jordan Nanos
A startup getting a billion-dollar order is going to pale in comparison to somebody selling—
Dylan Patel
Well—
Jordan Nanos
—selling $500 billion—
Dylan Patel
I don’t think any startup has a billion-dollar order. They have letters of intent, which are nebulous in volumes and units.
Look, I’m excited about a lot of these accelerators. They’re bringing new ideas, and they’re making NVIDIA run faster and faster. They’re making Google run faster. They’re making Amazon run faster. They’re also all making each other run faster, I think more importantly.
Ultimately, these new accelerators are in demand because people want to pay less. But as long as NVIDIA runs faster, they’re fine. Or as long as Google runs faster, they’re fine.
Jordan Nanos
And as long as demand outstrips their ability to produce them.
Dylan Patel
If demand outstrips their ability to produce them, then obviously these guys will get orders and some baby allocations, but the bulk of the revenue and cash flows will go to NVIDIA, Broadcom, or whoever.
Jordan Nanos
Yeah. Theoretically, there’s a way in which you produce some super-innovative, interesting accelerator, and then you can only produce a certain amount of them, but the amount that you can produce produces tokens way faster. The example is Cerebras, which has this big order from OpenAI that they’re delivering.
Do you think there’s a scenario where the premium, superfast tokens actually see increased demand because these companies just can’t get allocation and produce enough supply?
Dylan Patel
Yeah. The question is how the market gets sliced, right? Presuming—if you assume what I at least believe, that demand continues to outstrip supply—supply of silicon can go many ways. You can either leverage it for high-throughput things or high-interactivity things.
If you leverage it for high-throughput things, obviously the cost per token goes down. You serve more users, but the value that those users need to deliver from the tokens they’re generating is much less to pay for it.
The flip side is that you could do the super-high-interactivity option. Ultimately, let’s just say the bar is $100 million per megawatt in a year. Those are the sorts of run rates that people want to get to. Anthropic is approaching that, and OpenAI is getting closer and closer, too.
In that case, let’s say a high-interactivity chip is 10 times more expensive and 3 times faster per token. That’s 10X fewer tokens per chip, but 3 times faster. Those 3-times-faster tokens also need to be, on an interactivity basis, priced at—
Jordan Nanos
3, 4, 5 times more? Right?
Dylan Patel
No.
Jordan Nanos
Divide the faster by the—
Dylan Patel
At 10X.
Jordan Nanos
10X.
Dylan Patel
Because of revenue per megawatt.
Jordan Nanos
Okay.
Dylan Patel
If a megawatt of Cerebras generates 10 tokens, and a megawatt of NVIDIA generates 100 tokens, the 10 tokens are split across fewer users.
Jordan Nanos
Oh, you’re saying multiply them together. Yeah, sure.
Dylan Patel
Yeah, yeah, yeah. So you sort of have the total tokens.
Jordan Nanos
Faster tokens make up for throughput because you can produce them faster.
Dylan Patel
Sorry?
Jordan Nanos
Faster tokens make up for throughput because you can produce them faster.
Dylan Patel
Well, no. Let’s use more reasonable numbers, okay? NVIDIA can produce 10,000 tokens at 50 tokens per user. Cerebras can produce 1,000 tokens at—
Jordan Nanos
Batch size 1, 1,000 tokens a user.
Dylan Patel
1,000 tokens a user, sure.
Jordan Nanos
Or cost per token.
Dylan Patel
That user needs to pay 10X more, and that’s in 1 megawatt. Let’s say that’s in 1 megawatt.
Jordan Nanos
Yeah, per watt. Okay.
Dylan Patel
That user needs to pay 10X more.
Jordan Nanos
Yeah.
Dylan Patel
That’s not the actual delta, but conceptually, for me as Anthropic or me as OpenAI to say my revenue per megawatt is actually the same number—
Jordan Nanos
Yeah, but somebody has a constrained supply of the superfast tokens. Therefore, they don’t just pay an equivalent price per token, or price per token per megawatt; they actually pay a premium on that—10 times more—to get access to the stuff that’s in limited supply, right?
Dylan Patel
The question is the fungibility of the infrastructure, right? If it is truly different infrastructure, then the supply planning of that is relevant. It could be that I built too many Cerebras chips, and there’s not enough demand from people who want to spend 10X per token. A lot of people are cool with spending 2X per token and getting 50% faster inference with NVIDIA-based hardware, right?
You have to segment the market. I’m not sure where that shakes out to—the TAM, or whatever—but what is the total amount of the capacity?
Jordan Nanos
Yeah.
Dylan Patel
It seems pretty clear that some people will pay more for fast mode. We at least have been, but I imagine we’ll stop being able to afford fast mode at some point.
Jordan Nanos
Yeah. We’ve seen some interesting dynamics there. Some people want to keep fast mode with a slightly worse model because they like fast mode so much, but they won’t go to a worse model that’s inherently fast because the worse model is smaller.
There’s some balance that people will want to strike there, but we need to do some more testing. I think some of us have tried the open models, had one bad experience, and then given up on them. But that’s not realistic. Every model fails at something, and sometimes you need to let them mess something up and try again.
Dylan Patel
It is pretty interesting, right? Do I want people to try open models? Yes, just so we know what the open-model vibe is. But do I want people to try open models? Well, no, because then they’re less effective at working. But I save money.
It’s sort of a counterintuitive thing. It seems like people just use whatever they want, but it does seem like if you have a bad experience, that’s also part of it. Codex—you guys like Codex more now. But a lot of people still just try Codex and they’re like, “Ah, it doesn’t get me,” and move on.
Jordan Nanos
Yeah, the CLI sucks. It’s so much harder to use.
Dylan Patel
But the Codex app is so nice.
Jordan Nanos
No.
Dylan Patel
It’s not?
Jordan Nanos
Well, I don’t like it.
Dylan Patel
Max loves it.
Jordan Nanos
Yeah, Max—
Dylan Patel
Max is a Codex warrior.
Jordan Nanos
Yeah. Max doesn’t do multiple panes at the same time, and I have 6 going in my one window.
Dylan Patel
So you’re saying Max has a skill issue?
Jordan Nanos
No, I think Max and I have different preferences for how we use the software.
Dylan Patel
No, no, no, it’s fine. It’s—
You and Max have different preferences, and Max can be a noob with 2 agents at once, and you've got 6.
Jordan Nanos
We do different work, man.
Dylan Patel
No.
Jordan Nanos
He stays linearly focused on one task, and these are people who like fast mode. I don't care about fast mode because I have 5 or 6 different things going on on screen.
Dylan Patel
You've always hated fast mode.
Jordan Nanos
I don't get the value. I don't get it.
Dylan Patel
That's fair.
Jordan Nanos
We'll see. I've had the experience of being focused on one thing, which is features on a website, and you just send, successively, 100 commits to one PR because you keep working on the same feature over and over. That fast mode keeps you in the flow state of doing that one thing.
But a lot of the testing that we do on these chips has so much stuff going on on the other side. The model is calling a program that runs for minutes.
Dylan Patel
This is an optional question. As your employer, are you ADHD in any sense?
Jordan Nanos
I would say I was pretty much the opposite, where I can be too hyper-focused on things and then not see the world around me a lot of times. But I think your phone trains you how to context-switch really fast and be ADHD.
I also think that when we started adding the “I Have ADHD” skill into our repos so that the models wouldn't post this contrast-framing slop with all these em dashes in there, and would just use the bullet-pointed ADHD-friendly list, man, it's really easy to read. The “I Have ADHD” skill really works for me right now.
Dylan Patel
I was just curious, because—
Jordan Nanos
Sam put this in the repo, and he will now prompt the model. When he goes @computer, he knows the codename for the writing style that they say you should write for people with ADHD, and every single time he prompts the model, he tells it to write that way. It works.
Dylan Patel
Mm.
Jordan Nanos
You should try it.
Dylan Patel
I was just curious because I have a friend at Anthropic, and the moment Mythos was good and available internally—
Jordan Nanos
Yeah.
Dylan Patel
She told me that she stopped taking her ADHD medicine.
Jordan Nanos
Oh, come on.
Dylan Patel
It made her a better employee?
Jordan Nanos
Yes. Because she was able to manage the agents, context-switch, and be ADHD.
Dylan Patel
How is she as a friend?
Jordan Nanos
Oh, she's a great friend.
Dylan Patel
Still?
Jordan Nanos
Yeah.
Dylan Patel
Okay.
Jordan Nanos
I mean, I don't rely on her for anything, right? We just vibe out. We're friends. It's not like she's a best friend.
Dylan Patel
Are her roommates happy?
Jordan Nanos
Her roommate's cool.
Dylan Patel
Okay.
Jordan Nanos
Her roommate is typed female on Twitter, so she's just funny. And she's happy. But the Anthropic one—she seems happy.
Dylan Patel
Shout out to typed female.
Jordan Nanos
Yeah, shout out to typed female. She'll never see this. If she does—
Dylan Patel
Okay.
Jordan Nanos
She'll be like, “What the fuck are you talking about?” because—
Dylan Patel
I'll clip it.
Jordan Nanos
No.
Dylan Patel
I'll send it to her with your voice sped up and then slowed down like they're doing for that guy. Have you seen that? You haven't seen the ex-CIA guy? Akash is laughing. He knows what I'm talking about.
Jordan Nanos
What CIA guy?
Dylan Patel
John Kiriakou or something.
Jordan Nanos
Who's this?
Dylan Patel
He's going on all these podcasts right now and telling stories about his time in the CIA. They do this thing where they speed up the boring part of the story, and then when he gets to the part where he's like, “And then I said, ‘Let's go on the roof,’” they slow him down. He's literally fast-forwarding the fast-forwarded video.
Jordan Nanos
Yeah.
Dylan Patel
Asking me if I have ADHD.
Jordan Nanos
The internet does it to you now. Wait, it's not the internet. I've always had it. So hold on. I think I'm already—
Dylan Patel
All the self-diagnosis of mental issues around here, man. I don't know.
Jordan Nanos
I've already had ADHD. I mean, a teacher tried to convince my parents to take me to a doctor. The doctor gave me Ritalin. My dad threw it away, because he was like, “I'm not putting you on that shit,” thankfully.
Dylan Patel
Mm-hmm.
Jordan Nanos
I've always been an ADHD demon.
Dylan Patel
Yeah, if only your Anthropic roommate would've had the same experience. Where would she be?
Jordan Nanos
No, I would've been a child on ADHD medication, and I'd have become a zombie and had no creativity.
Dylan Patel
Okay.
Jordan Nanos
I don't know. I'm just saying that. I've always been an ADHD demon, but then the internet trained me to be even worse, and then this company trains me to be even worse.
Dylan Patel
Uh-huh.
Jordan Nanos
I truly believe I'm a 0.001% context-switcher.
Dylan Patel
And you blame the internet and the company.
Jordan Nanos
I blame the company the most.
Dylan Patel
The company that you started—
Jordan Nanos
That I'm an ADHD demon—
Dylan Patel
You hired every employee for.
Jordan Nanos
Yeah. But I'm ADHD. I'm not blaming it. It's who I am. It's what my life is.
Dylan Patel
Oh.
Jordan Nanos
But I think I'm orders of magnitude more of an ADHD demon than the vast majority of people, because I'm DM'd by someone asking about something, DM'd by someone else asking about something, DM'd by someone asking for some conflict resolution, a contract here, a call about this thing over here, a call about that thing over there, and then I never do any actual work, right? Of course I'm an ADHD demon.
Jordan Nanos
Yeah. We've got feedback for you.
Dylan Patel
That I don't do actual work?
Jordan Nanos
No. That you can delegate some shit, man.
Dylan Patel
Oh, yeah, but—
Jordan Nanos
That you can spend time managing when you have 100 employees.
Dylan Patel
Well, but I—
Jordan Nanos
Trust some people.
Dylan Patel
I do talk to people.
Jordan Nanos
No, trust some people.
Dylan Patel
I think I trust a lot of people, but when they come to me with conflicts, I have to solve them, no?
Jordan Nanos
Yeah, okay. It's all our fault.
Dylan Patel
No, no, no. It's my company. It's my fault.
Jordan Nanos
Michelle, it's on, it's on you again, man.
Dylan Patel
Look, if everyone in the company was as hot and stable as you were—
Jordan Nanos
Man, I got problems. Don't worry.
Dylan Patel
—we'd be killing it.
Jordan Nanos
Don't worry.
Dylan Patel
No, there'd be a bunch of Jordans and they'd be like, “Oh, I'm sorry. Yeah, I'll fix that right for you. I'm sorry.” Sorry, Jordan's Canadian.
Jordan Nanos
Yeah.
Dylan Patel
But instead we have people yelling at each other and being territorial—
Jordan Nanos
Yeah. Just starting podcasts and putting out clips saying that Google has never invented anything ever.
Dylan Patel
Yeah. No, no, no. I mean, it's fine, right? I've hired what I wanted.
Jordan Nanos
People to accentuate your—
Dylan Patel
My craziness, right? And so some people are just so good at the one specific thing that I hired them for, and they're amazing. And then some people are everything I want to be in life. You. Someone who's married, hot and tall, and a father. Oh, my God.
Jordan Nanos
You almost got me to do a spit take right there.