Speaker 1
Agents can do 3 things. They can access your files, they can access the internet, and now they can write custom code and execute it. You should really only let an agent do 2 of those 3 things.
If you can access your files and write custom code, you don't want internet access because that's 1 vulnerability, right? If you have access to the internet and your file system, you should know the full scope of what that agent is capable of doing. Otherwise, malware can get injected or something else can happen.
A lot of what we've been thinking about is: How do we enable this, because it's clearly the future, but also, what are the enforcement points that we can start to protect?
swyx
All right, welcome to the Latent Space podcast in the Chroma Studio. Welcome to all the guests here. We're back with our guest host, Alessio Fanelli. Welcome. Good to have you back. And our friends Nader and Kyle from NVIDIA. Welcome.
Speaker 2
Thanks for having us.
Speaker 3
Thank you.
swyx
I don't even know your titles. I know you're an engineering leader and architect of Dynamo or something.
Speaker 3
I'm 1 of the engineering leaders and architects of Dynamo.
swyx
And you're director of something developer-related?
Speaker 4
Yeah.
swyx
You're the “developers, developers, developers” guy at NVIDIA.
Speaker 4
I focus on open source, agent marketing, Brev, developer tools, and things like that.
swyx
We're recording this ahead of NVIDIA GTC, which is coming to town again and taking over the town. We'll all be there, and we'll talk a little bit about your sessions and stuff.
Speaker 3
We're super excited for it.
swyx
1 of my favorite memories is how you always do marketing stunts. When you were at Brev, you had this surfboard that you took down to GTC. NVIDIA apparently liked it so much that they bought you. What was that like?
Speaker 3
Our logo was a shaka, and we were always trying to stay true to who we were. So much of being a startup is pretending that you're a bigger, more mature company than you are. Evan Conrad from SF Compute was just like, “You guys are 2 dudes in a room. Why are you pretending that you're not?”
So we said, “Okay, let's make the logo a shaka.” We brought surfboards to our booth at GTC, and the energy was great. We had some palm trees, too.
The palm trees actually poked out over the walls, so you could see the Brev booth from very far away when no one else could.
Speaker 4
Yeah, I remember it pre-acquisition. I was like, “Oh, those guys look cool.”
Speaker 3
That makes sense, because we signed up really last-minute and had the last booth. It was all the way in the corner, and I was worried that no one was going to come. That's why we had the palm trees and brought in the surfboards.
We even had 1 of our investors bring her dog. She was just walking the dog around to try to bring energy toward our booth.
Speaker 4
Yeah, she's the best.
swyx
As a conference organizer, I love that. Everyone who sponsors a conference shows up with a booth saying, “We are changing the future of AI,” or some other generic bullshit. No—actually try to stand out and make it fun. People still remember it after 3 years.
Speaker 3
I'll send you this clip if you want to add it in. My wife—my fiancée at the time—was in medical school, and she came to help us because it was a big moment for us.
We bought this Cricut, which is a vinyl printer, because how else were we going to label the surfboard? We got the surfboard, luckily purchased it on the company card, and got a Cricut. We put “fine-tuning for enterprises” or something like that on the surfboard.
It was 1:00 a.m. the day before we went to GTC. She was helping me put the vinyl stickers on, and she goes, “You son of a bitch. If you pull this off, you son of a bitch.”
After the acquisition, I stitched that together with the news of the acquisition and sent it to our family group chat.
swyx
She made a good choice there. Was that basically the origin story for Launchables?
Speaker 3
Maybe we should explain what Brev is.
swyx
Yeah, we should.
Speaker 3
Brev is a developer tool that makes it really easy to get a GPU. We connect a bunch of different GPU sources, and the basic idea is: How quickly can we SSH you into a GPU?
Whenever we talked to users, they wanted a GPU. They wanted an A100. If you go to any cloud-provisioning page, it's usually 3 pages of forms, or somewhere in the form there's a drop-down with some weird code that you need to know translates to an A100.
I remember thinking that every time someone says they want an A100, the piece of text they're telling me they want is stuffed away in the corner. So we thought, “What if the biggest piece of text was what the user was asking for?”
When you go to Brev, you just see big GPU chips with beautiful animations that show the type of GPU.
Speaker 4
Animations that you worked on.
swyx
Back in the day, before you could just prompt it.
Speaker 3
Handcrafted artisanal code.
I was actually really proud of that. I made it in Figma and then struggled to figure out how to turn it from Figma into React. What it actually is is just an SVG with all the styles included.
When you change the chip—whether it's active or not—it changes the SVG code, which somehow renders as if it's animating. We just slowed down the transition. It's really just a JavaScript function that changes the underlying SVG, and that's how I figured out how to move it from Figma.
That's artisan work.
Speaking of marketing stunts, we used those SVGs to make these cards—a GPU gift card—that I handed out everywhere.
Speaker 4
That was actually my first impression of Brev.
swyx
I think I still have 1 of them.
Speaker 3
They look great. I still have a ton of them in our garage, but they don't have labels. We should honestly bring them back.
I found this old printing press just around the corner on Van Ness. It's a 3rd-generation San Francisco shop. I came in as an excited startup founder, and they had this crazy old machinery. The whole building was so physical: You could see the machines, and they had pedals to move the saws and everything else. I don't even know what some of the machinery was.
I saw all 3 generations—the grandfather, the father, and the son. The son was around my age.
Speaker 4
It's like a holy trinity.
Speaker 3
We took the same SVG and printed it using foil printing. They make a mold that's the inverse of the A100, put the foil on it, and press it into the paper.
I remember when we got them, he said, “Hey, don't forget about us.” Early Apple and Cisco business cards were apparently made there. He said they get the startup businesses, but as they mature, they go somewhere else.
I think we were talking with marketing about using them for—
Speaker 4
swyx
As a very, very small Brev investor, I remember thinking, “Why are we spending time doing these stunts for GPUs?” As a typical cloud-hardware person, you go into AWS, pick a p5.48xlarge or whatever from a list, and look at the specs. Why animate this GPU?
I do think it shows the level of care that goes throughout Brev and Dynamo, and NVIDIA as well. That's what struck me most when we first came in: the amount of passion that everyone has.
You talk to Kyle, or really any VP I've met at NVIDIA, and they go so close to the metal. Almost a year ago, my VP asked me, “What's Cursor? Are you using it, and if so, why?” I was surprised by that. He downloaded Cursor and asked me to help him use it, or at least show him why we were using it.
The amount of care, passion, and appreciation for the moment is really remarkable. This is a very unique time, and it's cool to see everyone appreciate that.
Before we move over to the research topics and the stuff that you're working on, I want to tell the story of the acquisition. Not many people have been through an acquisition with NVIDIA. What's it like? Anything you'd like to say?
Speaker 3
It's a crazy experience. The thing that was most exciting for us was that our goal was simply to make it easier for developers.
Speaker 1
We wanted to find access to GPUs, make it easier to do that, and then—oh, actually, your question about Launchables. Launchables was just making one-click deploys for any software on top of the GPU.
What we really liked about NVIDIA was that we felt like we just got a lot more resources to do all of that. I think NVIDIA’s goal is to make things as easy for developers as possible, so there was a really nice synergy there. I think, when it comes to an acquisition, the amount that the soul of the products align is going to speak to the success of the acquisition, and so it, in many ways, feels like we’re home. This is a really great outcome for us. I love brev.nvidia.com; you should use it.
Speaker 2
It’s a front page for GPUs. If you want GPUs, you go there, and it’s like—internally, it’s growing very quickly. I don’t remember; you said some stats there.
Speaker 1
Yeah. I wish I had the exact numbers, but internally and externally, it’s been growing really quickly. We’ve been working with a bunch of partners, customers, and ISVs. If you have a solution that runs on a GPU and you want people to use it quickly, we can bundle it up in a Launchable and make it a one-click run.
If you’re doing things and you want just a sandbox or something to run on—like OpenClaw, a huge moment, super exciting—we’ll talk about it more. Internally, people want to run this, and we know we have to be really careful from a security-implications perspective. Do we let this run on the corporate network? Security’s guidance was, “Hey, run this on Brev.” It’s a VM sitting in the cloud, off the corporate network, and it’s isolated.
That’s been our stance internally and externally about how to even run something like OpenClaw while we figure out how to run these things securely.
swyx
But I think you were almost the right team at the right time, when NVIDIA was starting to invest a lot more in developer experience—or whatever you call it, UX. I don’t know what you call it. NVIDIA has always invested in software, but this is a different audience.
Speaker 1
It’s a wider developer base.
swyx
Yeah, right.
Speaker 1
swyx
Yeah. And you know, it’s funny—it’s not—
swyx
So what is it called internally? What is this that people should be aware is going on there?
Speaker 1
Developers.
swyx
Yeah. It’s called just developer experience, or is there a broader strategy here?
Speaker 1
NVIDIA always wants to make a good developer experience. The thing is, a lot of the technology is just really complicated. I think AI is having a huge moment, not because data scientists in 2018 were quiet then and are much louder now. The pie is bigger: there’s a whole bunch of new audiences.
My mom’s wondering what she’s doing; my sister taught herself how to code. I actually think, just generally, AI is a big equalizer, and you’re seeing a more technologically literate society, I guess. Everyone’s learning how to code; there isn’t really an excuse for that. Building a good UX means that you really understand who your end user is, and when your end user becomes such a wide variety of people, then you have to almost reinvent the practice, right? You have to—
swyx
—and actually build more developer UX, right? Because there are tiers of the developer base that were added. The hackers who are building on top of OpenClaw, for example, have never used a GPU. They don’t know what CUDA is; they just want to run something.
Speaker 1
Yeah, right. You need new UX that isn’t just, “Hey, how do you program something in CUDA and run it?” And then we built Torch when deep learning was getting big, but recently, the amount of layers added to that developer stack has just exploded because AI has become ubiquitous. Everyone’s using it in different ways. It’s moving fast in every direction: vertical and horizontal.
swyx
You guys even take it down to hardware, like the DGX Spark. It’s basically the same system, just thrown up on a big GPU cluster.
Speaker 1
Yeah. Yeah. Yeah.
swyx
Blackwell.
Speaker 1
Yeah. We saw the preview at last year’s GTC, and that was one of the better-performing videos of our NVIDIA coverage so far.
swyx
Awesome.
Speaker 1
This will beat it.
swyx
That was actually—fingers crossed.
Speaker 1
Yeah. Even when the DGX Spark was first coming out, getting to be involved in that from the beginning of the developer experience, it just comes back to—
swyx
You were involved.
Speaker 1
Yeah. I mean, I just got an email; we were thrown into the loop. Suddenly, I was getting an email from a bunch of the engineering VPs about the new hardware GPU system—not chip, just the GPU system—that we were putting out, and I was like, “Okay, cool. Nat is now involved with this for the UX. What am I going to do here?”
I remember the first meeting. I was just kind of quiet as I was hearing the engineering VPs talk about what this box could be, what it could do, and how we should use it. One of the first ideas people were considering was, “The first thing someone’s going to want to do with this is get 2 of them and run a Kubernetes cluster on top of them.” And I was like, “Oh, I think I know why I’m here.”
The first thing we’re doing is easy SSH into the machine. The person who wants to run a Kubernetes cluster on top of Sparks has a higher propensity for pain than someone who buys it and wants to run OpenClaw right now.
swyx
If you can make sure that’s as effortless as possible, then the rest becomes easy.
Speaker 1
There’s a tool called NVIDIA Sync. It just makes the SSH connection really simple. If you think about it, if you have a Mac or a PC or whatever, if you have a laptop and you buy this GPU and want to use it, you should be able to use it like it’s a GPU in the cloud, right? But there’s all this friction of how you actually get into that. That’s part of Brev’s value proposition: there’s a CLI that wraps SSH and makes it simple. Our goal is just to get you into that machine really easily.
One thing we just launched at CES—it’s still in early access, and we’re ironing out some kinks, but it should be ready by GTC—is that you can register your Spark on Brev. And so now, if—
swyx
Like remote-managed local?
Speaker 1
Because Brev can already manage other clouds anyway, right?
swyx
Yeah. You use the Spark on Brev as well, right?
Speaker 1
Yeah, exactly. You set it up at home, you can run a command on it, and then it essentially appears in your Brev account. You can take your laptop to a Starbucks or a cafe, and you can continue to use your Spark just like any other cloud node on Brev.
swyx
It’s just like a pre-provisioned data center in your home.
Speaker 1
Yeah, exactly.
swyx
Yeah. Yeah.
Speaker 1
Tiny little data center.
swyx
One more thing before we move on to Kyle. I just have so many Jensen stories, and I love mining Jensen stories. My favorite so far is “SOL.” What is “SOL”?
Speaker 1
“SOL” is actually—I think, of all the lessons I’ve learned, that one’s definitely my favorite.
swyx
It can always stick with you.
Speaker 1
Yeah. When you’re a startup, everything’s existential, right? We’ve run out of money. We were at risk of losing payroll. We’ve had to contract our team because we ran out of money.
Because of that, you’re really always forcing yourself to understand the root cause of everything. If you get a date or a timeline, you know exactly why that date or timeline is there. You’re pushing every boundary, and you’re not just accepting a no just because. As you start to introduce more layers and become a much larger organization, “SOL” is essentially, “What is the physics?”
The speed of light moves at a certain speed, so if something’s moving slower, then you know something’s in the way. Before trying to layer reality back in about why something can’t be delivered by some date, let’s just understand the physics. What is the theoretical limit to how fast this can go? Then start to tell me why, because otherwise people will start telling you why something can’t be done.
Actually, I think any great leader’s goal is just to create urgency.
swyx
There are compelling events, right?
Speaker 1
“SOL” is a term at NVIDIA that’s used to instigate a compelling event. You say, “This is done.”
swyx
How do we get there? What is the minimum—as much as necessary, as little as possible—that it takes for us to get exactly here?
Speaker 1
It helps you just break through a bunch of noise.
swyx
Yeah. Instantly.
One thing I’m unclear about is, can only Jensen use the “SOL” card—like, “Get the hell out”—because obviously it’s Jensen? Can someone else be like, “No,” like—
Speaker 1
Frontline engineers use it?
swyx
Yeah.
Speaker 1
I think it's not so much about “get the [__] out.” It's more like, “Give me the root understanding,” right? If you tell me something takes 3 weeks, it's like, yeah, first principles: why is it 3 weeks? What is the actual limit of why this is going to take 3 weeks?
If you wanted to buy a new computer and someone told you it was going to be here in 5 days, what's the SOL? The SOL is, “I could walk into a Best Buy and pick it up for you,” right? Anything beyond that—is that practical? Is that how we're going to give everyone in the company a laptop? Obviously not. So that's the SOL. If we have to get more than 10, suddenly there might be some constraints. And so now we can piece the reality back together.
swyx
So this is Paul Graham's “Do Things That Don't Scale.”
Speaker 1
Yeah. And this is also what people would now call founder mode. It's actually really interesting because there's a second hardware angle to SOL that doesn't come up for the whole organization. SOL is used culturally at NVIDIA for everything.
swyx
I'm also mindful that this can be annoying sometimes, when someone keeps going, “SOL,” and you're like, “Guys, we have to be stable. We have to learn to [__] plan.”
Speaker 1
Yeah, I encountered that with Alec, right? We have a new conference, so we need to launch. We have goals for what we want to launch by the conference, and at the end of the day, it's GTC.
We did it for CES, we did it for GTC DC before that, and we're doing it for GTC San Jose. Every year, we have a new moment, and we want to launch something. We want to do so, and that does mean that some level of prioritization needs to happen.
It is difficult, right? I think you have to be careful with what you're pushing. Stability is important, and that should be factored in. SOL isn't just “build everything and let it break.” That's part of the conversation.
As you're layering in all the details, one of them might be, “Hey, we could build this, but then it's not going to be stable for XYZ reasons.” One of our conversations for CES was, “Hey, we can get this into early access, registering your Spark with Brev.” But there are a lot of things we need to do to feel really comfortable from a security perspective. There's a lot of networking involved before we deliver that to users.
So it's like, okay, let's get this to a point where we can at least let people experiment with it. We had it in a booth, we had it in Jensen's keynote, and then let's go iron out all the networking kinks. That's not easy, and so that can come later. That was the way that we layered that back in.
swyx
It's not really about saying you don't have to do the maintenance or operational work. It's more about saying that it highlights how progress is incremental, right? What is the minimum thing that we can get to? Then there's the SOL for every component after that, but there's the SOL to get you to the starting line. That's usually how it's asked.
On the other side, SOL came out of hardware at NVIDIA, right? SOL is literally: if we ran the accelerator, or the GPU, at basically full speed with no other constraints, how fast would we be able to make a program go?
Speaker 1
Yeah. Yeah. Right. In training, you work back to some percentage of MFU, for example. Yeah, that's a great example. So there's an SOL, there's MFU, and then there's what's practically achievable.
swyx
Cool. Should we move on to Kyle's side? Kyle, you're coming more from the data science world. Whenever I meet someone who's done work in tabular data, graph neural networks, or time series—basically, when I go to NeurIPS or ICML and walk the back halls, there's always a small group of graph people, a small group of tabular people, and no one else there. It's very niche work, you know what I mean? It's important, interesting work if you care about solving the problems that they solve.
Speaker 2
Yeah, but everyone else is just LLMs all the time.
swyx
Yeah. It's like the black hole, right? Has the event horizon reached this yet at NeurIPS? But those are Transformers too, and those are also interesting things. Anyway, I just wanted to spend a little bit of time on that background before we go into Dynamo proper.
Speaker 2
Yeah, sure. I took a different path to NVIDIA than that. I joined 6 years ago—7 if you count when I was an intern. I joined NVIDIA right out of college, and the first thing I jumped into was not what I had done during my internship, which was some work for autonomous vehicles, like heavyweight object detection. I jumped into something like recommenders; this was popular.
swyx
Yeah, you did RecSys.
Speaker 2
Yeah, RecSys. That was the tabular data at the time, right? You have tables of audience qualities and item qualities, and you're trying to figure out which member of the audience matches which item—or, more practically, which item matches which member of the audience.
At the time, we were trying to enable recommenders, which had historically been a CPU-based workflow, to run really well on GPUs. It's since been done: there are a bunch of libraries for XGBoost that run on GPUs. The common models, like the Deep Learning Recommendation Model, which came out of Meta, and the Wide & Deep model, which was released by Google, were very accelerated by GPUs using the fast HBM on the chips, especially for vector lookups.
It was very interesting at the time and super relevant because we were starting to get this explosion of feeds and things that required recommenders to be actively on all the time. I transitioned a little bit toward graph neural networks when I discovered them because I realized you could use graph neural networks to represent relationships between people, items, and concepts. That interested me, so I jumped into that at NVIDIA and got really involved for 2-ish years.
swyx
Something I learned from Brian Catanzaro is that you can just choose your own path at NVIDIA.
Speaker 2
Oh my god. Yeah.
swyx
Which is not a normal big-corporation thing. You have a lane; you stay in your lane.
Speaker 2
I think that's probably the reason why I enjoy being in a big company as a startup guy.
swyx
The mission is the boss.
Speaker 2
Yeah. Yeah. It also shows, right? NVIDIA is just releasing state-of-the-art stuff in every domain. You expect foundation models with Nemotron, and voice models are just randomly top-tier. Parakeet just comes out. The voice team has always been producing. In every other domain, there's always a paper that comes out, a dataset that comes out.
I mean, it also stems back to what NVIDIA has to do, right? You have to make chips years before they're actually produced. You need to know; you need to really—
Alessio Fanelli
The design process starts 3 to 5 years before the chip gets to the market.
swyx
Yeah. I'm curious more about what that's like, right? You have specialist teams. Is it just that people find an interest, go deep on whatever they want, and that feeds back into, “Okay, we expect predictions”? The internals at NVIDIA must be crazy, right? You must not even have to sell to people—you have your own predictions of where things are going, and they're very based, very grounded, right?
Yeah, it's really interesting. There are 2 things that does. One is that we really index into passion. There's a big organizational, top-down push to ensure that people are working on the things they're passionate about. If someone proposes something that's interesting, many times they can just email someone way up the chain who would find it relevant and say, “Hey, can I go work on this?”
swyx
I worked at a big company for a couple of years before starting on my startup journey, and it felt very weird if you were to email out of chain, if that makes sense. The emails at NVIDIA are like mosh pits.
Huge.
swyx
It's just 60 people, just whatever. And there's something messy about it, like reply all.
Alessio Fanelli
Oh, it gets insane. It's insane. It must help you manage the context.
swyx
But that's actually a weird thing. I used to be like, “Why would we send emails? We have Slack.” I'm the exact opposite. I feel so bad for anyone who's messaging me on Slack because I'm so unresponsive.
You're emailing. I'm email-maxing out.
Speaker 2
Email is different. Email is perfect because—
We can't work together on Slack. [laughter]
Speaker 2
Email is great because important threads get bumped back up, right? Slack doesn't do that. I just have this casino going off on the right or on the left, and I don't know which thread was from where. But there's the thread, and then there's also just the subject, so you can have working threads.
I think what's difficult is when you're small—if it's not 40,000 people—I think Slack will work fine, but I don't know what the inflection point is. There is going to be a point where that becomes really messy, and you'll actually prefer having email because you can have working threads. You can CC more than 9 people in a thread.
You can fork stuff.
Speaker 2
You can fork stuff, which is super nice. And so that's part of where you can propose a plan. You can also just start. Honestly, momentum is the only authority, right? If you can just start to make a little bit of progress and show someone something, then they can try it. That's, I think, been the most effective way to push anything forward, and that's both at NVIDIA and generally.
Yeah.
Speaker 2
There's another concept that's explored a lot at NVIDIA, which is this idea of a 0-billion-dollar business. Market creation is a big thing at NVIDIA.
Alessio Fanelli
You want to go and start a billion-dollar business.
Speaker 2
Jensen says we're completely happy investing in 0-billion-dollar markets. We don't care if this creates revenue. It's important for us to know about this market. We think it will be important in the future. It can be 0 billion for a while. I'm probably mangling his words here, but I'll give an example. NVIDIA's been working on autonomous driving for a long time.
Like an NVIDIA car.
Speaker 2
No, they use Mercedes, right? They're around the HQ, and I think it finally just got licensed out. Now they're starting to be used quite a bit. But for 10 years, you've been seeing Mercedes with NVIDIA logos.
If you're in Santa Clara, it's actually south. Yeah.
Speaker 2
So 0-billion-dollar markets are a thing. Jensen—
I mean, okay, look, cars are not a 0-billion-dollar market, but yeah. [laughter]
Speaker 2
I think he's messaging zero today, but—
Or even internally, right? An org doesn't have to ruthlessly find revenue very quickly to justify its existence, right? A lot of the important research, a lot of the important technology being developed—that's kind of where—
Speaker 2
Research is very ideologically free at NVIDIA. They can pursue things that they—
Were you in research officially?
Speaker 2
I was never in research officially. I was always in engineering. I'm in an org called Deep Learning Algorithms, which is basically just: How do we make things that are relevant to deep learning go fast?
That sounds freaking cool.
Speaker 2
And I think a lot of that is underappreciated, right? Like time series. This week, Google put out TimesFM, a new time-series paper. RecSys—semantic IDs started applying transformers and LLMs to RecSys—and when you think of the scale of companies deploying these, like Amazon recommendations and Google Web Search, it's huge scale, and you want fast—
Yeah, actually, there's a fun moment that brought me full circle. Amazon Ads recently gave a talk where they talked about using Dynamo for generative recommendation, which was super weirdly cathartic for me. I'm like, “Oh my God, I've supplanted what I was working on. You're using LLMs now to do what I was doing 5 years ago.”
Speaker 2
Yeah.
Alessio Fanelli
Amazing. Let's go right into Dynamo. Maybe introduce it sort of top-down.
Speaker 2
At this point, a lot of people are familiar with the term inference. Funnily enough, I went from inference being a really niche topic to being something discussed on normal people's Twitter feeds.
It's on billboards here.
Speaker 2
Yeah, very, very strange. Driving and seeing just an inference ad on 101. Inference at scale is becoming a lot more important. We have these moments like OpenClaw, where you have these agents that take lots and lots of tokens but produce incredible results. There are many different aspects of test-time scaling, so you can use more inference to generate a better result than if you were to use a short amount of inference. There's reasoning, there's re-querying, there's adding agency to the model, allowing it to call tools and use skills.
Dynamo sort of came about at NVIDIA because myself and a couple of others were talking about these concepts. You have inference engines like vLLM, SGLang, and TensorRT-LLM, and they have one single copy. They think about things as one single copy, one replica—one version of the model. But when you're actually serving things at scale, you can't just scale up that replica because you end up with performance problems. There's a scaling limit to scaling up replicas, so you actually have to scale out, to use some Kubernetes-type terminology.
We realized there was a lot of potential optimization we could do in scaling out and building systems for data-center-scale inference. So Dynamo is this data-center-scale inference engine that sits on top of frameworks like vLLM, SGLang, and TensorRT-LLM, and just makes things go faster because you can leverage economies of scale—the fact that you have KV cache, which we can define a little bit later, on all these machines, that is unique, and you want to figure out ways to maximize your cache hits. Or you want to employ new techniques in inference, like disaggregation, which Dynamo introduced to the world in March—not introduced it; there was an academic talk beforehand—but we're one of the first frameworks to start supporting it. We want to combine all these techniques into a modular framework that allows you to accelerate your inference at scale.
Speaker 1
By the way, Kyle and I became friends on my first day at NVIDIA, and I always love that he teaches me—
Speaker 2
New things.
Speaker 1
By the way, this is why I wanted to put two of you together. I was like, “Yeah, this is going to be good.”
Speaker 2
It's very, very different. We've talked to each other a bunch. Actually, you know—
Speaker 1
You asked, like, “Why can't we scale up?”
Speaker 2
Yeah.
Speaker 1
Model—you said model replicas.
Speaker 2
Yeah. So scale up means assigning more—
Speaker 1
Heavier.
Speaker 2
Yeah, heavier—making things heavier, adding more GPUs, adding more CPUs. Scale out is just having a barrier saying, “I'm going to duplicate my representation of the model, or representation of this microservice or something, and replicate it many times to handle the load.” The reason you can't scale up past some points is that there are hardware bounds and algorithmic bounds on that type of scaling.
So I'll give you a good example that's very trivial. Let's say you're on an H100. The maximum NVLink domain for H100s—for most DGX H100s—is 8 GPUs. If you scaled up past that, you're going to have to figure out ways to handle the fact that now, for the GPUs to communicate, you have to do it over InfiniBand, which is still very fast but isn't as fast as NVLink.
Speaker 1
Is it like 1 order of magnitude—like hundreds?
Speaker 2
It's about an order of magnitude. Um—
Speaker 1
Not terrible. Yeah, I need to remember the data sheet here. I think it's about 500 GB/s unidirectional for NVLink and about 50 GB/s unidirectional for InfiniBand. It depends on the generation.
Speaker 2
I just want to set this up for people who aren't familiar with these kinds of layers and transfer speeds. Maybe even just going a few steps back: most people are very familiar with seeing that you can use things like SGLang and vLLM on your laptop. You can just run inference. There's all—
Speaker 1
You can run it on that laptop.
Speaker 2
You can run it on a laptop, then you get to, “Okay, the model's got pretty big, right? GLM-5 doubled the size, so what do you do when you have to go from, ‘Okay, I can get 128 GB of memory; I can run it on a DGX Spark,’ then you have to go multi-GPU. Okay, multi-GPU—there's some support there. Now, if I'm a company and I don't have—I'm not hiring the best researchers for this, right?—but I need to go multi-node, right? I have a lot of servers. Well, okay. Now there are efficiency problems, right? You can have multiple 8-GPU H100 nodes, but is that efficient? How do you do that efficiently?”
Speaker 1
Yeah. How do you represent—how do you choose how to represent the model, right? That's a hard question everyone asks: How do you size—“Oh, I want to run GLM-5,” which just came out, a new model. There have been like 4 of them in the past week, by the way—a bunch of new models.
Speaker 2
You know why, right? DeepSeek.
Speaker 1
No comment. [laughter]
Speaker 2
Yeah, but GLM-5, right? We have this new model. It's of a large size, and you have to figure out how to both scale up and scale out, right? Because you have to find the right representation that you care about. Everyone does this differently. Let's be very clear.
Speaker 1
Everyone figures this out in their own way. I feel like a lot of AI, or even ML, is like this. I think people assume it should be an easy recipe to follow. There was a tweet a few months ago that said, “Why hasn’t fine-tuning as a service taken off?”
Speaker 2
That might be me. [laughter]
Speaker 1
It might have been you. Yeah. But people want it to be such an easy recipe to follow. But even if you look at an ML model—
Speaker 2
Specific to you.
Speaker 1
Yeah. And the model has so much tinkering, right? When you see a model that has however many experts in the MoE model, it’s like, “Why that many experts?” I don’t know. They tried a bunch of things, and that one seemed to do better. And I think when it comes to how you’re serving inference, you have a bunch of decisions to make. You can always argue that you can take something and make it more optimal, but I think it’s this internal calibration and appetite for continued calibration.
Speaker 3
Yeah. And that doesn’t mean people aren’t taking a shot at this, like Tinker from Thinking Machines—RL as a service. It also gets even harder when you try to do big-model training, right? We’re not the best at training when they’re pretrained. We saw this with Llama 3, right? They’re trained in such a sparse way that Meta knows there’s going to be a bunch of inference done on these, right? They’ll open-source it, but it’s very trained for what Meta’s infrastructure wants. They want to run inference on it a lot.
Now, the question to think about is, say you want to serve a chat application or a coding copilot, right? You’re doing a layer of RL, and you’re serving a model for a certain number of people. Is it a chat model or a coding model? So, Dynamo, back to that. It’s like—
Speaker 1
Yeah, sorry. We sort of jumped off and jumped back into that topic. Everyone has their own journey, and I like to think of it as defined by what model you need and what accuracy you need. Actually, I talked to N about this earlier. There are 3 axes you care about.
Speaker 3
What is the quality that you’re able to produce? Are you accurate enough, or can you complete the task with high enough performance?
Speaker 1
High enough performance, yeah.
Speaker 3
There’s cost: can you serve the model—or serve your workflow, because it’s not just the model anymore; it’s the workflow, the multi-turn interaction with an agent—cheaply enough? And then, can you serve it fast enough?
We’re seeing all 3 of these play out. We saw new models from OpenAI that are faster. You have these new fast versions of models. You can change the amount of thinking to change the amount of quality, right? Produce more tokens, but at a higher cost and a higher latency.
Really, when you start this journey of trying to figure out how you want to host a model, you think about 3 things: What is the model I need to serve? How many times do I need to call it? What is the input sequence length? What does the workflow look like on top of it? What is the SLA? What is the latency SLA that I need to achieve? Because there’s usually some constant—you know the SLA that you need to hit.
Then you try to find the lowest-cost version that hits all of these constraints. Usually, you start with those things and do a bit of experimentation across some common configurations. You change the tensor parallel size, which is a form of parallelism.
Speaker 1
I’d say it goes even deeper. First, you’ve got to think about what model you—
Speaker 3
It’s like a multistep design process because, as you said, you can choose a smaller model and then do more test-time scaling, and it’ll equate to the quality of a larger model because you’re doing the test-time scaling, or you’re adding a harness or something. So, yes, it goes way deeper than that.
But from the performance perspective, once you get to the model you need to host, you look at that and say, “Hey, I have this model. I need to serve it at this speed. What is the right configuration for that?”
Speaker 1
Do you guys see the recent paper I just saw a few days ago that said if you run the same prompt twice, you’re getting, like, double the success rate? The key thing there is that you give it the context of the failed try, right? So it takes a shot, and this has been basic guidance for quite a while: just try again because it tried—
Speaker 2
Just try again. Did you try again?
Speaker 1
It’s a paper from Google, if I’m not mistaken, right? I think it’s like a 7-page little short paper. The title is very cute, and it’s just like, “Yeah, just try again.” Give it the context of the failed attempt. You just say, “Hey, take a little bit more, take a little bit more information. Try and fail, fail—”
Speaker 3
That basic concept has gone pretty deep. There’s self-distillation RL, where you do self-distillation, you do RL, and you have past failure, and you know that gives some signal. People take “try it again”—not strong enough. [laughter]
Speaker 1
For listeners who listen here, Vivek and I run a second YouTube channel for our paper club, where—
Speaker 2
Oh, that’s awesome.
Speaker 1
Vivek just covered this self-distillation and all that. That’s why he’s so up to speed on it.
Speaker 2
I’ll have to check it out.
Speaker 1
Yeah, it’s just a good practice. Everyone needs a paper club where you read papers together, and the social pressure kind of forces you to—
Speaker 3
There’s a big inference reading group at a big—
Speaker 1
I feel so bad every time. He shared one of your guys’ pieces in that—I forget. He’s on my team, actually. Funny, there’s an employee transfer between us: he worked for Nate at Brev, and now he’s on my team.
Speaker 3
He was our head of AI.
Speaker 1
I’m always looking for, “Can I start another podcast that only does that thing?” Is there something here? I mean, I don’t think there are new inference techniques every day, so it’s like—
Speaker 3
You would actually be surprised by the amount of blog posts you see.
There was a period where it was like Medusa, Hydra, Eagle—you know.
Speaker 1
We have new forms of decoding. We have new forms of speculative decoding, and it’s exciting when you guys put out something like Nemotron, because I remember the paper on Nemotron-3. The amount of post-training, the amount of tokens that the GPU-rich can just train on—and it was a hybrid state-space model, right?
Speaker 3
Yeah, it’s co-designed for the hardware. One of the things was always that state-space models don’t scale as well when you do a conversion, or whatever the performance is, and you guys were like, “No, just keep training.” Nemotron shows a lot of that. Yeah.
Also, something cool about Nemotron is that it was released in layers, if you will, very similar to Dynamo. It was essentially released in an aggregated form. The pretraining and post-training datasets are released. The recipes for how to do it are released. The model itself is released, so you can benefit from us turning on the GPUs.
But there are companies like ServiceNow that took the dataset and trained their own model, and we were super excited and celebrated that work.
Speaker 2
Zoom is— [laughter]
Speaker 1
I think, just to add, a lot of models don’t put out base models. If that’s the case, why has fine-tuning not taken off? You can do your own training, but—
Speaker 3
Yeah, that’s true.
Speaker 1
You guys put out base models. I think you put out everything.
Speaker 3
I don’t know about base. Can base be cancelable?
Speaker 1
Base can be cancelable.
Speaker 3
Safety training. [laughter]
Speaker 1
Did we get a full picture of Dynamo? I don’t know if we—
Speaker 2
What I’d love is for you to break down the 3 axes you mentioned. What is prefill, what is decode, and what are the optimizations that we can get with Dynamo?
Speaker 3
Yeah, that’s a great point. To summarize that 3-axis problem, there are 3 things that determine whether or not something can be done with inference: cost, quality, and latency. Dynamo is supposed to provide you with the runtime that allows you to pull levers, mix it up, and move around the Pareto frontier, or the Pareto surface, that determines whether this is actually possible with inference and AI today.
Speaker 1
It gives you the knobs.
Speaker 3
Yeah, exactly. It gives you the knobs.
One thing that we use a lot in contemporary inference, and that is starting to pick up in general knowledge, is this concept of disaggregation. Historically, models would be hosted with a single inference engine, and that inference engine would sort of ping-pong between 2 phases.
There’s prefill, where you’re reading the sequence and generating the KV cache, which is basically just a set of vectors that represent the sequence. Then you use that KV cache to generate new tokens, which is called decode.
Some brilliant researchers, across multiple different papers, essentially made the realization that if you separate these 2 phases, you actually gain some benefits. Those benefits are basically that you don’t have to worry about step-synchronous scheduling. The way an inference engine works is you do 1 step, then you finish it, and then you start scheduling the next step.
Speaker 1
It’s not fully asynchronous. The problem is that prefill and decode are actually very different in terms of both their resource requirements and, sometimes, their runtime. You would have prefill that would block decode steps because you’d still be prefilling, and you couldn’t schedule because the step has to end. You remove that scheduling issue, and then you also allow yourself to split the work into 2 different types of pools.
Prefill is typically—and this changes as model architecture changes—compute-bound most of the time when a sequence is sufficiently long. It’s compute-bound on the prefill side because you’re doing a full pass over all the weights and the entire sequence every time you do a prefill step. On the decode side, because you don’t have the quadratic computation of the KV cache, it’s usually memory-bound because you’re retrieving a linear amount of memory and doing a linear amount of compute, as opposed to prefill, where you retrieve a linear amount of memory and then use a quadratic amount of compute.
swyx
You know what’s funny? Exo Labs did a really cool demo where, for the DGX Spark, which has a lot more compute, you can do the compute-hungry prefill on a DGX Spark and then do the decode on a Mac, and so—
Speaker 1
That’s faster. Yeah.
swyx
Yeah, so you can do machine stratification.
Speaker 1
With our future generations of hardware, we actually announced, with Rubin, this new accelerator that is prefill-specific. It’s called Rubin CPX.
swyx
I have a question. When you do the scale-out, is scaling out easier with Dynamo because, when you need a new node, you can dedicate it to either prefill or decode?
Speaker 1
Yeah. Dynamo actually has a Kubernetes component in it called Grove that allows you to do this crazy scaling specialization. It has this representation that I don’t want to go too deep into Kubernetes here, but there was a previous way that you would launch multi-node work. It’s called LeaderWorkerSet. It’s in the Kubernetes standard, and LeaderWorkerSet is great. It served a lot of people super well for a long period of time.
But one of the things that it struggles with is representing a set of cases where you have a multi-node replica that has a pair, right? You know, prefill and decode. Or it’s not paired, but it has a second stage that has a ratio that changes over time.
swyx
And prefill and decode are 2 different things. As your workload changes, the amount of prefill you’ll need to do may change. The amount of decode that you’ll need to do might change, right? Let’s say you start getting insanely long queries. That probably means that your prefill scales harder because you’re hitting this quadratic scaling growth.
Speaker 1
Yeah.
swyx
For listeners, prefill will be long input and decode will be long output, for example, right?
Speaker 1
Yeah. Decode scale—I mean, decode is funny because the amount of tokens that you produce scales with the output length, but the amount of work that you do per step scales with the amount of tokens in the context.
Alessio Fanelli
Yes.
swyx
So it both scales with the input and the output.
Speaker 1
That’s true. But on the prefill-decode side, if suddenly the amount of work you’re doing on the decode side stays about the same or scales a little, and the prefill side jumps up a lot, you actually don’t want that ratio to be the same. You want it to change over time.
So Dynamo has a set of components that tell you how to scale. It tells you how many prefill workers and decode workers it thinks you should have, and it also provides a scheduling API for Kubernetes that allows you to actually represent and effect this scheduling on your actual hardware, on your compute infrastructure.
swyx
Not going to lie, I feel a little embarrassed for being proud of my SVG function earlier. [laughter]
Alessio Fanelli
No, it was really cute. I like—
swyx
It’s all engineering. It’s all engineering.
I’m technical. One thing I’m curious about, seeing everything that’s going on here at a systems level, is that we’re scaling it up in distributed systems. I think one thing that’s kind of the moment right now is people are asking: are there any upper bounds in terms of—let’s just call it context length, for want of a better word—but you can break it down however you like?
Speaker 1
Yeah. I just think—well, clearly you can engage in hybrid architectures and throw in some state-space models in there all you want, but it still looks very attention-heavy.
Yes.
Speaker 1
Yeah, long context is attention-heavy. We have these hybrid models—
swyx
And most models cap out at 1 million context, and that’s it. For the last 2 years, that’s been it.
Speaker 1
Yeah, the model-hardware-context co-design thing that we’re seeing these days is actually super interesting. It’s my passion, my secret side passion. We see models like Kimi or GPT-OSS. I’m going to use these because I know specific things about these models.
Kimi K2 comes out, right? And it’s an interesting model. It’s a DeepSeek-style architecture. It is MLA. It’s basically DeepSeek scaled a little bit differently, and obviously trained differently as well. But they talked about why they made the design choices.
For context, Kimi has more experts but fewer attention heads and, I believe, a slightly smaller attention dimension—but I need to remember; I need to check that. That doesn’t matter, but they discussed this at length in a blog post on Juejin, which is—
swyx
In Chinese.
Speaker 1
Yeah. So it’s actually an incredible blog post. All the ML systems people I’ve seen on Twitter are very brilliant, but the creators of Kimi K2 actually talked about it in a blog post, and they say, “We actually did an experiment around attention.”
Attention scales with the number of heads, obviously. If you have 64 heads versus 32 heads, you do half the work of attention. You still scale quadratically, but you do half the work. And they made a very specific trade-off in their system, in their architecture.
They basically said, “Hey, what if we gave it more experts?” We’re going to use more memory capacity, but we keep the amount of activated experts the same. We increase the expert sparsity, so we have fewer experts active. The ratio of experts activated to the number of experts is smaller, and we decrease the number of attention heads.
swyx
And, for context, what we’d been seeing was that you make models sparser instead. No one was really touching heads. You were just having—
Speaker 1
Well, they implicitly made it sparser.
swyx
Yeah, for Kimi they did. They also made it sparser, but basically what we were seeing was people were at the level of, okay, there’s a sparsity ratio: you want more total parameters, less active, and that’s sparsity.
But what you see from papers from labs like Moonshot and DeepSeek is that they go to the level of, okay, outside of just the number of experts, you can also change how many attention heads and fewer attention layers, more attention layers—
Speaker 1
Yes, yes.
swyx
So that’s all basically coming back to, just to tie it together, hardware-model-code design, which is—
Speaker 1
Hardware-model-context-code design, right? Like, if you were training a model that was really, really short-context—
Alessio Fanelli
Or, like, really good at super-short-context tasks, you may design it in a way such that you don’t care about attention scaling because it hasn’t hit that turning point where the quadratic curve takes over.
swyx
How do you consider attention or context as a separate part of the code design? I would imagine hardware-model-code design would be hardware-model-context-code design, because the harness and the context that is produced by the harness is a part of the model once it’s trained in.
Even though towards the end you’ll do long context, you’re not changing the architecture through training—
Speaker 1
I mean, you can try. [snorts]
swyx
You’re saying everyone’s training the harness into the model?
Speaker 1
I would say to some degree.
swyx
Or there’s code—
Speaker 1
I know there’s a small amount, but I feel like not everyone has gone full send on this. I think it’s important to internalize the harness that you think the model will be running into the model.
swyx
Interesting. Okay.
Speaker 1
Bash is like the universal harness.
swyx
Yeah. I’ll give an example here. I mean—or just an easy proof, right? If you can train against a harness and you’re using that harness for everything, wouldn’t you just train with the harness to ensure that you get the best possible quality out of—
Speaker 1
Well, I can provide a counterargument, which is: you want to provide a generally useful model for other people to plug into their harnesses, right? So harnesses can be open source, right?
swyx
I mean, that’s effectively what’s happening with Codex. Yeah.
Alessio Fanelli
But you may want a different search tool, and then you may have to name it differently, or—
Speaker 1
I don’t know how much people have pushed on this, but can you train a model—would it—have people compared training a model for the harness versus post-training for—
I think it’s the same thing, just extra post-training.
I see. And so, I mean, Cognition does this, Cursor does this, where you just have to—if your tool was slightly different, either force your tool to be like the tool that they trained for, or undo their training for their tool and then retrain.
Speaker 1
Yeah, it’s really annoying, and, like—
I would hope that eventually we hit a certain level of generality with respect to understanding new tools.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
It's not AGI. It's just a really stupid "learn my tool" [__]. I don't know if I can say that [laughter], but my point is that I look at the slopes of the scaling laws, and this slope is not working, man. We're at a 1-million-token context. Maybe next year it's 2 million, but we're not going to 100 million.
Speaker 1
This doesn't work. This doesn't work. What's kind of funny is that we always want to see a trend that we can predict, but every time something comes, it's been a leapfrog. I don't know how we go from 1 to 2, but I imagine what's likely to happen is that we break through that with something new.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah. There's actually an interesting formalization of this. There's an essay—it's a pretty interesting essay—by Leopold Aschenbrenner called Situational Awareness.
Speaker 1
Okay. Yes.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
He introduces a concept called an unhobbling.
Speaker 1
Right. So Leopold, in this essay, details, "Hey, I want to get to this point in intelligence, and I think that it's 4 orders of magnitude worth of compute, data, and training away." He says, "I think data centers can scale up by about this much. I think you can scale up the data and some other things by this much." But one of the things that makes the rest of that order-of-magnitude growth possible is unhobbling: scientific discoveries made during model architecture search or training that really, really, really impact how you're able to scale.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
A good example of this might be that we see a lot of models trained with multi-token prediction natively during pretraining. In the DeepSeek paper, they say, "Hey, this actually helped us ensure more stable convergence." There are unhobbings like that, and then there are rather large unhobbings.
Architecturally, a lot of our models have different types of attention. One of the problems with attention is that you have a lot of KV cache, but people have found different forms of attention, like grouped-query attention and MLA in DeepSeek—multi-head latent attention—that decrease the burden that KV cache places on the model, which allows you to grow longer in context.
Speaker 1
Yeah, and that's very drastic for DeepSeek.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah. For context, the total context length of DeepSeek is 128,000 tokens, or it might be 256,000 with RoPE extension. That entire 128,000-token context fits into 8 GB. Previously, a context of a similar size on Llama 405B was 40 or 80 GB at the same precision.
Those unhobblings really decrease the cost of that size. I wouldn't be surprised if we see the ability to break through to 10 million, 20 million, or 100 million tokens of context through an unhobbling showing up.
Speaker 1
And it's just science. So more deep-learning algorithms is what it is.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
More deep-learning algorithms. [laughter]
I could actually give you an example of something theoretical—not a theory-theory, but something theoretical—
Speaker 1
An unhobbling that you're excited about?
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Well, an unhobbling that—I haven't seen it, so it could be a tarpit and could just not work. But I'd be really excited to see a model that does prefill and decode differently: a model that does prefill locally, document-wise, in chunks, and then does decode globally across the entire sequence.
Logically, it doesn't seem like you'd necessarily need KV to be associative between documents that have no mutual association. But that places a lot of burden on prefill—or, sorry, on decode, and on pure attention within the decode phase, to make those connections, since the KV is static at that point.
You see other techniques that are interesting like this, too. If prefill becomes local and decode is still global, you solve that prefill quadratic-scaling problem because you have a bunch of small chunks that you prefill independently.
Speaker 1
Okay. All right. Well, let's wait and see. But I think it'll be pretty exciting.
Speaker 2
Fingers crossed. Yeah. Yeah. Yeah.
Speaker 1
I'm excited for prefill and decode on separate hardware. With the Groq acquisition, can we decode on Groq? Can we get super fast? [laughter]
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
I don't think I'm allowed to comment on this.
Speaker 1
And Mark is going to shoot arrows at us.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, he's in the room. [laughter] I'm super excited to see the team come in. I've had the pleasure of working with some of the Groq people coming in, so I know—
Speaker 2
Sunny—we've had him at the same conference that you're at.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
On the Groq side, yeah, I use the associated NVIDIA U [?].
Speaker 2
Yeah, I think there's only one PM for Claude Code, and it's Sky Wu. The rest of them—there's Dev [?], there's Boris, maybe.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Exactly. I mean, let's go into agents. I think this was the last part of the discussion we planned.
Speaker 1
How have we not talked about it? [laughter]
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
You guys scheduled it. I was like, "Okay, let's have cohesive sections."
Speaker 2
I mean, there's big news, right? NVIDIA has a huge deployment of Codex.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
It uses everything, and we use Cursor, and we use—
Speaker 2
That's a pretty big deployment, right? That's tens of thousands of people.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Totally, yeah. I mean, it goes back to the mosh pit of emails we mentioned earlier, or just how fluid the organization feels. When there's new technology, people will email it out and everyone will try it. If it makes people's lives easier, it'll spread like wildfire.
A lot of times Jensen will get it and be like, "Let's make this work across the company. Let's make this work right now."
Speaker 1
Honestly, if I were a startup, I feel like a cool hack is that if you have something that's going to save an NVIDIAN's time, they'll spread it to a couple of people in the same way, right? It'll just spread like wildfire.
Speaker 2
Careful before your email blows up from startups. [laughter]
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Well, you've got to know the person, right? But no, I love using Codex. It's been a ton of fun. I've been using it personally and at work. It's been great to see the rollout.
Speaker 1
Oh, yeah.
NVIDIA's AI Engineers
Agent Inference at Planetary Scale and "Speed of Light"
Something really funny: on the access we got to Codex and Claude Code, I found this person at the company—his name is Carlos. He wrote an Outlook CLI, just a CLI for email. And this was—
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
I've been using that.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, like 4 or 5 weeks ago. Once I got Codex access, I installed the CLI. It had a skill, and I asked it to go through all of my emails, which are very messy. So if I don’t respond to your email, I’m really sorry, but I asked it to give me a summary, highlight any escalations that I should look at, put any thread that it thinks I should respond to in a folder, and then archive everything.
And it did. So if I missed your email, it’s because it didn’t get to it. [laughter]
swyx
So I should put a prompt injection in my emails?
Yeah, what you should do is just FaceTime. [laughter]
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah. My SLA is highest on FaceTime, but that was magic. I sent it in a big email thread to 500 people. A bunch of folks tried it out, and I started FaceTiming whoever I could at the company to get them set up with this.
swyx
That specific example—you guys deal with some pretty sensitive emails. Is there a security review with this? One guy made it for himself, but it’s not meant for all the—
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
The security team at NVIDIA is incredible. Shout-out to them. They’re trying to—we have an amazing security team because they’re progressive, and they know that this is really important technology and we have to bring it in.
If you think about working at a big company, your laptop is usually very locked down, and you can only access certain things. NVIDIA engineers don’t have those restrictions, so you’re expected to understand the risks when you try things out. Very quickly, we made sure to involve security in what we were doing.
There’s actually a lot that we’ve been thinking about, especially with OpenClaw. Agents can do 3 things: they can access your files, they can access the internet, and now they can write custom code and execute it. You should really only let an agent do 2 of those 3 things.
If it can access your files and write custom code, you don’t want it to have internet access, because that’s one source of vulnerability. If it has access to the internet and your file system, you should know the full scope of what that agent is capable of doing. Otherwise, malware can get injected or something can happen.
A lot of what we’ve been thinking about is how we both enable this, because it’s clearly the future, but also what enforcement points we can start to protect.
swyx
There’s certainly a directive like, “Hey, we have a company account or company agreement with OpenAI. We use OpenAI models here,” or choose whatever.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
No, no. I would never put any company data in a model that we haven’t vetted. It has to go through security.
swyx
Contrary to that, obviously you could run your own models. You have Nemotron, and you have an internal cluster.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
I think we’re Dynamo’s first customer. [laughter]
Actually, there’s a funny story about how I got the experience that informed what we needed for Dynamo. At one point, there was a website called build.nvidia.com, which allows people to try models. It gives you an API service, so you can call the model with a REST API and get a response.
I ran the model side for that, and at one point it was the largest inference deployment. It still may actually be the largest inference deployment. I’ve since handed that off to some people, and they’re doing a wonderful job.
This is an extremely underknown or less-known resource: build.nvidia.com. You can get any of these open-source models. It’s rate-limited, but it’s free, so it’s perfect for hackers.
swyx
The SLA on getting day-zero models up is about a day. They’re incredibly good at figuring out the right way to host the model and get it up there as soon as it comes out. You ran this?
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, I ran it a long time ago. It was originally called NVIDIA AI Playground. Then it was called AI Foundation, and then it was called Build. I ran the model side of it.
There was a large, multi-organizational team. I ran which models we should host, how we should host them, and what the proportion of them should be. Then, of course, there was an SRE team that made sure things ran well and scaled the models. I ran the model side: how do we get the model to silicon, and which models are important? I also worked with our product team to determine which models were important.
swyx
There’s also a middle ground in between, right? For the hacker who wants to try anything, there’s the Brev console, then there’s Dynamo. There was also NIM, right? I remember it had its little moment a year or 2 ago. Is it still—
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, NIM is inference software.
I think it’s no longer an acronym; it’s just a name. NIM is how enterprises can take any of this technology and run it with support and everything else. That includes Dynamo, as well as our other optimizations that are packaged up for enterprise.
swyx
Yep. Anyway, you got a bunch of experience running internal inference gateways and playgrounds.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah.
swyx
Bill also helped build NVIDIA’s first internal VS Code thing. We called it NVCode.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
It’s like the extension. Eventually, it was a fork of VS Code.
swyx
We joked a while back that we should have a fork-of-VS-Code hackathon, where you build the best fork of VS Code.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
We were doing a “Make a Billion Dollars”—
swyx
Someone from VS Code was there, and he was somewhat down to get involved. I was like, “Oh, you should do that.” Then the cool thing became a Fork Chrome hackathon.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Chrome and IDEs are not cool.
swyx
I was talking to Joseph from Roo Code, your partner in crime. We were talking about how, with the new Alpamayo model, NVIDIA just released an open-source autonomous-driving model. The Mercedes cars that you saw drive—sounds crazy.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, so we were thinking: could we hackathon a driverless car? I have my old car. Let’s just try it. [laughter]
swyx
We could even have a race. It’s the first person to automate their driving over a weekend.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
We could take it to Treasure Island in the middle of the day and just see. How many cameras do we need? 1, 2, 3, 4? I don’t know. I think we’re going to try it. You should do it with us.
swyx
We do have an autonomy track. It was at a fair, and Waymo was there.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, NVIDIA did send people. It was for GR00T, not because we didn’t have the driving thing yet.
swyx
Yeah, that’s cool. I think Wayve also has a version of this. They have open-source driving, and they’ve done a fun hackathon.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
I really want a Tesla with Tesla-level self-driving—
swyx
But as a Smart car—as a 2-seater that’s basically a wheelchair with a roof. The demand has always been there. They’ve been saying this for about 5 years.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
swyx
Really?
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, they were a different manufacturer.
swyx
I thought it was one of those things where we’d see someone buy the brand and revive it. I would buy it.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Someone hears this and goes, “Buy your car.” That’s crazy, because Mercedes—
swyx
Because I think Mercedes owns the brand, and they’re saying they’re going to make them.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
I don’t know. I feel like they own the brand, and your dream might come true.
swyx
Every time I try to park in San Francisco, I have to buy a Smart car, because 20% of the parking lots in San Francisco only fit Smart cars.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
swyx
Really?
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
That’s what I mean.
swyx
This comes from someone who basically does not drive.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
That’s where the Vespa was a life hack.
swyx
Yeah, exactly. What happened to the Vespa?
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
I used to have this yellow Vespa. I left it outside the hacker house when we moved out. It was always there, and then about a month ago it wasn’t there anymore. I’ve been meaning to—
swyx
You forgot about it.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah, and left it.
swyx
Yeah, no, this is probably hazardous.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Speaking of hackathons, I also wanted to give a big shout-out to the world’s shortest hackathon.
swyx
Let’s go. You did it twice—a handful of times.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Yeah. There’s going to be one at GTC. Oh, we’re doing L.A.
We have a bunch of challenges that we haven’t released, and you get to bring your agent to attempt to go through those challenges.
swyx
That’s like a zero-minute hackathon idea. You just bring your agent and press the go button. You’re not allowed to code; it’s just the agent doing the hackathon.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
It’s a good hidden eval, right?
swyx
Yeah. You make a repo, and I feel like this is something I would love to see from Cognition or someone else: “Come bring your agent. Drop it in.”
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
Because you don’t know what the task is. Let’s be—operate a browser, order a pizza, or play a Snake game. We’ll just see.
swyx
And you don’t know what the task is. You don’t even know what the judging categories are. Then we give it the judging categories: try to win as much as possible.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
It’s great, though. It turns into, “Let’s build something on Dynamo.”
swyx
It’s a great proposition. Anyway, funny story, actually.
NVIDIA's AI Engineers: Agent Inference at Planetary Scale and "Speed of Light"
We have a couple of people at NVIDIA. We've been working with security to bring agents really close to compute. So now we have stuff where we can tell Dynamo, “Go run some experiments with Dynamo on XCluster and try it right now.” Once you get queued, send this request load.
We've actually been able to one-shot problems. We used to have this problem where, with Dynamo, you had to find the right configurations. We do some of that automatically, but you still need a good initial configuration to use. We've had an agent completely one-shot that. It gets the compute, runs a couple of experiments, and says, “This is the best. These are part of the Pareto frontier. Go run this.” Then we give that to people, and it's faster than anything they have.
Agent UX and agent marketing are super important. There's something we've been thinking about a lot. Alec is redoing the entire Brev CLI so that you can fetch all the different compute types that are available. I don't know, it's going to be really soon, but then you can browse what GPUs are available, provision one, SSH to it right there, and pipe all the commands.
Speaker 1
But I think it goes back to the Alex CLI. Coding agents have been so much more effective than general-purpose agents, and I think a large part of that is that they just have access to the terminal, like you said. That means they have access to everything you've installed into your terminal.
They can run things. They can write code, compile the code, and, if there are errors, fix them. They can run your suite of tests because that's all just in your terminal.
Speaker 1
What got me really excited about the Outlook CLI is that we're now turning through building CLIs for the entire business suite—building CLI SAP, Go. I've also done that for myself. Really? Yeah.
We're going to open-source all of this. These are CLIs for the business applications, and we would love for someone to run with this and build an OpenCLI Foundation or something. NVIDIA would love to support anyone doing this.
Every dev tool should really have good CLI support at this point. At one point, it was, “You want your docs to be accessible by an LLM, right? You want LLM-friendly documentation.” No, everything needs some CLI support.
swyx
It's kind of funny, right? Computing began with a terminal, with a shell, but we said it wasn't empathetic to humans, so we built these nice user interfaces. Now we have LLMs navigating our user interfaces. Ironically, we're not empathetic to the machine anymore.
Yeah. Just give the LLM access to the shell.
swyx
One thing that slightly makes me uncomfortable is: why do we have to build CLIs? Why can't we just expose APIs?
I have an interesting answer to this. There are a couple of reasons. Portability is one issue. Sometimes APIs aren't discoverable or reachable by some types of systems. There's also an element of locality. The CLI is literally you interfacing with your local system, which is a little bit different.
You could still do it through an API, but this highlights the difference between a CLI and an MCP. They kind of occupy the same purpose: you call them, they do something on the system, and then they're done.
swyx
I think in pretraining there's just an enormous amount of command-line data.
Yeah.
swyx
Even if we ignore RL—let's say you're doing no harness or post-training—the amount of CLI versus API documentation for navigating the world of the CLI and your file system is enormous.
Alessio Fanelli
Yeah, right. I think there are a couple of things, too. Your intuition is right: the CLI is just wrapping the API, functionally.
swyx
Functionally, right.
I think it's nice because, first, you're being very specific and even pedantic about what you're doing, and that's really good. You're describing the problem space, so you know the surface area for vulnerability and what network calls you're making. It's not arbitrary or decided on the fly; it's predetermined, which is important from a security perspective.
If you were to write a bunch of API requests, how would you do that? Would the model use Python? I kind of like that everything with a CLI is just Bash because it's ubiquitous. It's just there, and you don't have to make sure certain environment variables are set up.
If your Python version is different from my Python version, we're using the same model to do the same thing. Is it going to write different code? It probably would. So it's nice to work with a human as well.
swyx
No, I think it's about making those decisions happen ahead of time.
One last thing on this sort of agent co-location, or whatever you call it: one pattern I'm tracking for this year is that I always try to think about what the theme of the year is going to be. Last year was definitely coding agents. This year is definitely coding agents breaking out of containment into broader ways.
It definitely has to rent a human.
Yeah.
swyx
Yeah, I'm on there. Are you really?
Alessio Fanelli
I'm like $5,000. I'll do anything.
swyx
Really?
I think so. I need my bowls from Costco.
swyx
But I think the best part is that only the agent can book me.
It's basically just another labor marketplace.
swyx
Mechanical Turk was this.
So I have a weird story about why I did it. Going back to your example of just giving an agent access to compute: you guys are GPU-rich at NVIDIA.
swyx
He's not shy about it.
I have a 24/7 agent running. I hooked it up to Runpod, and it doesn't shut down instances. I've tried prompting it and giving it instructions: “Shut down when you're done.” But it's like, “Keep it warm. I'll need it soon.”
It's horrible at time estimates, too, because it realizes, “I'll need it in 45 minutes.” Forty-five minutes of human time is actually 3 minutes of agent time. So I'm booting it up and waiting, and I'll just leave it on all night. Modal's good at shutting down after some inactivity.
I had it on my local server, a little dual-GPU thing, and it just stays on. I have a little space heater at home now. So basically, they don't care about the concept of money. Just burn it. “I need it. It's useful.”
Another DGX Spark would be really nice. I think it's super useful for agents because you buy it once, plug it in, and then it can rip.
swyx
I'm going to make an NVIDIA ad here.
Okay. The Blackwell RTX PRO 6000 cards are only, I think, $8,000.
swyx
PRO.
Alessio Fanelli
PRO are only, I think, $8,000. Slightly cheaper.
Yeah, it's much cheaper than the data-center card.
swyx
Yeah.
Alessio Fanelli
And it's got 96 GB of VRAM. So if you and your crew want to run a local agent in your home, it's got a significant amount of VRAM. I've thought about purchasing this and running it in my basement, except my neighbors would hate me.
swyx
It's just a single two- or three-slot GPU.
Yeah, it's a PCIe PC GPU. You can go buy that. The big difference compared with RTX gaming GPUs is that it's Blackwell, obviously, and it's a professional GPU. It has a lot of VRAM, which means you can run pretty large models on it.
swyx
You can stack 4 of them for the Max-Q in a system.
That's a beast.
swyx
It's beefy. You can run—what is it?—96.
You can run anything with 96 GB. You don't lose speed.
swyx
But they are slow. Their performance and speed will be somewhat slower compared with an API.
Oh, yeah, that's true. Again, a big fleet and economies of scale allow you to get both speed and throughput. You can run—I'll give an example—an optimization called WideEP. I'm not going to go into it fully, but it featured heavily in InferenceMAX for DeepSeek.
There's a great set of stories from NVIDIA and SemiAnalysis about why WideEP is important. For MoE models, it's basically essential. The level of parallelism and the level of scale-up parallelism used for it go beyond a certain barrier, and it really, really is important to have an NVL72 GB200 NVLink system to serve at scale.
I don't remember the cost improvement against Hopper, but with this NVL72 system, you're getting something like 35 times cheaper per token for a lot of the curve.
swyx
Yeah.
Alessio Fanelli
Which is crazy.
swyx
Yeah.
Alessio Fanelli
And that's normalized per GPU, obviously, because the GPU is part of the cost.
swyx
One thing I'm exploring is that this year is also the year of the subagent, where you have the main agent, but it also kicks off tools that are themselves agents with limited capabilities and different prompts.
For example, one thing Cognition does is, before you kick off a search, they use a fast context model. You kick off a process to search across the codebase. A lot of the time, that's better than indexing—not all the time. You should still index for some things—but the idea is that agents should be able to command subagents and probably run them close to inference as well.
swyx
I don't know if that's architecturally possible, or even—
Yeah, we're thinking about that for Dynamo. That's our big theme for the year, because if you can design that into your systems, a lot more people will use it. Right now, it's mostly theoretical because you pay a lot of back-and-forth coordination cost.
swyx
I think you'll net speed up, though, right? Even at a basic level, with speculative decoding, you're running a small model. You're running 2 instances, but it's a net speedup.
That is one example. Yes.
swyx
Yeah.
Alessio Fanelli
But this is a little different with agents.
swyx
Agents. Yeah. This isn't specific.
I think there's a summary of that trend that I like to give to my team: this is the year of “system as model.” Instead of having a single model be the thing, you have a system of models and components working together to emulate the black-box model. When you make an API call to something that's multi-agent in the background, it still looks like an API call to a model. You're still getting back—
swyx
But under the hood—
Yeah. Under the hood, it's a billion different models. That's a lot of complexity, right? With Dynamo and other libraries on NVIDIA, we're looking to help manage that complexity.
It's funny—we actually just released the model router for CES. With DGX Spark, you can have a local model running on the Spark and a foundation model, and then the model router decides when to send queries to which one. It's no longer either-or; it's about using the best of everything available to you.
swyx
You have a good post-trained model that's running on—
It also leads to the broader functionality of being able to manage the Spark.
swyx
Oh, that'd be cool. Oh, yeah—
As requested. There we go.
swyx
I'd like to extend and flip the question around. How much longer do you guys think agents are going to be running? That's one thing I've been throwing around: what happens when—
Speaker 1
I mean, always longer.
swyx
It even affects speculative decoding, right? Codex, compared to Claude Code, runs much longer tasks. That thing will run for 6, 7, or 8 hours. I'll run it overnight and come back. I have some crappy logging software I use, and there are times when it wants to go deep on research. It'll eat up 80,000 tokens, go on another run, go on another, and just eat through tokens. That's part of it: at the end, it does hit a long task.
Speaker 1
I think you only see that because there's insatiable demand for tokens, and every improvement that comes along just makes our demand even higher. It's kind of funny, right? If you have a teammate and ask them to do a task, are they going to save some effort and not think too hard about it? They're like, “Fuck no, I'm going to do my best.” You can have 4 shots, right? With the original Codex, before the app, why do 1 call? Give it 4 attempts. Just use all the tokens.
swyx
Try more.
Speaker 1
Try again. Try more.
Speaker 1
It's like the METR index, right? It's the thing that tracks how long models are able to run. I expect we'll see log-linear, if not log-superlinear, growth. Before the end of the year, we'll see an agent capable of running for longer than 24 hours with self-consistency the entire time.
Speaker 2
I would also point out that different domains have different desires. At a consumer level, I'm getting slightly frustrated at 20 minutes per basic query. Sure, you can optimize for 6- or 8-hour tasks, but I don't see myself launching many 1-week agents. If someone is doing GPU-kernel research, or medical or biological research, then sure, launch something that takes a long time. I think it will be somewhat domain-specific, because you also really need to train that into the system.
Speaker 1
What's doing your taxes?
swyx
Right. That's taxes—get it right.
Speaker 2
I wonder if that's what speculative decoding is: your agent figuring out what you might prompt it with the next day at night and prefetching it.
Speaker 1
Yeah, you can already do that. Branch prediction.
swyx
Oh, well, no. That's too low-level, but yes.
swyx
Sorry. One question I've got to get in. We actually did record a podcast with the METR folks right here. Their chart is about the human-equivalent hours of work, rather than how long the agents themselves are autonomous. There's a huge difference: 5 hours of human work versus 30 minutes of agent work.
Speaker 1
Yeah, 5 hours, right.
swyx
That chart is estimating the human-equivalent replacement. I think Anthropic released a more recent chart that showed Claude Code autonomy from their production traffic numbers, and that was 20 to 45 minutes.
Speaker 1
That's roughly where we are. That's the realistic thing. I do think there are experimental setups where you can use a Ralph Wiggum loop—just prompt it to keep going when it stops—and obviously that can go arbitrarily long. From my experience, around 20 to 40 minutes seems right for when I'm using Codex or Claude Code.
Speaker 1
When I want to spin up a new project, I'll often start with Replit, and it'll get into the project, I believe. With their new v3 agent, it'll spin up a web browser, click around, discover new bugs, and just keep churning. I think my longest run was over an hour.
Speaker 2
Before we see super-long-running agents, I think there's going to be an efficiency hit. Sure, you can take an hour and go down different paths, but you also want to be more efficient and smarter in your reasoning. I think that will actually go down before it goes back up. You don't want to scale nonoptimized systems just for the sake of it. As much as I love saying, “Use all the tokens,” they are expensive. Going from dense models to reasoning models adds cost. You're paying for a lot of tokens, and it doesn't make sense to scale systems that aren't optimized. There's always that balance.
swyx
Yeah.
Speaker 2
But I think you'll see both sides of it.
swyx
Yeah. So, 2023 was super exciting. If you were in SF, you were like, “Okay, I know this is going to be a huge, world-changing moment,” but it seemed like no one knew it yet. Maybe even before that—was it 2022? Maybe.
Speaker 2
Yeah. Roon had this tweet about how everyone who was in SF from 2021 to 2023 understood what it was like to be early.
swyx
Totally. Yeah, 2021—that's when I made my first OpenAI account. It was crazy. At the time, SF had not been doing well, so it felt like the concentration of founders in the city had risen. My neighbors used to be doing all sorts of things, but those people had all left. The only people who were still in the city were people who really wanted to build. It was cheap, too.
Speaker 1
Yeah, it was also way cheaper. I feel really bad for anyone who's trying to get rent now.
Speaker 1
Celo had a huge office. The blockchain company took over the old Casper building.
swyx
Yeah, they had the showroom and what I think was the back warehouse. It was a huge office, right across from OpenAI.
Speaker 1
Yeah, it was in the original arena.
swyx
I named the arena because of it.
swyx
Yeah. Rooflow [?], Mintlify, and Brev were there. You guys were there. I remember that was actually where you bought the AI.engineer domain.
Speaker 1
It was a really fun moment when we were all in this SoMa space. I don't know—it was a really cool community, especially being so early.
Speaker 1
And so you got me early Cruise access.
swyx
Oh, yeah.
Speaker 1
There was a long period of time when both Cruise and Waymo were just free.
swyx
Yeah. Always.
Speaker 1
I mean, they were so back.
swyx
So now Zoox is doing—
Speaker 1
Zoox's robotaxis. Yeah, totally. It's actually really cool that you guys have this studio so close to Celo, with this rock-climbing gym right around the corner. It's an awesome block.
swyx
Yeah. It's just a bit of San Francisco. I do think one thing I try to do with the podcast is bring what it's like to be in San Francisco to the rest of the world, and maybe also give a taqueria a shout-out.
Speaker 1
Yeah, my favorite tacos in the city: steak and shrimp. I know. They're very good.
Speaker 1
Yeah. And I guess what it's like to be in San Francisco is that everyone seems to be super supportive. Sometimes I feel like the city believes in you more than you do. I don't know if you remember, but I remember posting my first blog post. I had met you on Twitter, and you gave me an hour of your time, completely randomly. You coached me through writing content for developers. I was trying really hard not to come off as salesy or plug myself, so I stripped all the personality out of the blog post, and you brought that out.
Alessio Fanelli
You're like, “People don't care. It's okay to talk about what you're doing. You don't have to be weird about it.” I remember that really helped me figure out what our voice is and not shy away from it, so I'm always really grateful for you injecting your voice into everything. [Laughter] It's actually a huge advantage to be very genuine about what you care about.
swyx
Yeah, imagine some random person DMs you, like, “Can you give me feedback on this blog post?” It's pretty boring, and you're like, “Fine. He looks interesting. I'll just do a Zoom call.” Then you meet this guy. Yeah.
He's so energetic, just being right there. But I think people are trained to write a certain way in school, and they never—
swyx
Totally.
Alessio Fanelli
See, there's a broader world out—
swyx
To unlearn.
Writing is thinking, and everyone thinks differently, so you might as well—
swyx
Write your way.
Cool. Well, thank you for indulging us. Really broad-ranging discussion, but I love that you guys are sort of the young faces on video with so much energy, but also a lot of technical depth. I think people can learn a lot from this session. So, thank you. This is awesome. Thank you guys, and thank you for everything that you've done in the talk.
swyx
Yeah. The podcast, all the above, and CTCard[?] to it.
Yeah. Cool. Awesome. Thank you. Thank you.