Ben Thompson
And again, maybe you would go back and say, “I wish that had never happened.” That’s fair enough, but it has happened. What I think is better and healthier is to appreciate that all the things that bother you are ultimately human issues that have been exacerbated or accelerated by technology. By anthropomorphizing technology as good or bad and saying we need to make it good or bad, you’re assuming way too much power that you don’t have, and you’re actually skirting closer to even more disastrous outcomes than what may or may not have happened.
Andrew Sharp
Yes. And OpenAI should incorporate ads into the ChatGPT product at some point in 2026.
Ben Thompson
Yeah, it’s a big problem.
Andrew Sharp
Yes. I look forward to chronicling that particular adventure as the year unfolds.
1. Nvidia Licenses Groq Technology
For now, we will move to Concrete News. Exiting the dorm room here, from Bloomberg on Christmas Eve: Nvidia agreed to a licensing deal with artificial intelligence startup Groq, furthering its investments in companies connected to the AI boom and gaining the right to add a new type of technology to its products. The world’s largest publicly traded company has paid for the right to use Groq’s technology and will integrate its chip design into future products. Some of the startup’s executives are leaving to join Nvidia to help with that effort. The company said Groq will continue as an independent company with a new chief executive, it said Wednesday in a post on its website.
So, Ben, I just want to say: dropping this news at 4:00 Eastern on Christmas Eve is one of the greatest news dumps of my entire life. I still can’t believe it really happened.
It’s really a—what is the sort of meme about my T-shirt? It’s raising questions that are—
No kidding. And the timing of me putting on that T-shirt. TikTok did that a couple of years ago, where it was like 2 days before Christmas and they announced, “Oh yeah, our internal investigation—we did find out that we were actually spying on people.” But Nvidia beat TikTok here. They went on the afternoon of Christmas Eve.
A lot to talk about here. We don’t have that much time, but can you walk me through what Nvidia is acquiring with Groq and why Groq? We could cut the whole dorm room segment. I was just going to be backpedaling for the rest of this episode.
No, it was a wonderful—I enjoyed the dorm room segment. We’ll be back to the dorm room at some point in the next couple of weeks. But Groq and Nvidia: why does this make sense?
Ben Thompson
We’re slowly developing a nomenclature. We have the mullet section. I guess we have a dorm room section.
Andrew Sharp
Occasional visits to the dorm room. And, of course, we have TikTok, our parenting segment that we revisit too infrequently. But TikTok aside, talk to me about Nvidia here.
2. Groq Uses Deterministic Computing
Ben Thompson
Okay. What is Groq? Groq has a very different approach. Groq was actually started before the LLM AI explosion. It was focused more on machine-learning tasks generally, which are not as memory-intensive as LLMs are.
The idea with Groq is sort of software-defined inference. The first thing the Groq team built was not a chip; they built a compiler.
Andrew Sharp
Yeah.
Ben Thompson
A compiler is where you take code and change it into the 1s and 0s. It’s actually machine code—the 1s and 0s that actually run on the chip.
The idea was that if you have a super-deterministic algorithm that’s very simple—which these algorithms are, and that’s why they work better on GPUs than on CPUs—we’ve talked about this. GPUs are simpler calculation machines. They just do a bunch of calculations at the same time, as opposed to CPUs, which are much more complex. They have to handle if-then statements, branching, where stuff is, and all of that.
It’s a great irony that deterministic computing rests on probabilistic chips, which I would put CPUs in, particularly with things like branch prediction and all these bits and pieces, whereas probabilistic AI rests on very deterministic algorithms. You’re just doing a very deterministic calculation. It’s a weird, interesting paradox.
Groq takes that to the extreme. The idea is that if you know exactly what the calculation that you’re running is—there are no ifs, ands, or buts; you’re just running a very straightforward calculation—the way to make it run the absolute fastest is to define where every single thing is.
If you’re pulling into a gas station, there’s this game of dithering, right? You have the little arrow that points right or left on your gas tank. Apparently, the guy who invented that just died.
You pull up to a gas station, and you have to figure out: Where’s the pump? Which side am I on? What’s available? It’s going to take you a little bit of time to pull up to the pump, and then you have to get the gas thing, slide it over, put it in, and so on.
Contrast that with a Formula 1 car pulling into a pit stop. It’s super-defined: There are lines on the ground, people in place, and you know exactly where you’re going. That’s how you can get in, change tires, and be out in 2 seconds. They used to add gas, but even then it would still be a matter of seconds because they pull right up: gas right there, go in, boom, out.
Andrew Sharp
Mhm.
Ben Thompson
If you can predefine everything, it’s going to be way faster than dealing with uncertainty in the process of trying to execute something.
Even GPUs, which are very focused, are still somewhat probabilistic in nature and indeterminate in their calculations. The big thing is memory. They run on DRAM. There’s a variation of DRAM called high-bandwidth memory, or HBM, but it’s still the general concept: This is memory that has to be powered. You’re constantly having to supply a current to the gates to keep them where they are, and it’s a little indeterminate exactly what state they’re in at any one time.
You go to get information, and you have these abstractions of where it is, but not the exact location. You have to go and get it. You might have to wait a second for it to be refreshed, then you grab it and bring it back. It’s still incredibly fast—you can’t really grok, no pun intended, how fast it is—but it’s not as fast as it could be if you’re operating in a Formula 1 context, where you know exactly where the bit of data is. You go there, and it’s guaranteed to be there. It doesn’t need to be refreshed because it’s like a physical gate that’s set. It doesn’t need electrical current running through it.
Groq uses what’s called SRAM, which is RAM that’s actually on the die itself. You have the chip with all the logic gates on there; you’re not going off to a piece of memory. You’re actually on the same chip. It’s defined: It’s either on or off. It doesn’t need a continual current to refresh it. It’s kind of degrading. Then it goes back up.
You know exactly where every bit of data is. You go there and get it. What this means is, if you’re running a super-well-defined algorithm that is not branching—there are no if-then statements, you’re just calculating—and you can go and grab things and know exactly where they are, you can calculate much faster than if you’re inserting all these tiny bits of uncertainty throughout the process.
Andrew Sharp
Yeah.
3. Groq Makes Inference Faster
Ben Thompson
So, what this means is Groq is superfast. When it came out—
Andrew Sharp
And you can envision all sorts of different AI applications for it.
Ben Thompson
It was mind-blowing. 2 years ago, I embedded a clip of it. First off, you generate a response. Back then, it’s hard to remember because the stuff has gotten pretty fast, but it would spit out 1,000 words instantly, as opposed to watching it write down line by line, which was the common thing then.
They did a live segment on CNN that I embedded in the article when I first wrote about it, of someone talking to Groq as a real conversation. Again, it speaks to technological progress that we kind of get that now with the regular ones, but all that speaks to is the potential being even faster.
There are applications. Real-time speech, for example. Even in that demo, you can see it was a little slow, but you can imagine the speed that’s happened with regular AI models being applied to that model and making it faster.
Andrew Sharp
You and I are having a conversation here, right? There’s a very fast back-and-forth. When we’re in person, that back-and-forth is actually even faster. We have some latency here, right? Sometimes we speak over each other and go back and forth, which we don’t get when we’re in person, because in person you’re talking and I’m waving my arm. It’s just very apparent.
I want to talk right now to give you space to talk instead of me talking and having a meltdown if I don’t give you some space.
Ben Thompson
Well, no, I mean, that’s the perfect example. If you’re dealing with a chatbot customer-service agent, you want the fastest answers you can possibly get.
Andrew Sharp
Well, here’s another example. What is one of the AI things I’m super optimistic about? Because I’m a weirdo who thinks ads are good: personalized ads, right? You have stuff tuned to what you—
So, the whole ad ecosystem is actually pretty incredible, right? When you submit a webpage, it goes out.
Ben Thompson
It says, “Here’s who I think this person is,” or this sort of broad bucket of identifiers that I have. It runs an auction [laughter], then goes and gets an ad, puts it in, and serves you.
Andrew Sharp
In a matter of 3 seconds. To your point about us not properly appreciating how mind-blowing the progress of the internet is, independent of AI, that’s a great example of what we’re dealing with here in 2025.
Ben Thompson
Imagine you want that ad to be perfectly personalized to you. Now, again, I get how this sounds sort of terrifying and bad, so let’s just set that aside for a moment. But if you want real-time generative AI, so that it’s not just that generative AI lets you A/B test and find 1,000 interesting pieces of art instead of just the 10 that you generate on your own, what if it’s all generated in real time, immediately for you specifically?
Andrew Sharp
Once again, stipulate that Ben is a weirdo here, but this might be possible in the years to come.
4. The Inference Market Fragments
Ben Thompson
Basically, there’s going to be a market for massive speed and inference. That’s a market that Groq was going for, and it’s a market that Nvidia, with its architecture, was always going to be slower at.
Andrew Sharp
Now, there’s a huge trade-off.
Ben Thompson
You can’t fit very much memory on a die. Groq chips had 256 megabytes—I mean megabytes, not gigabytes. I think you had to link a gazillion of them together to even run a basic model, much less a high-end model. You definitely couldn’t run a high-end model with things like context—the context window—because that takes up a huge amount of memory.
Andrew Sharp
The way these things work is they generate a token.
Ben Thompson
Yeah.
Andrew Sharp
And then they run the exact same process to generate the next token. It’s hilariously inefficient. It’s like running these calculations again and again, just token by token by token. So you can see why something like Groq, even those little bits of being faster, manifests in being hugely faster.
Ben Thompson
But the point is, to generate a cohesive answer, you need to be storing every token that’s already been generated.
Andrew Sharp
And to do that takes a huge amount of memory, and so that’s why you still have Nvidia.
Ben Thompson
Well, say you want these thinking models that actually go through and go back and forth and evaluate different things and all these bits and pieces. All that’s doing is blowing up the context window and these memory requirements.
Andrew Sharp
So actually, one of Nvidia’s announcements with Vera Rubin this week is that a new part of a full-scale Vera Rubin server is that, within that whole multirack sort of unit, there are SSDs—actual solid-state drives—that are dedicated to storing this K/V cache.
You have these very large context windows, and you think about it from an agent perspective. If you have an AI that goes off and does work, it’s actually okay that the fetching is a little slow because you’re not part of the loop. The AI is just off doing work, right? Having larger context windows, being able to maintain more context and stay grounded, is more important than super speed. The point being, the higher-value the answer, the more willing you are to tolerate some latency in terms of what’s being served.
Ben Thompson
Right. But at the end of the day, if I’m calling a customer service agent and I have to wait 15 seconds for an answer—I’m making that up—the point being—
Andrew Sharp
The market is going to fragment. It’s not going to be just one way, one architecture, whatever it might be.
5. Nvidia Makes Groq More Valuable
Ben Thompson
So what Nvidia just acquired—or made a deal for; again, I keep saying “acquired.” They didn’t acquire anything.
Andrew Sharp
Get it right. Techmeme had to delete a tweet on Christmas Eve because they characterized it as an acquisition. This is a licensing deal. I think it was actually 90% of the employees, including the actual architects of this. Jonathan Ross in particular, who invented the first TPU, was the CEO of Groq, and they now have something in their portfolio to do this.
Ben Thompson
They licensed a product that will serve one of the fragments in the market of the future here. Now, Nvidia can make this way more valuable.
Andrew Sharp
This is going to be tricky: What problems go where? What should be served by a Groq-type model? What should be served by these very complicated SSD, hard-drive, east-west-maximizing-traffic models?
Ben Thompson
Wouldn’t it be nice to abstract that? Wouldn’t it be nice to have a software layer where you can easily say what you’re trying to do and it figures out the right way?
Andrew Sharp
A one-size-fits-all solution.
Ben Thompson
Well, something like CUDA, right? The idea of CUDA is that you don’t need to get down into the nitty-gritty. You don’t need to program individual shaders, and they’re going to have libraries that just do it for you. This fits into Nvidia’s approach of abstracting away some of that complexity into a higher-level sort of thing.
Nvidia is probably the most powerful company in terms of chip supply chains now—them or Apple. Groq doesn’t need to run on a super-old process, 14 nanometers from GlobalFoundries, which is terrible given that they need more space on-chip because that’s where the memory is also going. What would a Groq chip be like on a 2-nanometer TSMC process? It actually would be pretty freaking awesome.
Nvidia can get them in the door. They can make that happen. There are lots of things where it doesn’t matter. Nvidia is not buying the whole company because Nvidia uniquely can take their license and also all their employees and make it dramatically better than it’s going to be. And also, you’re going to pay through the nose for it. So that’s going to be the model. [laughter]
Andrew Sharp
Yeah. I mean, I was going to ask whether the world just belongs to Nvidia for the foreseeable future. It’s already the most valuable company in the world and certainly doesn’t look like they’re slowing down anytime soon. Do you think that this is bad for the tech ecosystem? If you were to argue that this is bad for the tech ecosystem, what would you say? A lot of people were concerned by the deal here and the way this entrenches Nvidia’s market dominance.
6. Antitrust Creates Strange Deals
Ben Thompson
Yeah. Well, there are 2 angles. Number 1 is just the structure of the deal itself. I wrote a big thing last summer about—
Andrew Sharp
Stinky deals.
Ben Thompson
Yeah. Stinky. It’s bad, right? In this case, it appears everyone was very well taken care of. The investors got paid out, and all the employees got paid out. Even if you just joined Groq last week, you were taken care of.
Andrew Sharp
Well, you know what’s a fun fact? I believe—I think I haven’t verified this. I just saw a tweet, so take it for what it’s worth.
Ben Thompson
Had Chamath put that money into Nvidia at the time he made the investment, he actually would have made— [laughter]
Andrew Sharp
Well, I saw that tweet. I think that tweet was debunked, but you could have done very, very well [laughter] investing in Nvidia at that point.
Ben Thompson
Okay, maybe debunked. I don’t stand by it. But if you joined Groq last week, your shares got vested, you got paid out; everyone got taken care of, and that’s good. It’s really important for Silicon Valley to maintain the incentives to join a startup, not just as the founder, but also—you need to go to work, right? You should reward employees who weren’t working at Nvidia for the last year.
Andrew Sharp
Right. And sometimes that reward is just having a stable job at Google for the next 40 years, and you can figure out what’s next. There’s something to be said for the fact that if you go join a startup, you’re not going to be unemployed. And that’s in danger.
Ben Thompson
Again, social pressure has kept it, it seems, to date with all these deals. We’re up to 8 or 9 of them. But it’s a big problem that there’s just not this assumption that if you join a startup and it fails, a big tech company will acquire you, you’ll have a job, and whatever—it’s going to be fine. This idea that we’re going to assume these are all bad and anticompetitive when the vast majority, again, are mostly failures has been, I think, very destructive.
Andrew Sharp
Yeah.
Ben Thompson
So I hate the structure in general just because it’s part of this trend, and I think Pandora’s box has been opened. If you’re a big tech company, why would you ever do a normal acquisition ever again? [laughter] Right? Because you don’t want to go through the regulatory review.
Andrew Sharp
And the irony is this is the one that probably needs the regulatory review. Here’s where, by virtue of valid acquisitions being forced to go through this structure just to avoid reviews that never should have happened—
Ben Thompson
Mhm.
Andrew Sharp
We created this structure for reviews to never happen, even when they should. And the reality is Groq’s approach was unique. Nvidia could not match it without building its own version of this sort of SRAM-dominated approach.
Again, Groq wasn’t going to take over the Nvidia market by any means. This is a niche part of the market, because there are huge parts of the market where you need an Nvidia-style approach, or you need a different ASIC approach, or whatever it might be. But it was a legitimate sort of competitor that Nvidia was just able to not buy, but effectively buy. And because they can bring so much to bear, it doesn’t matter that Groq is still a standalone entity. It’s not going to matter in the long run. And—
And because they’re printing so much money, it doesn’t matter that they paid an unbelievable premium for Groq as it exists today, right?
Ben Thompson
Which I don’t care about. Everyone’s like, “So much money.” I’m like, “I don’t care.” $20 billion is Nvidia’s quarterly revenue or quarterly profit.
Free cash flow. They made 23 billion in free cash flow last quarter. If the opportunity is as large as you think it is, counting dollars and cents—which, in a tech context, means counting tens of billions of dollars—who cares, right? So, I think it's a good move by Nvidia, and I think it's just funny that our regulators are never going to take accountability for the fact that they drove the creation of this structure of deals that has basically shut them off eternally from this. What are we going to do? Start saying Jonathan Ross can't go work for Nvidia? We're going to restrict the movement of individuals.
Andrew Sharp
Personal freedom. Sure. It'll be interesting to see how and/or if this new model is ever really addressed because, historically speaking, there's always been a bit of a cat-and-mouse situation with antitrust enforcement.
Ben Thompson
What are you saying, a company can't license?
Andrew Sharp
Maybe there are certain instances where licenses come under review. Yeah. Ben is silently shaking his head in disgust. I don't know. I'm not that concerned about it. It's more—
Ben Thompson
I'm concerned because people will latch on to Facebook acquiring Instagram, which, if you knew what was going on, you knew that was a problem, but regulators didn't know what was going on. So, their response was not to—
Andrew Sharp
appreciate their lack of understanding of technology. The response was to investigate every single small acquisition.
Ben Thompson
Mhm. That's how we got here. So, if we say, “Oh, we need to fix this. Let's start limiting employee movement. Let's start saying we're going to look into licensing deals. Are you serious?” That's just going to lead to even more bad outcomes and bad things and raising issues.
Andrew Sharp
Yeah. Well, it would be very difficult to look into. This is a non-exclusive licensing deal. So, NVIDIA is very, very—
Ben Thompson
And by the way, the way this was already decided, Google can make a non-exclusive deal with Apple or whatever it might be. So—
Andrew Sharp
It's been ratified. No problems there. Nothing to see there.
Ben Thompson
It's just an amazing example of an incredible regulatory own goal—
Andrew Sharp
Where—
Ben Thompson
by virtue of having any sort of humility about what they were trying to do, they basically just wiped out the possibility of doing anything ever. So, good job, guys.
Andrew Sharp
I also think fortuitous timing for NVIDIA. Jensen is pretty close with President Trump, so I'm not sure the DOJ would do anything. What are you going to do about the—
Ben Thompson
No. My position is, even if they were acquiring Groq, I don't know that there would have been an enforcement action. This started, to be fair, under the Trump administration. When I wrote about this, “First, Do No Harm,” I was at an antitrust conference with the Trump DOJ folks. Yeah, I went to dinner with them, actually, and I came away horrified. I'm like, “No, what are you guys doing?” So, yeah, it was not—