(Preview) Google Starts Dancing, The Winners and Losers of Gemini Week, OpenAI Has an Advertising Problem
- Gemini 3’s benchmark sweep is meaningful validation of Google’s comeback, not proof that every user will experience the best model. Ben Thompson’s base assumption is that the scores are “broadly correct” because Google has the resources, experience, and integrated stack one would expect eventually to win; he still says he is neither positioned nor qualified to certify “best in class.”
- Anthropic emerges as the week’s other winner because coding was the one benchmark Gemini did not lead. Thompson thinks Anthropic’s edge, dating to what he thinks was Claude 3.5, has been surprisingly persistent; Andrew Sharp says it has “a moat for now.”
- Benchmark leadership is hard to interpret when teams can optimize for or cheat the test. Llama 4’s benchmark-specific release is Thompson’s exhibit: disguising a mediocre model through benchmark gaming means “heads need to roll.” A friend’s characterization of Google’s bureaucratic tendency is that it “builds to the benchmark,” leaving real-world translation unresolved.
- Google’s strategic weapon is not merely Gemini 3 quality but a TPU-based cost structure that could be sustainably lower than rivals’. Model-chip co-design, cheaper specialized silicon, and freedom from NVIDIA margins compound Google’s enormous cash flow; Thompson calls that combination “pretty devastating,” though some scalability benefits remain theoretical here.
- ChatGPT’s installed base may blunt a technically superior Gemini just as Windows’ ecosystem blunted the better Mac. Thompson mentions 800 or 900 million ChatGPT users, then says not to quote him on whether that figure is weekly, monthly, or daily. Both stress that changing habits is harder than winning a fresh market: “Even if Gemini is better,” ChatGPT is already good enough and widely used.
- Release-week hype is outrunning users’ ability to distinguish increasingly capable models. In Thompson’s own comparisons, ChatGPT beat Gemini on UniFi bandwidth controls and turkey-brining timing, including the crucial skin-drying step; anecdotal as those are, they illustrate the difficulty of judging model differences alongside Sharp’s iPhone analogy of steadier, less spectacular gains.
1. Gemini 3 finally makes Google dance on its own terms
Sharp opens with the 2023 taunts that frame Gemini 3: Altman worried about a “lethargic search monopoly,” while Nadella wanted the “800-pound gorilla” to dance—and wanted credit for making it dance. Pichai’s mild reply, warning against “playing to someone else’s dance music,” has aged into Google’s comeback narrative.
Thompson’s benchmark problem is perceptual: AI is already so good for so many users that differences become difficult to see. “You can see down easier than you can see up,” and models may soon leave virtually everyone looking upward without a reliable sense of relative intelligence.
Capability also remains “very spiky”: a model can excel in one domain, hallucinate in another, or visibly pull from Reddit. Thompson therefore finds AI most interesting where no right answer exists and invention is valuable, or where correctness is objectively verifiable.
2. Coding preserves Anthropic’s lead while exposing benchmark games
On the reported suite, Gemini led every benchmark except coding, where Anthropic remained first. Thompson thinks Anthropic discovered something around 3.5 that “blew everyone’s mind,” producing a surprisingly durable lead and making it “one of the big winners of this week.”
Coding is unusually informative because “the code has to run”: it compiles or it does not, throws errors or it does not. That verifiability supports both credible measurement and reinforcement-learning loops that test outputs and improve post-training.
Public benchmarks invite distortion. Thompson cites Llama 4’s benchmark-specific version and condemns any team “disguising their failure by cheating”: the underlying mediocre model is one problem, but “the failure” and “the cover-up” together mean “heads need to roll.”
Google presents a subtler concern. A friend once characterized its bureaucratic weakness as building to benchmarks—“you get what you measure”—so dominant Gemini scores could reflect real superiority or KPI optimization. Thompson’s answer is not dismissal but continued real-world use.
3. Ordinary tasks puncture the launch-week narrative
Thompson’s model usage has already moved with experience: Grok 4 felt like “a big leap forward,” but browser friction and what seemed like declining quality pushed him back to a ChatGPT that seemed to be improving. His standard is sustained utility, not release-week excitement.
On a UniFi networking question, Thompson says ChatGPT explained how quality-of-service policies could limit guest bandwidth, then warned that software handling might slow the whole network. Gemini omitted that option and instead proposed separate access points connected through a 100-megabit router or switch—“number one, bizarre answer.”
Thompson considers turkey preparation a domain where he would like to think he is genuinely expert. In his comparison, ChatGPT covered thawing, wet brining, and leaving the bird uncovered so its skin dries and crisps; Gemini started too late and skipped drying, implying soggy skin. “You wanna talk about a confidence shaker.”
Sharp’s pushback is that these releases increasingly resemble mature iPhone launches: once-stunning leaps have become steadier and “less sexy.” Thompson agrees but adds that Google benefits from a powerful comeback narrative as well as structural reasons to expect technical leadership.
4. TPU economics turn Gemini’s quality into a larger threat
Thompson expected Google eventually to build the best model: it has worked on the problem longest, commands immense resources, and owns a fully integrated stack. Google invented the transformer, but OpenAI’s more modular Azure/NVIDIA setup nevertheless took the lead; Gemini now provides important validation for Google’s integrated approach.
TPUs are less programmable than GPUs but also simpler and cheaper to manufacture. Google can co-design training architecture and silicon, making architectural choices with TPU capabilities in mind; Thompson contrasts monolithic models with mixture-of-experts systems and the resulting communication and scaling tradeoffs.
Newer TPU designs may support stronger scaling across separate systems or facilities, although Thompson does not claim Gemini 3 validates that specific capability. What the release does validate is that coordination between Google’s model and hardware teams “is paying off.”
The economic consequence is a potential “sustainable cost advantage”: Google avoids NVIDIA’s margins, uses cheaper chips, and retains most of the secret sauce even with Broadcom’s assistance. Add cash flow that Thompson says is “better than ever,” and Sharp’s summary lands: more cash than almost anyone, potentially at lower cost than everyone else.
5. ChatGPT’s installed base may outweigh Google’s integration
Thompson corrects the standard Windows-versus-Mac history: DOS and its software ecosystem preceded the Mac, while Windows inherited compatibility with that installed base. The Mac introduced the earlier graphical interface and may have been better, but Windows was effectively first at the economically decisive software layer.
That history also revises the simplistic claim that modular systems always defeat integrated ones. Consumer markets reward integration because the user is the buyer and “there is no ceiling to the quality of the user experience”; enterprise purchasing gives spreadsheets and organizational constraints more power, even after SaaS improved bottoms-up adoption.
AI is not starting fresh. Thompson mentions 800 or 900 million ChatGPT users but immediately says not to quote him on whether the figure is weekly, monthly, or daily. Switching those users is harder than winning an unformed market, and Gemini being incrementally better may not be enough to change behavior.
Sharp’s pushback—worth keeping—is that Google’s true victories may be preserving Search while growing Cloud and YouTube, not converting ChatGPT users. Thompson’s closing analogy agrees: consumer AI may resemble the PC era, where the integrated Mac “was better, but it didn’t matter because everyone was already using Windows.”
Full transcript
Hello, and welcome to a free preview of Sharp Tech. Hello, and welcome back to another episode of Sharp Tech. I'm Andrew Sharp, and on the other line, Ben Thompson. Ben, how you doing?
You know, Andrew, there is—
Oh, boy.
Take, like, an NBA—
Deep sigh.
... take, like, an NBA basketball player. There are sort of 2 curves in their career. The first is athleticism. When you come into the league, you might still have a little bit more of a peak, particularly if you need to put on weight. Giannis would be a good example of this: He comes in very spindly, very thin, and puts on a bunch of muscle.
Mm.
But over time, you're going to peak relatively early in athleticism. Then it's going to go downhill, and yet we don't say a player's prime is when they're 23 or 24. We consider it to be usually when they're 27, 28, or 29.
28, 29.
The other curve is the increase in experience and understanding of how to actually play basketball.
The game slows—
The game—
...down the further—
That's right.
...you go in your career.
And it's good that the game—
Yes.
...slows down, because your body slows down as well. And so there's this period where the athleticism has declined, but not enough, and the experience has increased to more than make up for it. Eventually—
Mm-hmm.
...however, you can be as experienced as you want, and you just don't have it athletically anymore.
The lines on the graph diverge a little bit too much. Blake Griffin, on Amazon's new NBA coverage, was talking about the moment he knew he had to retire. He was running up and down the court with Evan Mobley, and he was playing for three or four minutes. This was Blake Griffin on the Boston Celtics, three or four years ago, and there was a foul. He was standing at the free-throw line next to Evan Mobley, and he looked up at Evan Mobley, who had been running up and down the court like nothing happened. Blake Griffin had his hands on his knees and was exhausted, and he was like, “Man, it's time to go.” So is that what you're confronting, that sort of moment right now?
No, no. The experience is you just learn how to do things, right? You learn how to write, to turn something out every day, to be a professional podcaster. Even when you're not feeling it, when you're feeling tired, you sort of get amped up.
Mm.
You're ready to do it. You learn how to do things like tease. There are segments in this podcast that I'm going to give a little hint toward. They involve getting old.
Okay.
And I... So I like to think that my experience is improving and is enough to compensate for the horrific reality that we have a question on this rundown that has been on the rundown for a couple of weeks. I was sure we had addressed this question previously. I got in an argument with you. I brought in Dummen [?] to review it, a professional podcast listener. He's like, “No, I don't think you talked about that.”
And you know what? My memory, on which I've depended to a near-absolute basis my entire life, is something I've just unknowingly relied on. There are so many things that are based on, “I just remember.” “How do you cite that article from 2017?” I'm like, “I don't know, I just thought of it and decided to link it.”
Yeah.
Fading fast, Andrew. I am facing my mortality. I am praying the experience. I am hoping that AI can save me.
Well, at least yours is a colorful, age-induced dementia, because you are effectively imagining podcast conversations that we did not have on the show.
It's true.
So—
So wait, I need the readers—if there is a question on this podcast that you think we covered before, you have to email in.
Email in for Ben's sanity. Exactly. Make Ben feel better over the Thanksgiving holiday.
As long as we're looking ahead, I have a topic on here that's going to look ahead to Thanksgiving.
Another tease.
Another tease.
Okay.
That's right. And we will do a Thanksgiving episode next week per tradition, unless there's big news.
Yep.
It was Thanksgiving three years ago when the OpenAI stuff went down.
That's right.
We know because it was the Las Vegas GP. It was the first one, and it was—
And we had to keep delaying the podcast. We were going to record in person in Las Vegas, and we had to keep delaying it because the news kept coming hot and heavy that week.
Right. I got back Sunday night, and I was delaying writing, which was good because Satya Nadella drops this tweet at around 2:00 a.m. “Oh, yeah, we're acquiring”—or whatever it might be—Sam Altman and Greg Brockman. Man, what a weekend that was.
So, barring something like that, we're going to have our silly, off-topic, massive mailbag. If it's about tech, it has to be clever and funny. We're not doing strategy breakdowns—X, Y, or Z. We're going to have a fun time. That's the plan.
Indeed, and if you've got questions for that episode, please submit them by 8:00 p.m. Eastern on Monday. Ben and I will record the episode on Tuesday afternoon, but I have to put the rundown together Monday night. And by the way, as bad as you feel for imagining podcast conversations that we didn't have, I am heartened to know that you actually read the rundowns before we come on to record every episode. This is evidence.
We did have this knockdown, throwdown fight approximately 8 minutes ago, so don't give me too much credit.
That's true.
But yes.
No, actually, you know what my thought was? My thought was I should not read the rundowns, because if I hadn't read this rundown previously and seen this question, which has been on there for a couple of weeks, I wouldn't have had this problem. So actually, I'm being punished for being prepared.
Okay.
1. Google Starts Dancing
Well, in any event, we are going to begin with Gemini and 3 quotes from 2023. Sam Altman in February 2023 said, “If I were sitting on a lethargic search monopoly and had to think about a world where there was going to be a real challenge to the way that monetization of this works and maybe even a temporary downward pressure, I would not feel great about that.”
And then Satya Nadella—
Where did he say that?
What's that?
Where did he say that?
He said that on Stratechery—in a Stratechery interview. That's right. I should credit the outlet here, okay?
Thank you.
Sam Altman to Stratechery in February 2023. Satya Nadella to The Verge in February 2023 said—
We don't need to cite everyone.
Fair enough. At the end of the day, Satya said, “They're the 800-pound gorilla in this. That is what they are. And I hope that with our innovation, they will want to come out and show that they can dance. And I want people to know that we made them dance.”
Look, I keep using the Altman quote because it's on Stratechery.
Uh-huh.
But this is the killer quote.
It's terrific.
So I guess we ought to credit The Verge. It's so good.
And then another quote that I had missed at the time, but I saw as I was preparing: Google CEO Sundar Pichai with a nice, mild-mannered response—“One of the ways you can do the wrong thing is by listening to noise out there and playing to someone else's dance music.”
I totally missed that at the time, but it's a wonderfully nerdy response from him. So there you have it, Ben. That's where we were 2 years ago: Google moving to its own beat. Two years later, Gemini is here. It's beating the competition on every model benchmark but one, and we'll talk about the implications for a variety of AI players.
One threshold question, though: As an analyst, how do you handle the release of a new model like this? How much stock do you put in various benchmarking tests?
2. Benchmarks Are Hard To Trust
It's really tough, and it's actually something that I've thought a fair bit about. This is sort of going to be an issue, challenge, question, and reason for optimism for OpenAI, actually, which I think we're going to get to in a bit, perhaps per our recycled question originally.
There you go. Another tease.
The reality is, AI is already so good for so many situations for so many people that it's harder and harder to perceive what differences there are.
Yeah.
Now, you could argue that's because you're too stupid, which very well may be the case. At some point, AI is like... It's easier—how to put this? If you're smarter than different people—
Mm-hmm.
...you can perceive the differences in intelligence. If you're not as smart as other people, they're all just smarter than you.
Yeah.
And one of the fortunes of being in Silicon Valley and getting to know lots of people is that I've definitely talked to people who are much smarter than me, and I'm not sure how much smarter they are than me.
They’re just smarter than me. And I’ve talked to other people that I’m like, “Eh, I can kind of rate you guys.” I know sort of who’s where. It’s an uncomfortable topic, but you can see down easier than you can see up.
Yeah.
And one of the challenges with AI is, if it’s not there yet, it’s going to get to the point where it’s sort of looking up for everyone.
Mm-hmm.
And so, what’s the difference between X, Y, and Z? Now, AI also tends to be very spiky. There’s some stuff it’s just really good at, and some stuff where you can tell it’s making it up. It will hallucinate much more in particular topic areas, or it will just get stuff wrong. Or you can really see the seams: you’re just pulling straight from Reddit.
Yeah.
Probably the most useful and important areas for benchmarking are areas like coding, for example, and the reason is that they’re verifiable.
Mm-hmm.
And this is a point that I’ve been making again and again. Andrej Karpathy had a tweet along similar lines this week, where AI is the most interesting at the 2 ends of the extreme. The one extreme is where there’s no right answer; this is where hallucination is actually a positive trait. It makes up new stuff. That’s actually incredible.
Yeah.
Computers have never done that before. The other extreme is where you know if something is right or wrong. This is why code, I think, has been a huge area, because the code has to run.
Mm-hmm.
It either runs or it doesn’t. It compiles or it doesn’t. It throws errors or it doesn’t. And in areas that are verifiable, not only can you know how good it is, but you can also do reinforcement learning, where you can do post-training to make the base model even better by running it through a bunch of tests and then learning from them and getting better over time. So what’s interesting about these benchmarks is, by the way, the one benchmark Gemini did not come in first in was coding.
Mm-hmm.
Anthropic is still number 1. Anthropic figured something out with, I think it was Claude 3.5, that blew everyone’s mind as far as coding, and it’s been a sustainable, persistent lead. Which has probably surprised me a bit. I don’t know what secret sauce there is and if it’s going to be figured out eventually.
Right, is it replicable?
But that’s a very positive sign. Yeah, I think Anthropic was actually one of the big winners of this week, where their bread and butter is the one thing that Gemini did not beat, at least according to benchmarks.
Yeah. They have a moat for now.
Now, the benchmark problem. Well, the other problem is the benchmarks are public. They do their best to keep the questions secret and adjust them and all these bits and pieces.
That’s part of what I wonder about: are models just being trained to the benchmarks at this point, and how much do we really learn from any of this?
Well, sometimes explicitly. Like, Llama 4 released this benchmark-specific version, and it scored super high.
Yeah.
I don’t understand the people who are critical of people in Meta AI being—well, set aside the superintelligence and who they hired. If you have a team that is disguising their failure by cheating, heads need to roll.
Yeah.
There are multiple issues going on, not just in the failure, but in the cover-up and all those sorts of things. And Mark Zuckerberg—one of the many reasons he looked bizarre, I think, in the interview I did with him last year, and kind of dumb in, I think, the interview with Dwarkesh in the same week—is sitting out there on a spit trying to defend the fact that he has a team that was cheating on benchmarks to cover up their failure.
Yeah.
That’s just not okay.
Their failure being effectively a mediocre model.
That LaMDA 4 wasn’t very good. Yeah.
Yeah.
It was a sort of poor model, and they’d sort of gone the right direction with this monolithic sort of approach.
Mm-hmm.
3. Google Builds To The Benchmark
Now, the funny thing about Google and Gemini—this is not an assertion about Gemini 3. All I will say is that there was a friend I was having a discussion with, and this was a while ago, I want to say 6 months or so, where they were tongue-in-cheek characterizing the various AI teams.
Yeah.
They were talking about the weaknesses in the teams. I think something about Anthropic was the goody-goody two-shoes bit: “Oh, we’re so special,” blah, blah, blah. The OpenAI bit—or they said that they reflect their founders. I can’t remember exactly what it was. I’m so sorry. I don’t want to speak for this person, who I’m not naming because I’m butchering whatever they said.
I was going to say—
The OpenAI one—
I have no idea what was trying to be said, even, but that’s fine.
Right, the OpenAI one was—the complaint about OpenAI is that it sort of kisses up to you all the time—
Yeah, cloying.
—and it’s just telling you what you want to hear.
Mm-hmm.
Right. And saying, “Yeah, that’s Sam Altman as well.” That was sort of the funny one as well. But Google was—Google builds to the benchmark. That’s part of the bureaucratic nature of a large company. You get what you measure.
Sure, KPIs.
4. Gemini Faces Real World Tests
And so it’s just interesting that they made this observation to me a while ago, and here we are today, and Gemini’s blowing the doors off in terms of benchmarks, but does that actually translate in the real world? I don’t know. And it’s really hard. From my perspective, what I do want to do is use Gemini regularly.
Mm-hmm.
When Grok 4 came out, it was a big leap forward, and I used it regularly. Then I got sick of using it in the browser on my desktop, where I use it the most.
Yeah.
It got a little worse. ChatGPT seemed to get a little better. I went back to ChatGPT. So I do need to use it more.
Yeah.
What’s interesting, for what it’s worth—and this is pure anecdata, which means it’s probably not worth anything—there were 2 extended discussions I was having with ChatGPT. I started on ChatGPT.
Mm.
One was about a networking issue and one was about turkey preparation. Now, the networking issue I know enough to be dangerous, but I wouldn’t say I’m an expert. What I will say is that the ChatGPT answer blew the Gemini answer out of the water.
Hmm.
It actually answered the question. It gave specific references, and it said why the solution I wanted was probably not a good idea. So I have all Ubiquiti UniFi equipment. I wanted to see: can I limit the bandwidth on the guest network? And you can do this through quality-of-service policies, but that’s going to kill your overall networking because it’s no longer on a dedicated hardware channel; it has to do it in software. Your whole thing will be slowed down. So it gives me this whole way out: “Yes, you can do it. This is how, but you might not want to do it because of why.”
Gemini didn’t even list the QoS option. Its option was to get completely separate access points for your guest network and put it on a 100-megabit router or switch.
Yeah.
Number 1, bizarre answer. Number 2, there was an answer it didn’t get. So anyhow, whatever. That happens.
Less comprehensive. Sure.
Yeah.
One example.
That happens. I don’t want to give one example. Now, what I am an expert at, I would like to think, is turkey preparation.
Great.
I’m a big believer in the wet-brine approach for turkey preparation. The problem with turkey is that people complain it’s dry. I have my own brine for poultry that I think is phenomenal. My chicken is world-class, and I’ve done turkeys in lots of different ways. I’ve done it in the oven. I’ve actually done it on the grill. It worked very, very well. I’m going to do a smoker this year, so—
Have you fried a turkey yet?
I have not. I might try that in the future. I’m going to try the smoker this year. That’s sort of a new thing, but I do have the brining process down pretty pat.
Mm-hmm.
And I was reviewing what’s the timeline from thawing the turkey—
Look at this.
—to doing—
Prepping. Love it.
Right.
Leaving no—
And so I was like—
—stone unturned before next week.
No, no, well, because you’ve got to start early. It takes 3 or 4 days to thaw, and then you would brine it. The key thing is that after the brine, you need to leave it uncovered in the refrigerator so the skin dries out, so you’ll get crispy skin.
Hmm.
ChatGPT covered all that and gave me a good timeline. I’m like, “Oh, I should put this in Gemini.” Gemini skipped the step of drying out the skin.
Oh, boy. Gemini forgot to throw it in the fridge to let—
Um—
—the skin dry.
Really, you want to talk about a confidence shaker.
Failed that benchmark.
We’re talking about Thanksgiving dinner, and it started out like, “Oh, start thawing on Sunday night.” I’m like, “What? That’s way too late. What are you talking about?” I go through it, and it’s like I’m going to pull it out of the brine. No, then you’re going to get soggy skin. Ugh.
Oh, boy. Wow.
Nothing to do with the economic implications of AI. Probably totally irrelevant and just bad luck for Gemini, but it speaks to why this is very difficult.
Yeah. It’s mystifying.
For me as a normie, it’s mystifying to know how much stock to put in any of these tests. Every model release is hyped for a week or so in advance. John Coogan of TBPN compared these releases to iPhone releases, where 10 years ago there would be mind-blowing leaps that would make the whole world stop, and now the improvements are steadier and less sexy than they were a year and a half ago.
That isn’t to say they don’t matter, but vibes-wise, it’s just sort of a shift as the ecosystem grows up here. Does that resonate with you at all?
Well, I think so, and it speaks to a couple of factors here. Number 1, Google does have the comeback narrative behind them.
Yeah.
Narratives matter, right? That’s how you framed this discussion from the beginning. As a longtime professional NBA podcaster who is one of the biggest proponents of narratives, that led to one of your really terrible takes that Russell Westbrook should have been MVP as a 6-seed a few years ago, just because you love the narrative.
Yeah.
This is a great narrative, right?
Yeah.
There are also systemic, structural reasons why you would expect Google to have the best model. They have the most resources. They’ve been working on it the longest. They have this fully integrated stack, which at the beginning you would expect the integrated player to actually do better.
The whole theory is that through integration, when stuff isn’t good enough, you can solve a bunch of small problems and do it better. Over time, as you get to the iPhone level where everyone finds it boring, it starts making more sense to have component pieces and people competing at different levels.
We had the opposite. Google invented the transformer, but then we had OpenAI running on Azure and using NVIDIA—a much more modular approach—coming in and being in the lead. Even with all my turkey problems and my networking issues, my base assumption is that the benchmarks are broadly correct.
Okay.
I would expect Google to eventually get the best model. If it’s now or it’s the next one, whatever, I’m not surprised.
Interesting. We’re going to go company by company. You wrote about winners and losers on Wednesday after Gemini week, but one specific question I had on Google and the integration: Google trained this model on TPUs, their Tensor Processing Units, not NVIDIA GPUs. For Google’s prospects specifically, how significant is that detail?
5. Google Has A TPU Advantage
Pretty significant. I think everyone had this sense that TPUs—which Google started developing a decade ago, particularly the latest versions—like NVIDIA GPUs, actually, to be fair, were not tailor-built for LLM workloads.
That’s more important for the TPUs because they’re less capable and less programmable than a GPU. It’s like a simplified version of what a GPU can do, but a GPU actually does more stuff. It’s more programmable and can do more things.
Part of the reason GPUs are more expensive and more complicated is not just because NVIDIA charges a lot; they’re more expensive to make. They’re more complex chips, and a larger chip means you have more yield issues, et cetera, et cetera.
Theoretically, they should be able to serve AI at a lower cost than everybody else with TPUs.
Well, yeah, especially once they get the right TPUs in place, in theory. Those TPUs can be developed, again at least in theory, hand in hand with their model development team.
The model, as they develop it, can inform how they decide to train it, for example. As I mentioned, Llama was monolithic. It was one big clump of digital neurons sort of figuring it out, as opposed to a mixture-of-experts approach, where you have different ones that light up in different places.
That presents its own problems in terms of crosstalk, traffic, communicating, and making sure you’re not overweighting one, et cetera, et cetera. This was in the context of DeepSeek, which used a very aggressive mixture-of-experts approach. They did it in a sort of unique way that let them do it with relatively low capabilities, which we should go back to in a moment.
For Google, there are architectural choices that go into training. Google can make those choices with knowledge of what TPUs are capable of. I believe that with the last generation of TPUs, they’ve been really focused on having much more scalability—not just one big system, but actually having much stronger assumptions around being separate, and even potentially being in separate facilities over time.
They’re building toward that, so you don’t need one monolithic, huge data center. You can actually do something that is in different locations, for example. All of which introduces a ton of cost and complexity, et cetera, et cetera. The more you can build with those things in mind, at least in theory, the better.
Yeah.
So this is a validation—not necessarily of the different-locations thing, but whatever they’re doing between the model building and the TPU design is paying off. The implication is that they have a sustainable cost advantage going forward.
Number 1, they’re not paying NVIDIA margins. Number 2, TPUs are cheaper to make than GPUs. The margin advantage is probably the more important one. They get to internalize all that.
Yeah.
Broadcom helps in some respects, but Google is doing most of the secret sauce. Broadcom’s margins aren’t going to be anywhere near what NVIDIA’s are, because Broadcom isn’t making the chip. Google is using some baseline. It’s more that Broadcom gets a percentage of the stuff, not that they get to charge whatever margin they want on top of it.
Mm-hmm.
So, yeah, this is the real economic threat here, and it’s substantial. Google has huge cash flow—better than ever, by the way.
More cash than anybody, potentially at cheaper costs than everybody else.
That’s right.
Yeah.
That’s right. Pretty devastating combination. To the extent that this is validated, it’s a big deal.
Mm-hmm.
We can get to what that means for the other companies, but, yeah, it’s validation. It’s not a surprise, but it is important validation.
Yeah. So, if Google’s model, Gemini 3, is in fact best in class, what are the opportunities—
And for the—let’s assume it is best in class. Again—
Yeah, for the purposes of this conversation—
I’m not in a position or qualified to say that for sure.
Let’s assume—
Yeah, but yes.
—that they killed it here. What opportunities for Google interest you most going forward in light of this success?
6. The Integration Advantage
What’s interesting is that, at least in past eras, the smartphone era is probably the best example. To go back to the beginning of Stratechery, I’ve told this story a million times, but at the time when I started, the broad sentiment was, “That’s neat that Apple invented the iPhone, but Android is going to obliterate them because modular beats integrated,” blah, blah, blah.
Mm-hmm.
Look what happens to Windows versus the Mac. As I’ve noted, the Windows-versus-the-Mac analogy is actually widely misunderstood, most prominently by—
Ahistorical, yeah.
—the late Dr. Clayton Christensen. But it’s just totally wrong. It wasn’t the case that the Mac was first and then Windows came in and ate its lunch. It was the case that DOS was first, rode on the coattails of IBM, built up a huge customer base and software ecosystem, and then the Mac launched.
Yes, the Mac was the first GUI, but Windows, which came later, ran all the DOS applications. So, actually, Windows was first when it came to the installed base and all those sorts of things. It’s not an applicable analogy.
Yeah.
The iPhone was different because the iPhone actually was first, and Apple built up a customer base and a developer base that was not going to just leave and go to Android.
It's funny how these myths about the past basically permeate everything. Everyone's analysis is totally wrong because of it.
Yeah.
I mean, it's the lowest-hanging fruit Stratechery ever had. If you actually change one of these facts about your understanding of what happened 20 years ago, your conclusions are going to be totally different. And boy, I could build a business just stating this very obvious thing.
And now here we are. No, I was not aware of that history until about a year and a half into hosting this podcast, when it came up, and I cannot believe that people basically just got the facts wrong in terms of the sequencing of what happened.
Yep.
And that informed—
Yep.
Well, because the sequence was that the Mac came before Windows, and the Mac famously sued Microsoft for stealing the concept, which of course Apple stole from Xerox—or borrowed. Great artists steal, whatever it is. It's understandable that this is the popular conception because the Mac did have the first GUI, the WIMP interface. Microsoft was second. But it misses the actual important layer: the software layer, third-party developers in particular, and the underlying OEM ecosystem. All that came before the Mac.
Uh-huh.
Okay.
So how does that relate to where we are and where Google might go?
Actually, it's interesting to contrast these two. And now—ugh—I'm actually regretting my update yesterday because I didn't dive into this. But that's okay. That's what Stratechery is for.
Exactly.
So if you look at the smartphone era, that's sort of what you would expect, or at least what I would expect. The big article I wrote about this—it was arguably the article I started Stratechery to write—was “Why Clayton Christensen Was Wrong.”
Mm-hmm.
It's a classic of the genre. You come up respecting someone, learn from someone; they're your hero. But to truly—
Take a shot at a legend.
Break through, you gotta take ’em down.
Make a name for yourself.
That's right.
Of course, yeah.
So what he was wrong about was the insistence that modularization would always beat integration in the long run. One of my takeaways in that article is that consumer markets are different because the user is the buyer. There is no ceiling to the quality of the user experience. This led to Eugene Wei really expanding on the invisible asymptote idea: you could just always get better and better and better in this respect.
Mm-hmm.
And so that meant that, actually, no, Apple wasn't going to be disrupted because consumers were not just going to decide, “Oh, this very abstract, spreadsheet-driven calculation of value relative to capabilities and suppliers and all this sort of thing—I'm going to switch to Android now.” All the examples Clayton Christensen used were in B2B businesses, where people are paid to be rational, to do spreadsheets, and to decide which one to choose.
Yeah.
I pointed to all these examples—I think it was autos or consoles, or lots of other things; consoles was definitely a big one—that stayed integrated because their focus was on elevating the user experience in a market where the buyer is the user.
Mm-hmm.
And so they're going to value that. Design matters in the consumer space in a way it doesn't always matter in the enterprise space. If you're the spreadsheet guy making buying decisions, one of your columns in your spreadsheet is user experience. It's not the driving factor.
Right.
When you have to use it, that's why people complain about their enterprise software, right? It's because there are lots of other factors.
It clearly has not been the driving factor for white-collar work over the last 30 years.
Right. Well, to some extent it did get better, and this is a great thing about the SaaS era and the whole bottoms-up philosophy: you win a team and then it spreads throughout the company. The good thing about that is that it necessitates it being good.
Mm-hmm.
So you get evangelists within your company that buy stuff and spread it around. The quality of enterprise software went up dramatically in the SaaS era relative to what was before, but still nothing like the consumer market.
I'm thinking of—
And so in that—
The battles I wage with Microsoft Teams on a weekly basis.
Yes. Uh—
That's what comes to mind when I think of—
Still pretty rough.
Good enough—
Let me tell you—
Solutions.
There's a reason we do use Teams for various things at Stratechery, and all our conversations happen in WhatsApp because it is much better.
That's right.
So generally, in the smartphone era, we saw that integration is particularly powerful in the consumer market because you can deliver a better user experience if you're just managing the whole thing from front to end. Even if you're giving up certain efficiencies along the way, those efficiencies—having modular interfaces—are costly in terms of the user experience.
Mm-hmm.
Because there are just rough edges that you're having to tie together. And as a rule, I think that's a useful thing to think about. The problem is that we're not starting from scratch today. ChatGPT is in the market. They have 800 or 900 million weekly active users. Don't quote me on weekly active users; it's monthly or daily, whatever.
Yep.
It's a much more difficult task to get people to change than it is to win a market fresh. So even if Gemini is better, how many people are actually going to even try it, given that ChatGPT is still pretty good and better for unified networking and cooking turkeys?
Totally. That's what I wonder about as everybody freaks out about Google. It seems like there's been an awakening of sorts to Google's dominance, but I'm unclear how much Gemini's model performance really moves the needle in that story. The big win here is preserving search dominance over the last couple of years and growing the cloud business and YouTube, et cetera. I understand why there's been an awakening to Google's dominance. It's just unclear what Gemini really means in that story because everyone I know still uses ChatGPT, not Gemini, and maybe that'll change, but I would guess not in the near term.
Well, this is where it's interesting to go back to the PC era and the whole point I just made: the PC era started out modular.
Mm.
And then the integrated player came in, got a rabid fan base that was also very small, and never really got large.
Yeah.
And that might be the better analogy for the consumer market in AI: the Mac was better, but it didn't matter because everyone was already using Windows.
All right, and that is the end of the free preview. If you'd like to hear more from Ben and I, there are links to subscribe in the show notes, or you can also go to sharptech.fm. Either option will get you access to a personalized feed that has all the shows we do every week, plus lots more great content from Stratechery and the Stratechery Plus bundle. Check it out, and if you've got feedback, please email us at email@sharptech.fm.