Kevin Roose
Casey, we have a special emergency podcast episode today about the launch of Gemini 3.
Casey Newton
Yes, Kevin, hotly awaited and much discussed among AI nerds here in Silicon Valley. We are finally about to get our hands on the genuine article.
Kevin Roose
Normally, we wouldn't break our Friday publication schedule to publish a special episode just about a new model coming out from one of the big AI companies. They're releasing models all the time, but there are a couple of reasons that we thought it was worth doing this week to talk about Gemini 3 in particular.
The first is that we got some time with Demis Hassabis and Josh Woodward, 2 of the leading AI executives at Google. Demis, of course, is the CEO of Google DeepMind, which is their in-house AI lab, and Josh Woodward is the VP of the Gemini team and some other stuff there at Google. So we are excited to talk to them and ask them about this big new model release.
But I think there are a couple of other reasons we were interested in doing this as well.
Casey Newton
One big thing, Kevin, is just that maybe more than other model releases, this one seems to have the attention of Google's competitors. We're hearing a lot of whispers from folks who work at other AI labs that it seems like Gemini 3 has managed to figure some things out in a way that may be bad for their businesses.
Kevin Roose
I think around the AI industry, there's sort of this feeling that Google, which struggled in AI for a couple of years there, had the launch of Bard and the first versions of Gemini, which had some issues. I think they were seen as catching up to the state of the art, and now I think the question is: Is this them taking their crown back?
So we'll get into all that with Demis and Josh. But let's just talk, Casey, about what we know about Gemini 3. They held a briefing early this week and told us a little bit about the new model and what it can do. So what did we learn about Gemini 3?
1. Gemini 3 Builds Custom Interfaces
Casey Newton
In terms of what it can do, which is always the most interesting to me, Google shared a few different things. In addition to saying all the things you would expect, like it's better at coding and it's better at vibe coding, it also is going to do some new things around generating interfaces for you when you ask it a question.
Nowadays, you ask most chatbots a question, and they'll spit back an answer in text. Maybe they'll show you an image. According to the Google folks, Gemini 3 is just going to start building custom interfaces for you. They showed an example where somebody wanted to learn about Vincent van Gogh, the painter, and Gemini 3 coded up an interactive tutorial that had all sorts of images and interactive elements.
They showed another example that involved building a mortgage calculator for buying a home over $1 million, which is the lowest amount of money that anyone at Google can imagine spending on a home. These are the kinds of things that you can expect to find in Gemini 3, Kevin.
Kevin Roose
I would say the theme of the briefing and of the materials that Google shared ahead of the Gemini 3 launch was: This is just better than their last model, Gemini 2.5 Pro, in basically all respects.
Some of the benchmarks that caught my attention: One was a benchmark test called Humanity's Last Exam, which is a very hard interdisciplinary exam that consists of a bunch of questions at basically a graduate student or PhD level. Their previous model, Gemini 2.5 Pro, got about 21.6% on that test, and Gemini 3 Pro gets 37.5% on that test.
That's basically the story of all of these benchmarks. They gave more than a dozen examples of various benchmarks where the new model just beats the old one handily. To a lot of people, I think that may not matter. Most people who are using Google's AI products are probably not out there trying to solve novel problems in physics.
But their basic pitch for this is just: This is a state-of-the-art model. Anything that you could do with ChatGPT or Claude or even the older versions of Gemini, you can do better with Gemini 3 Pro.
Casey Newton
They also talked about testing what they're calling the Gemini agent, which is going to be able to do one thing in particular that I've been waiting for somebody to do forever: look through your inbox, understand its contents, propose replies, organize emails together, and really help you get your inbox under control in a way that I personally have never been able to.
We basically only saw a few animated GIFs about that, but that will definitely be one of the first things that I try when I get my hands on Gemini 3.
Kevin Roose
They are not, we should say, rolling this out to everyone right away. It's going to be available this week for users in the Gemini app and also in AI Mode, which is the tab off to the side of the main Google Search engine.
It will also be available for developers in various products, but they're not saying when this will come to things like the Gemini integrations in Google Docs or Gmail, these very popular things that are used by billions of people a day.
But I thought it was interesting that they have brought this model to Google Search, albeit in this AI Mode that's not the main search bar. That, to me, suggests that they feel like they can serve this model cheaply enough to make it potentially something that billions of people could use, and that doing so would not melt their servers and incur billions of dollars of costs.
Casey Newton
So far, they say that usage keeps going up for AI Overviews, and every quarter they continue to make more money. It seems to be working out for them. Not working out for the rest of the web, but it's working out well for Google.
2. Google’s Distribution Advantage
Kevin Roose
I think that's obviously Google's big advantage here over its competitors: They have products that are used by billions of people a day, and they can shove Gemini 3 into those products over time, get more and more usage, get more data, and use that to improve their models.
Casey Newton
Which is why we always tell students when they ask us for advice: "Step 1, build an illegal monopoly."
Kevin Roose
Yes. And speaking of students, the other notable announcement that Google is making this week is that they are giving all U.S. college students 1 year of free access to a paid version of Gemini, which is, I think, a smart move.
I feel a little gross about it, like essentially telling students, "Hey, why don't you use this to maybe do some of your homework? Maybe help you with your exams? We'll give you the first hit for free."
Casey Newton
I was also struck during the briefing that we had this morning that I believe 3 different people used the phrase "learn anything." This seems like it has become a very prominent plank of Google's messaging. They're presenting Gemini as a learning tool, which maybe is just sort of a euphemism for a do-your-homework tool. I don't know.
Kevin Roose
Yes. Okay, so that is what we know about Gemini 3. We will be doing our own testing and reviewing of Gemini 3 once it is fully out on Tuesday. But for now, we wanted to just give you the basics and also bring you our interview with Demis Hassabis and Josh Woodward of Google DeepMind.
Casey Newton
Before we get to that, we should obviously make our AI disclosures. I work for The New York Times Company, which is suing OpenAI and Microsoft over the training of large language models.
Kevin Roose
And my boyfriend works at Anthropic.
Casey Newton
Demis and Josh, welcome to Hard Fork.
Speaker 3
Great to be here.
Speaker 4
Thank you.
Casey Newton
So 2 years ago, Sundar Pichai told us that Bard, rest in peace, was a souped-up Civic that was in a race with more powerful cars. What kind of car is Gemini 3?
Speaker 4
That's a good one. Demis, do you want to take it?
Speaker 3
I hope it's a bit faster than a Honda Civic. I don't really think of it in terms of cars. Maybe it's one of those cool drag racers.
Casey Newton
People are really excited about this model. We have been hearing from folks that have been early testing it. Obviously, you guys have shown off a lot of the benchmarks, which are very impressive. What can Gemini do on a concrete level that previous AI models couldn't?
Speaker 4
I'll jump in. Maybe a couple of things stand out. One, we're starting to see this model really excel at reasoning and being able to think many steps at the same time. Sometimes models in the past would lose their train of thought or lose track. This one's way better at that.
The other thing you'll see tomorrow as well is all kinds of new generative interfaces. This is our best model yet at being able to create new types of interfaces. It gives people a really custom design and answer to their questions.
The third thing I would say is that we've put a lot of investment into coding itself. A lot of the coding examples you'll see, along with some new products coming out, like Google Antigravity, will also showcase that.
Casey Newton
There’s been some discussion that, for average users, the chat use case can feel solved—that average users of products like Gemini almost can’t even think of a question to ask it that will generate something that feels meaningfully different from what they were able to get in the last model. To what extent does that feel true to you in Gemini 3, and to what extent do you think average folks are really going to notice a difference?
Speaker 4
Yeah, one of the things we’re seeing in some of the testing—and Demis, feel free to chime in too—is that this is a model that’s more concise and more expressive. It starts to present information in a way that’s much easier to understand, and I think for most people, that’s going to be a big immediate effect. Then I think what starts to get interesting is how these models start to interact with other types of information.
We talk a lot about how students are going to be able to learn with this model, or even how this model can connect to other types of data you might have in other Google products, with your permission. These are the ways I think we’re starting to show that it’s going beyond just the standard text Q&A back and forth.
Speaker 3
Yeah, I think I’d add to that. Its general reliability on things is incredibly high. You’ll notice that when you use it. I think we also worked quite hard on the persona, which we call internally the style of it. I think it’s more succinct, more to the point, and helpful. I feel like it’s got a better style about it.
I find it more pleasant to brainstorm with and use. Then I think there are various things where there’s almost a step change. I feel like it’s crossed a threshold of usefulness on things like vibe coding. I’ve been getting back into my game programming, and I’m going to set myself some projects over Christmas because I feel like it’s actually got to a point where it’s incredibly useful and capable on the front end and things like this—areas where perhaps previous versions weren’t so good.
3. AGI Timeline Holds Steady
Speaker 2
Demis, the last time we had you on the show, in May, you said that you think we’re 5 to 10 years away from AGI and that there might be a few significant breakthroughs needed between here and there. Has Gemini 3, and observing how good it is, changed any of those timelines, or does it incorporate any of those breakthroughs that you thought would be necessary?
Speaker 3
No, I think it’s dead on track, if you see what I mean. We’re really happy with this progress. I think it’s an absolutely amazing model and is right on track with what I was expecting, and with the trajectory we’ve been on for the last couple of years—since the beginning of Gemini, which I think has been the fastest progress of anybody in the industry. I think we’re going to continue on that trajectory, and we expect that to continue.
But on top of that, I still think there’ll be 1 or 2 more things that are required to really get the consistency across the board that you’d expect from a general intelligence. There are still improvements needed in reasoning and memory, and perhaps in things like world-model ideas that you also know we’re working on with Simmer and Genie. They all build on top of Gemini but extend it in various ways, and I think some of those ideas are going to be required as well to fully solve physical intelligence and things like that.
Both are true. I’m really happy with the progress of Gemini 3, and I think people are going to be pretty pleasantly surprised. But it’s on track with what we were expecting the progress to be, and I think that means still 5 to 10 years, with perhaps 1 or 2 more breakthroughs required.
4. Gemini’s Tool First Personality
Speaker 2
You mentioned Gemini 3’s style. There’s been a lot of discussion recently about AI companions and the relationships people are developing with them. How do you think about Gemini 3’s personality, and what kind of relationship do you want users to have with it?
Speaker 4
I would say that in the app itself, we see it on the team a lot as a tool, or as something you’re using to work through and cut through your day. Whether it’s helping with different types of questions you have or helping you create things, that’s really where we see it excelling and the direction we want to see it go.
If you zoom out and look at Gemini or some of our other projects, like NotebookLM or Flow, we’re really trying to think through how AI can be this superpower, this super-tool in your toolbox that you can use for writing, researching, creating films, or whatnot. That’s more where we’re focused. Over time, we’re really interested, as a team, in being able to track things like how many tasks we helped you complete in your day.
Speaker 2
Mm-hmm.
Speaker 4
That’s a new type of metric that I think we get excited about, and it’s a way that the original Google Search worked. You would come to it, try to get an answer or be sent to a page, and move on from there.
Speaker 2
Well, that all sounds very good and responsible, but I’m wondering about all the viral engagement you’re leaving on the table by not making this thing an erotic companion. Big oversight.
Speaker 4
No comment.
Speaker 2
Yeah. Some of your competitors have been very nervous in the days and weeks leading up to Gemini 3. I think they’ve started hearing the same rumblings that we have about this model being quite good, and maybe the narrative shifting from Google playing catch-up in AI to now being on top of the race, or at least in a leadership position there. Do you feel like Google is ahead in the AI race right now?
Speaker 3
Look, as you guys know very well, it’s a ferociously competitive environment—probably the most competitive there’s ever been. One can never say. The only important thing is your rate of progress from where you are, and that’s what we’re focusing on. We’re very happy about that.
I don’t really see it as, “We’re back in the lead,” or something like that. We’ve always pioneered the research part of this. I think we’re getting into our groove in making sure that’s reflected downstream in all of our products, and I think we’re really getting into our stride there. You saw that at the last I/O, I would say, and we’re getting better and better at that, with Google DeepMind being the engine room of Google.
Of course, there’s the Gemini app and NotebookLM, these AI-first products. But there’s also the powering up of all these amazing existing Google products, whether that’s Maps, YouTube, Android, or Search, with AI-first features—and, in some cases, reimagining things from an AI-first perspective, with Gemini often under the hood.
That’s going amazingly well, and I think we’re only midway through that evolution. It’s very exciting to see how much value and excitement our users are getting when they see each of those new features—for example, in Workspace and Gmail. There are almost endless possibilities there. We’re really excited about that, as well as all of these AI-first products that we’re imagining and prototyping.
Speaker 2
We had a historian on the show last week who was using an unreleased Google model in AI Studio, and it had sort of blown his mind with how it was able to transcribe these very old documents and reason correctly about what the measurements of the sugar were in this sort of 1800s fur trade in Canada. Do you think you can tell us once and for all: Was this man using Gemini 3?
Speaker 4
Not sure about that one.
Speaker 2
Okay. All right.
Speaker 4
I will say, though, the model is quite amazing at making these connections. I don’t know if the historian was using photos of old documents, diaries, or whatnot.
Speaker 2
Yes, that’s what he was doing.
Speaker 3
Yeah, that most certainly was.
Speaker 4
Okay, it’s very good at this. Someone like me, who has pretty poor handwriting, can take a page of notes and it’ll take that and run with it with no problem, no sweat.
Speaker 2
You mentioned on this call that you’re going to be integrating this into Search, in the AI Mode that’s a side tab on the main Google Search engine. Does that mean that you’ve found a way to serve this model more efficiently and cheaply than previous models?
Speaker 3
I think we’re always on the cutting edge. I feel like the thing we do really well, apart from the overall performance of our models and getting better and better at that, is the efficiency of our models—the distillation techniques and many other techniques that we created and pioneered, which we’re now putting to use.
Obviously, it’s necessary for us because we have extreme use cases of things like AI Overviews and others that we have to serve to billions of users. Then, of course, some of our cloud and enterprise customers really appreciate that cost efficiency too.
We’ve always tried to be on this Pareto frontier of cost to performance. Wherever you want to be on that frontier—if you value performance most, or if you value cost most—there’ll be a model in the model family for you.
So of course, we’re only announcing Pro today, but we’re also working on the other family of models for the 3.0 era. So you’ll see a lot more about that pretty soon.
Speaker 2
Yeah. It seems like every time we see the release of a new frontier model, we get to revisit the discussion about scaling laws and whether we’re beginning to see diminishing returns. I can predict a few Twitter accounts that will probably have something to say about this over the next few days. So I thought I would ask you before we have that discourse: How are you guys thinking about that in relation to Gemini 3?
Speaker 3
Yeah, we’re very happy with the progress Gemini 3 represents over 2.5. Referencing what we discussed earlier, the progress is basically what we’re expecting and on track, and we’re really pleased with it. But that’s not to say that there isn’t some kind of diminishing returns.
People, when they hear “diminishing returns,” think of zero or exponential, right? But there’s also something in between. So returns can be diminishing, but it’s not going to exponentially double with every era. It’s still well worth doing, right?
Speaker 2
Yeah.
Speaker 3
And there’s an extremely good return on that investment. So I think we’re in that era. As I said, my suspicion, although we’ll see, is that 1 or 2 more research breakthroughs are required to get all the way to AGI.
But in the meantime, you’re obviously going to need as-scaled-as-possible versions of these foundation models—multimodal foundation models that we’re building today and still seeing great progress on.
Speaker 2
Right. Which of the many benchmarks that you showed off today do you feel is going to matter most to the average user?
Speaker 4
Oh, that’s a good question. I think most people don’t look at the benchmarks as closely as we do, but the benchmarks are always a proxy, right? So you look at something like cracking 1,500 Elo on LMArena—that’s great, but what really matters is user satisfaction in the products, too.
What’s been encouraging to us is that these are still moving in the same direction. They’re good proxies for each other. Ultimately, we’ll put out all the benchmarks, and we’re very proud of them. They represent amazing progress.
But you also have to be able to translate that into product experiences that matter, so we try to do both with every one of these releases.
Casey Newton
Any new dangerous capabilities or safety concerns that come with the increased power of the model?
Speaker 3
I think we’ve taken quite a long time on this model because it’s frontier, has some new capabilities, and is very capable, as you can see from the benchmarks. And as Josh said, we make sure not to over-index internally on those benchmarks. They’re just a proxy for overall performance, and that’s why we care about them across the board, as well as how our users ultimately experience them.
We spend a lot of time on safety testing across all the different dimensions, with the safety institutes and also external testers that we work with, as well as doing a ton of internal testing. I would say this is our most thoroughly tested model so far.
Casey Newton
Do you want to mention any of those new capabilities that popped up, whether or not it was for a safety thing? Was there something in there where you thought, “Okay, yeah, we definitely need to make sure we’re sending this to a bunch of external researchers”?
Speaker 3
Yeah. Look, we’ve worked really hard on things like tool-call usage and function calling.
Casey Newton
Mm.
Speaker 3
These kinds of things are obviously super important for coding capabilities, and developers want that. It’s also very important in general for reasoning. But it also makes them more capable for riskier things, too, like cyber.
So we have to be doubly cautious as we improve those dimensions for all the good use cases. We’re continually checking all those kinds of measures to make sure they can’t be misused.
Casey Newton
Are we in an AI bubble?
5. The AI Bubble Has Real Value
Speaker 3
I think it’s too binary a question. My view on this, strictly my own opinion, is that there are some parts of the AI industry that are probably in a bubble. If you look at seed investment rounds being multi-$10 billion rounds with basically nothing, there are talented teams, but it seems like that might be the first sign of some kind of bubble.
On the other hand, I think there’s a lot of amazing work and value that we see, at least from our perspective. Not only are there all the new product areas, like the Gemini app and NotebookLM, but thinking more forward, there’s robotics and gaming. There are incredible uses of not just Gemini, but some of our other models, like Genie.
You can imagine my old game-playing background. I’m itching to think about what could be done there. And there’s drug discovery, which we’re doing with Isomorphic, and Waymo. So there are all these new greenfield areas.
They’re going to take a while to mature into massive, multi-$100 billion businesses, but I think there’s potential for half a dozen to a dozen of them. I think Alphabet will be involved with those, which I’m really excited about.
There are also immediate returns. We have the engine room, of course. This is the engine-room part of Google, where we’re pushing this into all of these incredible, multibillion-user products that people use every day. We have so many ideas; it’s just about execution. How would you reorganize Workspace around that? Android? YouTube? There’s just so much potential there.
I think a lot of that will also bring in near-term revenue and direct returns, while we’re also investing in the future—not to speak of cloud revenue, TPUs, and all of that, which I think is also going to be huge. So I feel really good about where we are as Alphabet, whether or not there’s a bubble.
I think our job is to be winning in both cases, right? If there’s no bubble and things carry on, then we’re going to take advantage of that opportunity. But if there is some sort of bubble and there’s a retrenchment, I think we’ll also be best placed to take advantage of that scenario.
Casey Newton
All right, let’s imagine Thanksgiving is coming up, and it’s the Bay Area. One of our listeners changes the subject from politics, which is upsetting everyone, to AI. Give people something to be excited about, and someone says, “Hey, I heard Gemini 3 just came out.” What can it actually do?
What’s the example that you would have our listeners show their friends, whether it’s on their phone or their laptop, to be like, “Get a load of this,” and save Thanksgiving?
Speaker 4
Yeah, I don’t know if it’ll save Thanksgiving, but it could probably provide some laughs. Our imagery models in Gemini are still the best in the world.
What I would say is, grab your phone. It can be an iPhone or Android; it doesn’t matter. Pull it out. You can take a selfie, put yourself in it, and edit it. People are still doing that in huge amounts, and it’s great fun.
Then I think you can show off any other capabilities in the new Gemini 3 alongside it. This is what we’re seeing: People are coming for a lot of these interesting use cases and then starting to try other parts of the app, too.
Kevin Roose
You heard it here: Nano Banana will save Thanksgiving dinner. Gentlemen, thank you. It’s great to talk, and thanks for making the time.
Speaker 4
Appreciate it.
Speaker 5
Thanks. Thanks for having us.
Speaker 4
Thank you all. Thanks, guys.