[BidClub_]
All-In · · 32 min

Inside Google DeepMind: AGI, Robotics, & World Models Explained - Demis Hassabis

Demis Hassabis

YouTube
TL;DR
  • Google DeepMind is now Alphabet’s centralized AI “engine room,” combining roughly 5,000 staff—80% or more, by Hassabis’s estimate, engineers and PhD researchers—with immediate distribution to billions of users. Gemini already powers AI Overviews, AI Mode and its own app, with Workspace and Gmail incorporation underway; the strategic advantage is a tight research-to-deployment loop across nearly every Google surface.

  • Genie 3 turns text prompts into controllable worlds whose pixels, objects and interactions are generated on the fly. The host contrasted it with conventional rendering engines; Hassabis said it was trained on video plus synthetic game-engine data and had reverse-engineered intuitive physics. It can maintain “a consistent minute or two” of interaction and preserve earlier changes. He sees these world models as foundational to AGI, robotics, smart glasses and assistants that understand physical context.

  • Google is pursuing both an “Android play” for robotics—a model layer spanning different machines—and vertically integrated model-and-hardware systems. Hassabis expects a “real wow moment” within a couple of years and eventually millions of robots, but says the algorithms need to become more reliable and hardware makers risk locking designs into factories just before a better generation arrives.

  • Hassabis rejects claims that today’s systems are broadly “PhD intelligences,” because isolated PhD-level abilities coexist with high-school math and counting failures. His AGI test is genuine invention: could a model capped at 1901 derive special relativity as Einstein did in 1905, or create a game as elegant as Go rather than merely discovering move 37? He estimates AGI is five to 10 years away and probably needs “one or two missing breakthroughs.”

  • Generative tools may commoditize production skills without commoditizing taste, vision or storytelling. Nano Banana’s key differentiator is consistency—changing the requested element while preserving everything else—while Veo collaborations suggest elite professionals could become “10x, 100x more productive.” Hassabis expects shared, professionally authored worlds to persist, but with audiences co-creating inside them.

  • Isomorphic aims to compress drug discovery from years or sometimes a decade to “weeks or even days” over the next 10 years. Its platform builds “adjacent AlphaFolds” for compound design, and Hassabis expects to enter preclinical work sometime next year, alongside Eli Lilly, Novartis and internal programs spanning cancer, immunology and oncology, with work involving places such as MD Anderson.

  • AI efficiency has improved roughly 10x—and in some cases 100x—for equivalent performance over two years, but frontier scaling means those gains have not reduced demand. Hassabis thinks AI will ultimately return more than it consumes through grid optimization, materials and new energy sources. If full AGI arrives within the next 10 years, he thinks it would usher in “a new golden era of science.”

Digest · the substance, structured for research

1. DeepMind has become Alphabet’s research-to-distribution engine

  • Hassabis learned of his Nobel roughly 10 minutes before it became public. He emphasized that the committee seeks both a scientific breakthrough and real-world impact, which can take “20, 30 years” to emerge—making AlphaFold’s early recognition genuinely surprising.

  • Google consolidated its formerly separate AI efforts into Google DeepMind, now roughly 5,000 people and 80% or more engineers and PhD researchers, by Hassabis’s estimate. He calls it “the engine room of the whole of Google and the whole of Alphabet.”

  • Gemini and DeepMind’s video and interactive world models now plug into nearly every Google product. Billions encounter them through AI Overviews, AI Mode and the Gemini app, while Workspace and Gmail are also being incorporated.

2. Genie 3 treats generated worlds as a route to physical intelligence

  • Genie 3 creates interactive environments from one text prompt: users navigate with ordinary controls while every visible pixel is generated on demand. In the demonstration, painted wall markings remained after the user looked away, while a chicken-suited person or jet ski could be inserted in real time.

  • The host contrasted the demo with Unity or Unreal and conventional rendering engines involving pre-created objects and programmed lighting and physics, describing it as 2D imagery generated on the fly. Hassabis said Genie 3 had “reverse-engineered intuitive physics.”

  • Trained on millions of videos plus synthetic game-engine data, Genie 3 can sustain many kinds of worlds for “a consistent minute or two.”

  • Hassabis’s framing: language and mathematics alone cannot produce truly general intelligence. AGI must understand “the physical world around us,” while generation itself is an expression of understanding the world’s dynamics—critical for robotics, smart glasses and context-aware assistants.

  • Gemini Robotics already converts language requests such as putting a yellow object into a red bucket into robotic-hand movements. Google is pursuing both a cross-robotics “Android play” and vertically integrated combinations of its newest models with particular robot designs.

3. Robotics will scale only after reliability and hardware converge

  • Hassabis has changed his mind on humanoids. He once expected mostly task-specific robots and still favors them for laboratories and production lines, but the ordinary physical world was designed around human bodies—“steps, doorways” and other existing features—making humanoid compatibility potentially important.

  • He remains “a little bit early” on the sector: algorithms need better reliability and physical understanding, though a “real wow moment” may arrive within a couple of years. His destination is millions of productivity-enhancing robots, not a firm five- or seven-year unit forecast.

  • The manufacturing dilemma is timing. Once factories commit to tens or hundreds of thousands of one design, rapid iteration becomes harder—and a more dexterous generation may appear six months later. The host compared robotics to 1970s computing; Hassabis agreed, except “10 years happens in one year probably.”

4. AGI requires invention, consistency and continual learning

  • AI-assisted science remains Hassabis’s career-long objective. DeepMind has applied related systems to protein folding, materials, fusion-plasma control, weather and Math Olympiad problems, but today’s AI still lacks “true creativity”: it may prove a supplied conjecture without originating a new hypothesis.

  • One proposed test is historical: cap a system’s knowledge at 1901 and ask whether it can discover special relativity as Einstein did in 1905. Another separates optimization from creation: AlphaGo invented move 37, but no current system could invent a game “as elegant, as satisfying, as aesthetically beautiful as Go.”

  • Claims that current models are “PhD intelligences” are “nonsense,” Hassabis argued. They display some PhD-level capabilities but are not generally competent at that level and can still fail high-school mathematics or simple counting when a question is phrased differently.

  • He places AGI roughly five to 10 years away. Scaling might close the gap, but his bet is that “one or two missing breakthroughs” are needed, including better reasoning, intuitive leaps, consistency and continual online learning. He said Google DeepMind is not seeing either convergence or an internal slowdown, while pointing to continued progress in Genie, Veo and Nano Banana.

5. Creative production gets cheaper while taste remains differentiated

  • Nano Banana’s key differentiator is not merely state-of-the-art image generation but controlled iteration: it follows a requested edit while keeping everything else consistent. Hassabis sees conversational “vibing” with tools replacing complex Photoshop-style interfaces.

  • Two effects coexist. Anyone can create without mastering elaborate software, while top filmmakers and artists can test ideas cheaply and become “10x, 100x more productive.” Collaborations with director Darren Aronofsky and others using Veo help DeepMind learn which professional features matter.

  • Hassabis expects neither purely personalized media nor unchanged one-to-many entertainment. Top visionaries may build dynamic worlds entered by millions, audiences may co-create selected elements, and the lead creator becomes “almost an editor of that world”—a new genre enabled by Genie-like systems.

6. Hybrid science models bridge scarce data while scale sustains energy demand

  • Isomorphic is building “many adjacent AlphaFolds” to design compounds that bind correctly without unwanted effects. Hassabis thinks discovery could fall from years—or a decade—to weeks or days over 10 years. He said the company expects to enter preclinical work sometime next year, has partnerships with Eli Lilly and Novartis, and is running internal programs while working with places such as MD Anderson.

  • Biology and chemistry often lack enough data for unconstrained learning, so AlphaFold mixes neural learning with known chemistry and physics, including bond angles and rules preventing atoms from overlapping. AlphaGo was similarly hybrid: a neural network recognized promising patterns while Monte Carlo search handled planning.

  • The long-run goal is to “upstream” handcrafted knowledge into end-to-end learning. AlphaZero followed that path by removing Go-specific knowledge and human game data, learning from scratch and extending beyond a single game.

  • Google’s serving requirements have driven roughly 10x—and sometimes 100x—efficiency gains for equal performance over two years. Those gains have not reduced demand because frontier experiments keep scaling. Hassabis thinks AI’s contributions to grid and electrical systems, materials and new energy sources will outweigh its consumption. If full AGI arrives within 10 years, he thinks it would usher in “a new golden era of science.”

Speaker 1

Welcome.

Demis Hassabis

Great to be here.

Speaker 1

Thanks. First off, congratulations on winning the Nobel Prize.

Demis Hassabis

Thank you.

Speaker 1

And thanks for the incredible breakthrough of AlphaFold. Maybe you've done this before, but I know everyone here would love to hear your recounting of where you were when you won the Nobel Prize. How did you find out?

Demis Hassabis

It was a very surreal moment, obviously. Everything about it is surreal. The way they tell you—they tell you 10 minutes before it all goes live. You can't really process it; you're shell-shocked when you get that call from Sweden. It's the call that every scientist dreams about.

Then there are the ceremonies and the whole week in Sweden with the royal family. It's amazing. Obviously, it's been going for 120 years. The most amazing bit is that they bring out this Nobel book from the vaults, from the safe, and you get to sign your name next to all the other greats.

It's quite an incredible moment, leafing back through the other pages and seeing Feynman, Marie Curie, Einstein, and Niels Bohr. You just carry on going backwards, and you get to put your name in that book. It's incredible.

Speaker 1

Did you have an inkling that you'd been nominated and that this might be coming your way?

Demis Hassabis

You hear rumors. It's amazingly locked down, actually, in today's age, how they keep it so quiet. It's sort of like a national treasure for Sweden.

You hear that maybe AlphaFold is the kind of thing that would be worthy of that recognition. They look for impact as well as the scientific breakthrough—impact in the real world—and that can take 20 or 30 years to arrive. So you just never know how soon it's going to be, or whether it's going to happen at all. It's a surprise.

Speaker 1

Well, congratulations.

Demis Hassabis

Thank you.

Speaker 1

And thank you. You let me take a picture with it a few weeks ago, so that's something I'll cherish.

What is DeepMind within Alphabet? Alphabet is a sprawling organization with sprawling business units. What is DeepMind? What are you responsible for?

Demis Hassabis

We now see DeepMind, or Google DeepMind as it's become. We merged all of the different AI efforts across Google and Alphabet, including DeepMind, a couple of years ago. We put it all together, bringing the strengths of all the different groups into one division.

The way I describe it now is that we're the engine room of the whole of Google and the whole of Alphabet. Gemini is our main model, but we also build many other models, including video models and interactive world models, and we plug them in all across Google now. Pretty much every product and every surface area has one of our AI models in it.

Billions of people now interact with Gemini models, whether that's through AI Overviews, AI Mode, or the Gemini app. That's just the beginning. We're incorporating it into Workspace, Gmail, and so on. It's a fantastic opportunity for us to do cutting-edge research and then immediately ship it to billions of users.

Speaker 1

How many people are there, and what's the profile? Are these scientists and engineers? What's the makeup of your organization?

Demis Hassabis

There are around 5,000 people in Google DeepMind, and it's predominantly—80% or more, I guess—engineers and PhD researchers. So, about 3,000 or 4,000 people.

Speaker 1

There's an evolution of models, a lot of new models coming out, and also new classes of models. The other day, you released this Genie 3 world model.

Demis Hassabis

Yes.

Speaker 1

What is the Genie 3 world model? I think we have a video of it. Is it worth looking at so we can talk about it live?

Demis Hassabis

Yes, we can watch it.

Speaker 1

Because I think you have to see it to understand it; it's so extraordinary. Can we pull up the video, and then Demis can narrate a little bit about what we're looking at?

Speaker 2

What you're seeing are not games or videos. They're worlds. Each one of these is an interactive environment generated by Genie 3, a new frontier for world models.

With Genie 3, you can use natural language to generate a variety of worlds and explore them interactively, all with a single text prompt.

Demis Hassabis

All of these videos and interactive worlds that you're seeing are environments that someone can actually control. It's not a static video; it's being generated from a text prompt. People are able to control the 3D environment using the arrow keys and the spacebar.

Everything you're seeing here is being generated on the fly. All of these pixels are generated as the player—or the person interacting with it—goes to that part of the world. All of this richness is generated in real time.

You'll see in a second that this is fully generated. This is not a real video. It's someone painting their room and painting some things on the wall. Then the player is going to look to the right and then look back. This part of the world didn't exist before, so now it exists. They look back, and they see the same painting marks they left earlier.

Again, every pixel you can see is fully generated. You can type things like “a person in a chicken suit” or “a jet ski,” and it will include them in the scene in real time.

Speaker 1

It's quite mind-blowing, really. I think what's hard to grasp when looking at this is that we've all played video games that have a 3D element to them, where you're in an immersive world, but there are no objects that have been created. There's no rendering engine. You're not using Unity or Unreal, which are the 3D rendering engines.

Demis Hassabis

Yeah.

Speaker 1

This is actually just 2D images being created and rendered on the fly by the AI.

Demis Hassabis

This model is reverse-engineering intuitive physics. It's watched many millions of videos—YouTube videos and other things about the world—and from that, it's reverse-engineered how a lot of the world works.

It's not perfect yet, but it can generate a consistent minute or 2 of interaction with you as the user in many different worlds. There are some videos later on where you can control a dog on a beach or a jellyfish. It's not limited to just human things.

Speaker 1

The way a 3D rendering engine works is that the programmer programs all the laws of physics: How does light reflect off an object? You create a 3D object, light reflects off it, and what I see visually is rendered by the software because it has all the programming for how to create and simulate physics.

But this was just trained on video, and it figured it all out.

Demis Hassabis

Yeah, it was trained on video and some synthetic data from game engines, and it just reverse-engineered it.

For me, this project is very close to my heart, but it's also quite mind-blowing because, in the 1990s, early in my career, I used to write video games, AI for video games, and graphics engines. I remember how hard it was to do this by hand—to program all the polygons and the physics engines.

It's amazing to see this do it effortlessly: all of the reflections on the water, the way materials flow, and the way objects behave. It's doing all of that out of the box. It's hard to describe how much complexity was solved by that model. It's really, really mind-blowing.

Speaker 1

Where does this lead us? Fast-forward this model to Genie 5.

Demis Hassabis

The reason we're building these kinds of models is that we've always felt that, while we should continue progressing with normal language models like Gemini, we wanted Gemini to be multimodal from the beginning. We wanted it to take any kind of input—images, audio, or video—and it can output anything.

We've been very interested in this because, for an AI to be truly general—to build AGI—we feel that the AGI system needs to understand the world around us and the physical world around us, not just the abstract world of language or mathematics. Of course, that's critical for robotics to work. It's probably what's missing from AI today.

Smart glasses are another example. A smart-glasses assistant that helps you in your everyday life has to understand the physical context that you're in and how the intuitive physics of the world works.

We think that building these types of models—these Genie models, and also Veo, our best text-to-video model—are expressions of us building world models that understand the dynamics and physics of the world. If you can generate it, then that's an expression of your system understanding those dynamics.

Speaker 1

That ultimately leads to a world of robotics—one aspect or application, at least. Maybe we can talk about that. What is the state of the art with vision-language-action models today?

I'm thinking of a generalized system—a box, a machine—that can observe the world with a camera, and then I can use language, text, or speech to tell it, “I want you to do it.”

Speaker 1

And then it knows how to act physically to do something in the physical world for me.

Demis Hassabis

That’s right. If you look at our Gemini Live version of Gemini, where you can hold up your phone to the world around you, I’d recommend any of you try it. It’s kind of magical what it already understands about the physical world. You can think of the next step as incorporating that into some sort of more handy device, like glasses. Then it will be an everyday assistant, and it’ll be able to recommend things to you as you’re walking the streets, or we can embed it into Google Maps.

With robotics, we’ve built something called Gemini Robotics models, which are sort of fine-tuned Gemini models with extra robotics data. What’s really cool about that—and we released some demos of this over the summer—is that you can have tabletop setups with two robotic hands interacting with objects on a table, and you can just talk to the robot. You can say, “Put the yellow object into the red bucket,” or whatever it is, and it will interpret that language instruction into motor movements.

That’s the power of a multimodal model rather than just a robotics-specific model: it will be able to bring real-world understanding to the way you interact with it. In the end, it will be the UI and UX that you need, as well as the understanding that the robots need to navigate the world safely.

Speaker 1

I asked Sundar this: Does that mean that, ultimately, you could build what would be the equivalent of a Unix-like operating system layer, or an Android for generalized robotics? At that point, if it works well enough across enough devices, there will be a proliferation of robotics devices, companies, and products that will suddenly take off in the world because this software exists to do this generally.

Demis Hassabis

Exactly. That’s certainly one strategy we’re pursuing: an Android play, if you like, across robotics—almost a kind of robotics OS layer. But there are also some quite interesting things about vertically integrating our latest models with specific robot types and robot designs, and doing some kind of end-to-end learning of that, too. Both are actually pretty interesting, and we’re pursuing both strategies.

Speaker 1

Do you think humanoid robots are a good kind of form factor? Does that make sense in the world? Some folks have criticized it as being good for humans because we’re meant to do lots of different things, but if we want to solve a problem, there may be a different form factor to fold laundry, do dishes, clean the house, or whatever.

Demis Hassabis

Yeah, I think there’s going to be a place for both. I used to be of the opinion, maybe 5 or 10 years ago, that we’d have form-specific robots for certain tasks. I think in industry, industrial robots will definitely be like that, where you can optimize the robot for the specific task. Whether it’s a laboratory or a production line, you’d want quite different types of robots.

On the other hand, for general use or personal-use robotics, and just interacting with the ordinary world, the humanoid form factor could be pretty important because, of course, we’ve designed the physical world around us for humans. Steps, doorways, and all the things that we’ve designed for ourselves—rather than changing all of those in the real world, it might be easier to design the form factor to work seamlessly with the way we’ve already designed the world.

I think there’s an argument to be made that the humanoid form factor could be very important for those types of tasks. But I think there’s also a place for specialized robotic forms.

Speaker 1

Do you have a view on hundreds of millions, millions, or thousands over the next 5 or 7 years? Do you have a vision in your head?

Demis Hassabis

Yeah, I do. I spend quite a lot of time on this, and I think we’re still a little bit early on robotics. I think in the next couple of years there’ll be a real “wow” moment with robotics, but I think the algorithms need a bit more development. The general-purpose models that these robotics models are built on still need to be better and more reliable, with a better understanding of the world around them. I think that will come in the next couple of years.

Also, on the hardware side, the key is that eventually we will have millions of robots helping society and increasing productivity. But the key question, when you talk to hardware experts, is at what point you have the right level of hardware to go for the scaling option. Effectively, when you start building factories around trying to make tens of thousands or hundreds of thousands of a particular robot type, it’s harder for you to update and quickly iterate on the robot design.

It’s one of those questions where, if you call it too early, the next generation of robots might be invented in 6 months’ time that’s just more reliable, better, and more dextrous.

Speaker 1

Sounds like, using a computing analogy, we’re kind of in the ’70s-era PC DOS kind of era.

Demis Hassabis

Yeah, potentially. But of course, I think the exception is that 10 years happens in 1 year, probably. So we’re in one of those years, right?

Speaker 1

Exactly. Let’s talk about other applications, particularly in science. True to your heart as a scientist—as a Nobel Prize-winning scientist—I always felt like the greatest things we would be able to do with AI would be the problems that are intractable to humans with our current technology, capabilities, brains, and whatnot, and we can unlock all of this potential. What are the areas of science and breakthroughs in science that you’re most excited about, and what kinds of models do we use to get there?

Demis Hassabis

AI to accelerate scientific discovery and help with things like human health is the reason I’ve spent my whole career on AI. I think it’s the most important thing we can do with AI, and I feel like if we build AGI in the right way, it will be the ultimate tool for science.

I think we’ve been showing at DeepMind a lot of the way forward with that—obviously AlphaFold most famously, but we’ve also applied our AI systems to many branches of science, whether it’s materials design, helping with controlling plasma in fusion reactors, predicting the weather, or solving Math Olympiad problems. The same types of systems, with some extra fine-tuning, can basically solve a lot of these complex problems.

I think we’re just scratching the surface of what AI will be able to do, and there are some things that are missing. AI today, I would say, doesn’t have true creativity in the sense that it can’t come up with a new conjecture or hypothesis yet. It can maybe prove something that you give it, but it’s not able to come up with a new idea or new theory itself. I think that would be one of the tests for AGI.

Speaker 1

What is that creativity as a human?

Demis Hassabis

Yeah.

Speaker 1

What is creativity?

Demis Hassabis

I think it’s these intuitive leaps that we often celebrate in the best scientists in history, and in artists, of course. Maybe it’s done through analogy or analogical reasoning. There are many theories in psychology and neuroscience as to how we as human scientists do it.

A good test for it would be something like giving one of these modern AI systems a knowledge cutoff of 1901 and seeing if it can come up with special relativity, like Einstein did in 1905. If it’s able to do that, then I think we’re onto something really important, where perhaps we’re nearing AGI.

Another example would be with our AlphaGo program that beat the world champion at Go. Not only did it win, back 10 years ago, it invented new strategies that had never been seen before for the game of Go—famously, move 37 in game 2, which is now studied. But can an AI system come up with a game as elegant, as satisfying, and as aesthetically beautiful as Go, not just a new strategy? The answer to those things at the moment is no.

That’s one of the things I think is missing from a true general system, an AGI system: it should be able to do those kinds of things as well.

Speaker 1

Can you break down what’s missing, and maybe relate it to the point of view shared by Dario Amodei and others about AGI being a few years away? Do you not subscribe to that belief? In your understanding of the structure, in your understanding of the system architecture, what’s lacking?

Demis Hassabis

I think the fundamental aspect of this is: Can we mimic these intuitive leaps rather than incremental advances that the best human scientists seem to be able to make? I always say that what separates a great scientist from a good scientist is that they’re both technically very capable, of course, but the great scientist is more creative. Maybe they’ll spot some pattern from another subject area that can have an analogy or some sort of pattern matching to the area they’re trying to solve.

I think one day AI will be able to do this, but it doesn’t have the reasoning capabilities and some of the thinking capabilities that are going to be needed to make that kind of breakthrough. I also think that we’re lacking consistency. You often hear some of our competitors talk about how these modern systems that we have today are PhD intelligences.

I think that's nonsense. They're not PhD intelligences. They have some capabilities that are at a PhD level, but they're not generally capable—and that's exactly what general intelligence should be—of performing across the board at the PhD level.

In fact, as we all know from interacting with today's chatbots, if you pose a question in a certain way, they can make simple mistakes with even high-school math and simple counting. That shouldn't be possible for a true AGI system. So I think we're maybe, I would say, 5 to 10 years away from having an AGI system that's capable of doing those things.

Another thing that's missing is continual learning: the ability to teach the system something new online or adjust its behavior in some way. A lot of these core capabilities are still missing. Maybe scaling will get us there, but if I was to bet, I think there are probably 1 or 2 missing breakthroughs that are still required and will come over the next 5 or so years.

Speaker 1

In the meantime, some of the reports and scoring systems that are used seem to be demonstrating 2 things. One, perhaps—and tell me if we're wrong on this—is a convergence of performance among large language models. And number 2, perhaps, is a slowing down or flatlining of improvements in performance with each generation. Are those 2 statements generally true, or not so much?

Demis Hassabis

No, I mean, we're not seeing that internally, and we're still seeing a huge rate of progress. But we're also looking at things more broadly. You see, with our Genie models and Veo models, Nano Banana is insane.

Speaker 1

Has anyone here used it? Can I see who's used it? Has anyone used Nano Banana?

Demis Hassabis

It's incredible, right? I'm a nerd who used to use Adobe Photoshop as a kid, and Kai's Power Tools. I was telling you about Bryce 3D. The graphic systems and recognizing what was going on there was just mind-blowing.

I think that's the future of a lot of these creative tools. You're just going to vibe with them or talk to them, and they'll be consistent enough. With Nano Banana, what's amazing about it is that it's an image generator—it's state-of-the-art and best-in-class—but one of the things that makes it so great is its consistency. It's able to follow instructions about what you want changed and keep everything else the same.

You can iterate with it and eventually get the kind of output that you want. I think that's what the future of a lot of these creative tools is going to be, and it signals the direction we're heading in. People love it, and they love creating with it.

Speaker 1

So, democratization of creativity, I think, is really powerful. I remember having to buy books on Adobe Photoshop as a kid, and then you'd read them to learn how to remove something from an image, how to fill it in, how to feather it, and all this stuff. Now anyone can do it with Nano Banana. They can just explain to the software what they want it to do, and it does it.

Demis Hassabis

I think you're going to see 2 things. One is this democratization of these tools, allowing everybody to use and create with them without having to learn incredibly complex UXs and UIs, like we had to do in the past.

On the other hand, we're also collaborating with filmmakers, top creators, and artists. They're helping us design what these new tools should be and what features they would want. People like the director Darren Aronofsky, who's a good friend of mine and an amazing director, have been making films with their teams using Veo and some of our other tools. We're learning a lot by observing and collaborating with them.

What we find is that it also superpowers and turbocharges the best professionals. The best professional creatives are suddenly able to be 10 or 100 times more productive. They can try out all sorts of ideas they have in mind at very low cost and then get to the beautiful thing they wanted.

I actually think both things are true. We're democratizing creativity for everyday use—for YouTube creators and so on—but at the high end, the people who understand these tools can get more out of them. There's a skill in that, as well as the vision, storytelling, and narrative style of the top creatives.

These tools allow them to iterate much faster, and they really enjoy using them.

Speaker 1

Do we get to a world where each individual describes what sort of content they're interested in? You say, "Play me music like Dave Matthews," and it'll play some new track.

Demis Hassabis

Yes.

Speaker 1

Or, "I want to play a video game set in the movie Braveheart, and I want to be in that movie."

Demis Hassabis

Yes.

Speaker 1

And I just have that experience. Do we end up there, or do we still have a one-to-many creative process in society? How important is this culturally? I know this is a little bit philosophical, but it's interesting to me. Are we still going to have storytelling where we have 1 story that we all share because someone made it, or are we each going to start to develop and pull on our own kind of virtual worlds?

Demis Hassabis

I actually foresee a world—and I think about this a lot, having started in the games industry as a game designer and programmer in the '90s—where this is the beginning of the future of entertainment. Maybe it will be some new genre or new art form involving a bit of co-creation.

I still think you'll have the top creative visionaries creating these compelling experiences and dynamic storylines. They'll be of higher quality, even if they're using the same tools, than what the everyday person can create. Millions of people will potentially dive into those worlds, but maybe they'll also be able to co-create certain parts of them. Perhaps the main creative person is almost an editor of that world.

Those are the kinds of things I'm foreseeing in the next few years, and I'd actually like to explore them ourselves with technologies like Genie.

Speaker 1

Right. Incredible. How are you spending your time? Maybe you can describe Isomorphic Labs. What is Isomorphic Labs, and are you spending a lot of your time there?

Demis Hassabis

I am. I also run Isomorphic Labs, which is our spinout company to revolutionize drug discovery, building on our AlphaFold breakthrough in protein folding.

Of course, knowing the structure of a protein is only 1 step in the drug discovery process. You can think of Isomorphic Labs as building many adjacent AlphaFolds to help with things like designing chemical compounds that don't have any side effects but bind to the right place on the protein.

I think we could reduce drug discovery from taking years—sometimes a decade—down to maybe weeks or even days over the next 10 years.

Speaker 1

It's incredible. Do you think that's in the clinic soon, or is that still in the discovery phase?

Demis Hassabis

We're building up the platform right now. We have great partnerships with Eli Lilly—I think you had the CEO speaking earlier—and Novartis, which are fantastic, as well as our own internal drug programs. I think we'll be entering the preclinical phase sometime next year. Then candidates get handed over to the pharmaceutical company, and they take them forward.

We're working on cancers, immunology, and oncology, and we're working with places like MD Anderson.

Speaker 1

How much of this requires—and I just want to go back to your point about AGI as it relates to what you just said—models can be probabilistic or deterministic. Tell me if I'm reducing this down too simplistically: the model takes an input and outputs something very specific, like it has a logical algorithm and outputs the same thing every time. It could be probabilistic, where it can change things and make selections: "The probability is 80%, I'll select this letter; 90%, I'll select this letter," and so on.

How much do we have to develop deterministic models that sync up with, for example, the physics or chemistry underlying the molecular interactions as you do your drug discovery modeling? How much are you building novel deterministic models that work with models that are probabilistic and trained on data?

Demis Hassabis

It's a great question. For the moment, and probably for the next 5 years or so, we're building what you might call hybrid models.

AlphaFold itself is a hybrid model. You have the learning component—the probabilistic component you're talking about, based on neural networks, transformers, and things like that—and that's learning from the data you give it, any data you have available. But in a lot of cases with biology and chemistry, there isn't enough data to learn from, so you also have to build in some of the rules about chemistry and physics that you already know about.

For example, with AlphaFold, you have the angles of bonds between atoms. You need to make sure that AlphaFold understands that you can't have atoms overlapping with each other and things like that. In theory, it could learn that, but it would waste a lot of its learning capacity. So it's better to have that as a constraint in the system.

The trick with all hybrid systems is how to marry a learning system with a more handcrafted, bespoke system and actually have them work well together. AlphaGo was another hybrid system, where a neural network was learning about the game of Go and what kinds of patterns were good, and then we had Monte Carlo search on top, which was doing the planning.

That's pretty tricky to do.

Speaker 1

Does that sort of architecture ultimately lead to the breakthroughs needed for AGI, do you think? Are there deterministic components that need to be solved?

Demis Hassabis

I think ultimately what you want to do is, when you figure out something with one of these hybrid systems, upstream it into the learning component. It's always better if you can do end-to-end learning and directly predict the thing that you're after from the data that you're given.

Once you've figured out something using one of these hybrid systems, you then try to go back and reverse-engineer what you've done and see if you can incorporate that information into the learning system. This is sort of what we did with AlphaZero, the more general form of AlphaGo.

AlphaGo had some Go-specific knowledge in it. But then with AlphaZero, we got rid of that, including the human data and human games that we learned from, and actually just did self-learning from scratch. Of course, then it was able to learn any game, not just Go.

Speaker 1

A lot of hype and hoopla has been made about the demand for energy arising from AI. This was a big part of the AI summit we held in Washington, D.C., a few weeks ago, and it seems to be the number one topic everyone talks about in tech nowadays. Where's all this power going to come from?

But I ask the question of you: Are there changes in the architecture of the models or the hardware, or the relationship between the models and the hardware, that bring down the energy per token of output or the cost per token of output? Could that ultimately mute the energy demand curve that's in front of us, or do you not think that's the case and we're still going to have a pretty geometric energy demand curve?

Demis Hassabis

Well, interestingly, I think both cases are true. Especially at Google and DeepMind, we focus a lot on very efficient models that are powerful because we have our own internal use cases, of course, where we need to serve, say, AI Overviews to billions of users every day. It has to be extremely efficient, extremely low latency, and very cheap to serve.

We've pioneered many techniques that allow us to do that, like distillation, where you have a bigger model internally that trains the smaller model. You train the smaller model to mimic the bigger model. Over time, if you look at the progress of the last 2 years, model efficiencies are 10 times, even 100 times, better for the same performance.

The reason that that isn't reducing demand is because we still haven't got to AGI yet. The frontier models keep wanting to train and experiment with new ideas at larger and larger scale, while at the same time, on the serving side, things are getting more and more efficient. Both things are true.

In the end, from the energy perspective, I think AI systems will give back a lot more to energy and climate change than they take, in terms of the efficiency of grid systems and electrical systems, material design, new types of properties, and new energy sources. I think AI will help with all of that over the next 10 years, and that will far outweigh the energy that it uses today.

Speaker 1

As the last question, describe the world 10 years from now.

Demis Hassabis

Wow. Okay. Well, 10 years—even 10 weeks—is a lifetime in AI. But I do feel like if we will have AGI in the next 10 years—full AGI—I think that will usher in a new golden era of science, a kind of new Renaissance. I think we'll see the benefits of that right across the board, from energy to human health.

Speaker 1

Amazing. Please join me in thanking Nobel laureate Demis Hassabis. Thank you. That was great. Thank you.

Inside Google DeepMind: AGI, Robotics, & World Models Explained - Demis Hassabis | BidClub