[BidClub_]
All-In · · 32 分钟

走进 Google DeepMind:AGI、机器人与世界模型解析——Demis Hassabis

Demis Hassabis

YouTube
TL;DR
  • Google DeepMind 如今已成为 Alphabet 集中的 AI“动力核心”,汇聚约5,000人;按Hassabis估计,其中80%或以上是工程师和博士研究员,并能直接触达数十亿用户。 Gemini 已经驱动 AI Overviews、AI Mode 和 Gemini 独立应用,Workspace 和 Gmail 也在接入;其战略优势在于,研究成果可以在几乎所有 Google 产品上快速部署,形成紧密的研究到落地闭环。

  • Genie 3 能把文字提示转化为可控世界,像素、物体和交互行为都按需实时生成。 主持人将其与传统渲染引擎作了对比;Hassabis 表示,Genie 3 的训练数据包括视频和合成游戏引擎数据,并已逆向还原出直觉物理。它可以让互动“稳定持续1至2分钟”,并保留此前做出的改动。在他看来,这类世界模型是 AGI、机器人、智能眼镜以及理解物理环境的助手的基础。

  • Google 正在同时推进机器人领域的“Android 路线”——覆盖不同机器的模型层——以及模型与硬件垂直整合的系统。 Hassabis 预计,未来几年内会出现“真正令人惊叹的时刻”,最终机器人数量将达到数百万台;但他也表示,算法可靠性仍需提升,硬件厂商则可能在更好一代产品出现前,就把设计锁定进工厂。

  • Hassabis 驳斥当今系统普遍具备“博士级智能”的说法,因为它们在拥有孤立的博士级能力的同时,仍会在高中数学和简单计数上出错。 他对 AGI 的测试标准是真正的发明:如果把模型知识截断在1901年,它能否像 Einstein 在1905年那样推导出狭义相对论;或者能否创造出像 Go 一样优雅的游戏,而不只是发现第37手这样的妙着?他估计 AGI 还要5至10年,可能需要“1至2个缺失的突破”。

  • 生成式工具可能让生产技能商品化,却不会让品味、视野或叙事能力商品化。 Nano Banana 的关键差异化在于一致性:只改变用户指定的元素,同时保留其他一切;而与 Veo 的合作表明,顶尖专业人士可能因此实现“10倍、100倍的生产力提升”。Hassabis 预计,由专业人士共同创作的共享世界仍会存在,只是观众也将参与其中、共同创作。

  • Isomorphic 的目标是在未来10年内,把药物发现从数年、甚至有时长达10年,压缩到“几周甚至几天”。 其平台正在构建用于化合物设计的“相邻 AlphaFolds”,Hassabis 预计公司将在明年某个时候进入临床前研究,并与 Eli Lilly、Novartis 合作,同时推进覆盖癌症、免疫学和肿瘤学的内部项目,工作也涉及 MD Anderson 等机构。

  • 在过去2年里,在实现同等性能的前提下,AI 效率提升了约10倍,部分情况下达到100倍,但前沿模型持续扩展意味着这些提升并未降低需求。 Hassabis 认为,AI 最终将通过电网优化、材料科学和新型能源等领域创造超过自身消耗的价值。如果完整 AGI 在未来10年内出现,他认为这将开启“科学的新黄金时代”。

摘要 · 为研究而整理的核心内容

1. DeepMind 已成为 Alphabet 连接研究与分发的引擎

  • Hassabis 在诺贝尔奖消息公开前约10分钟才得知自己获奖。他强调,委员会看重的不仅是科学突破,也看重现实世界影响,而后者可能需要“20年、30年”才能显现,因此 AlphaFold 如此早获认可确实出人意料。

  • Google 将此前相互独立的 AI 业务整合为 Google DeepMind。按 Hassabis 的估计,该部门如今约有5,000人,其中80%或以上是工程师和博士研究员;他称其为“整个 Google 和 Alphabet 的动力核心”。

  • Gemini 以及 DeepMind 的视频模型和交互式世界模型,正在接入几乎所有 Google 产品。数十亿用户通过 AI Overviews、AI Mode 和 Gemini 应用接触这些能力,Workspace 和 Gmail 也在接入。

2. Genie 3 将生成世界视为通往物理智能的路径

  • Genie 3 可以通过一条文字提示创建交互式环境:用户用普通控制方式移动,画面中每个可见像素都按需生成。在演示中,用户移开视线后,墙上的涂鸦痕迹仍然保留;穿着鸡服的人或水上摩托也可以实时插入场景。

  • 主持人将演示与 Unity、Unreal 以及传统渲染引擎作对比,后者依赖预先创建的物体,以及编程设定的光照和物理效果;他形容 Genie 3 是实时生成的2D影像。Hassabis 表示,Genie 3 已经“逆向还原出直觉物理”。

  • Genie 3 的训练数据包括数百万段视频和合成游戏引擎数据,可以让多种类型的世界“稳定持续1至2分钟”。

  • Hassabis 的框架是:仅靠语言和数学无法产生真正的通用智能。AGI 必须理解“我们周围的物理世界”,而生成本身就是理解世界动态的一种表现,这对机器人、智能眼镜和具备环境感知能力的助手都至关重要。

  • Gemini Robotics 已经能把“把黄色物体放进红色桶里”这类语言请求,转换成机器人手部动作。Google 正在同时推进跨机器人领域的“Android 路线”,以及将最新模型与特定机器人设计垂直整合的方案。

3. 机器人只有在可靠性与硬件 converging 后才能规模化

  • Hassabis 已经改变了对人形机器人的看法。他过去预计机器人将主要按任务定制,至今仍认为实验室和生产线适合使用专用机器人;但日常物理世界是围绕人体设计的——“台阶、门口”等既有设施都如此——因此人形兼容性可能很重要。

  • 他仍认为自己在这个行业的判断“稍微有点早”:算法需要提升可靠性和物理理解能力,不过未来几年内可能出现“真正令人惊叹的时刻”。他的目标是数百万台提升生产力的机器人,而不是给出5年或7年的确定装机预测。

  • 制造业面临的难题在于时机。一旦工厂承诺生产数万甚至数十万台同一设计的机器人,快速迭代就会变得困难;而更灵巧的新一代产品可能在6个月后出现。主持人将机器人产业比作1970年代的计算机业,Hassabis 表示认同,只是“10年可能1年就会发生”(“10 years happens in one year probably”)。

4. AGI 需要发明能力、一致性与持续学习

  • AI 辅助科学一直是 Hassabis 的长期目标。DeepMind 已将相关系统用于蛋白质折叠、材料、聚变等离子体控制、天气预测和数学奥赛题,但如今的 AI 仍缺乏“真正的创造力”:它可以证明别人给出的猜想,却无法自行提出新的假设。

  • 一个测试方案来自历史:把系统的知识截断在1901年,观察它能否像 Einstein 在1905年那样发现狭义相对论。另一个测试则区分优化与创造:AlphaGo 发明了第37手,但目前没有系统能创造出一款“像 Go 一样优雅、令人满足、具有审美之美”的游戏。

  • Hassabis 认为,宣称当前模型是“博士级智能”是“胡说”。它们确实展现出部分博士级能力,但并不具备普遍的博士水平;当问题换一种表述时,它们甚至仍可能答错高中数学题或简单的计数题。

  • 他认为 AGI 距今约5至10年。规模扩展可能缩小差距,但他的判断是仍需要“1至2个缺失的突破”,包括更强的推理、直觉式跃迁、一致性和持续在线学习。他表示,Google DeepMind 尚未看到收敛迹象,也没有看到内部放缓,同时 Genie、Veo 和 Nano Banana 仍在持续进步。

5. 创意生产成本下降,品味仍然稀缺

  • Nano Banana 的关键差异化并不只是最先进的图像生成能力,还在于可控迭代:按照用户要求修改,同时保持其他内容一致。Hassabis 认为,复杂的 Photoshop 式界面将被与工具进行对话式“试做”所取代。

  • 两种效果会同时出现:任何人都能创作,无需掌握复杂软件;顶尖电影人和艺术家则可以低成本测试想法,成为“10倍、100倍更高效”的创作者。与导演 Darren Aronofsky 等人使用 Veo 的合作,也帮助 DeepMind 了解专业人士真正需要哪些功能。

  • Hassabis 预计,未来既不会是完全个性化的媒体,也不会维持不变的一对多娱乐模式。顶尖创作者可能构建动态世界,让数百万人进入其中;观众可以共同创作部分元素,而主创者则“几乎成为这个世界的编辑”——这是 Genie 类系统催生的新类型。

6. 混合科学模型连接稀缺数据,规模扩展持续推高能源需求

  • Isomorphic 正在构建“许多相邻 AlphaFolds”,用于设计能够正确结合、同时避免不良副作用的化合物。Hassabis 认为,未来10年内,药物发现周期可能从数年甚至10年缩短到几周或几天。他表示,公司预计明年某个时候进入临床前研究,已与 Eli Lilly 和 Novartis 建立合作,同时推进内部项目,并与 MD Anderson 等机构开展工作。

  • 生物学和化学往往缺乏足够数据,无法进行不受约束的学习,因此 AlphaFold 将神经网络学习与已知化学和物理规律结合起来,包括键角,以及防止原子相互重叠的规则。AlphaGo 同样是混合系统:神经网络负责识别有潜力的模式,蒙特卡洛搜索负责规划。

  • 长期目标是把手工编码的知识逐步纳入端到端学习。AlphaZero 沿着这一路径前进:移除 Go 专属知识和人类棋局数据,从零开始学习,并将能力扩展到单一游戏之外。

  • Google 的推理服务需求推动了过去2年约10倍、部分情况下100倍的同等性能效率提升。但由于前沿实验仍在持续扩展,这些提升并未降低需求。Hassabis 认为,AI 对电网和电力系统、材料以及新型能源的贡献,最终将超过其自身消耗。如果完整 AGI 在未来10年内出现,他认为这将开启“科学的新黄金时代”。

Speaker 1

Welcome.

Demis Hassabis

Great to be here.

Speaker 1

Thanks. First off, congratulations on winning the Nobel Prize.

Demis Hassabis

Thank you.

Speaker 1

And thanks for the incredible breakthrough of AlphaFold. Maybe you've done this before, but I know everyone here would love to hear your recounting of where you were when you won the Nobel Prize. How did you find out?

Demis Hassabis

It was a very surreal moment, obviously. Everything about it is surreal. The way they tell you—they tell you 10 minutes before it all goes live. You can't really process it; you're shell-shocked when you get that call from Sweden. It's the call that every scientist dreams about.

Then there are the ceremonies and the whole week in Sweden with the royal family. It's amazing. Obviously, it's been going for 120 years. The most amazing bit is that they bring out this Nobel book from the vaults, from the safe, and you get to sign your name next to all the other greats.

It's quite an incredible moment, leafing back through the other pages and seeing Feynman, Marie Curie, Einstein, and Niels Bohr. You just carry on going backwards, and you get to put your name in that book. It's incredible.

Speaker 1

Did you have an inkling that you'd been nominated and that this might be coming your way?

Demis Hassabis

You hear rumors. It's amazingly locked down, actually, in today's age, how they keep it so quiet. It's sort of like a national treasure for Sweden.

You hear that maybe AlphaFold is the kind of thing that would be worthy of that recognition. They look for impact as well as the scientific breakthrough—impact in the real world—and that can take 20 or 30 years to arrive. So you just never know how soon it's going to be, or whether it's going to happen at all. It's a surprise.

Speaker 1

Well, congratulations.

Demis Hassabis

Thank you.

Speaker 1

And thank you. You let me take a picture with it a few weeks ago, so that's something I'll cherish.

What is DeepMind within Alphabet? Alphabet is a sprawling organization with sprawling business units. What is DeepMind? What are you responsible for?

Demis Hassabis

We now see DeepMind, or Google DeepMind as it's become. We merged all of the different AI efforts across Google and Alphabet, including DeepMind, a couple of years ago. We put it all together, bringing the strengths of all the different groups into one division.

The way I describe it now is that we're the engine room of the whole of Google and the whole of Alphabet. Gemini is our main model, but we also build many other models, including video models and interactive world models, and we plug them in all across Google now. Pretty much every product and every surface area has one of our AI models in it.

Billions of people now interact with Gemini models, whether that's through AI Overviews, AI Mode, or the Gemini app. That's just the beginning. We're incorporating it into Workspace, Gmail, and so on. It's a fantastic opportunity for us to do cutting-edge research and then immediately ship it to billions of users.

Speaker 1

How many people are there, and what's the profile? Are these scientists and engineers? What's the makeup of your organization?

Demis Hassabis

There are around 5,000 people in Google DeepMind, and it's predominantly—80% or more, I guess—engineers and PhD researchers. So, about 3,000 or 4,000 people.

Speaker 1

There's an evolution of models, a lot of new models coming out, and also new classes of models. The other day, you released this Genie 3 world model.

Demis Hassabis

Yes.

Speaker 1

What is the Genie 3 world model? I think we have a video of it. Is it worth looking at so we can talk about it live?

Demis Hassabis

Yes, we can watch it.

Speaker 1

Because I think you have to see it to understand it; it's so extraordinary. Can we pull up the video, and then Demis can narrate a little bit about what we're looking at?

Speaker 2

What you're seeing are not games or videos. They're worlds. Each one of these is an interactive environment generated by Genie 3, a new frontier for world models.

With Genie 3, you can use natural language to generate a variety of worlds and explore them interactively, all with a single text prompt.

Demis Hassabis

All of these videos and interactive worlds that you're seeing are environments that someone can actually control. It's not a static video; it's being generated from a text prompt. People are able to control the 3D environment using the arrow keys and the spacebar.

Everything you're seeing here is being generated on the fly. All of these pixels are generated as the player—or the person interacting with it—goes to that part of the world. All of this richness is generated in real time.

You'll see in a second that this is fully generated. This is not a real video. It's someone painting their room and painting some things on the wall. Then the player is going to look to the right and then look back. This part of the world didn't exist before, so now it exists. They look back, and they see the same painting marks they left earlier.

Again, every pixel you can see is fully generated. You can type things like “a person in a chicken suit” or “a jet ski,” and it will include them in the scene in real time.

Speaker 1

It's quite mind-blowing, really. I think what's hard to grasp when looking at this is that we've all played video games that have a 3D element to them, where you're in an immersive world, but there are no objects that have been created. There's no rendering engine. You're not using Unity or Unreal, which are the 3D rendering engines.

Demis Hassabis

Yeah.

Speaker 1

This is actually just 2D images being created and rendered on the fly by the AI.

Demis Hassabis

This model is reverse-engineering intuitive physics. It's watched many millions of videos—YouTube videos and other things about the world—and from that, it's reverse-engineered how a lot of the world works.

It's not perfect yet, but it can generate a consistent minute or 2 of interaction with you as the user in many different worlds. There are some videos later on where you can control a dog on a beach or a jellyfish. It's not limited to just human things.

Speaker 1

The way a 3D rendering engine works is that the programmer programs all the laws of physics: How does light reflect off an object? You create a 3D object, light reflects off it, and what I see visually is rendered by the software because it has all the programming for how to create and simulate physics.

But this was just trained on video, and it figured it all out.

Demis Hassabis

Yeah, it was trained on video and some synthetic data from game engines, and it just reverse-engineered it.

For me, this project is very close to my heart, but it's also quite mind-blowing because, in the 1990s, early in my career, I used to write video games, AI for video games, and graphics engines. I remember how hard it was to do this by hand—to program all the polygons and the physics engines.

It's amazing to see this do it effortlessly: all of the reflections on the water, the way materials flow, and the way objects behave. It's doing all of that out of the box. It's hard to describe how much complexity was solved by that model. It's really, really mind-blowing.

Speaker 1

Where does this lead us? Fast-forward this model to Genie 5.

Demis Hassabis

The reason we're building these kinds of models is that we've always felt that, while we should continue progressing with normal language models like Gemini, we wanted Gemini to be multimodal from the beginning. We wanted it to take any kind of input—images, audio, or video—and it can output anything.

We've been very interested in this because, for an AI to be truly general—to build AGI—we feel that the AGI system needs to understand the world around us and the physical world around us, not just the abstract world of language or mathematics. Of course, that's critical for robotics to work. It's probably what's missing from AI today.

Smart glasses are another example. A smart-glasses assistant that helps you in your everyday life has to understand the physical context that you're in and how the intuitive physics of the world works.

We think that building these types of models—these Genie models, and also Veo, our best text-to-video model—are expressions of us building world models that understand the dynamics and physics of the world. If you can generate it, then that's an expression of your system understanding those dynamics.

Speaker 1

That ultimately leads to a world of robotics—one aspect or application, at least. Maybe we can talk about that. What is the state of the art with vision-language-action models today?

I'm thinking of a generalized system—a box, a machine—that can observe the world with a camera, and then I can use language, text, or speech to tell it, “I want you to do it.”

Speaker 1

And then it knows how to act physically to do something in the physical world for me.

Demis Hassabis

That’s right. If you look at our Gemini Live version of Gemini, where you can hold up your phone to the world around you, I’d recommend any of you try it. It’s kind of magical what it already understands about the physical world. You can think of the next step as incorporating that into some sort of more handy device, like glasses. Then it will be an everyday assistant, and it’ll be able to recommend things to you as you’re walking the streets, or we can embed it into Google Maps.

With robotics, we’ve built something called Gemini Robotics models, which are sort of fine-tuned Gemini models with extra robotics data. What’s really cool about that—and we released some demos of this over the summer—is that you can have tabletop setups with two robotic hands interacting with objects on a table, and you can just talk to the robot. You can say, “Put the yellow object into the red bucket,” or whatever it is, and it will interpret that language instruction into motor movements.

That’s the power of a multimodal model rather than just a robotics-specific model: it will be able to bring real-world understanding to the way you interact with it. In the end, it will be the UI and UX that you need, as well as the understanding that the robots need to navigate the world safely.

Speaker 1

I asked Sundar this: Does that mean that, ultimately, you could build what would be the equivalent of a Unix-like operating system layer, or an Android for generalized robotics? At that point, if it works well enough across enough devices, there will be a proliferation of robotics devices, companies, and products that will suddenly take off in the world because this software exists to do this generally.

Demis Hassabis

Exactly. That’s certainly one strategy we’re pursuing: an Android play, if you like, across robotics—almost a kind of robotics OS layer. But there are also some quite interesting things about vertically integrating our latest models with specific robot types and robot designs, and doing some kind of end-to-end learning of that, too. Both are actually pretty interesting, and we’re pursuing both strategies.

Speaker 1

Do you think humanoid robots are a good kind of form factor? Does that make sense in the world? Some folks have criticized it as being good for humans because we’re meant to do lots of different things, but if we want to solve a problem, there may be a different form factor to fold laundry, do dishes, clean the house, or whatever.

Demis Hassabis

Yeah, I think there’s going to be a place for both. I used to be of the opinion, maybe 5 or 10 years ago, that we’d have form-specific robots for certain tasks. I think in industry, industrial robots will definitely be like that, where you can optimize the robot for the specific task. Whether it’s a laboratory or a production line, you’d want quite different types of robots.

On the other hand, for general use or personal-use robotics, and just interacting with the ordinary world, the humanoid form factor could be pretty important because, of course, we’ve designed the physical world around us for humans. Steps, doorways, and all the things that we’ve designed for ourselves—rather than changing all of those in the real world, it might be easier to design the form factor to work seamlessly with the way we’ve already designed the world.

I think there’s an argument to be made that the humanoid form factor could be very important for those types of tasks. But I think there’s also a place for specialized robotic forms.

Speaker 1

Do you have a view on hundreds of millions, millions, or thousands over the next 5 or 7 years? Do you have a vision in your head?

Demis Hassabis

Yeah, I do. I spend quite a lot of time on this, and I think we’re still a little bit early on robotics. I think in the next couple of years there’ll be a real “wow” moment with robotics, but I think the algorithms need a bit more development. The general-purpose models that these robotics models are built on still need to be better and more reliable, with a better understanding of the world around them. I think that will come in the next couple of years.

Also, on the hardware side, the key is that eventually we will have millions of robots helping society and increasing productivity. But the key question, when you talk to hardware experts, is at what point you have the right level of hardware to go for the scaling option. Effectively, when you start building factories around trying to make tens of thousands or hundreds of thousands of a particular robot type, it’s harder for you to update and quickly iterate on the robot design.

It’s one of those questions where, if you call it too early, the next generation of robots might be invented in 6 months’ time that’s just more reliable, better, and more dextrous.

Speaker 1

Sounds like, using a computing analogy, we’re kind of in the ’70s-era PC DOS kind of era.

Demis Hassabis

Yeah, potentially. But of course, I think the exception is that 10 years happens in 1 year, probably. So we’re in one of those years, right?

Speaker 1

Exactly. Let’s talk about other applications, particularly in science. True to your heart as a scientist—as a Nobel Prize-winning scientist—I always felt like the greatest things we would be able to do with AI would be the problems that are intractable to humans with our current technology, capabilities, brains, and whatnot, and we can unlock all of this potential. What are the areas of science and breakthroughs in science that you’re most excited about, and what kinds of models do we use to get there?

Demis Hassabis

AI to accelerate scientific discovery and help with things like human health is the reason I’ve spent my whole career on AI. I think it’s the most important thing we can do with AI, and I feel like if we build AGI in the right way, it will be the ultimate tool for science.

I think we’ve been showing at DeepMind a lot of the way forward with that—obviously AlphaFold most famously, but we’ve also applied our AI systems to many branches of science, whether it’s materials design, helping with controlling plasma in fusion reactors, predicting the weather, or solving Math Olympiad problems. The same types of systems, with some extra fine-tuning, can basically solve a lot of these complex problems.

I think we’re just scratching the surface of what AI will be able to do, and there are some things that are missing. AI today, I would say, doesn’t have true creativity in the sense that it can’t come up with a new conjecture or hypothesis yet. It can maybe prove something that you give it, but it’s not able to come up with a new idea or new theory itself. I think that would be one of the tests for AGI.

Speaker 1

What is that creativity as a human?

Demis Hassabis

Yeah.

Speaker 1

What is creativity?

Demis Hassabis

I think it’s these intuitive leaps that we often celebrate in the best scientists in history, and in artists, of course. Maybe it’s done through analogy or analogical reasoning. There are many theories in psychology and neuroscience as to how we as human scientists do it.

A good test for it would be something like giving one of these modern AI systems a knowledge cutoff of 1901 and seeing if it can come up with special relativity, like Einstein did in 1905. If it’s able to do that, then I think we’re onto something really important, where perhaps we’re nearing AGI.

Another example would be with our AlphaGo program that beat the world champion at Go. Not only did it win, back 10 years ago, it invented new strategies that had never been seen before for the game of Go—famously, move 37 in game 2, which is now studied. But can an AI system come up with a game as elegant, as satisfying, and as aesthetically beautiful as Go, not just a new strategy? The answer to those things at the moment is no.

That’s one of the things I think is missing from a true general system, an AGI system: it should be able to do those kinds of things as well.

Speaker 1

Can you break down what’s missing, and maybe relate it to the point of view shared by Dario Amodei and others about AGI being a few years away? Do you not subscribe to that belief? In your understanding of the structure, in your understanding of the system architecture, what’s lacking?

Demis Hassabis

I think the fundamental aspect of this is: Can we mimic these intuitive leaps rather than incremental advances that the best human scientists seem to be able to make? I always say that what separates a great scientist from a good scientist is that they’re both technically very capable, of course, but the great scientist is more creative. Maybe they’ll spot some pattern from another subject area that can have an analogy or some sort of pattern matching to the area they’re trying to solve.

I think one day AI will be able to do this, but it doesn’t have the reasoning capabilities and some of the thinking capabilities that are going to be needed to make that kind of breakthrough. I also think that we’re lacking consistency. You often hear some of our competitors talk about how these modern systems that we have today are PhD intelligences.

I think that's nonsense. They're not PhD intelligences. They have some capabilities that are at a PhD level, but they're not generally capable—and that's exactly what general intelligence should be—of performing across the board at the PhD level.

In fact, as we all know from interacting with today's chatbots, if you pose a question in a certain way, they can make simple mistakes with even high-school math and simple counting. That shouldn't be possible for a true AGI system. So I think we're maybe, I would say, 5 to 10 years away from having an AGI system that's capable of doing those things.

Another thing that's missing is continual learning: the ability to teach the system something new online or adjust its behavior in some way. A lot of these core capabilities are still missing. Maybe scaling will get us there, but if I was to bet, I think there are probably 1 or 2 missing breakthroughs that are still required and will come over the next 5 or so years.

Speaker 1

In the meantime, some of the reports and scoring systems that are used seem to be demonstrating 2 things. One, perhaps—and tell me if we're wrong on this—is a convergence of performance among large language models. And number 2, perhaps, is a slowing down or flatlining of improvements in performance with each generation. Are those 2 statements generally true, or not so much?

Demis Hassabis

No, I mean, we're not seeing that internally, and we're still seeing a huge rate of progress. But we're also looking at things more broadly. You see, with our Genie models and Veo models, Nano Banana is insane.

Speaker 1

Has anyone here used it? Can I see who's used it? Has anyone used Nano Banana?

Demis Hassabis

It's incredible, right? I'm a nerd who used to use Adobe Photoshop as a kid, and Kai's Power Tools. I was telling you about Bryce 3D. The graphic systems and recognizing what was going on there was just mind-blowing.

I think that's the future of a lot of these creative tools. You're just going to vibe with them or talk to them, and they'll be consistent enough. With Nano Banana, what's amazing about it is that it's an image generator—it's state-of-the-art and best-in-class—but one of the things that makes it so great is its consistency. It's able to follow instructions about what you want changed and keep everything else the same.

You can iterate with it and eventually get the kind of output that you want. I think that's what the future of a lot of these creative tools is going to be, and it signals the direction we're heading in. People love it, and they love creating with it.

Speaker 1

So, democratization of creativity, I think, is really powerful. I remember having to buy books on Adobe Photoshop as a kid, and then you'd read them to learn how to remove something from an image, how to fill it in, how to feather it, and all this stuff. Now anyone can do it with Nano Banana. They can just explain to the software what they want it to do, and it does it.

Demis Hassabis

I think you're going to see 2 things. One is this democratization of these tools, allowing everybody to use and create with them without having to learn incredibly complex UXs and UIs, like we had to do in the past.

On the other hand, we're also collaborating with filmmakers, top creators, and artists. They're helping us design what these new tools should be and what features they would want. People like the director Darren Aronofsky, who's a good friend of mine and an amazing director, have been making films with their teams using Veo and some of our other tools. We're learning a lot by observing and collaborating with them.

What we find is that it also superpowers and turbocharges the best professionals. The best professional creatives are suddenly able to be 10 or 100 times more productive. They can try out all sorts of ideas they have in mind at very low cost and then get to the beautiful thing they wanted.

I actually think both things are true. We're democratizing creativity for everyday use—for YouTube creators and so on—but at the high end, the people who understand these tools can get more out of them. There's a skill in that, as well as the vision, storytelling, and narrative style of the top creatives.

These tools allow them to iterate much faster, and they really enjoy using them.

Speaker 1

Do we get to a world where each individual describes what sort of content they're interested in? You say, "Play me music like Dave Matthews," and it'll play some new track.

Demis Hassabis

Yes.

Speaker 1

Or, "I want to play a video game set in the movie Braveheart, and I want to be in that movie."

Demis Hassabis

Yes.

Speaker 1

And I just have that experience. Do we end up there, or do we still have a one-to-many creative process in society? How important is this culturally? I know this is a little bit philosophical, but it's interesting to me. Are we still going to have storytelling where we have 1 story that we all share because someone made it, or are we each going to start to develop and pull on our own kind of virtual worlds?

Demis Hassabis

I actually foresee a world—and I think about this a lot, having started in the games industry as a game designer and programmer in the '90s—where this is the beginning of the future of entertainment. Maybe it will be some new genre or new art form involving a bit of co-creation.

I still think you'll have the top creative visionaries creating these compelling experiences and dynamic storylines. They'll be of higher quality, even if they're using the same tools, than what the everyday person can create. Millions of people will potentially dive into those worlds, but maybe they'll also be able to co-create certain parts of them. Perhaps the main creative person is almost an editor of that world.

Those are the kinds of things I'm foreseeing in the next few years, and I'd actually like to explore them ourselves with technologies like Genie.

Speaker 1

Right. Incredible. How are you spending your time? Maybe you can describe Isomorphic Labs. What is Isomorphic Labs, and are you spending a lot of your time there?

Demis Hassabis

I am. I also run Isomorphic Labs, which is our spinout company to revolutionize drug discovery, building on our AlphaFold breakthrough in protein folding.

Of course, knowing the structure of a protein is only 1 step in the drug discovery process. You can think of Isomorphic Labs as building many adjacent AlphaFolds to help with things like designing chemical compounds that don't have any side effects but bind to the right place on the protein.

I think we could reduce drug discovery from taking years—sometimes a decade—down to maybe weeks or even days over the next 10 years.

Speaker 1

It's incredible. Do you think that's in the clinic soon, or is that still in the discovery phase?

Demis Hassabis

We're building up the platform right now. We have great partnerships with Eli Lilly—I think you had the CEO speaking earlier—and Novartis, which are fantastic, as well as our own internal drug programs. I think we'll be entering the preclinical phase sometime next year. Then candidates get handed over to the pharmaceutical company, and they take them forward.

We're working on cancers, immunology, and oncology, and we're working with places like MD Anderson.

Speaker 1

How much of this requires—and I just want to go back to your point about AGI as it relates to what you just said—models can be probabilistic or deterministic. Tell me if I'm reducing this down too simplistically: the model takes an input and outputs something very specific, like it has a logical algorithm and outputs the same thing every time. It could be probabilistic, where it can change things and make selections: "The probability is 80%, I'll select this letter; 90%, I'll select this letter," and so on.

How much do we have to develop deterministic models that sync up with, for example, the physics or chemistry underlying the molecular interactions as you do your drug discovery modeling? How much are you building novel deterministic models that work with models that are probabilistic and trained on data?

Demis Hassabis

It's a great question. For the moment, and probably for the next 5 years or so, we're building what you might call hybrid models.

AlphaFold itself is a hybrid model. You have the learning component—the probabilistic component you're talking about, based on neural networks, transformers, and things like that—and that's learning from the data you give it, any data you have available. But in a lot of cases with biology and chemistry, there isn't enough data to learn from, so you also have to build in some of the rules about chemistry and physics that you already know about.

For example, with AlphaFold, you have the angles of bonds between atoms. You need to make sure that AlphaFold understands that you can't have atoms overlapping with each other and things like that. In theory, it could learn that, but it would waste a lot of its learning capacity. So it's better to have that as a constraint in the system.

The trick with all hybrid systems is how to marry a learning system with a more handcrafted, bespoke system and actually have them work well together. AlphaGo was another hybrid system, where a neural network was learning about the game of Go and what kinds of patterns were good, and then we had Monte Carlo search on top, which was doing the planning.

That's pretty tricky to do.

Speaker 1

Does that sort of architecture ultimately lead to the breakthroughs needed for AGI, do you think? Are there deterministic components that need to be solved?

Demis Hassabis

I think ultimately what you want to do is, when you figure out something with one of these hybrid systems, upstream it into the learning component. It's always better if you can do end-to-end learning and directly predict the thing that you're after from the data that you're given.

Once you've figured out something using one of these hybrid systems, you then try to go back and reverse-engineer what you've done and see if you can incorporate that information into the learning system. This is sort of what we did with AlphaZero, the more general form of AlphaGo.

AlphaGo had some Go-specific knowledge in it. But then with AlphaZero, we got rid of that, including the human data and human games that we learned from, and actually just did self-learning from scratch. Of course, then it was able to learn any game, not just Go.

Speaker 1

A lot of hype and hoopla has been made about the demand for energy arising from AI. This was a big part of the AI summit we held in Washington, D.C., a few weeks ago, and it seems to be the number one topic everyone talks about in tech nowadays. Where's all this power going to come from?

But I ask the question of you: Are there changes in the architecture of the models or the hardware, or the relationship between the models and the hardware, that bring down the energy per token of output or the cost per token of output? Could that ultimately mute the energy demand curve that's in front of us, or do you not think that's the case and we're still going to have a pretty geometric energy demand curve?

Demis Hassabis

Well, interestingly, I think both cases are true. Especially at Google and DeepMind, we focus a lot on very efficient models that are powerful because we have our own internal use cases, of course, where we need to serve, say, AI Overviews to billions of users every day. It has to be extremely efficient, extremely low latency, and very cheap to serve.

We've pioneered many techniques that allow us to do that, like distillation, where you have a bigger model internally that trains the smaller model. You train the smaller model to mimic the bigger model. Over time, if you look at the progress of the last 2 years, model efficiencies are 10 times, even 100 times, better for the same performance.

The reason that that isn't reducing demand is because we still haven't got to AGI yet. The frontier models keep wanting to train and experiment with new ideas at larger and larger scale, while at the same time, on the serving side, things are getting more and more efficient. Both things are true.

In the end, from the energy perspective, I think AI systems will give back a lot more to energy and climate change than they take, in terms of the efficiency of grid systems and electrical systems, material design, new types of properties, and new energy sources. I think AI will help with all of that over the next 10 years, and that will far outweigh the energy that it uses today.

Speaker 1

As the last question, describe the world 10 years from now.

Demis Hassabis

Wow. Okay. Well, 10 years—even 10 weeks—is a lifetime in AI. But I do feel like if we will have AGI in the next 10 years—full AGI—I think that will usher in a new golden era of science, a kind of new Renaissance. I think we'll see the benefits of that right across the board, from energy to human health.

Speaker 1

Amazing. Please join me in thanking Nobel laureate Demis Hassabis. Thank you. That was great. Thank you.