关于 Gemini 3,所有人都漏看了什么|Salim、Dave 与 Alexander Wissner-Gross 对谈 | EP#209
Peter Diamandis × Salim Ismail × Dave Blundin × Alexander Wissner-Gross
Gemini 3 让 Google 成为本期讨论中明确的前沿领先者,预测市场认为它以第一名结束今年的概率为91%,但到明年夏天仍居首的概率只有60%。 Alexander Wissner-Gross 称这是继4月 OpenAI 发布 o3 后最大的一次更新:覆盖面广、具备多模态能力,而且显然没有过拟合那些有利于宣传的基准测试。Peter Diamandis 更尖锐的判断是,自然语言软件意味着“从今天开始,世界已经与我们昨天生活的那个世界不同”。
代理正从助手跨入经济行为者阶段,而 Gemini 3 在模拟的 Vending-Bench 经济中取得近3,000%的利润优势,为这一转变给出了可量化结果。 该基准给每个代理 $500、运营工具和破产约束;Wissner-Gross 称,在那里取得成功“已经走到自主经营现实世界企业的一半”。Salim Ismail 的推论是:1年前讨论的3人初创公司,正变成零员工公司。
基准测试的饱和如今指向聊天机器人质量之外的数学、科学、工程和医学硬核研究。 根据讨论,Gemini 3 在 Humanity’s Last Exam 上接近50%,在 ARC-AGI-2 上大约是 GPT-5.1 的2倍,在 Humanity’s Last Exam 上实际上是 Claude 4.5 的2倍。Wissner-Gross 表示,如果到明年年底硬核研究问题还没有被这类模型攻克,他会“非常惊讶”,但同时对持续学习和超长上下文保留了限定。
Google 的护城河在于分发加整合,但真正解锁它的不是既有地位,而是竞争压力。 Gemini 可以跨 Google Workspace 执行任务、生成界面、致电商店并撮合交易;语音能力的提升也让 Diamandis 提到 Duolingo 较前一年下跌近50%。Blundin 的反向判断是,Google 内部早已拥有底层技术,只是 OpenAI 迫使它行动:“这就是 Google 会动起来的唯一原因。”
当超大规模云厂商可以复刻界面并向下延伸技术栈时,应用层估值仍然暴露在风险之中。 Cursor 在6个月内从约 $10B 上涨到 $30B,并融资 $2.3B;但 Blundin 表示,除模型接入外,Google 的 Antigravity 看起来“和 Cursor 一模一样”。他拒绝预测赢家:Cursor 的团队、资本和多模型架构都很强,但“它的核心定位极其脆弱”。
AI 的资本开支账单很可能主要由企业承担,因为企业会把昂贵的推理算力配置到高价值工作上。 GPT-5.1 的路由机制会给困难提示词分配更多算力,预示着这样一个市场:如果 AI 能让商品规划利润率提高20%,零售商可能会很乐意为它付费。Blundin 还指出一个罕见的定价优势:AI“把销售员直接内置进了自己的能力”,能够展示高端体验,然后告诉用户升级。
更便宜的智能可以降低生活成本,但丰裕不会自动到来:部署、监管、安全和社会分配将成为瓶颈。 Wissner-Gross 认为上游指标是每单位智能的美元成本,目前正“以同比约40倍的速度超通缩”;Ismail 则提出,抑郁症发生率可能比物质产出更早成为诚实的进步指标。生物安全暴露了其中的权衡:防御性 AI 可能与攻击能力同步扩张,但代价或许是无处不在的传感系统,以及一个类似“全球机场”的世界。
1. Gemini 3 让奇点显得出奇地平常
Peter Diamandis 开场直说:“Google 赢了”(“Google is winning”)。Gemini 3 只用了11个月就接替 Gemini 2,发布节奏快到足以让真正的不连续进步,听起来像“把模型名字填进去、再把编号填进去”的例行发布。
Diamandis 将今天的自然语言开发,与从 COBOL、0和1、十六进制一路发展到高级语言的大约40年编程史并置;Dave Blundin 补充说,人类最初甚至从汇编语言起步。直到现在,软件与机器之间的接口始终是代码;如今直接与机器对话,可能把这条边界从软件延伸到基因测序、白领自动化和机器人工业设计。
Wissner-Gross 的框架是:奇点可能只是“一种视觉错觉”,因为身处其中时,“时空感觉是平的”,每周发生的突破都会显得平淡。他将 Gemini 3 排在 o3 之后最大的发布,并认为 GPT-5 主要只是重新包装了 OpenAI 在4月已经交付的能力跃升。
2. Google 的存量用户如今拥有通用代理
Google 的演示让 Gemini 从回答请求转向执行请求:规划旅行、研究产品、执行多步骤操作并调用工具。生成式 UI 意味着答案可以变成带有图片、模拟和交互组件的定制界面,而不再只是“一堵文字墙”。
Wissner-Gross 认为 Gemini 与 Gmail、Calendar、YouTube 及整个 Google 环境的整合几乎无缝,但称这“可能是最不值得关注的部分”。更大的后果是,数十亿现有用户如今都有了“随叫随到的超级智能”。
他判断“大模型气味”的测试,是看能否跨模态生成:Gemini 只根据一张 MIT 的照片,就一次性生成了一个可交互的3D体素风校园。Antigravity 则提供了对应的软件开发界面——这是 Google 基于 Visual Studio Code 构建、由前 Windsurf 人才参与打造的代理式开发环境。
3. Vending-Bench 将自主经营变成量化测试
Andon Labs 的 Vending-Bench Arena 给代理一个模拟的 $500、电子邮件、互联网搜索、银行账户、库存和定价控制权。若连续10天未能支付每天 $2 的费用,代理就会破产;目标是最大化资本回报。
据称,Gemini 3 创造的利润比 GPT-5 或 Claude Sonnet 高出近3,000%。Wissner-Gross 喜欢这项基准,因为它把 AI 建模为“第一类经济行为者”,实际上是在执行中层管理者的工作,而不是回答孤立问题。
Diamandis 反驳称,模拟忽略了“员工的混乱”。Wissner-Gross 的回应是,自然语言交互的对手已经迫使代理与供应商谈判;加入绩效评估或员工互动,在技术上也不会难多少。
Blundin 估计,仅互联网广告就是一个规模达 $300B、且大部分由非人类完成的业务,并表示如果自动化经济规模不到 $1T,他会感到意外。Diamandis 没有解决的分配问题是:是否有人可以给代理 $10,000 稳定币并说“去帮我赚更多钱”,还是说这种接入只会扩大财富差距。
4. 一次生成正在瓦解软件技能门槛
Wissner-Gross 只给 Gemini 一条提示:“创建一个视觉效果惊艳、我可以直接玩的赛博朋克 FPS。音乐要好听,画面要丰富。”可玩的结果不到5分钟就完成,提示词大约140个字符——“这是我见过能力最强的一次性生成”。
如果写一条简短社交媒体帖子就足以生成一款游戏,Wissner-Gross 预计1年内会出现数十亿款游戏。Diamandis 对孩子的建议也随之改变:不要只是玩游戏,而是要设计和构建游戏,再修改生成出来的作品。
Blundin 将这一机会比作拍照手机让600万美国人得以全职从事网红工作。过去需要制作团队、摄像机和技术训练的生产能力,如今通过思想和语音即可获得,让那些前一天还不会写代码的人也能拥有职业机会。
5. 语音与现实世界交易强化 Google 的分发优势
Diamandis 说,过去 Gemini 的语音与 GPT-5 的 Ember 语音相比显得机械,但如今 Google 在自然度上已经领先。他将实时语言辅助与 Duolingo 面临的压力联系起来,并称 Duolingo 过去1年下跌了近50%。
Blundin 将功劳归于 OpenAI 带来的竞争压力:Google 开发了包括 Transformer 在内的大量底层技术,但一直不愿公开部署。“作为 CEO,说‘各位,赶紧动起来,这里有威胁’要容易得多。”
Wissner-Gross 认为,更丰富的口音和语音互动与其说是新能力,不如说是“解除束缚”(“unhobbling”)。从音频转文字再转音频的管线,转向直接的音频到音频模型,将解锁更细腻的对话;再结合 AR 眼镜,讨论者预计同声翻译将重塑国际交流。
Google 的购物代理可以致电附近商店询问库存和价格,这距离 Duplex 在2018年首次亮相已经过去7年。Blundin 希望保留人类对产品的主观判断;Ismail 则反驳说,AI 会知道得更多,还能即时展示图片。Wissner-Gross 的总结是:当正式整合不存在时,语音提供了一个出口——“一切都能有 API”(“APIs for everything”)。
6. 基准测试如今衡量的是距离经济有用研究还有多远
Wissner-Gross 为基准测试辩护,称其是文明进步的量化仪表盘。Humanity’s Last Exam 近似测试博士级问题解决能力,ARC-AGI-2 测试类似人类的视觉推理;它们的价值不在排行榜本身,而在于测试饱和意味着哪些现实问题已经变得可处理。
Gemini 3 在 Humanity’s Last Exam 上接近50%,在 ARC-AGI-2 上大约是 GPT-5.1 的2倍,在 Humanity’s Last Exam 上实际上是 Claude 4.5 的2倍。Wissner-Gross 表示,它看起来并不是狭窄的“刷榜”,而是一个均衡的通才模型,而非专门优化宣传友好的分数。
他的条件式预测异常具体:按照当前轨迹,如果到明年年底数学、科学、工程和医学领域的硬核研究问题还没有被 Gemini 这类模型攻克,他会“非常惊讶”。
限定出现在超长上下文下的持续学习能力上,Wissner-Gross 不会自动期待一次戏剧性跃升。检索能力则更强:Gemini 3 Pro 在测试大型上下文中隐藏事实的“大海捞针”任务上表现“惊人地好”。
7. Scaling 仍是反驳“需要新 AGI 范式”的核心依据
Ismail 问,在这种规模下保持连贯,是否意味着系统思维和世界模型已经出现。Wissner-Gross 给出了绝对判断:能够解决数学问题、编写源代码,并跨学科处理博士级工作的模型,已经在进行符号推理并建模系统;等待某种模糊的神经符号突破是“彻头彻尾的胡说”。
Diamandis 将 Gemini 3 描述为“7万亿参数级模型”,而前一年约为1万亿参数,并预计原始算力还会再增加10倍至40倍。他要求技术和医疗领域的领导者明确判断:基准测试何时足以允许模型自我改进,并可靠地治愈疾病。
Diamandis 提到 Yann LeCun 认为 LLM 是通往 AGI 的错误分支。Wissner-Gross 尊重另一条强调在具身空间中采取行动的路径,但认为可行道路不止一条:只要 scaling law 和能力仍在持续提升,就无需新范式,“也许我们真的可以继续 scaling”。
8. GPT-5.1 预示着按任务价值为智能定价的市场
Wissner-Gross 认为,GPT-5.1 的路由经济学比技术本身的变化更重要:困难提示词获得更多推理时算力,简单提示词获得更少。他将其与搜索广告相比较:间皮瘤诉讼搜索的价值,远高于一道算术题。
这种分配暗示了谁会为数万亿美元的数据中心资本开支买单。最典型的付款方可能是把大量资金投入高价值工作的企业——Blundin 举的例子是,如果更好的商品规划能让利润率提高20%,Target 就会购买算力——而不是每个消费者每月支付数百美元。
Blundin 不接受通过降低半个模型档位来节省成本,因为前沿模型的表现“就是不一样”。AI 也是历史上最奇怪的产品发布:它一边销售自己,一边与用户对话,提供令人信服的体验,还会亲自解释为什么更多能力需要升级。
9. 防御性协同扩展是应对 AI 生物武器的方案
OpenAI 向 Red Queen Bio 投资了 $15M,该公司结合 AI 与实验室测试来识别生物脆弱点。Wissner-Gross 借用红皇后的赛跑——只能通过奔跑来维持原地不动——解释“防御性协同扩展”:安全能力必须与模型能力同步提升。
Ismail 将2024年约 $34B 的生物防御市场,与预计在2034年或2035年翻倍的规模,以及一次可能只需约 $1,000 就能发起的数万亿美元级极端攻击进行对比。Diamandis 提议在机场和交通系统部署传感器,在本地测序空气中的生物因子,并以光速分发对策,而病原体只能以飞机速度传播。
Diamandis 认为,生物武器风险“正是美国开源 AI 已死的原因”,并称美国实验室不再发布前沿模型,而中国实验室仍在发布。他警告,本地模型可以绕过按查询触发的拒答系统,让一个原本无能为力的攻击者拥有“天才级 AI 作为副手”。
Wissner-Gross 将 Linus 定律推广为:“拥有足够的超级智能后,所有隐藏代理都会变得浅显。”代价是监控;Ismail 想象人类生活在“全球机场”里,而 Diamandis 保留了 EFF 的反驳:安全并不总要以牺牲隐私为代价。Wissner-Gross 提醒,不应过度聚焦风险,因为防御性 AI 也会同步扩展。
10. Cursor 的增长惊人,但在战略上极其脆弱
Grok 4.1 在文字排行榜上的第一名大约只维持了1周,随后领先优势就消失了。对 Wissner-Gross 来说,这段短暂统治说明,通用模型可以每周、甚至随着前沿加速而每天抹平针对基准测试优化的领先。
Cursor 的估值在6月至11月间从约 $10B 增至 $30B,翻了3倍,同时融资 $2.3B。它让编程变得代理化且易于使用,并将工作路由给 Claude 4.5、OpenAI 和 Grok 等外部模型。
Blundin 给出了一个令人不安的视觉比较:把 Cursor 和 Antigravity 并排打开,“你甚至不知道自己在哪个里面”。Cursor 提供多种模型,Antigravity 则绑定 Gemini 3;但 Diamandis 拒绝选择赢家,因为 Cursor 既有天才般的团队又资本充足,同时在结构上暴露无遗。
Wissner-Gross 预计,开发环境会引入自有模型,并掌握更大部分供应链。软件工程是第一个被自动化的主要高生产率劳动类别;雇主已经在非正式地区分两类工程师:一类在代理式编程出现前受训,另一类可能已经因代理而“萎缩”。
11. Project Prometheus 将资本从超级智能转向产业
据报道,Jeff Bezos 以 $6.2B 启动 Project Prometheus,并从 OpenAI、Google、Meta 等公司招募了近100名研究人员。Diamandis 强调,一家初创公司在第0天资产负债表上就拥有数十亿美元,这种情况既前所未有,也令人畏惧。
Wissner-Gross 认为,资本市场正从资助超级智能,转向资助超级智能之后的事情:解决数学、科学、工程和医学中的未决问题。他估计,这一机会是超级智能本身市场的10倍至100倍,并预计将有数万亿美元资金流入;招聘动向显示,生物学获得的重视可能高于报道所暗示的程度。
讨论者将 Prometheus 描述为 AI 从办公室走向工厂、物流、实体测试,最终进入太空自主产业的过程。Wissner-Gross 说,他职业早期构建的一个基础模型耗时5年,如今2个月就能重建;再过1年,这一时间还可能进一步缩短5倍至10倍,因为“你可以用 AI 构建下一个 AI”。
12. 丰裕取决于抬升底部,而不是缩小差距
Elon Musk 在录音中的判断是,AI 和人形机器人“基本上是让所有人变得富裕的唯一途径”。讨论者重新定义了成功:即使超级富豪仍然存在,真正重要的里程碑也是每个孩子能否获得充足的食物、水、能源、医疗和教育。
Diamandis 以2004年海啸后的印度尼西亚为例说明信息如何实现民主化。渔民获得手机并报告危险后,只用2个月收入就增加了30%,因为他们可以查看鱼价,并选择何时、何地出售;Ismail 补充说,他们还能识别出哪个港口给价更高。Diamandis 认为,AI 可以把这种信息优势延伸到个性化教育和早期医学诊断,而早期发现可能让治疗成本降低约100倍。
Diamandis 将美国平均家庭每年 $77,000 的预算拆分为:住房33%、交通17%、食品13%、保险和养老金12%、医疗8%、娱乐5%,教育约2.5%。他将远程办公、打印式住房、自动驾驶电动交通和 AI 医疗分别对应到这些支出上;他还表示,垂直农场可以实现7倍产量,同时节省99%的淡水。
Ismail 提议用抑郁症发生率,而不是单纯的物质丰裕,作为早期有意义的基准。Wissner-Gross 将上游变量定义为每单位智能的美元成本,并称其每年下降约40倍;剩余约束则是监管、社会凝聚力,以及能够分配“便宜到无需计量的智能”的安全网。
People who are already in the ecosystem now have a superintelligence at their beck and call. That’s probably the least interesting thing. When they’re on the cusp of the singularity, they’ll start soft-selling it. Gemini 3, which has just today climbed all the way up the third-party AI rankings. Let’s break down what this so-called Gemini leap means. This will change the game completely for everything, everywhere. Why is this not just another little faster, little better capability? We have a way of measuring progress in our civilization. AI is imminently, I think, well-positioned now that these benchmarks are saturating to start solving the hardest problems on Earth in math, science, engineering, and medicine. All of a sudden, you can build software by talking to the machine. This is like a different world starting today from the world we lived in yesterday. Now, that’s a moonshot, ladies and gentlemen.
You know, the hardest thing for me when I’m going over the slides is what to cut out. It’s all so good, right? Every one of them could be an entire hour-long conversation. The question of how we group it and how we actually make it a fun conversation is so challenging. There’s so much going on. We almost have an episode on robotics, an episode on energy, and an episode on AI.
Yeah, but then if we do that, we’re publishing more than once a week, which is a lot. Sometimes we do, but if you’ve gone 3 weeks without covering one of the fields, it’s disruptive shock therapy.
The world is over.
Well, the audience has a limited amount of time, too, so we’ve got to try to help them as much as possible.
Yeah. Twice a week, basically. That’s all you can do.
I hope you guys have as much fun as I do on this.
Oh, yes.
It’s awesome scanning all the breakthroughs and looking at how fast it’s all moving. It’s really about trying to figure out what this really means. Besides yet another benchmark, or yet another number that’s greater than another number, what does it mean for everybody?
For me, I get so buried in the day-to-day. There’s just so much going on, and if it weren’t for the podcast pulling me out of the weeds, I would miss all kinds of things. I get really frustrated when people don’t know what’s going on and aren’t reacting to it. I’m like, “The only reason I know what’s going on is because we do the podcast, and that prep time pulls me up out of the weeds.” So I love this time.
For sure. I told my son, “Hey, Gemini 3 is out, and it has amazing benchmarks.” He goes, “Yeah, insert name of model here, insert number here. Every week you tell me that.” I said, “Yeah, you’re right.”
Yeah.
We’re the antidote for that because, as we’re always saying, people get inured to things so quickly and miss the implications. That’s true even at MIT, where I’ve been for the last 3 days. But this is just not true. This is step-function, life-changing stuff week by week.
Yeah.
I think we should just jump in because there’s a lot. If you guys are ready, let’s go. I’m here with DB2, AWG, and Mr. ExO—your call signs—and let’s jump in. So—airports. They’re all three-letter airport signifiers. Okay, let’s get going here. So welcome to Moonshots, everybody. This is another episode of WTF Just Happened in Tech.
The real news, and for us the only news, is the implications: What does it mean, and what does it mean for you, your family, your business, your company, and your country? We’re going to open up with the hyperscalers: Google, xAI, and OpenAI. The TL;DR for this episode is that Google is winning. There’s a lot going on at Google. We just saw the release of Gemini 3 yesterday, which is why we’re recording today—trying to be right here, right now.
I’m going to share a video from Josh Woodward. Josh is a friend; I had him on the Abundance stage a year ago. He now heads Gemini and Google Labs. He’s a brilliant presenter, and we’re going to have him on this podcast in the new year. I’m excited for that.
Hey, everyone. My name is Josh, and I lead the Gemini app, Google Labs, and AI Studio. Today is the day: Gemini 3 is here, and it’s in the app. It’s our smartest model ever, and you can try it right now.
We have this new feature called Agent, and you can go into Gemini, describe a task, and it’ll get to work for you. You can plan a trip, research products, and do all these things. It acts on your behalf, takes multistep actions, makes tool calls—all of it.
The other thing I’m really excited about is that we’re entering a new era where you can create UI dynamically. The model creates these generative UIs, so when you ask a question, Gemini won’t just respond with a wall of text. It’ll pull in images and different interactive widgets, giving you a much more customized experience based on what you’re looking for. All of this gives you a more helpful response.
One more video here from Gemini, and then we’ll discuss it. This is their official “Introducing Gemini” video. Again, congratulations to Josh for taking the lead there and crushing it. We’ll talk about the benchmarks with AWG in a little bit, but before then—
Gemini 3 is the strongest model in the world for multimodality and reasoning. It’s our most intelligent model, helping you bring any idea to life.
In Google Search, Gemini 3 enables new kinds of generative user interfaces. It codes interactive simulations like this one, custom-built for your search.
In the Gemini app, you can supercharge how you learn, create, plan, take action, analyze complex videos, and more. We’re even introducing a new platform, Google Antigravity. It’s our vision of software development at the frontier of model intelligence. It lets you use Gemini 3’s agentic coding capabilities to accelerate how you build. This is just the beginning of our Gemini 3 series.
Okay, who wants to dive in first? Dave, do you want to jump in? What does this mean to you? Why is this not just another little faster, little better capability?
Dave is a full kid in a candy store here. This is great. I can’t wait to hear Alex’s take on this, too.
It’s at almost 50% on Humanity’s Last Exam. It’s such a step-function change in history. I was over at MIT last night talking to a bunch of undergrads, and I was trying to tell them, “Look, you don’t know this, but 40 years ago we started writing code as a species. We started with COBOL, and we started with ones and zeros and hexadecimals.”
That’s true. We started with assembly.
I swear to God, if you look at what happens today when you write code versus 40 years ago, it’s identical. It’s like a higher-level language; nothing’s really changed. All of a sudden, you can build software by talking to the machine. It is such a different world starting today and moving forward.
I’m hoping they can then generalize and say, “Well, it’s coding today, gene sequencing tomorrow, and all white-collar automation the day after that. Then all industrial design of robotics is done by voice.” This is like a different world starting today from the world we lived in yesterday, and it’s really hard to get people to fully understand the implications. I can’t tell you how big this is.
Alex, what’s your takeaway, buddy?
I’ve said in the past here that I think the singularity is probably an optical illusion. When you’re in the midst of it, space-time feels flat. Every time I hear the question, “What else is new? The benchmarks are going up and to the right, but it doesn’t feel really transformative,” that, to me, is a sign that when you’re in the midst of a singularity, space-time feels flat and breakthroughs that are happening essentially every week or every day feel prosaic.
There are so many transformative aspects of Gemini 3. Just walking through those 2 videos, starting from perhaps the least transformative aspects, the Gemini app itself—which is how many people are likely to first encounter Gemini 3—is now integrated with all of the other Google properties.
There’s been a lot of bellyaching over the past year: Why can’t I agentically have Gemini write my Gmail for me, organize my calendar, or interact with YouTube? I’ve been playing with Gemini Agent, the Agent mode part of Gemini 3, and that’s seamless at this point. It’s literally a single click to get Gemini 3 to order your entire Google-platform-based existence, or Google Workspace-based existence.
That’s probably the least interesting thing.
But it’s a powerful driver for people to switch to Google as an all-in platform, right? That’s really the situation they’re striving for.
Google has billions of users across all of its products already, so I’m not sure that, at the margin, the greatest impact on humanity is getting people to switch to Google. I think it’s more that people who are already in the ecosystem now have a superintelligence at their beck and call.
And again, that's the least interesting thing. A couple of more interesting things come from interacting with the model itself. Again, this isn't focusing yet on the benchmarks; this is just on interacting with the client. It smells. People in the community sometimes refer to something as “big-model smell”: a model that has certain types of capabilities that can't be arrived at through extended reasoning or through other, smaller-footprint attempts to extend a model.
Gemini 3 has what I think can be fairly termed big-model smell. You can ask it to do cross-modal or multimodal tasks that are very challenging to do elsewhere. One of my first tasks was to feed it a photo of the MIT campus and ask it to generate a 3D voxel, block-world-type rendering that I could interact with. In one shot—basically zero-shot—it produced an interactive 3D rendering of the MIT campus.
There's also—I don't want to let this point drop—Antigravity, the code-development environment, the integrated development environment focused on Gemini 3. My understanding is that many of the core members of the Windsurf team—we've talked about Windsurf in the past, a Cursor competitor—joined Google DeepMind and built Antigravity as a result. I was interacting with Antigravity, and it was a very impressive Visual Studio Code-derived experience for code development. So there are many pieces here, and that's before we get to the truly interesting stuff, in my mind, which is the benchmarks.
Yeah. You know, one of the things we said a while ago is that when they're on the cusp of the singularity, they'll start soft-selling it. You noticed Google put out all these mind-blowing benchmarks, and the only thing they put out in terms of content is that Josh Woodward clip from a second ago. Contrast that to the OpenAI GPT-5 release, which was a special hour-long presentation by Sam and so forth. This was, like you said, a very soft sell.
One thing I found fascinating is the speed at which we're upleveling the models. Gemini 2 was December of last year—11 months ago—and now we've got Gemini 3 coming out. We're seeing an increasing speed at which we're deploying them. We're seeing that across the board with the hyperscalers.
To comment narrowly on that from my perspective, Gemini 3 is the biggest model release since OpenAI's o3 in April—only 7-ish months ago. To the extent GPT-5 may have felt slightly underwhelming, I would argue it's because almost all of its raw capability jumps actually happened a bit before, in the form of o3.
Maybe think of GPT-5 as o3, which was actually o2 because o2 was trademarked, so it had to be called o3. GPT-5 was actually like o2.1. So I think we can't take credit away from OpenAI for the achievement that was o3, which was then partially repackaged as GPT-5.
For me, this is seeing Google go from a reactive assistant, where you're asking it for something, to an autonomous agent handling complex, real-world data. We're going to see that in the next slide. Let's go there. Gemini 3 delivers breakthrough profitability in an AI-run mini-economy. This is the Vending-Bench benchmark, which I love. Gemini 3 outperforms Grok, Claude, and ChatGPT in long-term business-management tasks. To explain to us what this means—the king of benchmarks—Alex, let's go to you.
I love benchmarks. I love this benchmark in particular. This is a benchmark, Vending-Bench Arena, maintained by a company named Andon Labs. It's derivative of another benchmark they maintain, named Vending-Bench 2.
The basic premise is that AI agents are given a simulated $500 to start. They're put in charge of a simulated vending machine, and they're given tools that they can manage. They have the ability to send and read emails—real, full natural-language emails. They're given the ability to search a simulated internet. They have a simulated bank balance. They can send money, receive money, stock and restock the vending machine, set prices, check inventory, collect cash, and so forth.
So this is really performing the role of almost a middle manager in charge of a vending machine. If the simulated agents maintaining the vending machine fail to pay a $2 daily fee for 10 consecutive days, they go bankrupt. The goal of the game is to maximize the return on investment for that initial simulated $500.
I think this is such a lovely, self-contained proxy for AI agents as first-class economic actors. If AIs can do a spectacular job of managing this pretty rich, simulated vending-machine world, then I think they're halfway to autonomously running their own real-world businesses and becoming AI entrepreneurs, at which point we get zero-human startups.
Wow. It's amazing, right? We talk about this: Gemini 3 is delivering almost 3,000% more profit than GPT-5 or Claude Sonnet. And you're right, we've talked about going after stablecoins and agents together, spinning up new businesses faster than you possibly can.
Now, the one thing this doesn't do is account for the messiness of employees. It would have to be a nonhuman business that it's running in order to really maximize profitability without dealing with employees.
Yeah. Go ahead.
I would actually argue that the email functionality built into the benchmark makes this less of a limitation. When it sends and receives emails, there's a large language model counterparty at the other end writing full natural-language emails. I could imagine a generalization—maybe a future version 3 or 4 of Vending-Bench—that takes into account performance reviews and interacting with employees. All of that, I think, is not technically that much more difficult to test.
If you can manage vendors and suppliers, then email communication with employees isn't that much harder.
Interesting, Dave.
Well, I was in the internet advertising business. It's $300 billion a year and completely nonhuman. The whole thing is automated bidding and automated placement. I'd be surprised if the nonhuman economy is anything less than $1 trillion already. The parts of the economy where you can just deploy this are going to grow very rapidly.
But did you notice how Alex has a lot more emotion in his voice as the AI is getting more sophisticated? [Laughter.] Is that an improvement in the algorithm, or is that just enthusiasm? His true identity is being revealed. [Laughter.]
If personhood is granted and I get to be a real person, a real boy as it were, then I get to run my own business too, I guess. [Laughter.]
Well, on this topic, I completely agree. We need many, many, many more benchmarks. The more real and practical they are, and the less technical they are, the more they open people's eyes to what's possible.
I think we desperately need more benchmarks in the medical area. Peter, you're the top guy on the planet in this, but we're getting so close—so close—to being able to first extend people's health spans, delay cancer, delay heart disease, and then cure them. If we do that quickly, I think we can save 30 million lives. There's 10 million a year.
This is very, very important to me personally, just because of some friends that I have in this situation. And I swear to God, this step-function improvement today puts that right in front of us.
Dave, imagine this in the future: instead of AI agents managing vending machines, you're going to be a part of a population and the agent's going to manage you. It'll be like, “Go outside, take a walk right now. Drink another glass of water. Go take these…”
That's the promise of the Jarvis thing you keep talking about here.
Yeah, it's coming, buddy. I know you've got a pesky leaf blower outside. I keep saying to Elon, “Would you please make electric leaf blowers? Just make them quieter.”
Nat Friedman has a $100,000 prize for anyone who can create a silent electric leaf-picking-up machine.
Oh, crazy, right? We're going to elevate it to an XPRIZE and put $10 million behind it. [Laughter.]
Let's do it. That's a great idea. I've got a couple of thoughts. One is that the entire stack of society can now be AI-mediated, which is kind of an incredible thing to be able to say. The second part of this is a really important point that Alex made: you can now build a company with literally zero employees.
We were talking about 3 employees a few months ago, Peter, and a year ago. Right now, it's down to 0. This is going to change the game and absolutely will happen. As Dave says, there's already a trillion-dollar-or-so economy out there, and this is going to get automated very quickly.
All right, so keep your eyes on this. As an entrepreneur, I think about this. When can I start spinning up companies? Can I give $10,000 in stablecoins to my AI agents and say, “Go make me some more money?”
And now the question is, is that available for everybody? Can anyone and everyone spin up an agent that is going out there and generating revenue for them? Because if it isn't, then we're beginning to have a widening wealth gap.
All right, let's go to our next story here. This is a story about a one-shot cyberpunk first-person shooter that I think you made, Alex.
That's right. I see the comments sometimes. I've remarked in the past that one of my favorite evals for a fresh model is to ask it to generate a cyberpunk first-person shooter, and some folks in the past have suggested that's nonsense.
So I thought it might be instructive, given the strength of Gemini 3, to ask it to one-shot the generation of a cyberpunk first-person shooter. The prompt that I gave it—the only prompt—was, “Create a visually stunning cyberpunk FPS that I can play. It should have nice music and rich visuals.”
Let's play the video. If you're watching on YouTube, enjoy this. If not, go to YouTube.
So, Neon Protocol. I do like the music. Actually, I immediately copied Alex's prompt and extended it, and my music came out absolutely nauseating. I said, “Make it even faster action and make it a deeper-pumping bass.”
And my version was just nauseating beyond the
Okay. Listen, this is what I keep telling my kids: instead of playing video games, at least design them and build them. This is just making it so much easier.
For everybody listening, you can do this. This is not something you have to have special access to. You can do exactly what Alex did in less than 5 minutes. So go ahead and try it, and then modify it. It's super fun. Also, Google has a limited amount of compute.
Everybody can do this for free, but after you hammer it for a few hours, it'll throttle you.
So take advantage of your first few free hours and have some serious fun and learn a lot. I was with Jack Hidary at FII, and one of the conversations I had with Jack—and I respect this very much—he says, “Instead of waking up in the morning and consuming, just scrolling through everything, get up in the morning and create something. Build something.”
So go on, Alex.
And to that point, it's never been easier. That was probably 140 characters or fewer. If you can post on X or post a short social media message, you can create a game on demand, which means that I think we should expect to see billions of games created in the next year because it's now so easy.
It's the most competent one-shotting I've ever seen.
Gaming slop.
Just to echo the conversation from last week with 140 characters and flying cars, it'll be amazing when the inner loop gets to a point where you can just use 140 characters to say, “Build me a flying car.” Correct? Yeah, it goes and does it.
You can do that right now. With 140 characters, you can create a simulated flying car with Gemini 3.
There are 6 million people in America whose full-time job is influencer.
And that was enabled by the camera phone. Prior to that, you needed a production crew and heavy cameras. You couldn't be an influencer. All of a sudden, because there's a 4K camera on every iPhone and there's great editing, 6 million people shift to influencer as a career. This is at least as big a shift.
If you say, “Video games are generic right now. Let me make something custom to my community, custom to people,” you can actually create it. Even if you couldn't code yesterday, today you can create something just using your thoughts and your voice. And so it opens up career opportunities.
Let's take a listen to this. This is the next article here: Is Gemini Live a more natural voice?
Is there any fish on this menu?
Yes, there's a sea bass.
Yeah, I love sea bass. Can you help me order that in Spanish?
Of course. Try, “Me gustaría.”
How's this? “Me gustaría la lubina, por favor.”
That sounds great.
Yeah. I think they made a nice move forward here. I used to love my GPT-5 voice. I use Ember when I'm talking to it, and Gemini felt stilted and not natural. So they really did a great job moving us forward. Super excited about that.
Interesting on the translation side. We talked in one of the previous pods about Duolingo being disrupted. Well, over the year now, it's down almost 50% in the last year. So, a lot of challenges there. They're going to have to reinvent their business model, which I'm sure they will. Dave, what are your thoughts?
Yeah, I'd like you to remember what Peter just said for later in the pod because I had the exact same experience. The OpenAI version of the voice was much more engaging. I can talk to it while I'm driving. It's great.
And then the Google version was stilted and robotic and just no fun. So now Google has leapfrogged, and it's actually better. But they did it under competitive pressure from OpenAI. I think you're going to see that theme throughout everything that we see on this pod: OpenAI will hopefully catch up and leapfrog again. But that's the only reason Google moves, because of that pressure. Otherwise, things just stall.
I mean, Dave, we had that conversation, and you noted it in our chat. A lot of the AI capability and a large number of the large language models were developed at Google, but until OpenAI released them onto the open web, Google was holding back.
It was the responsible thing to do: don't allow it to code itself; don't put it on the open web. That was the basic thesis of the last decade. And when OpenAI moved, Google had no other option but to move as well.
It's just a big company, you know? But I get it because I've run companies with hundreds or thousands of employees.
It's hard to make your company move. But then you get competitive pressure from a little, nimble company—
—and it's much easier as a CEO to say, “Guys, get your asses in gear. There's a threat here.” It's the kind of dynamic that makes America and the global economy move forward at all.
But all this technology, like you said, Peter, was originally invented inside Google. The transformer algorithm was invented inside Google, and it was just sitting there, literally not coming out the door at all. We could go through all the reasons. We've talked about them before, but sorry, Alex, you were going to say—
I would perhaps go even further and argue that many of these underlying capabilities are not just available, but they're available in the underlying data distribution that these models are trained on.
Exposing, for example, different accents is probably more of an unhobbling, as they would say, than anything else. It's not so much that capabilities are being added as restrictions are being removed.
Frontier models, in particular, when we see live-audio-type engagement, are moving from what they've been in the recent past—which is audio to text to text to audio—just directly audio to audio, which enables much richer audio interactions, including accents.
Yeah. And where we're going here with the next generation of AR glasses everyone's developing, we're basically plugging into your auditory and visual input. It's simultaneous translation. It is going to change how we communicate with people around the world in an extraordinary fashion.
This was a fun one. Again, continuing on the Google theme, the TL;DR: they really have gone and won hands down. I know, Dave, you and I are looking at the prediction markets. Google has literally skyrocketed to be the contender that's going to be the winner by the end of the year, and I think they got that mantle.
“Google AI helps users shop, compare, and call stores for the holidays.” New agentic features can call your nearby stores, check stock and pricing. Gemini apps add built-in shopping tools.
I mean, this is like, “Hey, call 20 stores within 10 miles of me and find out who's got the cheapest prices and put it on hold—or, better yet, purchase it for me and have it delivered tomorrow.” Holy cow. A lot to unpack there.
So I had to check on this one, Peter. It was all of 7 years ago that Google launched Duplex, their AI store-calling functionality, at I/O. 7 years ago—2018—the year after “Attention Is All You Need.”
It's been 7 years for this to make it into some fully realized format, but I think this is finally the beginning of AI starting to autonomously index the physical world. If you can have AI call stores autonomously, you can send AI-powered robots out into the physical world to index everything that's going on as well.
I'm curious what consumer behavior is going to be like. What I'm really interested in is what it's like on the other end, when you're in the store and you're getting all of these inbound calls. At what point is more than 50% of the calls AI calls?
Well, you have AI answer the AI calls, obviously.
I mean, is it going to be that you have to identify yourself as an AI? Probably.
That is what Duplex has historically done. It announces itself as an AI assistant.
Yeah, so far it's going to be state by state. The AIs are not announcing themselves, and we do a lot of this inside our lab here. About half the time, people are like, “Am I talking to an AI?” and the other half, they have no idea.
And so, do you have to answer if it asks?
You don't have to. In most states, you don't have to. But regulatory considerations are moving so slowly that it's completely ambiguous. As of right now, you don't have to, but it doesn't hurt to say, “Yeah, I'm an AI,” or even declare it up front. It's not hurting call-performance rates at all, so you might as well just say, “Hey, I'm an AI, but I'm so much more helpful than the guy you were going to talk to.”
My new business idea, then, is a little button on your phone. When an AI calls you, you flip it over to your AI, because when I'm calling a store, I want to speak to a human. I want to ask the human at the store, “What do you think about that product? How good is it? Are people returning it?” That interaction is a proper human-to-human interaction, but I'm not going to have that tolerance with an AI.
AI.
Wait, I want to challenge you on that.
If you call a store, why do you want to talk to a human? An AI is going to know way more about the inventory and the situation than a human.
No, yeah, exactly right, Salim. And not just that, the AI can pull up images in real time, show you the product, spin it around, and stuff. So it's not at all like talking to a human in a store. It's far, far more engaging.
I'll tell you what else. The VoiceRun guys here in the lab are doing OpenTable, doing restaurant bookings and stuff.
You wouldn't believe the fraction of restaurant bookings that are made by non-English-speaking people.
Or going the other way if you're traveling internationally. It's a lifesaver to be able to talk in a different language and do your full booking, and then the AI just translates it.
Fascinating.
I do think this is how we get to APIs for everything. Right now, there's a need for an escape valve for surfaces, for business interactions that don't support APIs. With an AI that can make voice calls and have arbitrary, unstructured interaction, we get APIs for everything.
Yeah, we do. Okay, one more article on the Gemini front: Gemini 3 benchmarks. We should probably skip this. I don't think anybody's interested in it, but okay.
You can't. You're teasing. Good.
Alex has just sent a drone to your house there. Watch your roof.
To take me out. I'm going to send my Duplex AI to give you a phone call.
All right, Alex, clue us in here. Gemini 3 benchmarks: How good are they? At the end of the day, what do they really mean? I mean, for the people watching and listening to our Moonshots program, I hear you talking about benchmarks every time, right? We're going to talk about some more benchmarks in a little bit, but what does it really mean? What does it mean to me?
Sure. I guess there's the headline: The numbers are going up and to the right, but who cares? We have a way of measuring progress in our civilization, and this is a precious moment when, with raw numbers, day by day, we can track progress toward solving some of the hardest problems that our civilization faces.
Humanity's Last Exam—say what you like about it. Some like it, some like it less, but it's an attempt, as are all of these benchmarks, to encapsulate, in a measurable, quantitative way, progress by AI toward solving hard problems. In the case of Humanity's Last Exam, it's an attempt to measure AI's ability to solve PhD-level problems. In the case of ARC-AGI-2, it's an attempt to model human-level ability to visually reason.
The so-what is that these benchmarks are all saturating, which means that AI, at this point, has the ability to perform PhD-level research. For the so-called average person, the implication is that AI is eminently well-positioned, now that these benchmarks are saturating, to start solving the hardest problems on Earth in math, science, engineering, and medicine. That's the so-what.
We spoke about that last episode with Sam Altman, speaking about science breakthroughs coming on GPT-6. That's his expectation, and here the numbers are impressive. If we're looking at GPT-5.1, Gemini 3 is basically doubling the ARC-AGI-2 benchmark. It is effectively doubling Claude 4.5 on Humanity's Last Exam. These are not incremental moves; they're significant step-ups.
Critically, it's not benchmark-maxing that we're seeing. There are some labs that have been accused of optimizing their AIs to do well on 1 or 2 of the benchmarks, and then, when you ask them something out of distribution, they fall over. That doesn't appear to be the case here. It feels like the team behind Gemini 3 really did a professional job of not overoptimizing toward narrow, spiky intelligence on any of these benchmarks to do well in a press release.
This feels like a well-rounded, generalist AI model. Given the trajectory toward saturating these benchmarks, I'd be very surprised if, by the end of next year, we're not seeing hard research problems succumb to AI models like this one.
Do you remember, 2 podcasts ago, Alex, we had that paper that came out on how to measure AGI, defining it in terms of—I don't know if it was 10 or 12 different dimensions? I wonder how Gemini 3 does on that. I'm sure we'll know soon enough.
I would expect it to do generically well on the spikes where models historically were doing well. As I recall, one of those dimensions where models historically did poorly was continuous learning with ultra-long context. Off the cuff, I wouldn't expect Gemini 3 Pro to do amazingly better on ultra-long context, but it does really well on retrieval scores.
I don't think it's shown in this slide, but there are other needle-in-a-haystack-type benchmarks that attempt to measure how well models are able to retrieve tiny facts of information buried in their context window. Gemini 3 Pro does amazingly well at retrieval as well. I think almost everything is going up and to the right at this point.
Yeah, there was 1 observation I had, and I wanted to check with you guys what you think of this. When you have coherence at this scale, it implies we have systems-level thinking inside these models. Is that accurate?
Could you say a little bit more, Salim, about what that means?
Well, because you've got, essentially—systemic thinking is one of the holy grails of deep, deep reasoning, right? You can look at the entire patterns of things shifting. It feels to me like, at this level of AI competency, you can get to that kind of systems-level thinking.
That means you can do world modeling in a really powerful way, using—almost, you shift the whole thing into symbolic reasoning when you can think in those concepts. So don't we get to that level very quickly?
I have so many thoughts, but the first thought that immediately jumps out at me is: Of course these are world models, and of course they're able to symbolically reason. They're solving math problems and writing source code.
I would argue that in the past, you've seen some commentators argue that there's some sort of nebulous neurosymbolic advancement that's waiting to drop. I think that's utter nonsense. Of course they're able to reason symbolically; the tokens are in some discrete space. And of course there are systems-level thinkers. They're able to solve PhD-level problems across dozens of disciplines. That requires understanding the world as a system. So, yes.
Yeah, I agree with that, and it turns into a philosophical debate, and nothing great usually comes out of it. But I will say that this is a 7-trillion-parameter-class model, and last year all the naysayers were saying, “Well, there's evidence that things will slow down, because last year we were at a trillion parameters,” and they were clearly wrong.
When you went from 1 to 7, we know next year is at least a 10× and up to a 40× step-up in raw horsepower. The naysayers are saying, “Well, things are going to level off unless we crack through some other level of System 2 thinking.” But they're clearly not leveling off.
I would challenge the technical audience out there looking at these benchmarks: You're almost obligated to think about 2 things if you're AI-inclined. One of them is: Where on these benchmarks does it become self-improving? Read all of Ray Kurzweil's work and really have an opinion on that, because that's tied heavily to benchmarks 1, 4, and 6 on this slide.
Have an opinion about where you need to be on 1, 4, and 6 in order for this thing to improve its own algorithm. That's a critical point. And then the other one is: Where do you need to be on the benchmarks to start proposing cures to diseases and being right?
If you work anywhere in health tech and you have no opinion on that topic, you're doing a disservice that's bordering on—in my opinion, bordering on—negligent homicide. This can save lives if you work on it, if you apply it to whatever you're doing in health tech.
You're obligated to get your head out of the sand, study the numbers, and at least have an opinion. Even if that opinion is, “No, it's not going to work.”
That's fine. I'm okay with that. But to say, “I don't know” or “I didn't listen to the pod,” that is absolute negligence.
Can I ask a question of you, Dave and Alex? Yann LeCun comes out saying we've gone down the LLM rabbit hole, and that's the wrong direction. We're optimizing on that, but we need to go through a different evolutionary tree to really get to AGI. What are your thoughts?
All the old people say that, and all the young people don't. When that tells you something out of the gate, you know you're sorting yourself into an age bucket just by saying it. There's definitely a philosophical divide in there.
But the question I would ask isn't whether there's another innovation that we need; it's whether a human will have that innovation or whether this exact AI scale will have that innovation. I would bet on the AI. Either way, we have so much to absorb just from where we are now. Forget everything else that may come along later.
Yeah.
I think there are also many paths to AGI, and I know and respect Yann LeCun's work. I know he favors an approach toward AGI that's more focused on actions in an embedded space rather than in terms of autoregressive models. That may be a perfectly legitimate approach as well.
But when I see the scaling laws continue to hold and capabilities continue to go up and to the right without any new paradigms, it makes me think maybe we really can just continue scaling and don't need to worry as much about yet another paradigm shift.
And let AI do that. All right, let's go on here.
Insert my normal rant about AGI here, and we can move on.
Okay. Yeah. So noted and approved.
All right, next story. OpenAI introduces GPT-5.1 for developers. Again, this is a benchmark question. First of all, this was announced before Gemini 3 came out, so I'm curious, AWG, whether this is still the case and why this matters.
Yeah, I think the economics of this—the microeconomics—are maybe even more interesting than the technical side. We're starting to see, and this is somewhat visualized in the chart you're showing, inference-time compute beginning to conform to the economic productivity of queries.
You know how, in Google Search, for example, if you search for mesothelioma litigation, you're going to see a bunch of very expensive AdWords ads. It's a very economically valuable query. On the other hand, if you search for an arithmetic query, you'll see no ads or almost no ads because it's not that economically valuable.
We're starting to see that same dynamic emerge here, where certain queries require lots of inference-time compute. What we're seeing at the routing layer with GPT-5.1 is even more compute being allocated to queries and prompts that really require a lot of compute. For the lighter, easier queries or prompts, we're seeing less compute get allocated.
I think this is actually pretty profound. It's not just a matter of moving around the deck chairs in some sort of zero-sum game. I think this is almost a premonition for what the economics of post-superintelligence will look like.
One of the things I think about most is who's going to pay, at the end of the day, for the trillions of dollars of capex in data-center buildout. Who's going to pay for it? Is it going to be the consumer? Will consumers, on average, be spending hundreds of dollars per month on consumer subscriptions for AI, or will it be enterprises that are spending billions of dollars, in some cases, for enterprise-level tasks?
I think what we're starting to see here is that, modally, probably it's going to be the enterprises paying lots of money for the most valuable tasks, in the same way we're seeing right now, in microcosm, some of these harder tasks and harder prompts get allocated a lot more inference-time compute at the expense of easier queries.
I would totally bet on that direction. If you're, say, Target, you can manage merchandising and get 20% extra margin on something, then it's worth the extra compute on the back end, and we'll see a lot of that.
But there are places where consumers will spend hundreds of dollars a month on their iPhone, on their plan, because it enables them in an extraordinary fashion. Remember that the money to be made here is on the margin, from persuading people to switch their behavior from what they otherwise would have done.
If they were going to spend the money anyway, that money doesn't go to the AI. It goes to the entire value chain underneath the phone manufacturer.
Mm-hmm. All right.
Well, I can tell you from my experience that you have to operate at the margin, at the extreme end of what these are capable of. I've tried to either save money or get more speed by dumbing it down by a half step, and it just isn't the same.
It feels like everybody wants to be at the forefront. This is the weirdest product that's ever been launched on humanity in that it's talking to you as it's selling to you.
And so you start with a subscription. They give you this incredible experience, and then it tells you, “Well, you want more of that? You need to upgrade.” But it's actually telling you—it's talking to you—about upgrading.
No product, no cable company, no iPhone has ever done that before. So it's a salesman baked into its own capabilities. It's kind of creepy, actually. It's very weird.
All right, let's stay on the OpenAI theme. This is a fascinating story, and it's an important one: OpenAI-backed startup aiming to block AI-enabled bioweapons.
This is a startup called Red Queen Bio, and they received a $15 million investment from OpenAI, which sounds really small compared to all the $100 billion and trillion-dollar investments being made. But Red Queen is using advanced AI plus lab testing to spot vulnerabilities in biological systems.
They're basically saying, “Hey, we want to stop people from using these AI models to create bioweapons.” Super important. Who wants to jump in first?
I'd love to speak to this one, maybe starting with the literary reference. For those not tracking, Red Queen in this case is a reference to a scene in Through the Looking-Glass where Alice and the Queen are constantly running just to stay in the same place.
The Red Queen's race, in general, is used as a metaphor for cases where a lot of effort is required basically to maintain a standstill. In this case, I think the other key concept that is ultimately quite profound about what Red Queen Bio has announced, and the reason why they're taking funding, is that we've just spent quite a bit of time talking about how, as you pour more compute onto these models, the capabilities keep increasing.
Inevitably, you have to worry about alignment and safety as well. In society, if you're growing a city and you double the population, you're going to approximately want to double the police force or the safety force.
Wouldn't it be wonderful if, as the capabilities of AI keep scaling and increasing, the safety measures, the alignment, and other properties that make them safe for humanity also benefited from scaling with more compute?
Seeing scaling laws and Red Queen Bio's announcement that they've uncovered scaling laws for biological safety measures, I think this is the way we achieve alignment. Again, the scaling law for police forces in a city is a little bit sublinear relative to population. Same idea here, but nonetheless close.
As capabilities increase, we want to live in a world where we achieve so-called defensive co-scaling, where the resources and capabilities of safety measures scale close to proportionally with the resources and capabilities of the underlying models.
Yeah, let me add some data to that. In 2024, the biosecurity and biodefense market was $34 billion, and it's expected to double a decade from now, by 2034 or 2035.
But here's the quote that really hits me: An extreme bioattack scenario could have a multi-trillion-dollar global loss. The notion is, could you create such a bioweapon for $1,000? It's the asymmetric situation where a small amount of money, using complex models, could do a lot of damage.
So there has to be this layer of defense. It's critical. When I talked to Eric Schmidt a couple of years ago at FII, I remember that the number-one scenario of greatest concern was bioweapons—something where you take an existing virus, change its viral payload, make it much more infectious, and release it.
You and I have had this conversation: one of the most important things is going to be setting up biosensing capabilities at train stations, airports, and bus stations that filter the air, look for, and rapidly sequence everything they come across. The majority of the bioweapons that are concerning are airborne, right? A person coughs or sneezes, and it's there. One thing in our favor is that these viruses, these bioweapons, can only move at the speed of an airplane. That's the fastest they can go, right? We saw that with the release of COVID. So if you can detect it at an airport, sequence it on the spot, develop an antiviral, and transmit that at the speed of light—not at the speed of 600 nautical miles per hour—then you have a chance of battling it.
This is also exactly why open-source AI is dead in America. Meta decided, "Okay, we're not open-sourcing," so now none of the U.S. labs are open-sourcing anymore. The only open-source models are coming from China. But if you're a U.S. company, usually a terrorist in a basement in some jurisdiction somewhere in the world isn't the sharpest tool in the shed, and you're counting on them not knowing how to build the weapon. When you give them genius-level AI as a sidekick, suddenly they're empowered to build virtually anything in that basement. That's the risk. No U.S. company wants to be responsible for that.
So they're trying to cut it off at the query level, saying, "As soon as you ask the AI to help you create a bioweapon, it stops." Open source would be a huge leak in that. The U.S. labs don't do open source anymore. The Chinese still do. Alex, comment on that. How do you deal with that if it's a model running on my laptop and somehow it contains enough knowledge to do this? I can query my laptop, and no one ever knows the query I've made. It's just resident there. How do we deal with that?
I think ultimately it all reduces to co-scaling. If you imagine having a fully self-contained facility, hypothetically, in your basement, the ultimate societal protection will be having lots of sensors and, more importantly, lots of AI screening—superintelligent AI screening—that can spot hidden agents.
I have this dictum that I think is super important on so many different levels in the software engineering world. Linus Torvalds, who created Linux, has this so-called Linus's law—and I'm going to butcher this slightly—that with enough eyeballs, all bugs become shallow. I would propose a sort of generalization: with enough superintelligence, all hidden agents become shallow.
To the extent that we have hidden agents in their basement building superweapons, I would expect that with enough superintelligence, defensively co-scaled, they become shallow. I've made the comment before that privacy is an illusion. This is just going to shatter even that illusion, because if you want safety, you're going to want agents listening and watching everything all the time.
This is an arms race. I think what we've seen throughout history is that we thought, "Oh my God, email's going to crash because of all the scams," and then we thought we had phishing and couldn't solve for that. We've used AI consistently in that sense, because people forget that the bad actors may use AI, and they will, but the good actors can also use AI. Therefore, you just have to be one step ahead. The question is, if that gap gets too big, one of the challenges with what you were saying earlier, Peter, is that you may not know what to look for in some of these, and that's the danger point.
An interesting little case study, too, because if you rewind the clock before Gmail took over, Microsoft had Outlook and Hotmail, and Google launched Gmail. The two promises were very different. Microsoft said, "We will never read your email."
Google said, "We will read every word of every email that you receive, but it's going to be read by an AI and not by a human. So we won't let human eyes look at your email, but we're going to do all kinds of things based on the information in your email read by the AI." People didn't care, and so everybody moved to Gmail. You have an interesting case study in how this plays out—just the human behavior.
So here, I think the equivalent is, "Hey, I'm talking to AI about my most personal things in the world." Peter, I think you're right: the AI is going to listen to every single word. If you're designing a bioterror weapon or a cyberattack, it's going to flag it and escalate it. If you're talking about your virtual girlfriend or whatever, it's going to be fine; it's just going to kind of hide that.
I remember talking to the head of one of the major intelligence agencies, and they had a very clever thing. They said, "Look, when there are known things like nuclear weapons or whatever, we put eyes on it. We try to watch it. When you have something like this that could be developed in secret, we've been actively opening up these communities and actually funding the biohacking movements, because then you can see things earlier."
But if you can do open-source bioweapon development in a lab in a bunker, that really causes a huge issue. We're going to have to rethink an approach, something along the lines of what Alex said. You remember when you gave the Ayn Rand Award to Michael Saylor? I don't know if you did the keynote. Michael gave this incredible speech, but those people are probably vomiting right now based on how this is evolving.
Well, listen, the bioweapon—I mean, you're not going to create a novel virus that has zero history involved. There are extensive registries of every virus that's ever been mapped. So when, at an airport, you identify and sequence something and it's not on that registry, you can then look at it, and LLMs—or future bio-LLMs—will be able to look at, "Okay, this is an infectious agent. This is something that's able to be airborne or waterborne." When you look at the proteins, you can tell what kind of virus or protein it's generating.
So you're going to be able to learn instantly when you sequence it, and rapid sequencing is here. But we're going to need this, and I think giving up privacy to a large degree—which you've talked about, Salim—when you're in an airport, you've basically given up your privacy right there.
You're being surveilled, and your rights can be taken away at any time. One framing of our era is that we're living essentially in a global airport. I think that continues to some extent. I don't see a way of coming back from that.
Well, good luck, although Brad and the EFF folks say there is a way of doing it. You don't have to compromise privacy for security. There are lots of mechanisms for solving this in other ways. That's their complaint: governments kind of go after the surveillance side just because, "Oh, this is great. We can surveil people under the excuse of security." But many times, you don't have to.
If I might close the discussion on this, I want to make sure we don't over-index on safety concerns, or so-called safety. I think these are very important concerns, but I also think that AI can be scaled to combat them, just as one might naively expect, prior to the development of modern cities, that crime would be overwhelming and that humanity would not be able to support itself in urban environments at scale. It turns out that we are able to.
Although we're not doing a book corner this episode, I would encourage everyone to read Vernor Vinge's Rainbows End, which does a glorious job of depicting what the future of AI-enabled biosafety looks like.
Amazing. Well, I'm the eternal optimist here, and I'm absolutely clear we're going to be able to overcome this. Let's move on to one more benchmark here: xAI releases Grok 4.1, ranking number 1 in major leaderboards for reasoning and writing. Back to our resident leaderboard expert.
My comment on this one is short. This lead in the Text Arena benchmark lasted approximately 1 week and was over. [laughter] My short comment here is that the race for the frontier is so intense that even if a frontier lab is perhaps benchmark-maxing toward a singular benchmark, generalist models seem to be able to push the frontier at this point on a weekly basis. I can only imagine, as timelines progress, what this is going to look like when these benchmarks are being toppled on a daily basis.
I'm sure Grok 4.5 and Grok 5 are around the corner. Let's move on to Cursor. Cursor triples its valuation in just a few months, from June through November, going from roughly $10 billion to roughly $30 billion in 6 months, and raised $2.3 billion. There's Michael Truell, the CEO of Cursor. Who wants to jump in here? I mean, this is a hot race between a whole slew of different coding tools out there.
This seems to be in Dave's wheelhouse.
Dave, yeah. What do you think?
Well, I'll tell you, I think this team is phenomenal, and most of the people around here think they'll rise to the occasion and succeed. But I also think that Antigravity looks exactly like Cursor. [laughter] I actually have both open on my laptop side by side, and other than a little cosmetic difference here and there, you don't even know which one you're in.
Then you look under the covers, and it's like, well, I can access all the models through Cursor, and I can only access Gemini 3 through Antigravity.
So there's a difference right there. But then the bet at Cursor is that Anthropic and the other models will be worth having, and Gemini 3 doesn't just run away with it anyway. It's really an interesting horse race right now, and I'm not going to make any prediction on it because you can't make a prediction on it. Their core positioning is incredibly vulnerable, but the team is brilliant. They're well capitalized.
Back up for those who don't know what Cursor is or what it does.
Fair.
Let's do that basic 101 right now. Dave or Alex?
Yeah. Cursor, I think everyone around here that I know uses it every day. It's the best, or has been the best, coding assistant that uses AI. It's fully agentic now, so you can just type in a prompt. You can talk to it now, too, and it'll just build things for you.
Under the covers, though, they don't own their own foundation model. It goes out to either OpenAI or Grok, and it has all of them in there. Anthropic is what I usually use—Claude 4.5. It organizes everything, cranks out the product, and configures your laptop for you. It just makes coding trivially simple. Anyone can do it, and it's pretty universally used.
It was early to market. When I think about the value of AI in this world, where does value aggregate? My list is data, scaffolding, user experience, and integration and customization, and then the models themselves. So where would you put Cursor in those categories? The scaffolding?
Anything other than the models themselves. Yeah, it's all of the above other than the models themselves.
Compared to Replit—we've talked about Replit a bunch—and Lovable, how do they compare to Cursor?
Replit and Lovable are much more for your mom and pop who want to build something like a video game quickly, or an invite to a birthday party with moving graphics, or whatever. You can build something while you're flying your plane, Peter, like you did. It's super, super easy to onboard. Cursor is more for hardcore engineers who are moving to AI and trying to get 10× more performance out of their engineering.
I would just note that, for what it's worth, all of these, or almost all of these, integrated development environment companies, including Cursor, are rolling out their own first-party models. It's almost inevitable that they want to climb down the stack to own more of their software supply chain.
I think the success that we're seeing from Cursor, which is of course very exciting, is a reflection that software engineering is probably the first high-productivity labor category that's being automated by AI. It won't be the last, but it's the first big one that we're seeing.
Surely AI-driven software development is now the default, right? I mean, you couldn't do it without it now, already, after just a few months.
To the point where this is crazy: I see companies that are almost treating potential software engineering hires by vintage. Did they get their degree and their experience prior to AI coding or not?
Are they spoiled?
Have they been ruined?
Basically, yes. Did they get their skills? Did they learn? Did they have lots of experience prior to the atrophying that comes, perhaps, with agent coding?
All right, I'm going to move us forward to another incredible article. This is a new startup funded by Jeff Bezos called Prometheus. Jeff put in $6.2 billion. By the way, the ability to start a company with $6 billion on your balance sheet has got to be frightening for a number of startups, and it's got to be incredibly accelerating. We've never seen this kind of thing—starting with billions, multiple billions of dollars, on day 0.
So what is Project Prometheus? It's AI-enabled engineering and manufacturing. It's basically learning from real-world experience so they can manufacture efficiently and focus on physical testing and simulations. I love this other bullet point here: Prometheus has hired nearly 100 researchers from OpenAI, Google, Meta, and other labs. They're just feasting on each other. They're stealing each other's well-trained talent.
If the going rate is $1 billion per researcher, then this is really underfunded.
They've got $6 billion on this. But I find that the 2 things I found fascinating right off the top—we'll talk about the meat of what Project Prometheus is in a second—are starting with that much money and that they're basically stealing from each other. Dave, what do you think?
Well, it's funny. I've had probably 12 meetings with different MIT teams in the last week, 30, 40, 50 at a time, and about half of them are computer science. The other half are not. The half that are not are saying, “How do I get involved? What do I do? What's my AI role?”
When MicroStrategy started, Michael Saylor was an aero/astro guy, and all the rest of the guys were in computer science. The company took off under Michael's leadership; it didn't matter what he studied. AI is like that. There's nothing in the computer science curriculum that teaches you much of anything anyway. Don't be intimidated.
Yeah, and so what you're pointing out here on this slide, Peter, is, okay, they stole another 100 people. Clearly, the industry wants 100,000 more people to come in. Why are you letting this guy get a $1 billion signing bonus? Why don't you get into the market, learn this stuff, and be there for $100 million? Just get in the game.
People get intimidated away from it because they feel like it's all geniuses, and they think, “I'm going to get crushed.” It's just not true. Just get in the hunt. Get into the game. This is the thing happening in the world now.
There's usually only 1 thing driving all change in the world. This is that thing. Just get into the middle of it. And then, Jeff, the other thing I'll point out in this is that there's a tendency to be intimidated by Elon Musk spending $6 billion or $7 billion building a massive data center in record time. How am I going to compete with that?
But the foundation models that will do parts creation or robotics simulation, or whatever, are different enough from a large language model that you can build a great foundation model company in parallel with OpenAI, Grok, Meta, and Gemini. It's okay. You shouldn't be intimidated by that either. And that's, I think, what Jeff is saying here.
Jeff Bezos bought all the robotics companies, put them into warehouses, and just ran away with warehouse automation, which created a whole litany of new startups working for Walmart, Target, and everyone else, like Symbotic, where Daniela Rus is on the board and does the robots now for Walmart's warehouses. Here, Jeff is saying, “Okay, Amazon is big enough that I'm actually going to be able to build a multibillion-dollar company within our own universe, our own channel.”
But that creates an opportunity for somebody to be outside the Bezos universe, doing it for everybody else.
And so, the design space is wide.
We're going to have Jeff Wilke on stage at the Abundance Summit this year. Jeff was the CEO of Amazon Worldwide. There were 2 divisions: one was AWS, and one was everything else, and Jeff Wilke ran everything else. He's actually super excited about this because this is what he's doing. He's got a company called Re:Build Manufacturing, which is working in this area, too.
Alex, let's get into the nitty-gritty here, right? Prometheus is building physical AI. It's world models again, like Fei-Fei Li and a little bit like Genie 3. These are world models understanding the laws of physics, chemistry, and engineering. So you can actually do real optimization. What are your thoughts here?
Yeah, I think we're starting to see the pivot of the capital markets from funding superintelligence to funding that which comes after superintelligence, which is, as I've argued in the past, solving math, science, engineering, and medicine. I think it's a 10× or 100× larger market opportunity—a larger addressable market—solving basically everything else after solving superintelligence than solving superintelligence itself.
$6.2 billion is a drop in the bucket. I would expect it's going to cost many, many trillions of dollars in funding to solve all outstanding problems in math, science, engineering, and medicine. There's been relatively thin reporting on what Project Prometheus is particularly focusing on. I've taken note that it seems to be absorbing a lot of old biology friends of mine, so it's possible maybe it ends up focusing a little bit more on biology and a little bit less on manufacturing. But I think this is where the action is after superintelligence.
Yeah, I have 3 points I want to make here. One, this is kind of a shift from chatbots to industrial agents, right? AI for the office is what we've had. This is AI for the factory floor, where there are physical consequences, where the systems are able to operate the factories because they understand the physical constraints, situations, and logistics.
The second thing is, I met Jeff in college. I was the chairman of SEDS Worldwide at one point, and Jeff was the president of SEDS at Princeton University when I was at MIT. Space has always been his passion. Congratulations to Blue Origin for its recent launch and landing.
We talked about that last time. But this kind of physical AI system is exactly what you need to operate heavy industry in space—to build factories in orbit, to build factories on the Moon, and to have them fully autonomous and capable.
The final thing I would say is that this is going to change. We've seen companies like Lila Sciences and other companies out there that are going to move from invention that happened through serendipitous human creation to invention coming from a computational process. That's when it gets super interesting, and that's what you've been talking about, my friend Alex.
That's right. What's neat about this is it felt to me like he's creating a backbone AI for everything in his world—Amazon, space, logistics, et cetera. This will service all of those.
And Elon will do the same, of course.
Mhm.
Like electricity, it's going to run through everything.
Yeah. Yeah.
The foundation model that I built early in my career took 5 years from the day I started writing the code until it was done. I can recreate it now in about 2 months, which I just did. If you look forward a year, that'll come down another 5 to 10 times. So you can use AI to build the next AI, which is essentially what I just did.
The same applies in mechanical design. If you said, “Wow, building an entire AI platform that designs rockets or robots is really hard,” well, it would have been, but now you can use the current AI to build that AI.
It cuts the time down tremendously. If you just look forward a year to where the existing AIs will be, that time is actually not intimidating at all. It's a good reason to get into the game and build these parallel AIs that work on very specific problems, whether it's biotech, mechanical design, futures trading, or whatever it is.
Build it from the old AI to the new AI.
The last time, I asked our subscribers to post questions. I took all the comments, put them into ChatGPT, and asked it to summarize the most important questions. There was a critical question that was asked, and I want to take a second and read it because I want to have an AMA about it.
It said, “What concrete milestones should people expect to see that prove abundance is coming? In other words, lower costs, new industries, accessible AI tools, and how do we ensure these benefits reach everyone rather than concentrating wealth among a small AI-augmented elite?”
I want to play a video that was posted on X today, and then we're going to talk about this question.
AI and humanoid robots will actually eliminate poverty. Tesla won't be the only company that makes them. I think Tesla will pioneer this, but there will be many other companies that make humanoid robots. There is only basically one way to make everyone wealthy, and that is AI and robotics.
All right, so that's Elon's thesis. I posted the question here again, and it's a real concern. Are we going to have runaway wealth concentration? Honestly, if you want me to believe in this future of abundance you keep talking about, guys, what are the concrete milestones, and how do we ensure these benefits reach everyone? How do I know it's actually coming? Let's jump into this.
Can I throw out a couple of points?
There's an important framing here. Let's not talk about the wealth gap, because the richest people in the world are always going to keep getting richer. The issue is more: can you lift the bottom if you care? You make this point all the time.
A thousand years ago, the king and queen on the hilltop would live below the poverty line today, and there were thousands of serfs who supported them.
They died of a tooth infection at age 22.
Or they were bled by leeches, as the king and queen sat up there. What we've done is, yes, we're heading toward a world where there are trillionaires living on Mars. But if every man, woman, and child has access to all the food, water, energy, healthcare, and education they could possibly want, we've lifted the bottom of humanity to a point where mothers can believe their children have access to everything they need. That's the world I want to live in. That's the world I want to create.
Let me speak just to that for a second. We forget, because we see all this—we see people getting richer, et cetera—but we have to remember the unbelievable benefits occurring at every level.
I'll give you a concrete example. When the tsunami hit Indonesia in 2004, all the ship-to-shore communications were wiped out. The government gave cell phones to all the fishermen, saying, “If you're out fishing and you see another tsunami, text it in,” et cetera. They found, to their surprise, that their incomes had increased by 30% over the next 2 months.
They looked into it, and all they were doing was texting to find out what the market price of the fish was. Should they stay fishing? Should they come in and sell?
Or which port they should go to, and who was paying more.
Yes. That little hint of what Alex would call the inner loop allows you to increase income pretty radically by having democratized access and demonetized access to cell phones, smartphones, and now AI. This will change the game completely for everything, everywhere.
I'll touch on 2 areas. One is education. You can now sit a child down with a smartphone and say, “Create a lesson plan for grade 7 algebra,” and they're going to learn 10 times faster than all the kids stuck in elementary schools in the West that, by law, have to go to these things.
The second is healthcare, where every single medical condition can now be diagnosed instantly. When you catch something early, the cost of treating it drops by something like 100 times. Those are 2 very concrete areas where AI will make a massive difference—2 areas that were traditionally inaccessible, hard to get, and expensive.
Let me read the numbers here. The U.S. average expenditures for a family in 2023 were $77,000. The number-one cost was housing: 33% goes to housing, 17% to transportation, 13% to food, 12% to insurance and pensions, 8% to health, 5% to entertainment, and about 2.5% to education.
Let's knock these down. Housing, number 1: now you can live outside of the city, where it's cheaper, and be able to telecommute in, reducing your housing cost. There is a future—it's not here yet—where we're 3D-printing houses, reducing the cost. What we saw on stage a couple of years ago, if you remember, Salim, was 3D-printed houses that were the cheapest per square meter, but also the most beautiful and luxurious, because you could get the greatest designers to create a standardized print file for people to use.
Transportation is 17%. Well, guess what? An autonomous electric Cybercab is 4 to 5 times cheaper than owning a car. It's going to be cheaper than an UberX and cheaper than a bus. So we're going to solve that.
Food—we've got to solve food better. We need, basically, vertical farms and stem-cell-grown meats.
Let me give you the statistic on vertical farms.
We've been doing horizontal farming since the beginning of time. Vertical farming is just crossing over now into economic viability. You can drip-feed water to the plants and know what nutrients the plants need because the sensors know it. You get about 7 times the yield of horizontal farming by doing things vertically because you have the right frequency of light hitting it.
You save 99% of fresh water. By the way, we use 70% of our fresh water globally for agriculture. The best calculation we've seen is that if you took 35 skyscrapers in Manhattan and turned them into vertical farms, they would feed the entire city sustainably.
Just think about that from a logistics, food security, pesticides, and fertilizer standpoint. There are massive changes coming down the pike, and this is before we apply AI to the whole mix. The radical changes coming are going to be so huge that the cost of everything should drop to near zero.
The amount of energy you need to feed 1 person is the amount of sunlight hitting 1 square meter, and that energy would feed somebody for a year. All we have to do is figure out a better loop for converting that energy into consumable foods, and we've got a long way to go.
Healthcare is 8% of our costs. You said it already. We know that an AI physician-diagnostician is significantly better than even the best physicians, and an autonomous robot eventually will be the best surgeon. The cost of that will be capital expenditure and electricity.
It's hard for people to believe this stuff now because it's on the bleeding edge—literally—but we're going to get there. Entertainment is 5% of our costs. Well, guess what? YouTube—what else could you want? Education, as you mentioned before: AI, YouTube, all these things.
So, we're demonetizing and democratizing this stuff. It's just hard for people to realize it. I think the challenge is that we compare ourselves to the Kardashians, right? We compare ourselves to people that we see on TV and on the internet all the time, versus comparing ourselves to what it was like for our parents or grandparents.
Yeah, I think that last point is the key one because we've had dirt-cheap food for a long time, but everybody still wants a $14 Starbucks latte, which you don't need to pay for, but there it is. Why do I feel that need? So, the metric I'd be tracking is actually depression rates, because I think AI properly deployed can hit that much more quickly than it can hit robotic automation that creates new homes for everybody—ones that are 10 times larger.
That's a great point.
And so I'd be looking at that as an early indicator that we're on the right path. It's not a no-brainer; you've got to really think it through because you mentioned rent is at the top. 33% of household income gets spent on housing, on average. But when you look below the poverty line, I think spending on drugs, alcohol, gambling, and pain relief is 3.
Yeah. The opioid addiction alone is a trillion-dollar—
Error, I guess.
And it's about 5 times more collectively than rent.
Go ahead.
Well, no. So, I'd be attacking that. If you want to come bottom-up and say, “Look, we want to create universal happiness with AI,” we've never had a tool that could attack it before, right? You can attack manufacturing automation. You can make food cheaper. You can have harvesters that mow down half the Midwest to create wheat. But all that does is create more of what's already abundant.
Yeah.
Alex, this is all about benchmarks. We've talked about this. You and I have been working on a paper on this subject. Can you speak to that?
I think it's so simple. I think what's upstream of all of these other milestones is the dollar cost per unit of intelligence. As we've discussed previously, right now that's hyperdeflating by something like 40x year over year. To keep the party going, and to make sure that all of these downstream considerations—cost of living, healthcare, housing, and so forth—all hyperdeflate ultimately alongside the cost of intelligence, I think it's largely a regulatory and social concern.
We've spoken previously about, for example, the difficulties of getting Waymos in Boston. That's a regulatory consideration. The cost of intelligence needed to autonomously drive cars around is making excellent progress. But ultimately, in order to provide essentially free autonomous on-demand transit to everyone, there's a regulatory bottleneck.
In order to ensure that the benefits of intelligence too cheap to meter become evenly distributed, I think it's going to require some revision of social cohesion and the social safety net to make everyone comfortable with the downstream consequences of intelligence too cheap to meter, including healthcare, housing, energy, and utilities too cheap to meter.
Yeah.
I did a calculation. If you wanted to have a reasonable life, you could do it for $20 a day in Bali. Housing costs about $10 a day, and your meals are literally about $2 a day, and then a bit extra. So, for about $20 a day, you could do it.
If you had 0.5 Ethereum, which is about $2,000, you can put it into DeFi trading pools and earn about 1% a day, which is about $20. So, 0.5 Ethereum of capital allows you to live crudely, but it allows you to live in a very lovely spot in the world at a very low cost.
Think about just that feedback loop, because as you double that, triple that, or 10x that, all of a sudden you get into a really great place. You can survive today on a very small amount.
My feet are in the sand. My feet are in the sand already. And, of course, that Ethereum comment was not investment advice, just to let everybody know. But it is interesting that Harvard has doubled down on Bitcoin.
Now that we're in the Bitcoin doldrums, it's nice to see the institutions. I remember when we went from wacky individuals buying crypto to institutions, financial institutions, sovereign funds, and so forth, but—
Countries also.
Not investment advice. All right, what an amazing episode. We’ve actually just gone through half of our stories. But to make this consumable, because the feedback we’ve gotten is please keep the episodes under an hour and a half, we’re listening to the comments and trying hard. So we’ll have to spin up another conversation on everything going on in data centers, energy, space, and so much more. It’s hard during the singularity to keep up with everything. The mind-blowing stuff from Gemini 3 was worth covering properly.
Just a reminder: last summer, not that long ago, Polymarket had said that the top 5 had an equal shot at being the best AI model by the end of the year. Now it's 91% Google, but by next summer that's down to 60%. So, it's sort of 50/50 that someone else will take the lead by next summer.
Well, that's what we should hope for because Alex said the key point, as usual: 40x is what you should expect next year. People really struggle with 40x in anything. So, if the cost per unit of intelligence comes down by 40x, or just raw intelligence goes up by 40x next year, you should expect that.
It's very hard to visualize all that means. So, we'll do everything we can on the podcast to try and make that tangible for people, but really try and digest that coming out of this incredible Gemini 3 breakthrough.
And just hats off to Josh Woodward, to Sundar Pichai, and to Demis Hassabis for an extraordinary job on Gemini 3. Just so proud of what they've been able to create. And, of course, a lot more coming.
I have one announcement. Sometime in December, we’re going to do a Meeting of Life session online. I’ve had enough clamoring from my community and other people and Peter’s people to go to Abundance that people want to do it, so stay tuned; we’ll get more details next time around.
We’ll do it also at the Abundance Summit on Wednesday night. This is Salim waxing poetically and philosophically for about 5 hours straight. It’ll be like a late-night French salon-type discussion—alcohol or equivalent mandatory—on the metaphysics, philosophy, and what it means to be alive today.
It starts at 10 p.m. What time does it end? Dawn. It depends on the audience, but the crazy ones have gone till dawn because we never get a structured conversation on the meaning of life. We never get that, so let’s have that conversation.
We’ll do it—you’ll do it—and I’ll join you until my bedtime at 9:00, then I’m exiting the building. But I’m going to do it online in about a month, so we’ll do it earlier.
Last time we talked about the potential for a Moonshot gathering. We’ve had 500 of you email us. If we get to 1,000, if you’re interested in a Moonshot gathering next fall, send an email to moonshots@diamandis.com and let us know you’re interested in having these conversations and gathering with other Moonshot listeners.
Once again, we put out our call for outro music. This is a piece by John Novotney called “Moonshots Metal Version.” You need to see this. This is not just music; this is a fun video. See, you look so sexy, Dave, and I love your ponytail, Alex. AWG’s got a ponytail in this and he’s rocking it.
On our outro, let’s watch and listen to this heavy-metal Moonshot music.
Oh my God, I haven’t seen this. Oh, that’s a good one. Very gentle. Oh, not to miss. Oh my God. This is amazing. Cool. Bottling the lightning.
I’ve got to get the ponytail on. Peter and Salim, your look in that video was really good. You should just do that. I love Salim with the sunglass move. Dave, you on the guitar, and AWG, you on the keyboards, and the ponytail was you, buddy. You’ve got to grow that ponytail, apparently. Go some lesson.
Well, thank you, John. That was amazing.
DB2, AWG, and Mr. ExO, have a fantastic week. I love doing this, and thank you to all our listeners.
Great episode. Take care, guys. Take care, guys.