NVIDIA 的 Jensen Huang 谈推理模型、机器人与驳斥“AI 泡沫”叙事
Huang 对2025年的判断是:具备现实 grounding、重推理的 tokens 已经到位,客户愿意为其支付足够高的价格,并让供应商实现盈利。 接入搜索的模型和基于置信度的路由器显著提升了准确率;他听说 OpenEvidence 的毛利率据报达到90%,并称 Cursor、Claude 和企业 OpenAI 工作负载同样拥有强劲利润率。关键变化在于,tokens 的价值已经“足够好,好到人们愿意花大价钱购买”。
AI 需求远不止聊天机器人,因为每个新生成的 token 都需要一座覆盖芯片、超级计算机、能源和熟练劳动力的“AI 工厂”。 Huang 看到3类工厂同时扩建——芯片工厂、新型计算机工厂和 AI 工厂——由此带来对建筑工人、电工、水管工、技术人员和网络工程师的巨大需求。他和 Guo 将短期劳动力影响定义为扩张,而非替代。
衡量就业的正确单位是工作的目的,而不是 AI 能自动化的某项任务。 Huang 提到 Geoffrey Hinton 曾预测放射科将由 AI 驱动,但表示放射科医生的数量反而增加了,因为更快的扫描分析带来了更多诊断、研究、患者和医院收入。他将同样的判断标准应用于律师、工程师和服务员:“技术往往解决的是任务,而不是目的。”
算力成本下降削弱了前沿 AI 必须永久集中在少数资本雄厚实验室手中的观点。 主持人提到,2024年 GPT-4 等效 tokens 的成本下降了超过100倍;Huang 预计硬件性能每年提升5–10倍,并表示10年内 token 生成成本下降10亿倍并不会让他感到意外。由于硬件、算法和模型的协同创新正推动成本“每年下降远超10倍”,落后6个月或1年的竞争者仍可能保持接近。
开源是创业公司、科学、教育和工业 AI 的战略基础设施,而不只是另一种聊天机器人商业模式。 Huang 称 DeepSeek 的论文可能是“去年对美国 AI 最大的单项贡献”,因为美国实验室和基础设施公司都从中获得了启发。他反对等待一个单一的“God AI”统治一切——他将其规模比作“圣经级”或“银河级”——因为现实产业现在就需要可适配的领域模型。
下一层可投资机会是数字生物学和 Physical AI 的垂直化。 Huang 预计多蛋白模型、蛋白质和化学品生成、具备推理能力的车辆以及多形态机器人将催生新的应用市场;未来5年,“真正令人兴奋的将是垂直化”。通用模型或许能提供99%的能力,但工业服务商必须把可靠性推近99.99999%,这为领域专家留下了大量价值空间。
Huang 反泡沫的核心依据是产能稀缺,以及远大于 OpenAI 收入的可服务市场。 他提到 NVIDIA 的自动驾驶业务正接近100亿美元,金融服务、机器人和数字生物学领域也在涌现数十亿美元机会,并用全球每年约2万亿美元的研发支出,说明研发方式正在转向 AI 驱动。无论是创业公司、大学还是产业界,他观察到的信号都十分明确:“所有人都快被产能需求逼疯了。”
1. 推理让 tokens 变成可盈利产品
Huang 并不意外于 scaling 延续,但对 grounding、推理以及模型接入搜索能力的提升感到欣喜。基于置信度的路由器如今能够识别何时需要外部研究,显著提高准确率,而不只是生成流畅文本。
2025年最出乎他意料的,是推理量——尤其是推理 tokens——在“多个指数级增长同时发生”的情况下快速攀升。更重要的是,这些 tokens 变得有价值:Huang 听说 OpenEvidence 的毛利率达到90%,并指出 Cursor、Claude 和企业 OpenAI 工作负载同样具备强劲的经济性。
主持人将这一质量跃迁与专业领域采用联系起来:Elad 表示,医生越来越把 OpenEvidence 当作可信资源;Harvey 则正在成为可信的法律工具或交易对手方。Huang 的总结既是商业判断,也是技术判断——AI 如今生成的 tokens“价值足够好,好到人们愿意花大价钱购买”。
2. AI 工厂先创造劳动力需求,再改变办公室工作
AI 与 Excel 这类预先录制好的软件不同,因为它会根据用户上下文和当前信息重新生成每一个 token。Huang 因此把生产 tokens 的计算机称为“AI 工厂”:它们持续制造供全球消费的智能。
这套生产体系需要3类工厂:更多芯片工厂,围绕 Grace Blackwell 等系统建设的新型超级计算机工厂,以及大规模 AI 工厂。Huang 将整个机架称为“一块 GPU”,凸显这类系统已经与传统计算机有多大不同。
就业方面,最直接的影响是对建筑工人、水管工、电工、技术人员和网络工程师的需求增加。Huang 表示,一些电工的工资已经翻倍,还要为项目出差,“就像我们一样——我们也会出差”,说明实体基础设施正直接受益于 AI 投资。
主持人更广泛的观点是,潜在需求远未耗尽:社会还没有用尽医疗、软件或工业产出的有用空间。Huang 认同,NVIDIA 生产率提高并不会导致裁员;它会带来更多想法、增长和利润,而这些利润又会为更多工作提供资金。
3. 自动化改变任务,但扩大工作的目的
Huang 提到 Geoffrey Hinton 曾预测放射科应用将由 AI 驱动,但表示 Hinton 当时关于不应进入这一职业的警告并未兑现:大约8年后,放射科医生的数量反而增加了。
关键在于“任务与目的”的区分。研究扫描结果是放射科医生的任务;诊断疾病和开展研究才是工作的目的。AI 让放射科医生能够深入查看更多扫描、要求补充影像、服务更多患者,并改善医院的经济效益——这些条件都可能推高对放射科医生的需求。
同一框架也适用于律师和工程师。阅读或起草合同不是律师的目的;解决冲突和保护客户才是。编码是软件工程师的一项任务,而工作的目的在于解决已知问题、发现新问题;Huang 表示,没有什么比工程师少写代码、多解决问题更让他高兴。
自动化可以帮助填补工厂工人和卡车司机的缺口,Guo 则指出会计和护理也是 AI 可以缓解的短缺领域,尤其是在老龄化加剧的背景下。部署10亿台机器人还将需要 Huang 所说的“地球上最大的维修产业”,正如汽车催生了修理工,如今的 robotaxi 也需要维护团队和运营中心。
4. 完整 AI 技术栈让开源不可或缺
Huang 的组织框架是一块5层蛋糕:能源;芯片;软硬件基础设施;多样化 AI 模型;以及 OpenEvidence、Harvey、Cursor、自动驾驶或具身机器人等应用。AI 自动化的是智能,但它覆盖生物、化学、物理、金融和医疗信息,而不仅是人类语言对话。
闭源前沿实验室可以选择任何能够赚取足够回报、继续支持投资的商业模式。但如果没有预训练开源模型和可复用的推理技术,Huang 表示,创业公司、大学、研究人员,以及制造业或医疗领域拥有百年历史的企业,在来得及将 AI 适配到自身领域之前,“就会被扼杀”。
因此,他传递的政策信息非常具体:“无论你决定什么、做什么,都不要忘记开源。无论你决定什么、做什么,都不要忘记生物学。”限制可见的模型层,可能损害其下方更广泛的创新飞轮。
Huang 否定了一个单一巨型模型很快吞并所有领域的叙事。一个同时掌握人类、基因组、分子、蛋白质、氨基酸和物理语言的系统“根本不存在”;“God AI 下周不会出现”,但每个产业现在都需要切实的计算能力进步。
5. 末日叙事可能阻碍安全所需的技术
Huang 称关于大规模失业或全能 AI 即将到来的科幻式描绘“极其有害”,尤其当受尊敬的 CEO 和研究人员把这类观点讲给不了解技术的政府时。被问及动机是否是监管俘获时,他拒绝猜测,但表示,那些主张限制竞争对手的公司拥有“严重冲突”的利益。
他的历史反例是,过去两三年的快速投资带来了 grounding、推理和研究能力。人们担心的终局并未到来;系统反而变得更实用,也更能完成用户要求的事情。
Huang 的安全层级首先从性能开始:“安全的第一部分是性能”——汽车或模型必须按宣传运行,或许要达到99.999%的可靠率。合成数据、网络安全、监控、grounding 和偏差削减公司都说明,安全本身需要持续的技术投资。
他还反转了边际成本的质疑。如果 AI 的成本大幅下降,一个关键代理就不必独自运行;它可以被数百万个监控 AI 包围。成本下降可能让监督无处不在,就像降低“每个角落都部署警察”的成本。
6. Token 经济学侵蚀静态资本护城河
Elad 引用了团队成员的一项分析:2024年 GPT-4 等效推理成本下降了超过100倍。Andrej Karpathy 的开源项目把过去看似需要数十亿美元和一台超级计算机的工作压缩成一个周末的练习——不过讨论中纠正了“个人电脑”的说法,认为更接近 Spark。
NVIDIA 从 Volta、Ampere、Hopper、Blackwell 到即将推出的 Rubin,持续叠加架构改进、更多晶体管和更大产能。Huang 将每年5–10倍的计算能力提升称为常态,而 Moore 定律大约是每18个月提升2倍;10年下来,单靠 AI 硬件就可能提升10万倍至100万倍。
再叠加模型和算法改进,Huang 表示,10年内 token 生成成本下降10亿倍并不会让他意外。训练成本下降得没那么激进,但他称,如果一次训练需要1亿美元或5亿美元,“明年就会便宜10倍”。
主持人提出相反的观点:头部实验室可以把每次效率提升都投入规模扩大10倍的训练,继续保持规模优势。Huang 的回应是,计算负担并不会同比增加,因为硬件、训练算法、模型架构和跨实验室学习会共同累积,使落后6个月或1年的参与者仍能“保持接近”。
7. 研究与专业化扩大竞争范围
DeepSeek 是 Huang 眼中共享学习的最佳案例:他称其论文可能是硅谷研究人员过去几年读过的最重要论文,也可能是“去年对美国 AI 最大的单项贡献”。Guo 认为,这是多年来唯一一项让人感觉处于前沿的开源工作;Huang 则表示,它帮助了美国实验室、创业公司和基础设施公司。
Guo 引用了 Ilya 关于 AI 重新进入“研究时代”的表述,同时指出 scaling 仍在多个维度延续。Huang 预计模型会出现分化:一个用于编码,另一个面向消费者易用性,其他模型服务细分领域,因为企业已经不再需要“煮干整片海洋”。
预训练并未结束;Huang 的语义重点是,预训练是在为“真正的训练”做准备,而后者如今被称为 post-training。训练模型所需的数据可能很少——也许只需要可验证的结果——但算法训练仍然计算密集,并且可以打造无需掌握一切、却在特定领域极其出色的专业模型。
NVIDIA 自己采用 Cursor 的情况进一步印证了任务与目的的区别:每名工程师都在使用它,但公司仍在大量招聘。Guo 预测,编码将成为首个达到10亿美元ARR的 AI 原生应用业务,这同时反驳了“一种模型吸收一切”和“开发者工具只能做小生意”两种观点。
8. 当模型架构持续变化,可编程性保护 NVIDIA
Huang 回忆,市场曾反复声称专用 CNN 芯片、transformer 芯片或其他 ASIC 会让 NVIDIA 过时。专用硬件可以高效执行单一工作负载,但 transformer 本身正通过新型 attention 机制、diffusion、自回归、混合 SSM 等架构持续演进。
在 Moore 定律只能带来有限晶体管增量的情况下,Huang 希望从算法和架构中获得“数百倍”的提升。可编程平台能够吸收最终胜出的任何技术路线,而固定硬件则可能在快速变化的研究领域中优化了错误的阶段性方案。
兼容性也会激励研究人员只需针对庞大的存量平台优化一次:FlashAttention、SSM、diffusion、自回归、CNN 和 LSTM 都可以跨代运行。Huang 表示,NVIDIA 最新系统 NVL72 是全球成本最低的 token 生成器,“优势大到难以衡量”,尤其适合困难的推理工作负载。
9. 生物学、车辆与机器人正在接近应用突破
Huang 预计,多模态、长上下文、推理和合成数据将为数字生物学带来“ChatGPT 时刻”。蛋白质理解正迈向多蛋白表示和生成;他提到一个称为 LA prina 的开源模型,Guo 则指出 Chai 等公司正在推进端到端分子设计。
与人类语言相比,生物学的瓶颈在于真实世界数据稀缺,因此实验基础设施和合成数据十分重要。Huang 所期待的拐点,是出现一个蛋白质基础模型和另一个细胞基础模型,随后理解、生成以及相关数据飞轮可以共同加速。
推理能力也应当把汽车带出感知和规划阶段。Huang 将自动驾驶的发展脉络概括为:从智能传感器和人类设计的“数字轨道”,到感知、世界模型和规划的模块化系统,再到端到端模型,最终进入能够将陌生情境拆解为已知组成部分的推理系统。
他承认:“我们开始得太早了。”如果自动驾驶只在3年前启动,行业或许已经处在大致相同的位置。机器人应当发展得更快,因为这些基础已经具备,但人形机器人仍面临机电和安全问题,包括一台300磅机器人自身的重量、跌倒风险,以及与儿童互动的风险。Huang 还表示,NVIDIA 的自动驾驶技术栈刚刚获得全球最高安全评级,Tesla 位居第二。
10. 垂直服务商、能源与产能支撑反泡沫论
Huang 眼中的机器人市场远不止人形机器人:“一切会移动的东西都将成为机器人。”通用 AI 可以被装进汽车、挖掘机、履带式设备、单臂或六臂系统;软件平台可以服务多个垂直领域,设备专家则负责将其落地为真正可靠的产品。
消费级 AI 达到90%就能让用户惊喜,达到80%就能让用户满意;工业 AI 则会让用户把注意力全部集中在失败上。核心平台或许能达到99%,但垂直解决方案提供商必须将其推向99.99999%。最终的系统集成正是 Huang 未来5年判断的基础:“真正令人兴奋的将是垂直化。”
能源是眼下的物理约束。Huang 表示,如果不是政府扭转了“能源增长”叙事,美国就会把工业革命拱手让给别人。他希望电网和表后都获得更多能源,包括天然气、核能、风能和太阳能,但认为未来10年可能只有天然气能真正推动前进。Guo 表示,2027–28年的近期发电需求仍难以解决,但认为 AI 需求正在催化气候创新,包括新电池公司和太阳能聚光器。
谈到中国,Huang 预计2026年会更加建设性:中国既是对手也是伙伴,完全脱钩是“天真的”。他主张制定更细致的出口管制政策,同时兼顾国家安全、技术领导力和繁荣,并指出中国已经生产许多芯片,在军事和国家安全需求上也可以依赖 Huawei。主持人提出防火墙和历史上的就业结构变化;Huang 则从技术栈层面回应,认为中国互联网增长曾让 Intel、AMD、Micron、SK hynix 和 Samsung 受益,而中国的开源贡献也帮助了美国创业公司。
Huang 的反泡沫论同样从聊天机器人以下的层级出发:即使没有 OpenAI、Anthropic 或 Gemini,他仍认为,从 CPU 转向加速计算也足以让 NVIDIA 成为一家市值数千亿美元的公司。AI 随后叠加自动驾驶——NVIDIA 相关业务正接近100亿美元——以及金融服务、机器人和数字生物学领域的数十亿美元机会。
他的自外而内测算从全球100万亿美元 GDP 出发,并粗略假设其中2%,即2万亿美元,是每年的研发支出。当湿实验室、量化金融、车辆和其他研究活动转向超级计算机时,这些研发活动需要大规模基础设施;与此同时,在 Huang 看来,OpenAI 如果拥有2倍产能,收入可以翻倍;拥有10倍产能,收入可以增长10倍。
Elad 对一项被广泛引用的 MIT 研究提出质疑,该研究称大多数企业 AI 部署并没有多大用处:实施、工作流整合、组织重组和企业规划周期,都可能超过研究窗口。他会去观察3万–4万家创业公司——这些采用者行动更快——因为在研究人员和建设者之中,“所有人都快被产能需求逼疯了”。
Jensen, thanks so much for joining us today.
So great to have you guys. What an amazing year.
What a year.
Happy Hanukkah, Merry Christmas, and happy New Year coming up.
Yep. Happy holidays.
With everything that’s happened in 2025, and being in the middle of the vortex with it, what do you reflect on and say, “This surprised you most,” or, “This is the biggest change”?
Let’s see. There are some things that didn’t surprise me. For example, the scaling laws didn’t surprise me because we already knew about that. The technology advancement didn’t surprise me. I was pleased with the improvements in grounding. I was pleased with the improvements in reasoning. I was pleased with the connection of all of the models to search.
I’m pleased that there are now routers in front of these models, so that depending on the confidence of the answers, they can go off and do the necessary research and generally improve the quality and accuracy of answers. I’m hugely proud of that. I think the whole industry addressed one of the biggest skeptical responses to AI, which is hallucination and generating gibberish and all of that stuff.
I thought that this year, the whole industry—from every field, from language to vision to robotics to self-driving cars—made big, big leaps in the application of reasoning and the grounding of the answers. Would you guys say this year?
Huge. Things like OpenEvidence for medical information, where doctors are now really using it as a trusted resource, and Harvey for legal—you’re really starting to see AI emerge as one of these things that’s become a trusted tool or counterparty for experts to actually be able to do what they do much better.
That’s right. In a lot of ways, I was expecting it, but I’m still pleased by it. I’m proud of it. I’m proud of all of the industry’s work in this area. I’m really pleased and probably a little bit surprised, in fact, that token generation rates for inference, especially reasoning tokens, are growing so fast—several exponentials at the same time, it seems.
I’m so pleased that these tokens are now profitable, that people are generating them. I heard somebody say today that OpenEvidence, speaking of them, has 90% gross margins. Those are very profitable tokens.
Yeah.
They’re obviously doing very profitable, very valuable work. Cursor’s margins are great. Claude’s margins are great. For the enterprise use of OpenAI, their margins are great. Anyway, it’s really terrific to see that we’re now generating tokens that are sufficiently good—so good in value—that people are willing to pay good money for.
I think these are really great foundations for the year. Some of the things in the narrative, of course, the conversation with China, really occupied a lot of my time this year. Geopolitics, the importance of technology in each one of the countries—I spent more time traveling around the world this year than just about any time in all of my life combined.
My average elevation this year is probably about 17,000 feet, so it’s nice to be here on the ground with you guys. I think geopolitics and the importance of AI to all the nations are all worth talking about later. Of course, I spent a lot of time on export controls and making sure that our strategy is nuanced, really grounded, and promotes national security, while recognizing the importance of various facets of national security.
There were a lot of conversations around that. Of course, lots of conversation about jobs, the impact of AI, energy, and the labor shortage. We covered everything, did we? Everything was AI.
Everything was AI. It was incredible.
AI was definitely the center of the storm for every one of those themes. Maybe one we can start with is jobs and employment, because when I look at the traditional AI community—even before things were scaling and before AI was really working—there was a strong doomsday component among the people working on AI. Oddly enough, the people who were most trying to push the field forward were often the people who were most pessimistic, which is very odd. Why would you do both at once?
I feel like that narrative has taken over some subset of the media or other things, despite all the things that we think are very positive about what AI has done. It’s going to help with health care, education, productivity, and all these other areas.
In general, whenever we have a technology shift, you have a shift in terms of the jobs that are important, but you still have more jobs.
That’s right. Could you talk about how you think about employment and jobs—what people are saying and what you think the real narrative is there?
Maybe what I’ll do is ground it on 3 points in time: now, the very near future, and then some point out in the distance, along with maybe some counter-narratives.
Something else to think about with respect to jobs in the near term is that AI is not just AI—it’s software. But it’s not prerecorded software, as you know. For example, Excel was written by several hundred engineers. They compiled it, and it’s prerecorded. Then they distribute it as is for several years.
In the case of AI, because it takes into account the context, what you asked of it, what’s happening in the world, and contextual information, it generates every single token for the first time, every time.
Mm-hmm.
Which means that every time you use the software, and everything that we do, AI is being generated for the first time ever. Just like intelligence, our conversation today relies on some ground truth and some knowledge, but every single word is being generated for the first time here.
The thing that’s really unique about AI is that it needs these computers to generate these tokens every single time. I call them AI factories because they’re producing tokens that will be used all over the world.
Some people would say it’s also part of infrastructure. The reason why it’s infrastructure is because it affects every single application. It’s used in every single company, every single industry, and every single country. Therefore, it’s part infrastructure, like energy and the internet.
Because of that, and because of the number of computers necessary to generate these tokens—and because this has never happened before—we need these factories. Three new industries have emerged. Three new types of plants have to be created. Number one, we have to build a lot more chip plants.
Mm-hmm. TSMC is building, right? SK Hynix is building a lot more plants.
And so we need more chip plants. We need more computer plants. These computers are very different. These are supercomputers the world has never seen before. Grace Blackwell looks like a very different type of computer than anything that’s ever been made. An entire rack is one GPU.
And so we need new supercomputer plants. Then we need new AI factories. These 3 types of plants are currently being built in the United States at very large scale, quite broadly, all over the United States, for the very first time.
The number of construction workers, plumbers, electricians, technicians, and network engineers—the number of skilled laborers necessary to support this new industry in the near term will be enormous. Let’s just face it.
I’m so excited to hear that electricians are seeing their paychecks double. They’re being paid to travel, like us. We go on business trips; they’re going on business trips. It’s really terrific to see that these 3 types of plants and factories are creating so many jobs.
The next part is the near-term impact of AI on jobs. One of my favorite examples is—I love Geoffrey Hinton. He said 5, 6, or 7 years ago that, in 5 years’ time, AI would completely revolutionize radiology; that every single radiology application would be powered by AI; that radiologists would no longer be needed; and that the first profession he would advise people not to go into was radiology.
He’s absolutely right. One hundred percent of radiology applications are now AI-powered. That’s completely true, and in some 8 years’ time, it has now completely pervaded radiology. However, what’s interesting is that the number of radiologists increased.
And so now the question is why. This is where the difference between the task and the purpose of a job comes in. A job has tasks and a purpose. In the case of a radiologist, the task is to study scans, but the purpose is to diagnose disease.
And to research.
That’s exactly right. They’re doing research. In their case, the fact that they’re able to study more scans more deeply, request more scans, and do a better job diagnosing disease means the hospital is more productive. It can have more patients, which allows it to make more money, which allows it to want to hire more radiologists.
The question is: What is the purpose of the job versus what is the task that you do in your job? As you know, I spend most of my day typing. That’s my task, but my purpose is obviously not typing. So the fact that somebody could use AI to automate a lot of my typing—I really appreciate that, and it helps a lot.
It hasn't really made me less busy. In a lot of ways, I become busier because I'm able to do more work. I think the second part to consider is the task versus the purpose of the job. This example really strikes home because my sister-in-law, Erin, actually leads nuclear medicine at Stanford, right? She's in radiology, and with all the technology advancements that are coming—
These doctors really welcome it, and they are working 20 hours a day trying to do more research and serve more patients.
Exactly. I think one thing that is often missed, beyond the diversity of jobs being created by this investment in infrastructure, is actually how much latent demand there is for different goods that we need in society, like better healthcare. I don't think anybody feels like, "You know what? We've reached the tiptop, the mountaintop, of what American healthcare or global healthcare could be." The more we can make these people productive, the more demand there will be.
That's exactly right. If NVIDIA were more productive, it wouldn't result in layoffs; it would result in us doing more things.
I met your new-hire class today. You seem to be hiring every week, anyway.
That's exactly right. The more productive we are, the more ideas we can explore. As a result, the more growth—and the more profitable we become—which allows us to pursue more ideas. I think you're absolutely right that if the job, if your life, if the world—the problems—are literally already specified and there's no other problem to solve, then productivity would actually reduce the economy. But it's clearly going to increase the economy.
I think the next part that I would consider is that people say, "Gosh, all of these robots that we're talking about are going to take away jobs." As we know very clearly, we don't have enough factory workers. Our economy is actually limited by the number of factory workers we have. Most people are having a very hard time retaining their workers.
We also know that the number of truck drivers in the world is severely short. The reason for that is people don't want those jobs where you have to travel across the country and live in different parts of the country every single night. People want to stay in their town and stay with their families. I think the first part is that having robotic systems is going to allow us to cover the labor-shortage gap, which is really severe and getting worse because of an aging population. This is not only the United States; it's all over the world, as you guys know.
We're going to cover the labor shortage. But the second part that people forget—and, as a result, we'll go there—is that there are shortages in other places where people talk about AI being relevant. Accounting would be an example where there are shortages. Nursing is another example. You can go through multiple other industries and say, "Okay, there are gaps," right?
And AI is trying to help fill those gaps.
That's exactly right.
Automation is going to help us increase and solve the labor gap. Now, people also don't remember that when we have cars, we need mechanics to take care of our cars. If you look at the robotaxis that are even on the streets today, it's taken 10 years for that to happen. Look at all the maintenance crews and all of the various hubs where you have to take care of these robotaxis, and just imagine we have 1 billion robots.
Mm-hmm.
It's going to be the largest repair industry on the planet. I think a lot of people don't think this through.
This is the part where you said that when we create this type of automation, we create another job. Right now, AI is creating so many jobs. The AI industry is creating a boom of jobs.
I think one of the core challenges here is that it's very easy to draw a straight line of extrapolation from, "Oh, there are tools that help lawyers be more productive. They're going to replace the lawyers." But it takes an incremental step of reasoning to say there's a sucking sound in the economy for everything in AI infrastructure. There's actually a sucking sound toward all of this latent demand in the places where we have gaps, where a lot of policymakers have focused on, "We can't replace or reduce what we have," when there's really far more demand in what we actually are not—
And in the case of a lawyer, what's the purpose of the lawyer versus the task of the lawyer?
Reading a contract and writing a contract are not the purpose of the lawyer. The purpose of the lawyer is to help you resolve conflict, and that's more than reading a contract. It's more than writing a contract. The purpose is to protect you. That's more than reading a contract; it's more than writing a contract. I think it's really, really important to go back to what is the purpose of the job versus the task that we use to perform that job. That changes over time.
Yeah. The other big theme of the year that you mentioned, which I think is really important to touch upon, is both China and the rise of Chinese open source in particular. Some of the highest-scoring models against benchmarks now are Chinese models. On the open-source side, on the closed side, it's still a lot of the U.S. models, but things like Qwen and DeepSeek are doing very well.
You've long been a proponent of open source in general. Could you share your views about China emerging in AI and open source, and what the U.S. should be doing in terms of both open source and its own industries?
When you think about these complicated, interconnected, dependent networks of problems—this big goop, this mesh of problems—it's always good to go back and find a framework for what it is that we're talking about.
In the case of AI, what is AI? Of course, the technology and capabilities of AI are about automation. It's about the automation of intelligence for the very first time. You could combine it with mechatronics technology to embody that mechatronics and make it perform tasks. So that's what AI is: automation.
But what is the stack that makes AI possible? What's the technology stack, the functional stack? The easiest way to think about that is that it's kind of like a five-layer cake. At the lowest level is energy. It transforms energy into the output that I just described. The next layer is chips.
The next layer is infrastructure, and that infrastructure is both hardware and software. This is where land, power, and shell come in—this is where the construction of data centers happens—and the software stack for orchestrating them. The layer above that is what everybody thinks about, which is AI, the models.
We know this, but it's really helpful to understand that AI is a system of models. AI is a technology that understands information, and there's human information. We often think about AI as a chatbot, but remember, there's biological information, chemical information, physical information of all kinds, financial information, healthcare information—information of all modalities and all kinds.
Human language is at the foundation of many things, but it's not the essence of everything. Biology molecules don't understand English. They understand something else. Proteins don't understand English; they understand something else.
I think the next layer—the important thing—is where the AI models are. But AI is very, very diverse. The layer above that is applications, and it depends on the industry. You already mentioned OpenEvidence. You mentioned Harvey. There's Cursor. There's all kinds of applications. Full Self-Driving is really an application, an AI application embodied into a mechanical car, and Figure is an AI application embodied into a mechanical human.
You've got all these different applications. This five-layer stack is one way of thinking about it. The next way of thinking about it, as I just mentioned, is that AI is really diverse.
When you now have this framework of what the technology capabilities are, how to build the technology, and how diverse it is, you can come back and think about the question: How important is open source?
Without open source, today, of course, the frontier models—the leading labs—have chosen to use a closed-source application approach, which is just fine. What people decide to do with their business models is, in the final analysis, their business. They have to calculate the best way for them to get a return on investment so that they can scale up and make better advances. However they made that calculus is fantastic.
On the other hand, without open source, startups would be challenged. Companies in different industries, whether it's manufacturing, transportation, or healthcare, would be challenged. Without open source today, all of that AI work would be suffocated. They need to have something that's pretrained. They need to have some fundamental technology about reasoning.
From that, they could all adapt, fine-tune, and train their AI models into exactly the domain and application they want. What people really miss is the incredible pervasiveness and importance of open source to all of these industries. Large companies—some of the 100-year-old companies that I work with—would be suffocated without open source.
Open source at this point is driving all of our data centers. It’s driving a big chunk of telephony in the world, in terms of Android and other devices. It’s driving a lot of the industrial applications that you were talking about.
It’s already pervasive, and I think the big question is open source.
Without open source, higher ed wouldn’t happen.
Education, research—
Startups—the list goes on. We talk all day long about the tip, the most visible part of it, the part that’s most newsworthy maybe, but underneath that is such an important space of open-source AI. Whatever we decide to do with policies, do not damage that innovation flywheel.
I spend a lot of time educating policymakers to help them understand: Whatever you decide, whatever you do, don’t forget open source. Whatever you decide, whatever you do, don’t forget biology.
I think the counter-narrative here that is worth addressing is that there should essentially be a monolithic vertical player and a monolithic asset—a single model that does it all—and that we can’t give away that crown jewel to other countries or non-American companies. Your argument is that we actually need this huge diversity of AI applications, and the American advantage—or any sovereign advantage—is actually in the whole stack, right? It’s the capability to deliver any piece of it.
I guess someday we will have God AI.
But when is that day?
That someday is probably on biblical scales—galactic scales. I don’t think it’s helpful to go from where we are today to God AI. I don’t think any company practically believes they’re anywhere near God AI, nor do I see any researchers having any reasonable ability to create God AI.
The ability to understand human language, genome language, molecular language, protein language, amino acid language, and physics language all supremely well—that God AI just doesn’t exist.
And yet we have a lot of industries that need AI.
AI is, if you will, at the simplistic level, just the next computer industry. Give me an example of a company, an industry, or a nation that doesn’t need computers.
Mhm.
We all don’t have to wait around for God AI for us to advance, right? God AI is not showing up next week. I’m fairly certain of that. God AI is not going to show up next year, but the whole world needs to move forward next week, next year, and next decade.
I think the idea of a monolithic, gigantic company, country, nation, or state that has God AI is unhelpful. It’s too extreme. If you want to take it to that level, then we ought to just all stop everything. What’s the point of even having governments? Why are they doing policies? God AI is going to be smart enough to avert or work around any policy, so what’s the point?
We ought to bring things back to ground level and start thinking about things practically and use common sense.
This seems to be a big theme in general in this conversation. There’s been a lot that’s been put out there that seems very extreme if you actually think about it: jobs and employment, nobody being able to work again, God AI solving every problem, or saying we shouldn’t have open source for some reason despite open source already powering much of our industries.
That’s right.
One of the themes of 2025 seems to be that a lot of extremes were painted in public around AI that, if you look at them very closely, don’t really follow a logical path toward happening anytime soon.
Yeah.
It sounds like it’s really important to have this conversation.
It’s extremely hurtful, frankly. I think we’ve done a lot of damage with very well-respected people who have painted a doomer narrative, an end-of-the-world narrative, or a science-fiction narrative. I appreciate that many of us grew up with and enjoyed science fiction, but it’s not helpful. It’s not helpful to people, the industry, society, or governments.
There are many people in government who obviously aren’t as familiar or as comfortable with the technology. When PhDs in this and CEOs of that go to governments and explain these end-of-the-world scenarios and extremely dystopian futures, you have to ask yourself: What is the purpose of that narrative? What are their intentions, and what do they hope? Why are they talking to governments about these things—to create regulations to suffocate startups?
For what reason would they be doing that? Do you think that’s just regulatory capture, where they’re trying to prevent new startups from showing up and being able to compete effectively? What do you think is the goal of some of these conversations?
I can’t guess what they have in mind. I know that the concern is regulatory capture. As a policy or as a practice, I don’t think companies had to go to governments to advocate for regulation of other companies and other industries. In practice, their intentions are clearly deeply conflicted, and they’re clearly not completely in the best interest of society. They’re obviously CEOs, they’re obviously companies, and they’re obviously advocating for themselves.
I think if we can all come back to where we are today and think about where the technology is going to be, then literally, in 1 year’s time—as we were talking about in the beginning—some of the proudest moments are when the industry was able to invest very aggressively in advancing AI technology instead of being slowed down.
Remember, just 2 years ago people were talking about slowing the industry down. But as we advanced quickly, what did we solve? We solved grounding, reasoning, and research. All of that technology was applied for good, improving the functionality of AI.
Yet the end has not come.
Yet the end has not come. It’s become more useful, more functional, and more able to do what we ask it to do. The first part of the safety of a product is that it performs as advertised.
The first part of safety is performance. The first part of the safety of a car isn’t that some person is going to jump into the car and use it as a missile. The first part of the car is that it works as advertised—99.999% of the time, working as advertised.
Mhm.
It takes a lot of technology to make that car or make that AI work as advertised. I’m really glad that in the last 2 or 3 years, the industry has invested so much in enhancing the functionality of AI as advertised.
I think if we were to look at the next 10 years, we have so much work to do to make it work as advertised. Meanwhile, as you know, you both invest so much in the ecosystem. You see so many companies being built for synthetic data generation so that the AIs could be more grounded, more diverse, less biased, and safer. You’re investing in a whole bunch of companies in cybersecurity using AI for cybersecurity.
People think that because the marginal cost of AI is going to go down significantly, AI is going to be dangerous. It’s exactly the opposite. If the marginal cost of AI is going to go down significantly, one AI is going to be monitored by millions of AIs.
Mhm.
More and more AI is going to be monitoring each other. People can’t forget that an AI is not going to be an agent by itself. It’s likely that an AI is going to be surrounded by agents monitoring it.
It’s no different from saying that if the marginal cost of keeping society safe were lower, we would have police on every corner.
One thing that we were talking about a little bit earlier was just the cost of AI and how it’s been coming down. In 2024, the cost of GPT-4-equivalent models, if you look at a million tokens, came down by over 100×. Somebody on my team did this analysis to show that. The costs are dropping pretty dramatically and very rapidly, partly because of all the advancements you’ve been driving at the NVIDIA level, but also because of efficiency gains across the stack.
At the same time, model companies are talking about how the costs are rising, how there are enormous capital moats to building these things out. How do you think about the cost of training and the cost of inference over time, and what does that mean for the average end user or the average startup company trying to compete or people trying to do more in this industry?
I forget the statistic, but Andre Karpathy estimated the cost of building the first ChatGPT, I think.
Versus now, I think you could do that on a PC.
Yeah. Yeah. It's probably tens of thousands of dollars at this point, or maybe even less.
Right. And so it costs nothing.
Mhm.
He has an open-source project that you can do in a weekend.
Oh, is that right? Okay. That's incredible. Right. We're talking about 3 years.
Mhm.
Mhm.
What people said cost billions of dollars—supercomputers built, raising billions of dollars in order to do all that—now costs something that you can do on a weekend on a PC.
Or a Spark—sorry, probably not quite a PC.
Okay. Not quite a PC. Yeah. We're improving our architecture and performance every single year. The first GPT, I think, was trained on Volta. And then Ampere, and I think the first breakthroughs, none of it included Hopper.
Mhm.
Of course, Hopper has been around for the last 2 or 3 years, and we're on Blackwell for the last year and a half or so. Every single one of these generations, the architecture improves, and of course the number of transistors and the capacity go up every single generation, very easily every single year from a computing perspective. The combination of all that getting 5 to 10x every single year is not unusual. And here comes Rubin just around the corner.
We're seeing 5 to 10x every single year. Compounded, it's incredible. Moore's law was 2x every year and a half, and over the course of 5 years, that's 10x; over the course of 10 years, that's 100x. In the case of AI, over the course of 10 years, it's probably 100,000x to 1,000,000x. And that's just the hardware.
Mhm.
Then the next layer is the algorithm layer and the model layer. The combination of all that—the fact that if you were to tell me that, in the span of 10 years, we're going to reduce the cost of token generation about a billion times—I would not be surprised.
Okay. And so that's the tokenomics of AI. On the training side, it's not quite as aggressive in cost reduction, but it's close. If you were to say that every single year we're increasing by 2 or 3x, over the course of 10 years, that's incredible. But the important idea is, when somebody says it cost $100 million to train something or half a billion dollars to train something, well, next year it's 10 times less. Next year, it's 10 times less.
For people to scale these things up, though, right? So the counterargument is, well, we'll just get bigger every year by 10x or 100x, or we'll try to offset that decrease in cost by scale.
And others can't keep up. Yeah. But really, what's happening is—and this is where the economics come in, as you know—the scale went up by a factor of 10, but the computational burden did not go up by a factor of 10 because you're getting the compounded benefits of all 3 things. The hardware is going up, the algorithms of the training models are going up, and of course the model architecture is going up, and we're getting the benefit of learning from each other.
This is, let's face it, DeepSeek was probably the single most important paper that most Silicon Valley researchers read in the last couple of years.
It was the only thing that felt frontier that was open in years. The value of open source is, again, putting out these papers.
That's right. Literally, DeepSeek benefited American startups, American AI labs all over, and infrastructure companies all over. It was probably the single greatest contribution to American AI last year.
If you said this out loud, of course, people would kind of shudder that American AI is actually learning from and benefiting from AI from other nations. But why would that be surprising? AI researchers all over America are Chinese natives and come from different countries. We benefit from every country. We benefit from every researcher, and not all of the world's ideas have to come from the United States.
So I think, back to your original question, it is the case that some of the narratives around the cost of AI are about scaring everybody out of the market: nobody ought to do pre-training but us; nobody should do training these frontier models but us. Because of innovation in models, algorithms, and the computing stack, the cost of AI is actually decreasing by well more than 10x every single year. And so if you're just 1 year behind, or even 6 months behind, you could really stay close.
And I think one thing that felt very different to me about 2025 is Ilya said recently that we're in the age of research again versus an age of scaling. I think both things are happening, by the way. Everybody is also trying to scale on multiple dimensions.
Yeah, exactly. Both are happening.
Being 6 months behind, or being at a 100K versus a 200K cluster, I think matters if you are competing symmetrically. But now you have people from frontier labs, or at the very top of the game, who have very different ideas about how to progress from here or who are working on a diversity of problems. And I think that felt different from 2024, maybe, where there was a lot of energy focused on just pre-training scale and LLMs.
And several other dynamics. As the market grows, each one of these models could choose to have verticals or segments where they want to differentiate. Somebody could decide to be better at coding. Somebody could decide to be easier to access so that it could be a greater consumer product.
The diversity of these models means that you could probably make a niche leap without having to be great at everything else and still be super valuable to the market. It's no longer necessary to boil the entire ocean.
Two years ago, because it was called pre-training, people said, "Pre-training is over." First of all, pre-training is not over. But the point of pre-training is to train yourself for training. That's why it's called pre-training: to prepare yourself to do the real training. And now we call it post-training. It's kind of weird. I think it's just training, but pre-training is pre-training and therefore it's training.
Training, as we all know, is where compute scaling directly translates to intelligence. The data necessary to train a model is actually pretty small. Maybe it's just the verifiable results. Now it's really algorithmic and very compute-intensive.
You don't have to be good at everything in life, as you know. Just like all of us, we don't have time to learn everything equally well. We decide to choose a specialty and focus all of our energy on it, and we become superhuman or incredibly good at something that other people are not. And so I think AI labs are going to start doing the same. They're going to start bifurcating into various segments, and over time you're going to see startups do the same. They'll find a micro-niche and take something open and then be incredibly good at it.
Well, I think one of the most optimistic views here is actually that these micro-niches are quite valuable, right? I was talking to Andrej because I've been talking to a lot of people about their predictions for next year. We'll ask you yours as well, of course. But he asked, "What is an example of a prediction that would have been prescient last year?" And my answer—everything's easy in retrospect—is that coding would be the first application-level business that gets to $1 billion of ARR as an AI-native app, right?
And I think if you had taken an old-world view of this, you would have believed one of 2 narratives. One is a single model does everything, and it'll all just be subsumed into something monolithic.
And 2 is that developer tools never get very big, right? Well, it kind of depends on how valuable the developer tool is. Now, I think many more people understand that software engineering is a niche and there's more demand than ever for it, but I think we'll see more like that.
Also interesting, we use Cursor here, and we use Cursor pervasively here. Every engineer uses it, and the number of engineers—we just mentioned it—the number of people we're hiring today is just incredible.
Yep.
Right. Monday is "Come to Work at NVIDIA Day," and why is that? This is now the purpose and the task.
The purpose of a software engineer is to solve known problems and to find new problems to solve. Coding is one of the tasks. And so if the purpose is not coding—if your purpose literally is coding, somebody tells you what to do, you code it—all right, maybe you're going to get replaced by AI. But all of our software engineers, their goal is to solve problems.
And it turns out we have so many problems in the company, and we have so many undiscovered problems. And so the more time they have to go explore undiscovered problems, the better off we are as a company. Nothing would give me more joy than if none of them were coding at all. They're just solving problems.
You see what I'm saying? And so I think that this framework of purpose versus task is really good for everybody to apply. For example, somebody who's a waiter: their job is not to take the order. That's not their job.
As it turns out, their job is to help us have a great experience. If an AI is taking the order or even delivering the food, their job is still to help us have a great experience. They would reshape their jobs accordingly.
I think the question about the cost of compute is really important. Let me come back to why we are so dedicated to a programmable architecture versus a fixed architecture. Remember, a long time ago, a CNN chip came along and they said NVIDIA was done. Then a transformer chip came along and NVIDIA was done. People are still trying that.
Yes.
Yeah. NPUs—and the benefit of these dedicated ASICs, of course, is that they can perform a job really, really well. Transformers are a much more universal AI network, but the space of transformers, as you know, is growing incredibly: the attention mechanism and how it thinks about context, diffusion versus autoregressive models, and these hybrid SSM transformers. For example, Nemotron—we just announced a new hybrid SSM.
The architecture of transformers is changing very rapidly, and over the next several years, it’s likely to change tremendously. We dedicate ourselves to an architecture that’s flexible for this reason, so that we can adapt.
Remember, because Moore’s law is largely over, the transistor benefit is only 10%, maybe once every couple of years, and yet we would like to have hundreds of × every year. The benefit is actually all in algorithms, and an architecture that enables any algorithm is likely going to be the best one, right? The transistor didn’t advance that much.
I think our dedication to programmability is, number one, for that reason. We have so much optimism for innovation in algorithms and iteration in software that we protect our programmability for that reason.
The second thing is that by protecting this architecture, our installed base is really large. When a software engineer wants to optimize their algorithm, they want to make sure that it doesn’t run on just one little cloud or one little stack. They want it to run on as many computers as possible.
The fact that we protect our architecture compatibility means FlashAttention runs everywhere, SSMs run everywhere, diffusion runs everywhere, autoregression runs everywhere. It doesn’t matter what you want to do. CNNs still run everywhere. LSTMs still run everywhere.
This architecture, which is architecturally compatible, gives us a large installed base that’s programmable for the future. That’s really important in the way that we help to advance the field.
As a result, all of this drives the cost down. I’m super proud that our latest innovation, the NVL72, is the lowest-cost token-generation machine in the world by enormous amounts. The reason for that is that inference is really, really hard.
People didn’t expect that. For us, it’s probably easier to train, but for inference, it’s incredibly hard to generate tokens. As costs drop, you usually open up new applications or new verticals that become more and more accessible.
We talked a little bit about coding—Cursor, Cognition, and other companies that have really benefited from that in the last year. Do you have any thoughts or predictions in terms of what the next breakthrough industries will be, or new applications or areas you’re most excited about coming in 2026 in particular?
Because of 2 or 3 things, I think several industries are going to experience their ChatGPT moment. I believe that multimodality and very long context are going to enable really, really cool chatbots. But that basic architecture, in combination with breakthroughs in synthetic-data generation, is going to help create the ChatGPT moment for digital biology.
That moment is coming.
By digital biology, do you specifically mean other aspects of protein folding, protein binding, or protein diagnosis? I see proteins.
I think we’re good at protein understanding.
Mm-hmm.
Now, multiprotein understanding is coming online. We recently created a model called LA prina. It's open. It’s for multiprotein understanding, representation learning, and generation.
I think protein understanding is advancing very quickly. Protein generation is going to advance very quickly. ChatGPT moment for proteins.
There are a lot of interesting companies working on molecule design in an end-to-end way, like Chai.
Exactly. Then, of course, chemical understanding and chemical generation, and then protein-chemical conformation understanding and generation.
Is that right?
That combination—the ChatGPT moment, the generative AI moment—all of that stuff is coming together for digital biology.
To your point about new industries, the way I think about it is investing in the inputs for this AI as well. All of these things around biology, chemistry, and materials science require real-world data generation and experimentation, right? That’s new infrastructure too.
New infrastructure. Synthetic data is going to be really important because they just have such sparsity of data, and they don’t have as much as human language. The real breakthrough is going to be when we can train a world foundation model—a foundation model for proteins, a foundation model for cells.
I’m very excited about both of those things. Once we have a foundation model, our understanding capability and our generative capability—the data flywheel—is really going to take off.
The second area that I’m excited about is robotics. Reasoning made huge breakthroughs in language, but because of reasoning, cars are going to be able to perform better. Instead of just perception cars and planning cars, they’re going to be reasoning cars. These cars are going to be thinking all the time.
When they come up to a circumstance they’ve never encountered before, they can break it down into circumstances they have encountered before and construct a reasoning system for how to navigate through it. The out-of-domain, out-of-distribution part of AI is going to be very much addressed by reasoning systems.
As a result, we could do more things than we were taught to do. Between generative AI, multimodal vision-language-action models, and reasoning systems, I think we’re going to see big breakthroughs in humanoid robots or multi-embodiment robots.
What do you think is a time frame for that? If you look at the self-driving analogy, self-driving technologies were based on very different types of neural networks than what we’re using today. There’s been a big switchover over the last 2 or 3 years in terms of how we do a lot of that.
We started too soon. Self-driving cars really had 4 eras. The first era was smart sensors connected into a car—the Mobileye era—and even the earliest days of Waymo.
Yeah. You’re talking about using smart sensors, a lot of human-engineered algorithms, and extensive mapping, and then different systems for planning and perception.
Exactly. You’re essentially creating a car that is driving on digital rails, right? It’s no different from the rails at Disneyland. There are digital rails.
So that’s the first generation. The second generation: during that generation, you have perception, a world model, and planning. These modules each have the limits of their technology. Perception was first affected by deep learning, and then it propagated through the pipeline.
Mm-hmm.
That system was too brittle, and it only knows how to perform what you taught it. Now, where we are, are end-to-end models, and where we’re going to go next are end-to-end reasoning models. There you go. Those are the 4 eras, in a lot of ways.
If we had started self-driving cars 3 years ago, we would probably be in exactly the same place.
All our poor friends who were working in self-driving.
I don’t mind it. I’ve been working on it for 10 years. NVIDIA’s self-driving car stack, by the way, has the number-one-rated safety in the world today. We just got that rating last week. Number 2 is Tesla.
I’m very proud that 2 American companies are at the top.
So, from a robotics perspective, do you think that because we’ve already built all these sorts of technologies in the modern era, robotics won’t take the same 10 or 15 years?
That’s right. I’m much more optimistic about robotics because we’ve kind of advanced the foundational technology.
Now, people are thinking about humanoid robotics. Humanoid robotics has a lot of challenges. There are all the mechatronics challenges. For example, it’s not helpful if the robot weighs 300 pounds. What happens if it falls over? What about interacting with kids, and so on and so forth?
And so, you have all kinds of challenges to deal with. I'm certain that we're going to solve those. But remember, the fundamental technology that goes into a humanoid robot can go into a pick-and-place robot.
One thing I've been curious about for robotics in particular is: if I look at who won, or who's perceived as winning, in self-driving, it's largely incumbents, right? It's Waymo, it's Tesla. You mentioned the safety rating NVIDIA's gotten. And so it's people who've been working on this for a long time.
It took a lot of capital; it was really intensive to get there. You have supply chain, hardware, and all this extra complexity. Do you think the same thing will be true in robotics? Are the winners basically going to be Tesla with Optimus and other people who have both been in the industry for a while but also have all those incumbent effects? Do you think there's room for startups?
Tesla will be one of the leaders—one of them, and surely a major one. But everything that moves will be robotic. Everything that moves is a very large space. It's not all humanoid robots. And yet every AI will be multi-embodiment.
Just like a human, we're multi-embodiment AI ourselves: we could sit in a car and embody that; we could pick up a tennis racket and embody that; we could pick up a chopstick and embody that.
People are general-purpose, right? They can do all these things.
Exactly. And so AIs are going to become general-purpose. You have one arm doing pick and place, maybe it's 2 arms doing pick and place, maybe it's 6 arms doing pick and place. So I think you're going to have all kinds of different sizes and shapes. It could be a caterpillar. It could be an excavator. It could be all kinds of stuff.
AI will embody those, just as a construction worker embodies an excavator, embodies a tractor.
Could there be a small number of companies then that do the embodiment for everything, or are you saying more that there are going to be niche applications?
You should definitely see a lot of software companies, and then that software company could serve a lot of different verticals. But each one of the verticals will still have solution providers that ground it all and turn it into something that works perfectly. Does that make sense?
Because in the case of AI for consumers, if it works 90% of the time, you're delighted—you're mind-blown. If it works 80% of the time, you're satisfied. In the case of most industrial and physical AIs, if it works 90% of the time, nobody cares about that. They only care about the 10% that it fails—basically, 100% dissatisfaction. And so, you've got to take it to 99.99999%.
The core technology might be able to get you to 99%.
And then a vertical solution provider like Caterpillar or somebody could take that core technology and make it 99.999% great. Do you think that's what happens earliest on? Because in markets that are this immature, it seems one of the fastest paths to market could be full verticalization, right? You just have control of iteration speed.
The difficulty of verticalization for technology that is general-purpose is that you don't have the R&D scale to build a general-purpose technology. Now, of course, open source helps that tremendously, which is the reason why you're going to see a big surge of vertical opportunities in AI in the next several years.
My prediction would be, over the course of the next 5 years, the excitement is going to be verticalization.
Notice we're excited about OpenEvidence, we're excited about Harvey, we're excited about Cursor. Cursor is horizontal, but it's kind of a horizontal vertical.
And so I'm super excited about all the verticals. A lot of people said, "AI is going to get so good that all these wrapper companies are going to be obsolete." It just misses the big point.
The reason why somebody can talk about the life of a surgeon is because they've never been a surgeon. The reason why somebody who builds AI can talk about the life of an accountant and a tax expert is because they've never been a tax expert. The reason why somebody could talk about being a busboy without being a busboy is that they've never been a busboy.
And so I think you've got to be a little bit more empathetic about the depth of the complexity of the work and try to truly understand the purpose of the work. Oftentimes, the technology addresses the task; it doesn't address the purpose.
So, I guess one of the other narratives—from the narratives we're looking at that are true versus not true for 2025—one other narrative that's come up has been more about energy and energy utilization, and whether we'll have enough energy to support AI. How do you think about that?
In the first week of President Trump's administration, he said, "Drill, baby, drill." He got so much flack for that. If not for this entire change in sentiment about energy growth in our country—
We can all concede now that we would have handed this industrial revolution to somebody else. And we're still power-constrained.
We're still power-constrained. Yeah.
Without energy, there can be no new industry.
Mhm.
And of course, we've been energy-starved now for what, a decade? If not for the fact that President Trump reversed that narrative, we would be completely screwed.
Without energy, you can't have industrial growth. Without industrial growth, the nation can't be more prosperous. Without being more prosperous, we can't take care of domestic issues. We can't take care of social issues, and on and on and on.
So the fact of the matter is, we need energy to grow. We need every form of energy. We need natural gas. We need more energy on the grid; we need more energy behind the meter. We're going to need nuclear. Wind is not going to be enough. Solar is not going to be enough. Let's just all acknowledge that we'll take it—we'll take everything we can.
But the fact of the matter is, I think, for the next decade, natural gas is probably the only way to go forward.
What's really interesting is, I agree the timeline is too far out to address people's power-generation issues in 2027 and 2028, where large players building clusters are very concerned. But the biggest drivers of climate innovation in the U.S. have actually been a result of this AI infrastructure problem, because people look at the demand—
Finally. That's right: demand.
They look at the demand, and the demand is driving people to create massive new battery companies, solar concentrators. It's put new energy—new energy, like, you know, willpower—behind—
The AI industry is driving all of that sustainable-energy industry.
Yeah. Because people see that there is going to be demand for it, right? So even if—and I think there is no practical answer in the small-number-of-years time frame versus natural gas, right?—it still drives climate innovation.
Yeah, no question about it. No question about it. And I think that's exactly right: doomer messages cause policy, and that policy may affect the industry in some way. But there's nothing more powerful than demand.
Look at all the jobs being created. Look at all the industries being formed around it. Sustainable energy, likely—and when history rewrites it, Sarah, I think you're going to be absolutely right that, if not for AI—well, AI is probably the biggest driver for sustainable energy ever.
A friend of mine has a saying: doomers are the people who sound smart at dinner parties, and optimists are the people who drive humanity forward. And I think that's very true for all these things we've talked about. Yeah, so—
Yeah, it's really true.
Yeah. Well, that's one of the big takeaways from this last year: the battle of narratives.
And it's too simplistic to say that everything the doomers are saying is irrelevant. That's not true. A lot of very sensible things are being said. It is too simplistic to say that when somebody is optimistic, they're just naive.
It needs to be grounded in reality. Yeah, that optimistic people are just naive, you know—
And that's obviously not true.
But I think we just have to be mindful of the balance of it. When 90% of the messaging is all around the end of the world and doom and pessimism, I think we're scaring people from making the investments in AI that make it safer, more functional, more productive, and more useful to society.
So we just need to make sure that it's more secure. All of that takes technology. Security takes technology. Safety takes technology. I appreciate that my car is safer today because it has better technology than a car 50 years ago.
And so I think it takes technology to be safe, technology to be secure. I'm delighted to see that the advancement of technology is still accelerating and ongoing. We just have to make sure that the policymakers around the world, the governments, are able to think about balancing these 2 ideas.
So, I guess we've talked a lot about 2025 and the narratives of 2025. How do you think about 2026? What are you excited about? What do you see coming? What do you think are big changes that we should be aware of?
I am optimistic that our relationship with China will improve.
Mhm.
President Trump and the administration have a really grounded, common-sense attitude and philosophy about how to think about China: They're an adversary, but they're also a partner in many ways. The idea of decoupling is naive, and the idea of decoupling, for whatever reason—philosophical reasons or national security reasons—is not based on common sense. The more deeply you look into it, the more you see that the 2 countries are actually highly coupled.
Both countries ought to invest in their own independence. When you depend too much on someone, the relationship becomes too emotional, as you know. [laughter] So it's good to have some independence, or as much independence as either would like, but to recognize that there's a lot of coupling, a lot of dependence between the 2 countries. I think there needs to be a nuanced strategy, a nuanced attitude about how to manage this relationship in a productive way for all the people of both countries and for all the people around the world.
Everybody depends on a productive, constructive relationship between the 2 most important nations, and the single most important relationship for the next century. We have to find that answer. I'm delighted that President Trump is looking for a constructive answer. I think next year will be a much better year than the last several.
I'm happy with what the administration was able to suggest: an export-control policy that is grounded in national security, recognizing that they already make so many chips themselves and can depend on Huawei themselves for their military, for their national security. They have ample technology to do that. American technology, although general-purpose, is unlikely to be used by their military because their military is too smart, just as our military is too smart to use their technology. It's grounded in national security, technology leadership, and national prosperity.
One of the things we always have to remember is that the world's mightiest military is supported by the world's mightiest economy. The wealth that we generate brings jobs home, creates prosperity in the United States, provides tax revenues, and ultimately funds the mightiest military on the planet. That circular system, that interconnected system, requires a nuanced strategy. I'm pleased to see some of the progress in that area that allows American technology companies to keep America first and keep America ahead, and to support American technology leadership, on the one hand, to win globally.
And then China, of course, is sorting itself out—well, not sorting itself out, but sorting out its attitude about how to think about American technology. The historical argument has been that, if you look, for example, at the internet, there was what was known as the Great Firewall, right? China basically prevented U.S. competition from entering China, while the opposite wasn't as true. There was mass expatriation of U.S. jobs and industry to China as part of the development of the 1990s and 2000s.
I think a lot of the things that people have brought up from a China-U.S. policy perspective, besides just the military adversarial relationship, spheres of influence, and all the various things like that, also include the economic imbalances that are perceived to exist between the 2 countries.
The way that I would think through that is to go back to the first principles of technology again. Let's say the internet: You have the chip industry, the systems industry, the software industry, and the services industry on top. Remember, China's internet growth has been a boon for Intel and AMD selling CPUs, Micron selling DRAM, SK hynix and Samsung selling DRAM. It is the second-largest internet market for the American technology industry.
Maybe it wasn't helpful to some layer of the stack—the Googles of the world—but don't exclude every layer of the stack. Always come back, every single one of these things, and take a step back and look at the whole stack. Maybe that's a theme for today as well. Technology is actually not just the internet software application layer that's been very dominant for 2 decades; it's the whole stack.
Remember, as Intel and AMD prospered with the internet industry in China's growth—the China industry growth—don't forget China also contributed tremendously to open source. No country in the world contributes more to open source than China. Look at all the startups here in America that were able to benefit from that open source to create the new startups that are here. You can't look at one area in isolation. You have to look at the whole life cycle of the technology and look at every layer of the stack. Does it make sense?
China's internet industry generated enormous prosperity for America.
Mhm.
Just not at the internet company per se.
Jensen, my other investor friends will not forgive me if I don't ask you about 2026 on the business side. Are we in an AI bubble?
AI bubble. Yeah, there are a lot of ways to reason through that.
When asked that question, my mind goes to: What is AI, and where are we in that? There's AI, then there's computing. As you know, NVIDIA invented accelerated computing. Accelerated computing does computer graphics and rendering; AI doesn't. Accelerated computing does data processing, SQL data processing; AI doesn't. Accelerated computing does molecular dynamics and quantum chemistry; AI doesn't. All these are things that people could say someday AI will, but it doesn't today.
Accelerated computing is really essential for classical machine learning, XGBoost, recommender systems, the whole process of feature engineering, extract, load, and transform. That entire data science and machine-learning life cycle uses accelerated computing.
The first thing to go to is, in the context of NVIDIA, what we see is the shift from general-purpose computing to accelerated computing because Moore's law has largely ended. You can't use CPUs for everything anymore like you used to, and so it's just no longer productive enough. It's not deflationary enough. We have to move toward a new computing model, and that's where accelerators come in.
If generative AI—excuse me, if chatbots, let's just go with OpenAI, Anthropic, and Gemini—if none of that existed today, NVIDIA would be a multihundred-billion-dollar company. The reason for that is because, as you know, the foundation of computing is shifting to accelerated computing. That's the first thing to realize: Take a step back and ask yourself what is actually happening now.
Now, the next layer up, the question about AI becomes: What is AI? We ask the AI bubble question, and we always go back to OpenAI's revenues, 100%, don't we?
Mhm.
You ask somebody, “Hey, is there an AI bubble?” Everybody goes directly to OpenAI's revenues. First of all, if OpenAI currently has twice the capacity, its revenues would double. You guys know that if they have 10 times the capacity, their revenues would be 10 times greater. I really believe their revenues would be 10 times greater. They need capacity.
This is no different from NVIDIA needing wafers from TSMC. Just because NVIDIA exists and we're doing great doesn't mean we don't need capacity. We need capacity. We need capacity of DRAM. In our world, it's sensible to everybody: We need capacity. Well, in their world, they need factories.
If they don't have factory capacity, how do they generate tokens, which is where we started our conversation today? They need factory capacity in order to increase their revenue growth.
Nonetheless, we also said that AI is more than chatbots. It includes all these different fields of science. NVIDIA's AV business is coming up on $10 billion. Nobody ever talks about that. You have to train world models. You have to train these AI AVs, and it's happening—robotaxis are happening all over the world.
Our AI work with digital biology, our AI work in financial services—the whole industry of quants, quantitative trading, is moving toward AI. They used to be classical machine learning. A whole bunch of humans—they call them quants, right? These specialized mathematicians were trying to figure out what the predictive features are. Now we use AI to figure it out.
Financial services is one of our fastest-growing segments. Billions of dollars in quants, in financial services; billions of dollars in AV; billions of dollars in robotics coming up; billions of dollars in digital biology. And so how big can all that be? Well, simple logic is this—simple math. Whether you think that AI is going to replace a labor shortage or workforce shortage of any kind, let's ignore that for a second.
The world is at $100 trillion in GDP. Out of that, let’s just say 2% annually is R&D. Let’s go back 5 years ago: if you were to take the largest drug company in the world, where were all of its R&D wet labs? Today, what are they doing? Building supercomputers.
There’s a fundamental shift in how they think about that $2 trillion. It used to be $2 trillion for the old way of doing things. It’s now going to be $2 trillion in the AI way of doing things. Well, $2 trillion is going to need $2 trillion of R&D, and that R&D is going to be powered by a whole bunch of infrastructure.
That’s the reason why we’re building supercomputers everywhere around the world. I think if you reason about it from the outside in—from the foundation up, from the outside in—you come to the conclusion that what we’re experiencing, what all 3 of us are experiencing, is that the amount of computing demand is insane.
Give me an example of a startup company that goes, “No, we’re good.”
They are all dying for computing capacity. Give me an example of a researcher in any university, a scientist in any company, who says, “Got plenty of capacity.” Everybody is dying for capacity.
We have a global, multicompany, multi-industry shortage. It’s not just about OpenAI, even though OpenAI could use a lot more capacity as well. I think the narrative is not helpful, and it’s a little too superficial to say, “How do you prove there’s an AI bubble? $12 billion of revenues, hundreds of billions of dollars of infrastructure being built.” That’s a little too simplistic.
Yeah. The other thing people tend to point out is the MIT study. There’s some study that I think came out of MIT that claimed most enterprise deployments of AI weren’t that useful. And you’re like, well, did you do the change management? Did you do a reorg? Did you integrate it into tooling? How long did it even take to implement it?
If a planning cycle in an enterprise is a year and something is implemented in 6 months, it feels like there are a lot of these overstated things that get a lot of attention, but then you map it against what’s actually happening.
Yeah.
And the growth of these companies using AI is just a completely different world. If you want to find out where the world's innovation is happening, I would not go find out at an enterprise. Would you guys agree?
Yeah.
Enterprise is the slowest adopter of new technologies. I would go talk to all of the startups, the 30,000 or 40,000 startups that are currently doing this stuff. I would go talk to OpenEvidence. How’s it working? I would go talk to Cursor. How’s coding working, by the way? I would just go talk to these people.
Healthcare—the most conservative of all. But guess what? They are so concerned about getting the right answer—
That’s the value of having something like OpenEvidence.
Yeah.
You can do grounded, high-quality research and get that research as information to you. Nobody wants to do research. They want answers. Nobody wants to do search. They want answers. Is that right?
Abridge is a great example of that, too, where they’re basically making it really easy to do physician notes instead of the physician sitting there and doing it. Back to your point on task versus—
Task versus purpose. Exactly.
And I think a different way to think about the demand is that there are so many jobs where the work is actually an impossible ask, right? Of a doctor or a radiologist: keep up with the world’s biomedical knowledge in R&D, which is accelerating in computing and otherwise.
Like arXiv papers.
There was a time when you and I read—
You and I both used to do that. I don’t do that anymore, but now I just load it all into ChatGPT.
ChatGPT. Now I just load it all in with all of the ones that are interesting, and I make it learn it.
And then I make it summarize, and then make another summary, and I interact with it.
But the point is, we used to do search. We don’t do that anymore. I don’t do search. We used to do research.
The goal is to get answers. The goal is to get smarter. These AIs allow us to do all that.
All of it comes back to this. It’s all more helpful if you come back to the framework that says AI is a multi-layer cake. AI is not just a chatbot. AI is very, very diverse in all of the industries, modalities, information, and applications that it addresses.
When you think about wanting to win—
When you think about America winning AI, it should not just be about having this company win AI. We should try to win across the board and across domains.
Across domains.
Exactly. When we think about open source, all of a sudden this is a helpful framework. When we think about winning, it’s a helpful framework. When we think about energy, it’s a helpful framework, because we need factories. Factories need energy, and without energy, we have no factory. Without factories, we have no AI. That’s a helpful framework.
If we have a better understanding—a system, a framework for understanding what AI is—I think the narratives will become more common sense. The narratives will become more pragmatic and more balanced.
We want to keep people safe, but one of the best ways to keep people safe is advancing technology quickly. I think the industry is doing that, and I’m very proud of the industry for doing that.
No one wants to drive a car from the first decade of cars. And so I think—
ABS is a really good thing.
Yes.
ABS is a really good thing. Lane keeping is a really good thing. There’s no question FSD is a really good thing.
And I think people will be excited about the third or fourth year of AI.
Yeah. No doubt. I say with great pride that the industry made tremendous strides this last year—all the technologies we’ve mentioned. The scaling laws are so intact that we now know that more compute means more intelligence.
The innovations in one sector diffuse and spread across all of the other sectors so fast. I’m so happy to see all that. I think the next 5 years are going to be extraordinary. No doubt about it. I think next year is going to be incredible.
Amazing. Well, we’re excited to talk to you at the end of next year, too.
Yeah. Looking forward to it. Thank you guys for all the work that you do. Congratulations. What a great year.
Wow. Amazing year.
Yeah. A lot. Thank you.
Yeah. Thank you. Happy New Year.
Find us on Twitter at No Prior Pod. Subscribe to our YouTube channel if you want to see our faces. Follow the show on Apple Podcasts, Spotify, or wherever [music] you listen. That way you get a new episode every week. And sign up for emails or find transcripts for every episode at no-briers.com.