为什么拥有无限 GPU 的 AI 实验室仍会失败——Anjney Midha,AMP
Anjney Midha 认为,资金和算力并不能保证产品交付,真正的约束是执行协同和组织文化。 他表示,Google 将约 95% 的节点利用率视为接近宕机,而业内最佳 MFU 只有 60–70%;协调不佳的实验室会让微小的规划误差在组织层级间不断累积。“AI 扩容的本质,应当是提高常识的价值。”
Amp 给出的答案,是建立一个中立的算力网格,汇聚分散在不同云和芯片上的供给。 目标是让“megaflops 像 megawatts 一样流动”,为参与者保证基础容量,同时动态分配峰值需求。Amp 已开始锁定明确披露的 1.3GW 规模需求,相当于约 400 亿美元云支出;但 Midha 估计,未来 4 年其团队可能需要 6GW 的峰值容量。
数据中心必须与所在社区达成明确的利益交换,否则许可风险就会变成基础设施风险。 Swyx 转述一项估算称,今年美国最多 20% 的数据中心可能缺乏足够的社区支持,但他也提醒这一数字可能被高估。Scott Nolan 提议的“新 AI 契约”是:将算力价格从每小时 4 美元提高到 4.50 美元,并通过现金把增加的 0.50 美元直接返还当地居民;Swyx 还建议提供更便宜的电力。作为算力客户,Swyx 表示自己很乐意买单。
替代性 AI 芯片可以扩大供给,而无需让整个技术栈碎片化。 Matrox 选择 NVIDIA 的参考架构和机架规格,再将创新集中在逻辑 die 与系统协同设计上:“你不可能在每条战线上同时作战。”更深层的约束是信任——芯片从设计到流片约需 2 年,因此设计者需要尽早了解模型架构的变化。
垂直一体化实验室内部囤积研究,为独立资本和基础设施创造了切入口。 Midha 表示,DeepMind 的 6 个月内部禁发期,在研究显现商业价值后可能永久化,导致公开发表的内容出现逆向选择。Amp 的 Foundry 因此为前沿团队提供支持,包括今年早些时候向 Anthropic 投资数亿美元;同时,Amp 也将多余算力捐给非营利组织和大学实验室。
临终预测是 Midha 所举最清晰的、兼具公共价值的 outputmaxxing 案例。 他说,Stanford 的纵向数据集覆盖至少 1200 万名患者,而当时 Medicare 和 Medicaid 超过 30% 的支出用于临终关怀;即使是回归模型和简单神经网络,看起来也具备技术可行性。真正卡住项目的是监管,而不是建模,因为错误诊断的责任无法从医生转移给 AI 系统:“过去 14 年里,我没有一天能把这件事从脑中移开。”
文化是 Anthropic 编码能力突围背后脆弱、却能不断复利的优势。 Midha 不接受“只是掷中了幸运骰子”的解释:Anthropic 在资源稀缺中花了 4 年做好准备,将编码设为 P0,因为计算机使用能力可能推动 AGI 进展,并承受了“21 次拒绝”,由此逼出了清晰方向。“文化不是一套信念,而是一套行动”(Culture is not a set of beliefs. It’s a set of actions)——它需要每天浇灌,而不是一座永久护城河。
1. 未充分利用的集群暴露协同失败
Midha 将节点分配与模型 FLOP 利用率区分开来:约 96% 的节点使用率应当是标准,Google 据报将 95% 视为宕机,而当前最佳 MFU 约为 60–70%。他说,他讨论的大多数集群连第一项测试都通不过。
他的诊断指向组织:出资方、集群部署者、调度器和产出测量者之间隔着太多层级。就像两条起点夹角只有细微偏差的直线,微小的失配到了规模化阶段就会变成巨大的浪费。
解决方案是借鉴半导体行业的迭代式 bring-up,并非 AI 专属发明。新能力的出现,不足以成为放弃运营纪律的理由:“这次不一样”或许适用于能力,但不适用于基础设施常识。
2. 算力池化可能优于全栈自有
Swyx 提出的问题值得保留:像 xAI 或 OpenAI 这样的垂直一体化运营商,难道不是最有动力提升效率吗?Midha 的系统性回答是,利用率既可以通过一体化提升,也可以通过池化提升;Amp 有意选择后者。
Amp Grid 覆盖多云和多种芯片,在调度和经济层协调闲置算力。它的目标是让“megaflops 像 megawatts 一样流动”,将原本不可互换的容量变成共享资源。
Midha 将 Amp 定义为独立系统运营商,而不是 neocloud。历史上的电网之所以有效,是因为中立协调者汇聚了发电机组和需求不相关的客户——钢厂在夜间达到峰值,鞋厂在白天达到峰值——同时不拥有底层资产;PJM Interconnection 是他所说的成熟类比。
Amp 表示,正在围绕 1.3GW、4 年期规模组织供需,但强调全部容量尚未锁定。理想稳态是 1.3GW 的基础负载容量,加上未来 4 年约 6GW 的峰值容量。
3. 动态分配仍需要指挥权
Amp 的调度模型呼应了 Google 的 Borg/GQM 实现,以及 Midha 在 a16z 的 Oxygen 项目:团队获得有保障的基础容量,将可中断任务排队,并花费 credits 重新排序。Midha 举例称,Nano Banana 2 获得 10 个 credits,Genie 3 获得 5 个,因此前者可以中断后者;他特别说明这是虚构示例。
Swyx 提供了关键反例:Brain Marketplace 确实存在,但 Midha 表示,Google 内部的 credit 机制最终导致 Google 错失 GPT。Midha 更广义的控股公司结构,则是将池化基础设施与 Foundry 资本业务结合起来,用于孵化和支持前沿实验室。
4. 社会许可正成为算力投入
Swyx 转述一项估算称,今年美国最多 20% 的数据中心可能难以获得社区支持,同时明确警告这一数字可能被高估。仅有就业岗位并不足够,因为居民承担着电网、环境和许可成本。
General Matter 创始人 Scott Nolan 的提议把利益交换具体化:如果算力成本为每小时 4 美元,就收取 4.50 美元,并将多出的 0.50 美元转给所在社区。Swyx 表示,作为算力客户自己会乐意支付,因为更便宜的本地电力或直接现金能让居民成为合作伙伴,也能让最终获得的算力供给更可靠。
Midha 预计,审计和调查最终会追上那些“快速行动、打破东西”的运营商。他偏好的合作方,是拥有 20 多年经验的数据中心老兵:他们有土地、电力、信用记录,也经历过互联网繁荣与萧条周期;他们可能不会赞助 NeurIPS 的庆功酒会,“但他们是成年人”。
5. 独立基础设施可以释放被囤积的研究
Midha 认为,大型实验室内部存在市场失灵:具有商业潜力的 DeepMind 研究可能进入 6 个月禁发期,并无限期无法公开。由此产生逆向选择——公众看到的,可能恰恰是企业认为价值最低的工作。
Foundry 作为 Amp 的资本部门,负责孵化和资助独立的前沿实验室;Midha 表示,今年早些时候 Amp 向 Anthropic 投资了数亿美元。Amp 也将多余算力捐给非营利组织和大学实验室。
他用一个刻意夸张的规模检查开玩笑:“我们只有 1.3GW 的算力,这根本不算什么。”但这个规模相当于约 400 亿美元的云支出。即便如此,如果团队同时推进超导体、医疗模型和其他前沿项目,这个算力池仍然不够。
6. 临终 AI 展示了 outputmaxxing 如何进入公共政策
Midha 读 Stanford 生物信息学研究生时,加入了 Nigam Shah 的临终预测项目,因为在他的估计中,Stanford 的纵向数据集覆盖至少 1200 万名患者;他说,唯一更大的数据集来自 VA。据报道,当时 Medicare 和 Medicaid 超过 30% 的支出用于临终关怀。
实际问题从不确定性开始:终末期患者可能听到“6 个月到 6 年”,而医疗事故风险让医生不愿给出更明确的建议。患者于是“什么都试一遍”,把最后几周耗在医院里,而不是在知情状态下决定如何使用剩余时间。
Midha 的文化对比来自个人经历,而不只是技术判断。他在 Chennai 长大,看到死亡被视为人生的另一阶段,并通过庆祝性的游行来纪念;而在美国医疗体系中,不确定性最终默认指向延后死亡。
建模本身看起来并不难——“连回归模型都能用”,简单神经网络也可以,如今 RL 能做到更多。尚未解决的障碍是监管责任:错误临床诊断的责任无法从医生转移给 AI 系统。14 年过去,Midha 仍把患者赋权称为“我愿意为之死守到底的阵地”。
7. Outputmaxxing 意味着扩张而不丢失协同
Midha 将这套纪律称为“Outputmaxxing 部”;如果在更正式的场合,他可能会称之为 Alignment 部。向一个次优扩展方案投入 500 个或 50 万个 GB300,并不等于遵循 Bitter Lesson;Anthropic 决定统一采用 Transformer 架构,则说明聚焦如何释放速度。
组织规模扩大后,专业分工、接口和抽象层随之增加,每一层都会引入沟通损耗。他的工程问题是:系统能否在纵向和横向扩张的同时,避免层与层之间的有损传输。
他认为有两条出路:将协议标准化,直到协同接近无损;或者创造足够的新供给,让损耗变得不再重要。室温超导体是他的极端案例——实现无损能源传输,并且按他的说法,几年内就能拥有“飞行汽车”。
8. 兼容型硅片把创新推向真正的瓶颈
Swyx 问及非 NVIDIA 加速器。Midha 举出的反例是 Matrox:它采用 NVIDIA 的参考架构和物理 I/O 规格,因此芯片可以接入已经按 NVIDIA bring-up 方案建设的场地。
这一选择让 Matrox 可以专注于逻辑 die 和系统协同设计的改进。NVIDIA 发布的标准因此成为赋能基础设施,而不只是竞争壁垒,尤其是因为在 Midha 看来,NVIDIA 无法满足全部生产需求。
真正的依赖转向信任。一块芯片需要约 2 年才能完成流片;如果模型架构在产品上市前发生变化,设计就可能被搁置。Google 内部团队可以与 Gemini 或 PaLM 研究人员并肩协同设计,独立创始人却失去了这种优先可见性。
Midha 认为,Amp 与不同实验室的关系可以跨过这道信任边界。算力网格也可以与 SF Compute 等交易所双向连接,但短期供给已经极度紧张:他说,过去 6 周已经消灭了 Amp 原本预计年末会有的多余容量,而那些已经融资数十亿美元的创始人正发短信询问“15 个节点”。
9. 科学家创始人靠精准度和持续维护文化取胜
Midha 不认同投资人把人当作静态类别的习惯。在 Berkeley 等机构达到学科顶尖水平的研究者,都是“思想上的运动员”,已经练就了争取资源、赢得组织信任、面对技术判断、发表成果并将其推出去的能力。
CEO 还必须具备超越科学本身的对抗能力,能够处理招聘、团队、招募和客户等问题。他提到 LM Arena 的 Anastasios、Black Forest Labs 的 Robin、Mistral 的 Guillaume 和 Anthropic 的 Dario,作为顶尖研究能力可以迁移到公司领导力的例证。
精准度也能消解虚假的竞争。“世界模型”掩盖了多个不同问题,而“实时动作预测模型”则揭示了真正的工作内容。以这种分辨率开展工作的研究者往往会寻求合作,反倒是商业抽象制造了谁在获胜的简单故事。
Anthropic 的编码突围是 Midha 最后的样本。它可能受益于运气,但“运气偏爱有准备的头脑”(luck favors a prepared mind):4 年来在稀缺、效率和安全信念下的准备,加上“21 次拒绝”,最终迫使编码进入 P0,成为通向 AGI 的路径。
因此,文化不是静态护城河,而是“一种非常脆弱、极易破碎的东西”(a very brittle, fragile thing),需要通过行动每天维护。Midha 表示,现金充裕、从未经历困难的实验室,可能永远无法确定自己要守护的阵地;在 Periodic Labs,他反对重新聘用一名因更高薪酬而离开、直到取得 SOTA 突破后才回归的人,因为这一决定本身就会重写文化。
Anshumita
There are so many AI labs today that have all the cash they need. They have all the compute they need, and they're still not able to ship anything. Then you start seeing people leave and so on. My diagnosis is: it's the culture. If you stop taking the actions that demonstrate the mission alignment with what you've said to your team and to the world matters to you, then your culture starts to fray.
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads. And we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you and it means absolutely everything to me and my team that works so hard to bring the week. If you do it, I promise you we'll never stop working to make this show even better. Now let's get into it.
We're in Periodic Labs with Anshumita, CEO and founder of Amp. Welcome.
Thanks for having me.
At Google, there are 2 types of utilization usually, right, that you're measuring in these clusters. One is node utilization, and the other is MFU. Node utilization is usually what percentage of cards in the data center are used. If it's not at 95%—
There's no excuse.
There's no excuse, right? I think 95% at Google, which is where my co-founder Seb came from—he built the Borg X Borg GQM scheduler at Google—and there, I think 95% is considered an outage. So, 96% node utilization should be standard. Most single-down clusters are not running at that. That's one. And then MFU utilization should be, I would say, best-in-class today, somewhere between 60 and 70%.
I think it's a leadership question, right? Fundamentally, it's an alignment question: are the people who are funding the cluster and then deploying the cluster actually aligned? Sometimes, theoretically, they are, but in practice, the number of people in the chain—the supply chain between the capital and whoever's managing the cluster, and then whoever's measuring what the output is—means there are just so many degrees of separation.
Have you ever heard that sort of radian metaphor? At the beginning of an arc, if you have 2 lines that are just off by a few degrees, they spread out at scale. I think what's happening is that, in a lot of cluster implementations and infrastructure at a lot of frontier labs and other teams, they initialize the plan, which is kind of like a North Star, with a team that wants to do good. But then they're required to scale so fast instead of iteratively that the wastage compounds really quickly at scale.
I think we know the answer, which is just to do iterative bring-ups. If you spend time with people who've been in the semiconductor industry or the DSA industry for a long time, this is not new. I don't think AI should be an excuse. Sure, something—what is new? Okay, we have a lot of new capabilities, but that doesn't mean you just abandon common sense. Common sense should always be in fashion.
What is new?
AI scaling doesn't change the need for common sense. In fact, if anything, AI scaling should be putting a premium on the value of common sense and infrastructure, because the margin of error now is so much lower and the cost of wastage is so much higher.
The cost of wastage, by the way, is not just economic. I'm obviously an investor, or I'm an investor by background. Over the last few years, now we're running an AI infrastructure business called Amp, and I think it's okay to say this time is different on the capabilities front. We're genuinely getting capabilities of a kind we haven't had before. That doesn't give you an excuse to say this time is different for everything, especially infrastructure.
Look, I love the hacker mindset and the hustler mindset, and that's great for the startup mindset. But you remember this moment when Zuck went from saying, "Move fast, break things," to "Move—"
Fast with stable infrastructure.
Move fast with stable infrastructure. I think now we need to move fast with responsible infrastructure.
Yeah. They're going to say, "Where is the impact?" There was a really interesting discussion in our class yesterday. Scott Nolan, who's the founder of General Matter, came by Stanford to speak about energy bottlenecks, and he had a phenomenal idea.
He said, if you look at the marginal unit economics of compute per hour, let's call it $4 an hour. If you're having to bring up a new data center in a new community, why not just say, "We're going to charge $4.50 an hour," and take that marginal impact, or that marginal increase, and literally give it to the local community as cash? I can tell you, as a customer of that compute, I would love that. I'd be happy to pay an additional $0.50 per hour at scale.
Wow, yeah.
Because if that means the public benefit is so clear to the communities that the data centers are coming up in, I'm going to feel like that compute is much more reliable. Up to 20% of all data centers this year in the U.S., my understanding is, are at risk.
Of community backlash?
Of not getting the community support they need to be brought up.
Wow. That's a huge number.
Yeah. Now, we—I think we should dig into what that number is. I think it's a little bit overstated. These things can get overreported. But—
They don't just care about jobs. They care about all the other stuff around it, right? They care about the power grid, and they care about the environment.
Permitting and so on. Imagine if you said, "There's a new AI deal. If we're bringing up a data center in your community, we're actually going to reduce the cost of your electricity bill." Okay, now we're talking.
Yeah.
Right? The community is going, "Okay, now this is a deal. I feel like a partner in this."
Anshumita
Yeah.
Right now, that's not happening.
Anshumita
Mm-hmm. There will be audits. There will be investigations. When the regulators come—I don't know when it's going to be—the folks who are moving fast and breaking things in the name of AI progress had better be prepared.
That's certainly not how we're procuring compute. We're trying as much as we can to work with partners who have long-term track records, many of whom, by the way, are not AI providers. I think this whole idea of neo-clouds being somehow this new category is a lot of marketing speak.
There are really good, reliable, trusted data center providers in America who've been around for 20-plus years. I love those folks. They know how to—sure, are they sponsoring happy hours at NeurIPS? No. Are they legibly Bitter Lesson-built? No. Are they hanging out at Situational Awareness parties? No. But they're adults.
Yeah.
Anshumita
They can run land. They can run power. They have credit histories. We sit down and have conversations. Many of them live in Silicon Valley. They've had to deal with the boom-and-bust cycles of the internet, and I love those folks. They're stable infrastructure partners and thinkers.
I think there's a lot of short-term thinking going on in the compute layer, and it's going to catch up to us. It's not going to be good.
You talk about aligning incentives. I would think that aligning incentives means you have the full stack in 1 company, which is xAI and OpenAI, right? So, as a standalone infrastructure layer, why are you somehow more aligned to your portfolio companies than people who just own the whole thing?
In systems design, there are 2 regimes of architecture, right? You have integration, and then you have pooling and utilization. Or rather, the way to increase utilization often is that you can do systems integration, where you collapse a lot of process into 1 node, or you can pull out a process from a node and share that resource amongst several different nodes.
We see the Amp Grid, which is the system we're building here, and which is basically a compute grid. We're trying to do for compute what the electric grid did—
Yeah, what the power grid did for electricity.
This is a pooling and utilization layer across clouds. We're actually the opposite of a full-stack integration. It's much more horizontal. It's multicloud, and it's multisilicon. The goal is to try to make megaflops flow like megawatts.
That is very hard to do today for many reasons. There are stranded pools of compute all over the place, and there's no fungibility. Right now, we do it at the level of scheduling, and we often do it at the economic layer.
But as we start to announce what we're working on, it's extraordinary how many folks are coming out of the woodwork saying, "Hey, I'm actually working on a way to make compute fungible at this part of the stack and that part of the stack." As a grid, we'd like all of these folks to participate on the grid.
People will often ask me, "Andre, you're a neo-cloud?" And I go, "No, actually, neo-clouds are suppliers."
Sometimes they laugh at how, as a venture capitalist, I go, “No, actually, they are demand, sort of off-takers of the grid.” We see ourselves as what’s called an independent system operator.
If you study the history of the electric grid, once it became legible to a lot of factories and industrial participants that, “Hey, actually, it turns out pooling is a good idea. We should pool our generators instead of all having half a generator running at half capacity in our backyard,” there was a need for an independent entity that could coordinate all these parties: power-generation facilities, transmission lines, and factories.
That neutral coordination mechanism is very critical. If you study the history of grids, the most enduring ones were those that never owned their own assets. They were ones that had, or often started with, long-term anchors that were uncorrelated sources of demand: a steel factory or shoe mill, or whatever, in a particular town.
They weren’t competitive. The steel factory wanted to spike up at night, and the shoe mill wanted to spike up during the day. So then you pool and share, right? Each of you is guaranteed some base load, but then you schedule your spikes to drive peak utilization across the town.
The gold standard, so to speak, historically has been utility companies like PJM Interconnection in the Northeast of America, where, over many years, they became what’s called an ISO, an independent system operator of the grid. So that’s how we see ourselves.
Yeah.
The Professor of Outputmaxxing
Economically, that’s what we are. From a technical perspective, we started at the scheduling layer because Sab and Mihai, who run engineering here, built that scheduling system. They did that at Google.
And you have infrastructure folks from Discord as well. I don’t know if Discord is the primary identity or whatever. I’m just—
The Professor of Outputmaxxing
No, Discord was—
Choosing a well-known name.
The Professor of Outputmaxxing
Well, I was running the developer platform there. I was not responsible for the internal infrastructure. That was actually a guy by the name of Mark Smith, who was extraordinary.
Discord is actually a counterexample. I had the chance to learn a lot about full-stack infrastructure there.
It’s the other architecture, which is that Discord built its own WebRTC, so its voice and video infrastructure.
Mhm.
The Professor of Outputmaxxing
Discord did not use third-party infrastructure for communication. It was all built in-house. The way you maximize utilization was by pulling demand from the world’s 200 million-plus monthly active gamers, right? That’s how those stacks were constructed.
Again, in system design, the 2 concepts that keep coming up over and over again are abstraction and composition.
Right—and bundling and unbundling.
The Professor of Outputmaxxing
Abstraction and composition,
like verticalization and horizontalization. In that sense, Amp is an independent system operator of the grid. We pull demand and supply from a number of partners we trust at about 1.3-gigawatt scale over 4 years. Then we pull demand from some of the world’s best research labs and so on.
We’re working with Periodic Labs, which needs extraordinary long-term demand. The idea is that each of them is guaranteed base load on the grid, but they can spike up and down flexibly on compute on much shorter timelines as needed.
That was roughly the design of the program I came up with at a16z called Oxygen. It was also the design of the Borg X Borg GQM implementation at Google that Mihai and Sevin built: How do you allow teams inside Google, on the internal infrastructure, to be guaranteed capacity for their base workloads, but when they need to spike up for research, ensure that capacity is available?
Of course, the big innovation that was implemented in the infrastructure space maybe 3 or 4 years ago at Google was the idea of interruptible demand, where you queue up a bunch of jobs and, through a sort of credit system, there can be a bidding mechanism. It’s dynamic prioritization, basically, and jobs can get interrupted based on somebody else saying, “You know what? I have 10 tokens, 10 credits, that I want to spend on this job.”
Another team lead or research lead is like, “Genie 3 or whatever is only worth 5 credits, and Nano Banana 2 is worth 10 credits.” So the Nano Banana job gets priority. That’s a made-up example.
It’s very real. Brain Marketplace was real.
The Professor of Outputmaxxing
Yeah.
And we’ve covered this on the pod with David, who was—
The Professor of Outputmaxxing
Oh, great. Okay, awesome.
The criticism is that sometimes you need central commands to go all in on something. Sometimes capitalism via credits doesn’t work.
The Professor of Outputmaxxing
It’s not a criticism of Amp. I’m just saying this is a thing that has been tried internally within Google, and it led to Google missing GPT.
We structured ourselves essentially very similarly to Google. We’re structured as a holdings company. Alphabet Holdings is Alphabet Holdings, and then they have subsidiaries called Google, Other Bets, and so on.
We have Amp Holdings, and we have our infrastructure business. Then we have a capital business called Foundry that incubates new frontier AI labs and invests in them as venture capital. Like Periodic, we put a few hundred million dollars into Anthropic from our fund earlier this year.
Wherever we feel like teams are making progress, especially researchers and others who are pushing the frontier inside existing labs like DeepMind, I find that there comes a point where they feel misaligned with the dictatorship of Alphabet Holdings.
At that point, sometimes the dictatorship doesn’t want them anymore, and they’re like, “Thank you. You’ve done your job here. You’ve helped us through the zero-to-one phase, and for whatever reason, we’re going to deprioritize your amazing omni-model or whatever it is. Instead, we’re going to prioritize coding.”
I think that’s a tragedy, but I get it. Sergey and Demis are running their own business there, but that doesn’t mean the rest of us should sit around waiting for that progress to get unlocked for the rest of the world and humanity.
If you think about how much extraordinary research has happened inside DeepMind over the last 10 years, I mean, Demis and Sergey and those guys did such a great job. But at the end of the day, so much of that has never seen the light of day.
Or they’re papers only, but they never actually shipped into production—
The Professor of Outputmaxxing
I mean, what’s worse is that the paper is actually not even being published anymore because there’s a 6-month embargo inside of DeepMind, right? We’ve heard about this: A paper comes out, and then I think it has a 6-month embargo window. If anybody on the business team says, “This could be interesting,” it’s embargoed for life.
Exactly.
The Professor of Outputmaxxing
So the stuff that gets published is the stuff that’s not good enough. There’s an adverse-selection problem, basically.
It’s a common complaint at NeurIPS, by the way. It’s like, “Why would I look at the papers that are the trash of GDM?”
The Professor of Outputmaxxing
Again, I think it’s a tragedy. I get it; they’re running their business. But for the rest of the space, I think there are negative externalities from research being hoarded. That’s a market failure, and somebody needs to unlock that research.
We can’t do it on our own. We only have 1.3 gigawatts of compute. That’s nothing. That’s about $40 billion of cloud spend. We’re going to need a little more.
That’s a new number I haven’t come across—the gigawatt number. That’s huge.
The Professor of Outputmaxxing
To be clear, we haven’t secured all of it. That’s how much demand we’ve started to secure. I think publicly, we haven’t actually confirmed how much we have for this year.
Where do you want to get to?
The Professor of Outputmaxxing
I think the steady state would be that we have a base-load pool of 1.3 gigawatts of capacity at all times. For spike capacity, right now my estimate is that we need roughly 6 gigawatts over the next 4 years for all our teams to feel like they’re able to keep moving the frontier, whatever they’re working on.
That could be superconductor discovery over here. There’s a new investment we’re working on right now in the end-of-life prediction space in healthcare. It’s extraordinary how much you can give people.
This was actually my graduate-school work. I went to graduate school for bioinformatics at Stanford Med.
Yes.
The Professor of Outputmaxxing
I was this really weird cat. I was never satisfied with my major options. At one point I was an economics major, then I was a computer science major, then I was an MCS major, called Mathematical and Computational Science. They decided they were going to end that major, so I took all that coursework and applied it to my graduate degree in bioinformatics, which was a master’s program.
Then I thought I was going to do a PhD. I never ended up doing it. I dropped out and went work at Cliner, but I was lucky enough to apprentice with a professor at Stanford Med. His name is Nigam Shah, and he was working on end-of-life prediction.
Stanford is one of the only research facilities in America that has a longitudinal patient data set that’s large at scale.
The Professor of Outputmaxxing
I think it’s at least 12 million patient lives. The only larger data set is at the VA, the Department of Veterans Affairs. To do research, like deep learning and so on, on that data set—it was called the STRIDE data set at that time—you had to be a Stanford Med School affiliate, which is why I went and enrolled in the bioinformatics department.
Wow.
The Professor of Outputmaxxing
Deep learning was early. Nigam Shah had the vision to see that you could do end-of-life prediction to help palliative care. In America, over 30% of all Medicare and Medicaid spending, at least at that time, was spent on end-of-life care.
We grew up in Asia, so we all have a very different relationship with death than I find folks who grew up in America do. At least, I won’t speak for you. In America, spiritually and culturally, especially in Western societies where the Christian tradition frames death as this terminal point, there’s often a judgment day and so on. The way we view death is with a finality.
In Indian culture, in Hindu culture, death is one—
Buddhist as well.
The Professor of Outputmaxxing
You’re a Buddhist, yeah. So it’s one step in a journey of many lives, right? I grew up in the city called Chennai in the south of India, and when people die, you dance on the street. There’s a procession where your body is carried to be cremated, and your family celebrates. There are drums and so on. It’s because the idea is that you’re going to be reincarnated. You’ve been liberated from the responsibilities of this life, and now you’re on to your next. It’s like going off to a new college or whatever, right?
It was so alien to me when I got here as an undergrad that the medical system works backward from that assumption: We have to view death as this terminal thing and delay it, postpone it. It’s a bad thing. At the time, clinical decision support in the United States was a very primitive field.
Even to this day, physicians in the United States will often tell you when you have a terminal disease, “We’ve diagnosed you, which is great. Our ability to diagnose you is extraordinary. You have somewhere between 6 months to 6 years to live.” What do you do with that information?
Yeah.
The Professor of Outputmaxxing
The error bars are so high that, in times of uncertainty, we default to culture. When the culture is, “This is a bad thing. I’ve got to prolong my life,” then you start doing things like a whole regimen of drugs and therapies. You often spend weeks and weeks in the hospital, and that deteriorates your quality of life.
When that deteriorates your quality of life, instead of spending your last few days doing the things you love with your family, you’re spending them on a hospital bed. That ends up being 30% of Medicare and Medicaid spending. From a systems perspective, physicians often feel like they need to provide higher certainty because there’s always some uncertainty in end-of-life diagnosis.
If you provide the wrong diagnosis or recommendation to your patient, you can be sued for medical malpractice. Then your license can be taken away. It can be catastrophic for your career. In contrast, in countries where that’s not the case, what you often observe is that physicians are quite prescriptive with their recommendations.
They say, “Hey, this is your condition. The literature says that you probably have this much time on Earth left. My expert opinion is that you are an outlier,” or whatever. They try to be more prescriptive, and that empowers a patient, right? Then a patient can say, “I trust my doctor. They said on average I have 6 months to live, but if I do these things, I may have a shot because of my particular predispositions or my genetic history,” or whatever.
That empowers you to go about your life in a more scientific way than leaning on religion, culture, spirituality, and so on. In contrast, here, because of that medical malpractice thing looming over your head, a physician never gives you a clear recommendation.
Right.
The Professor of Outputmaxxing
Instead, you say, “Okay, doc. Well, let’s try it all.” Then you start a whole regimen of drugs and therapies. You often spend weeks and weeks in the hospital, and that deteriorates your quality of life. Instead of spending your last few days doing the things you love with your family, you’re spending them on a hospital bed.
That ends up being 30% of Medicare and Medicaid spending. It’s worse for the patients. The doctors feel terrible. The American taxpayers pay a huge amount of money.
This is why Nigam Shah, who is a professor at Stanford, said, “Honestly, the big problem is end-of-life care.” I kind of sat down with him. I was this young guy—I was 21—and I was like, “I want to work on a big problem.” He’s like, “The big problem is end-of-life care.”
We started trying to use deep learning on these STRIDE patient data sets to ask, “Could you have an AI system make a recommendation that is orders of magnitude more precise about how much time you have left once you’ve been diagnosed with a terminal condition than a human?” If we can get that precision high enough, then you can empower the patient.
It turns out the tech works. Once you get the data set, RL works. Honestly, even regression models work. You don’t need to get that fancy. Half the time, we were just running very simple neural nets. Today, what we can do with RL is extraordinary.
The problem, then and now, is regulatory, because you actually can’t shift the burden of a wrong clinical diagnosis from the physician to the AI system. At that time, I got quite disillusioned—10 years ago, or 12 years ago—because I felt I just didn’t have the resources to influence regulation.
Today, I’m very lucky. I’m in a different place. I’m a lot older, and I’ve been spending a lot of time on my next incubation: How can we unlock patient empowerment by training AI models to do end-of-life prediction with much more precision?
Oh, wow. And you’re still focused on this?
The Professor of Outputmaxxing
I haven’t been able to get this out of my mind a single day for the last 14 years. This is the hill I would like to die on. This, too, I would say. You know what? I actually prefer not to die.
Yeah, exactly.
The Professor of Outputmaxxing
I think 2 issues that should be bipartisan in America are, first, how we empower patients to make the right clinical decisions at the end of their lives such that we’re reducing the taxpayer burden with science. It’s just good old science, and AI can help here.
The second is net-positive data centers, because I think that’s the biggest critical bottleneck on training good enough AI models to help people at the end of their lives. They’re sort of 2 sides of the same scaling-bottleneck curve. For those 2 problems, we formed Amp as a public benefit corporation.
My wife and I, whom you’ve met—her passion is education. Her family is a long line of educators and physicists. This class is my attempt to stop being the black sheep of the family and be an educator. But if I’m not educating, the thing I would be doing is working on these 2 problems, whether on the political spectrum or as a researcher back in some lab.
And my hope is, if anyone's listening to this podcast, if they're passionate about either of those two topics, I'd love to hear from them. We can share the contact in the show notes, but we're looking for people to join both of those missions, on the political side as well as on the medical research side.
You said this is a discipline that you want to form. You call it variously “frontier systems.” It’s variously called a one-person frontier lab. What is the ideal name or shape of this? What is the mission?
The Professor of Outputmaxxing
Of the class?
Of the discipline that you’re exploring. The class is called “Frontier Systems,” but for me, maybe one phrase is that you’re just anti-waste: wasting GPUs, wasting human life, and wasting Medicare. Is there a broader theme that I’m missing that you can encapsulate more succinctly?
The Professor of Outputmaxxing
From an engineering perspective, it’s basically outputmaxxing. It’s the Department of Outputmaxxing—
Of what we have.
The Professor of Outputmaxxing
Exactly. I’m a huge believer in optimal outcomes. I think both in America and other countries, we’re losing our appreciation for nuance, and this is the thing with AI—it’s the same case, right? “Oh, the Bitter Lesson holds.” Okay, fine. But that doesn’t mean you just throw 500 GB300s—or 500,000 GB300s—at your suboptimal model scaling and waste a bunch of compute.
It also doesn’t mean that the most optimal approach is to have 50 different architectures where there isn’t enough standardization. One of the reasons Anthropic has had extraordinary velocity is because they picked the Transformer architecture and said, “This is simple. Let’s double down on it,” right? Luckily, there’s enough investment going into the space that we can afford other architectures, but at the time, investment was too fragmented into other architectures. That arguably unlocked scaling.
I think there’s a philosophy. We all owe it to ourselves to do outputmaxxing with a new capability called AI on a global level. If I was starting a new department at Stanford, depending on how fuzzy or technical I wanted to be, I’d probably call it the Department of Alignment.
It’s an overloaded term.
The Professor of Outputmaxxing
But alignment really is a hard problem.
Mhm.
The Professor of Outputmaxxing
And I think when you unlock it, full-stack alignment is super hard in any organization and in any system. In a venture capital firm, if you can have full-stack alignment between your limited partners, the founders who are creating the value, and ultimately the public that owns the IPO stock, that is a gift that keeps giving.
When you study the history of these systems, when they start off, they usually start out small-scale, where the feedback loop is actually so tight that there’s alignment. Then, the more you try to scale, the more division of labor happens, the more specialization happens, and at each step you add abstractions. Wherever there’s an API interface, there’s loss. There’s communication loss.
So I think a really cool thing for us to figure out is whether there’s a way for us to have our cake and eat it, too. As an engineering discipline, is there a way to actually scale up and scale out without losing any alignment, without lossy transmission?
Standards.
The Professor of Outputmaxxing
Standards are one way. The other way is that you just have net-new capability. What we’re trying to do here is discover new superconductors. A room-temperature superconductor would be a lossless transmission mechanism for energy. I mean, we would have flying cars.
Yeah.
The Professor of Outputmaxxing
Yeah, right—within a few years of having a new room-temperature superconductor. So I think those are the two options. You either have to standardize on protocols or API specs that allow lossless communication, or you can come up with a whole new capability that unlocks so much abundance that standardization doesn’t matter, because you just unlock net-new capacity.
The Professor of Outputmaxxing
This is what I spend my days thinking about these days.
I think every infrastructure person who wants scale and wants to output-max eventually ends up thinking about this. We don’t have time to go into it, but we did an episode with SF Compute that’s trying to standardize the futures contract for compute. I don’t know how that’s going, by the way.
The Professor of Outputmaxxing
I think Evan is awesome, and SF Compute is the kind of effort that I hope we can accelerate, because what often happens is these exchanges are really hard to get—it’s hard to bootstrap them. They often require many inefficiencies between parties. There are trust-boundary inefficiencies in infrastructure, because you don’t trust one part of the stack to give another part of the stack visibility. There are capital-markets inefficiencies and operational inefficiencies.
If you can inject a single shock to the system—a ton of compute demand or supply—then you can accelerate these new flywheels. My hope is that one day, or soon, if SF Compute has excess capacity, they just hook it up to the grid and get flooded with demand from us.
On the other side, if they have a ton of demand but don’t have supply, they again hook up to the grid. It’s a two-way protocol where they can just hook up to our capacity. I don’t think we’re too far from that today.
Our working implementation of it is mostly through a group of labs, universities, and a few trusted parties who all feel like they’re in alignment, to borrow an overused word. But our hope is to just have it be an open protocol that anyone can hook up to.
Hook up for demand or hook up for supply? Primarily demand, it sounds like. You wouldn’t want to offer demand.
The Professor of Outputmaxxing
Both. Unfortunately, what’s happened the last 6 weeks is that we thought we’d have a bunch of excess capacity by the end of this year. It’s all gone.
It’s exploding.
The Professor of Outputmaxxing
Yeah, it’s all gone. My text messages are full of friends—I mean, we know many of these people. These are founders who have raised billions of dollars in San Francisco, going, “Any chance you have 15 nodes in the next few weeks?”
What is the scope for non-NVIDIA? You have Lisa Su coming and Reiner Pope as well, so there’s a lot of demand for more performance, alternative architectures, and all that. At the same time, this hurts your standardization.
The Professor of Outputmaxxing
I don’t think so. Actually, Reiner’s a great example. Reiner’s the CEO and founder of Matrox. I had him by for office hours in the class earlier today, and he brought up an insight that I hadn’t considered before.
When they decided to pick the standard for their data center, they picked the NVIDIA reference architecture. So the Matrox chips just plug into any site that has an NVIDIA bring-up plan.
It’s just software, then. It’s not the hardware.
The Professor of Outputmaxxing
Well, from an input/output perspective, it’s the same footprint as an NVIDIA rack.
That makes sense.
The Professor of Outputmaxxing
Where they have innovated a bunch, from what I can tell, is on systems co-design, which is where a lot of the gains are to be had. He picked it—he was like, “Andre, you know, there’s just so much work to do when you’re building a new chip company.”
Can’t fight on every front.
The Professor of Outputmaxxing
You just can’t fight on every front. My question to him was, “You’re working on this new chip. Their tape-out is next year. Who are you going to partner with to host the chips?”
He said, “Whoever will host them. That’s just not my focus.”
I asked, “But how did you decide?” Back to the earlier systems-design question, he decided that he didn’t want to be a fully integrated chip provider. The bottleneck they’re focused on is the logic die, and he feels they can crank out a ton of performance gains through co-design there.
But then that means you delegate. The data center provider is a different part of the stack, and so he’s dependent on that part of the ecosystem to host his chips and get the performance gains to the customer. So now you have another abstraction, and you might have loss.
I asked him, “How do you prevent loss?” Back to your point, he said, “I just picked the NVIDIA standard because I wanted to piggyback off an existing protocol. What’s great about NVIDIA is that the reference architecture is known. It’s open. They’ve published it.”
Jensen has actually enabled someone like Reiner to build a chip company like Matrox, and I don’t see them as competitive. Compute demand is so high. I don’t think NVIDIA is able to meet the production demand, so we just need more chips.
I think what Matrox has done is very smart: they’re just not going to innovate on the data center design. “Thank you, Jensen. You’ve done all the hard work.” Where they can innovate is somewhere else. I think that’s very healthy. I think that’s how we unlock new bottlenecks.
My view is that these chip teams, like Matrox, who have arrived at the insight that co-design is the way, have a primary bottleneck: the trust boundary. To do co-design well, you need visibility into the next model generation as soon as possible, because it takes 2 years to tape out.
If, by the time I bring my chip to market, your model architecture has changed, I’m hosed. When he was inside Google, he was sitting next to the Gemini team. He was on PaLM or whatever.
His co-founder was one of the PaLM guys, I think.
The Professor of Outputmaxxing
Yes, exactly. When you’re inside the trust boundary of Google, your systems co-design loop is super tight. When you leave as a founder, one of the biggest risks you take is that now you’re outside the trust boundary.
What I love doing is helping chip teams that can help us unlock more capacity for the independent ecosystem access trust. If I’ve been involved with a lab from day one, and I was lucky enough to work with Anthropic, then be on the board of Mistral, and help Black Forest Labs get started, I think at this point I’m on 6 or 7 different teams.
Only 6? I feel like my mental number was going to be 13, but yeah, it’s—
The Professor of Outputmaxxing
No, I go deep with one at a time.
You were founding CEO of Arena?
The Professor of Outputmaxxing
No, that wasn’t an—
Administrative CEO.
The Professor of Outputmaxxing
It was an administrative 5-month gig while Waylon and Anastasia were graduating from their PhDs. They didn’t need a product team, so I helped recruit the heads of engineering, product, and design. But Anastasia has always been the CEO of that company. I played a pinch-hitting job. I was an intern—I was CEO intern for 5 months.
I interviewed him. He’s very well-spoken. I think he’s a former debate champion, but he’s also very quantitative and mathematical, which is such a unicorn.
The Professor of Outputmaxxing
You look at his output—he’s an output-maxer. By the time he was graduating from his PhD, which he only graduated from last year, he had published more work, with a higher citation count, than people twice his age. At the same time, he’d already started a project called LM Arena that was being used by millions of people as a side project.
Time and time again, what I’ve realized is that venture capitalists suck at seeing human beings as dynamic agents. They want to put you in a box: this is your thing.
The first time I got introduced to Anastasios, somebody had told me, “He’s amazing, but he’s a researcher.”
The Professor of Outputmaxxing
Yeah.
I was like, “What do you mean, he’s a researcher? He’s not a CEO—not a founder.”
The Professor of Outputmaxxing
Not a CEO, exactly. I was like, are you crazy? Have you met Dario? Dario is a scientist. He's gone from zero to what will soon be a trillion-dollar company in 4 years. Being a CEO, nominally speaking, is not that hard. Being a good CEO is hard. Being a great CEO actually requires a level of performance that scientists who have already published at the top of their field have accomplished.
It is super hard to be a competitive scientist. To publish in academia over the last 20 or 30 years, to make it to the top of your discipline at a place like Berkeley, you were a star athlete. You are an athlete of the mind, and you perform at the highest levels. To get there, whether you're Anastasios or Wei-Lin at Berkeley, Robin, who with Black Forest Labs created Stable Diffusion, or Guillaume at Meta, who created LLaMA before he started Mistral, the amount of human leadership you have to demonstrate to get the resources, get the trust of the organization, publish it, put it up—I mean, I would just fund researchers all day.
Right? If they've put SOTA out there, they're star athletes already. If they haven't done SOTA, they can still be good CEOs, but then I find the failure mode is that they just don't want to be CEOs. They primarily want to publish, and that's okay, too.
One of the things we do with the Ampere Grid is donate the excess compute we have to nonprofits, like university labs. We carved out a couple thousand H100s, but I do think there's extraordinary research being done on university campuses. My father-in-law is a physicist. He's a professor. There's extraordinary work in physics, and we need that.
But if you want to be a CEO, what you need to be willing to do is be super confrontational outside of science. Within the scientific community, some of the best researchers are very confrontational about their convictions, right? This architecture is right. To be a great CEO, you basically have to be willing to be confrontational up and down the stack.
To your own team.
The Professor of Outputmaxxing
To your own team, hiring, recruiting, customers. Well, I would say, yeah, pretty much to everyone.
Yeah, yeah. I see a little bit of that in my own work, but I can't imagine the stakes that Dario has had to go through.
The Professor of Outputmaxxing
It's pretty good.
I don't think the stakes are that different from how you're feeling them, right? Stakes are personal scaling vectors. The stakes that seem so low to you, like having this podcast, where you can talk to somebody and just have a conversation—I mean, you're an extraordinary communicator. Already in this conversation, you've pulled more out of me than most people. You know, I've been on 12 podcasts in the last 2 weeks.
We've just seen each other enough that there's some base trust. I know that you know I've done my homework, and I know that trust is a big deal for you, so—
The Professor of Outputmaxxing
Right.
Yes. I think trust is about consistency, and you and I have seen each other in the community for years, right? I don't know—the first time we met was at NeurIPS in New Orleans. I don't know if you remember that.
The Professor of Outputmaxxing
I remember it.
God. Raiko had set up this—
The Professor of Outputmaxxing
Yeah.
You know, Raiko's amazing, and he set up this luncheon, and—
The Professor of Outputmaxxing
Yeah, I was like, “Who's this Discord guy?” I'm like, okay, but—
No, you weren't like that. You made some investment. You were much less polite. You were like, “Who's this VC?” No, I was—oh my God, I'm so sorry. I didn't know who you were.
The Professor of Outputmaxxing
I'm so sorry, but the introduction was bad. I didn't know who you were.
But this is the thing about context, right? There are people in the scientific community for whom the stakes are sometimes very high because they haven't had the emotional—we know what it's called—EQ coaching and mentorship. To have scientific impact, you often need to be an extraordinary, emotionally in-tune person with the folks you're trying to influence.
What comes so naturally to you is actually a super high-stakes thing to other people. I wouldn't assume that Dario is more stressed out than you. You'll be surprised how similar and small the problems are that some of the world's biggest leaders are facing compared with yours.
That's what I've learned from this class. The guest speakers are Sam, Satya, and Jensen. They are Coachella. They are Coachella, right? We got all the headliners, and I'm very lucky that some of these people have either mentored me over the years or I've done business with them.
When you take the performative stuff out, and any assumptions you may have about these people that you read in the press or on Twitter, we're all just humans. We're all trying to get along. What's so special about this moment is that AI is forcing a lot of people to revise their assumptions about how the world works and go back to first principles, or go back and educate themselves.
I won't name who this person is, but I was at an event last week in Texas and ran into somebody who said, “And I came across the class. What do you think about real-time action prediction models?” You don't know how happy it made me feel when they asked me that question. I know they've done the work. They've challenged the assumption.
They didn't ask me, “What do you think of world models?” They said, “What do you think of real-time action prediction models?” World models, don't get me wrong, are cool and everything, but you and I both know that they're a layer of abstraction that is sometimes not usefully precise enough.
The Professor of Outputmaxxing
Right? Arguably—
There's like 4 different kinds of world models.
The Professor of Outputmaxxing
Exactly.
We've done the part with General Intuition, by the way.
The Professor of Outputmaxxing
Oh, cool. Yes, I love him. He's great. What I love about people who've done that level of work is that they realize they're not in competition with the people the rest of the world thinks they're in competition with.
Yeah, because they're not in the category. They're in the specific thing they're trying to do.
The Professor of Outputmaxxing
They're focused on their mission, and they have a systems understanding of the problem they're trying to solve. When somebody else says, “I'm working on real-time action prediction models, too,” Pim goes, “Oh, I love that person. I want to learn from them.”
But the minute they're like, “Oh, that person's a world model person,” it's like, “Which type of world model person?” Mostly, they're just trying to figure out if they're wasting their time, because we don't have enough time.
Pim, for example, really loves this other company I work with that we've talked about, called Black Forest Labs. He's mentioned to me multiple times that he thinks what Flux is doing is really cool. Andy Blatman came by and spoke in the class.
What I find over and over again is that, for people who do the work and can be usefully precise enough about what is actually going on in the world of frontier research, the sense of camaraderie is still real and alive. But it gets lost sometimes when you have to abstract the technical complexities into business terms and then the VCs are like, “How are you different from that world model company?”
Yeah. Where do I even start to explain this stuff? And then there's the misalignment. I think people listening get a sense of what it's like to operate at a real level, where you're yourself, rather than at the journalist level, where you have to put everyone into a rough category and create a narrative of competition: who's winning today, who's behind.
The Professor of Outputmaxxing
Yeah. This idea of winning is so weird to me.
You do want to win. You want to compete. You want to compete with us.
The Professor of Outputmaxxing
No, I think you want to lead. You want to push the frontier. You want to push the state of the art. You want to do something that hasn't been done before. You want to capture value, but you don't want to capture so much value that people think you're out of line with your mission or not trying to do what's best for the world.
You want to capture enough value that you can keep innovating, right? I think that people want to lead. They don't really want this idea of winning and losing. Again, I love Jensen. He's a leader. The mindset that he talked about on Dwarkesh's podcast was, “I didn't wake up with a loser mindset.” I think that was awesome, right? Because he's an engineer.
Dwarkesh has done the work, so even though it was very obvious to me that they were talking about the same thing, they were talking past each other. Jensen has this 5-layer-cake abstraction of how the industry works, and Dwarkesh had, I think from that podcast, more of a pre-training, mid-training, post-training systems-loop concept.
It's just a factor of who he talks to, right? It's pretty clear.
The Professor of Outputmaxxing
It’s the abstraction, the mental models—it’s the whole deal. So much of the problem in the world is reasoning by analogy.
Mhm.
The Professor of Outputmaxxing
And then the assumptions that are held invisibly.
Yeah, I’ve said this is actually the best time in human history for first-principles thinkers, because everything you think will happen is actually now coming true.
The Professor of Outputmaxxing
Correct. [laughter] And the venture capital community is notorious for this, where people, in times of uncertainty, cling to axioms that end up being true from the previous era. They proclaim them with confidence as if they’re truths, but they’re not. It’s very important to see the distinction between a heuristic and an axiom.
Yeah.
The Professor of Outputmaxxing
An axiom can be proven—
Like, from an internal-consistency point of view.
The Professor of Outputmaxxing
A heuristic is where you use a shortcut. And my God, the number of people I have had to put up with over the last few years who use heuristics as axioms to judge people, to judge which companies are going to succeed. I mean, the number of people who are like, “Oh, yeah, yeah, yeah, Anthropic—they’re just training models right now. But this won’t continue.”
They’re going to be the B2B SaaS.
The Professor of Outputmaxxing
Yeah, which, over the fullness of time, if you squint at it, maybe. But the way you arrive there is so important that you can just dismiss people. Here’s what happened. Anthropic basically achieved takeoff in October of last year. That training run—
Whatever F37?
The Professor of Outputmaxxing
I forget the numbers now, but whatever that checkpoint was—
We saw it at Cognition.
The Professor of Outputmaxxing
Yeah, right? Probably, to those of us in the community, especially once post-training was done and it was released in December—
Yeah, can I sneak a sneaky question in there? I don’t know if you have a perspective; maybe you don’t. The number-one question is: How did Anthropic crack coding, right? Because Claude 1, Claude 2—okay, it was part of it, but it wasn’t a big deal. And the leading hypothesis is that it was a lucky dice roll that was then compounded, right? It was mildly better, but then they saw it and they were like, “Okay, let’s really invest.”
The Professor of Outputmaxxing
I had this very annoying teacher. I went to this boarding school called Rishi Valley in India, which is a bird preserve—like 350 acres of bird preserve in rural India. And there was no technology for 7 years.
There was this teacher—I won’t name them—but they would have this phrase. I hated it every time he said this to me. He was like, “Luck favors the prepared mind,” which is a common saying, but the way he delivered it always grated on me because I was always one of those kids who got a good grade without trying very hard. High school and middle school aren’t that hard if you’re generally paying attention and so on.
I would get an 80% grade, and he would keep pushing me, saying the reason I didn’t get the 95-plus percent was because I wasn’t that lucky. And I would say, “What do you mean?” because I would think that I deserved that grade, and I would sometimes argue with him. And he’d say, “You didn’t have a prepared mind.”
There was basically one time where I got a 95 or 96 on a subject, and I felt entitled. I was like, “Okay, I’m going to keep doing this.” And I didn’t. Then he was like, “Luck favors a prepared mind. You got lucky last time, but you have to stay prepared.” And I didn’t understand what he meant. Now, as I’m older, I’m like, “Okay, these adults actually knew a thing or two.”
Anthropic has been the most prepared company for 4 years. So when the right context data comes in and the right developers start sending in the right context diffs, sure, you could say they got lucky. But if you ask me, they’re pretty damn prepared, with paranoia, for 4 years.
And if you remember, it was so hard for them to get going early on that they had to do so much more with so much less. You just have to be prepared to be so efficient.
Yes. There are numbers on their burn compared to OpenAI. I’ve written about it, but they are so much more efficient in the way—
The Professor of Outputmaxxing
It’s not even—
Not even close.
The Professor of Outputmaxxing
Yeah, but it’s so clear, right? How to output-max for the world. They have been prepared. And you could call that luck, but—
The Professor of Outputmaxxing
Luck favors a prepared mind.
This is one of the things that I was going over in some of your old lectures, and you were like, “Data—people think it’s a moat, and actually it’s culture. And actually, it’s the team.” There are different levels of moats, and this is the ultimate one that determines everything else, which you can then compound.
The Professor of Outputmaxxing
You’re saying culture is the ultimate one.
Yes.
The Professor of Outputmaxxing
Yeah, but the thing about culture is that it’s very fragile. I don’t think there are very many moats that are actually moats. It’s a nice concept, but in reality, you have to replenish your culture.
Ben Horowitz was a speaker in CS183 on Tuesday, and I asked him this question about the culture bottleneck in—
Teams, because there are several AI teams.
His book, The Hard Thing About Hard Things?
The Professor of Outputmaxxing
The Hard Thing About Hard Things. But more concretely, there are so many AI labs today that have all the cash they need. They have all the compute they need, and they’re still not able to ship anything. Then you start seeing people leave and so on. My diagnosis is the culture.
I asked him, “Ben, you know, he’s been one of the most aggressive investors in AI labs.” He goes back to this thing, which resonates in my mind a lot.
When I used to work at a16z, I would book a conference room, and right outside the conference room that was closest to the toilet—because it was the fastest way for me to go use the bathroom between Zoom meetings—
Oh my God, I’ll put my sink-by-toilet optimization.
The Professor of Outputmaxxing
Okay, never mind. It was not healthy—
[laughter]
—in hindsight, but maybe this is TMI. Anyway, outside that conference room, on the wall, was this quote that was printed that said, “Culture is not a set of beliefs. It’s a set of actions.” And it's by Bushido, a Japanese philosopher.
If you stop taking the actions that demonstrate the mission alignment to what you’ve said to your team and to the world matters to you, then your culture starts to fray. So it’s not actually a moat, I would say. It’s a very, very brittle, fragile thing that requires daily tending to, like a garden.
But if you figure out the system to keep that garden tended, which I think ultimately comes down to knowing yourself, because most naturally, if you’re authentic and so on, you’ll naturally make trade-offs that seem effortless to you but that reinforce your culture, then that becomes this very hard thing for other people to catch up to.
At Anthropic, from day 1, there was this mission-like, missionary-like zeal and belief that, hey, these capabilities will scale. These systems are stochastic, not deterministic. There will be error bars. And until we crack interpretability, there’s risk.
Yeah.
The Professor of Outputmaxxing
At some point, people will stop using Claude just for coding. They’ll use it in some mission-critical context where it’ll throw off a bug. And then people are going to come blame them.
They want to be on the right side of history, where they said, “Yes, this is a powerful technology. We think it’s going to change the world. And we want to be very measured and scientific about the fact that, hey, guys, these are statistical models. That’s how statistics works. Ultimately, when you’re training neural nets, it is just a statistical system.”
And I think that belief that safety is important—and that it might seem toy-like in the early days, and sometimes there you could say Anthropic totally exaggerated the risk, like 2 years ago when they said, “Let’s not launch Claude 1” or whatever—well, okay, maybe in hindsight. But hindsight is 20/20.
At the time, they didn’t know how that model would be used. And to them, it felt existential if somebody came and said, “You were responsible if this wrote a bug.” The liability associated with that is massive.
So how do you prevent against that? Well, day in, day out, you say, “Safety, safety, safety, safety.” And when you start deviating from that, you have the team hold you accountable. You have the world hold you accountable. And I think that becomes a moat over time.
At some point, that moat will get challenged and so on, and then it will become fragile. I hope it endures, because that’s the beauty of having founders run the show. They can make really hard trade-offs to do mission alignment.
The hardest part is in the earliest days, when you don’t have a group of people who are going through difficulty, stress, and crisis together. Then your culture doesn’t get defined sharply enough. And that’s what I’m worried about right now. There’s so much money going to these labs; there’s no hardship. There are no—
21 nos.
The Professor of Outputmaxxing
There were 21 nos. And that, in hindsight, was a feature, not a bug, for Anthropic. The number of people who said no, the number of people who said, “Sorry, we’re already investors in OpenAI”—that is competitive defense.
It forces you to really understand what the hill is that you want to die on at the expense of everything else. What’s the P0? Their P0 from day 1 was coding.
The reasoning mechanism there was: if we crack coding, then we will crack AGI. Our mission is AGI. We want to get there safely. If we focus on coding, it’s such a generally powerful capability.
The Professor of Outputmaxxing
That it can accelerate all kinds of work on a computer. And if we can accelerate all kinds of work on a computer, we can get to AGI. You know, as a result, they've had to say no to so much other stuff.
Here, superconductivity is the mission. Coding is not the mission. So we use Claude. We'll use Claude. We don't care about that. The mission defines everything.
And I think teams who can raise too much money too fast, too early, who don't have to define what the P0 is, because that's the only thing, when you have scarce resources, you've got to—
The Professor of Outputmaxxing
Invest in. Those cultures end up being the most fragile and brittle. And they almost don't even make it to takeoff.
So, let's apply this to Periodic, since we're here. Sure. What is the constraint or the hardship that they were forcing themselves to go through?
The Professor of Outputmaxxing
Dude, here? Here? Are you crazy? No. Well, yeah, okay. On a technical level, it's physics. It's literally reality.
I mean, is there another one that's like company-building?
The Professor of Outputmaxxing
Yeah. Liam was a co-creator of ChatGPT. And Doge was skip level from Demis at DeepMind. He had created Gnomics, one of the most important tools to come out of DeepMind.
At the time, I was a visiting scientist at the Stanford physics department, and we had started benchmarking frontier models on physics and science capabilities. They were not very good. They were good at doing things like summarizing papers, but if you said, “Hey, could you analyze the scientific data coming out of a condensed-matter physics lab?”—I was in the condensed-matter physics group at Stanford—they were terrible.
So, it was not popular 12 months ago. I won't go into details, but there were people who said, as recently as a few months ago, that they wanted to join the company. For whatever reason, they took a job elsewhere. They kind of reneged on their commitments and took a job elsewhere that offered more money.
Then we had a technical breakthrough and created a SOTA system.
Okay.
The Professor of Outputmaxxing
I was excited.
Yeah, we'll cover it. We'll be doing a separate part on Periodic.
The Professor of Outputmaxxing
And then they wanted to come back. And I said, “No. No way. If you come here, you had your shot. You had your shot.”
Because it's actually about culture.
The Professor of Outputmaxxing
Of course.
Principles, yeah.
The Professor of Outputmaxxing
You know, and look, I believe in second chances and so on, but time will need to heal. Some of those wounds will leave deep scars. But because I started my company at 24, 25, I went through the whole cycle of betrayal and drama, and so you realize that Silicon Valley is both a very missionary place and a very mercenary place.
Sometimes people lose their minds when big money gets involved, which is, in the grand scheme of things, quite small money.
Life-changing to me, maybe less to you, but a lot of people have not been taught how to deal with money. And, yeah, we didn't come from that kind of privileged background, right?
The Professor of Outputmaxxing
I'm a street dog, man. Look, I grew up in Rishi Valley. We didn't have much. This was enforced brutalism. Jiddu Krishnamurti started the school, and it was like, “You will sleep on a hard slab of stone.” The mattress was this thin.
Then you go up to Singapore. When I got to Singapore, I was part of the scholarship program, which was amazing. I'm very grateful to the Singaporean government. But I was at St. Andrew's JC, and our dorm, which was by Boon Keng MRT—
Which is not a prestigious neighborhood.
The Professor of Outputmaxxing
It was a transition dorm because you were building this beautiful residential campus on-site at SAJC in Potong Pasir. We were the last—I think the second-to-last—batch to be at the transition site, which was some old immigrant-labor housing.
Yes. That's where we keep the people who work in the factories and stuff.
The Professor of Outputmaxxing
Right. So I lived there during my 11th and 12th grades. I slept in a bedroom the size of this, literally from there to here.
Yeah.
The Professor of Outputmaxxing
There were bunk beds: one bunk bed here, one bunk bed there, one on top, one on top, one more here, and then here was where we kept our toiletries and clothes and stuff. And when one guy would climb onto his bed there, this one would shake.
Oh my God.
The Professor of Outputmaxxing
It was amazing. I loved every minute of it. One of my roommates was a top-ranked Dota player from BRC, from China. He didn't speak a lick of English. I loved him. Amazing guy.
I mean, all the Singapore scholars are fantastic. And honestly, we should treat you guys better because of what you're going to do, but it's cool to know.
The Professor of Outputmaxxing
No, I mean, what I'm saying is I don't need much to be happy in life. When you've lived through that, money is a way we sometimes measure ourselves. But when it stops becoming just a byproduct and becomes more of a measure, it stops having meaning. It invokes Goodhart's law.
You use it to do more meaningful things. It's a resource to pursue your mission. I've kept you longer than I was supposed to, but we should continue this in Part 2.
You know, I really enjoyed this. Yeah, I mean, you're so inspirational, and there's more I want to dig into about how you've set everything up, every single one of your investments, how AMP is going, but we're running out of time for that. Thank you so much for joining us.
The Professor of Outputmaxxing
Good to see you, man. Let's get chicken rice sometime.
Yes. Actually, tomorrow, I'll send you the details. I'm hosting a birthday party.
The Professor of Outputmaxxing
I don't get an invite?
To a Singaporean birthday party? Yes. You get an invite right now.
The Professor of Outputmaxxing
Okay, cool.
All right, thank you.
The Professor of Outputmaxxing
All righty, thanks, man.