将算力推向物理学极限
Verdon 的核心判断是,AI 的概率型工作负载运行在为压制噪声而消耗能量的确定性硬件上,最后又由软件把随机性加了回来。 Extropic 的混合信号芯片则利用电子的随机动力学,加速离散变量和连续变量上的马尔可夫链蒙特卡洛。其主张是“放松对电子的控制”(“loosen our grip on electrons”),把熵从负担变成计算资源。
规模化的逻辑是物理性的:功耗越低,需要的电荷越少;但当电荷变得稀缺,其离散性会产生不可避免的噪声。 传统机器将错误率控制在约 (10^{-15}),为在“热危险区”内维持确定性而承担不断上升的能耗。Verdon 因此认为,热力学计算不可避免:“你迟早得走向热力学计算。”
Extropic 称,公司已从 3 个 pBit 的超导原型,推进到拥有约 300 个自由度的硅芯片,并计划明年扩展到数百万个自由度。 据称,其可控 pBit 只需“数百阿焦耳”即可产生熵;相关结果已提交同行评审,公司承诺在节目录制后提供私测 alpha 硬件、论文和开源软件。公司目标是实现芯片级 1,000–100,000X 的能效提升,但 Verdon 明确将其与整机的制冷成本区分开来。
一个对投资者最有参考价值的预测是晶圆级估算:20瓦功耗下,约可容纳15亿个 pBit 和200亿参数。 Verdon 将其与当前功耗20千瓦、甚至可能达到100千瓦的晶圆级系统对比,并认为,一个拥有数千亿参数的热力学模型,能效或许可以做到接近人脑的10X以内,而不是今天所称的1亿倍差距。这些都是前瞻性预测,并非已经验证的系统级基准。
近期的逻辑是叠加,而不是彻底替代 GPU。 Verdon 预计,确定性处理器将继续承担经典功能,概率芯片负责采样,量子硬件则补充涉及量子系统的工作负载。他表示,现有建设“在可预见的未来”仍然安全,因为模型迁移需要时间;即使出现“好1,000X”的技术,也只是改变未来的电力与算力比。
改变底层介质,可能改变最终胜出的模型架构,因为“transformers 并非神圣不可侵犯”。 Softmax、扩散模型、思维树搜索、蒙特卡洛树搜索、强化学习 rollout 和测试时推理,已经暴露出对采样密集型算力的需求;高效的物理采样可能让能量模型重新获得发展空间,并引发一场针对概率硬件优化的架构“寒武纪大爆发”。
Verdon 将去中心化算力与一项政治目标联系起来:个人应当拥有那些持续在线、能够延伸自身认知的系统。 他拒绝把仅仅获得一个集中控制的“唯一上帝模型”视为真正的自主,而认为那只是被同化进“Borg 思维”。在神经硬件出现之前,他预测人类与 AI 会先通过共享感知和行动实现一种柔性的融合。他更广泛的 e/acc 主张——“加速,否则死亡”——偏好增长、方差和反垄断式试验;Ramstead 则追问癌症式增长、威权主义局部最优和灾难性尾部风险。
1. 还原论在复杂系统成为研究对象后失效
Verdon 回溯了自己的求学路径:7岁时就着迷于亚原子粒子,之后在 McGill 学习数学和物理,又在 Waterloo 和 Perimeter Institute 研究理论物理。他原本期待找到一个简洁理论,或许再加上一种奇异推进方式,让人类扩张变得可行。
思想上的断裂来自凝聚态和量子引力问题:它们无法被还原成少数几个可解释参数。即使掌握微观方程,预测涌现行为也可能需要“规模类似整个宇宙的计算量”;许多系统无法通过解析方法简单地重整化成干净的高层规律。
Verdon 将接受不透明的计算表示称为一次“自尊死亡”:人类物理学家可能并不是“故事的主角”。张量网络和深度学习系统可以像自然本身一样难以解释,但可编程的复杂性仍然能够学习压缩表示,并提供预测能力。
2. 量子信息搭起从物理到 AI 的桥梁
“it from qubit”计划是对 Wheeler“it from bit”的推广,将宇宙视为量子计算机,或视为一个自我模拟系统。无论 AdS/CFT 是否真正带来了统一,Verdon 最终都认为,复杂主义,以及能够匹配自然复杂性的计算,是更有生产力的框架。
他的第一个答案是“以火攻火”:用可调控的量子复杂性去建模量子复杂性。他参与开发了早期的参数化量子程序,即量子神经网络;后来成为 Rigetti 的早期用户,再加入 Google 开发 TensorFlow Quantum,最终在 Alphabet X 负责量子机器学习。
其中可复用的原则是底层介质匹配:受物理约束的表示,应当运行在其原生动力学实现相同物理过程的加速器上。量子表示应放在可控的薛定谔演化上,随机表示则应放在可控的随机演化上。
3. 量子制冷让更热的计算机变得有吸引力
在量子计算领域近8年后,Verdon 认为进展过慢,于是转向研究容错的热力学负担。量子机器试图维持接近绝对零度、接近零熵的状态,并通过他所谓的“算法式制冷”持续抽走噪声。
他的思路反转很简单:让环境噪声进入系统,直到量子相干性消退,有用动力学转为随机动力学。与量子计算不同,他认为热力学加速并不存在复杂度类上的分离,只有速度和能耗上潜在的巨大常数级收益,足以让数字仿真在规模化时变得不切实际,但并非不可能。
Ramstead 进一步明确了定义:基于物理的计算,是利用组件真实的物理属性,而不是对其进行数字仿真。Verdon 则保留了宽泛的分类,包括量子、随机、光子,甚至包括通过“水洼”训练出来的神经网络;同时,他区分了基于物理的执行与受物理启发的软件。
4. 确定性有能耗代价,Extropic 选择利用噪声
经典硬件通过让信号远大于电子抖动,迫使晶体管进入绝对的开或关状态。Verdon 援引麦克斯韦妖认为,“知识是有成本的”:降低熵、在每个时钟周期维持确定性,必然需要消耗能量。
他的替代方案是“学会放手”(“learn to let go”):不再紧紧拽动信号,而是“温和地引导”噪声动力学走近平衡。这个哲学类比是有意为之:软件已经从命令式编程交给了梯度下降,硬件也应停止把每一次波动都当作敌人。
从技术上看,Extropic 正在打造用于马尔可夫链蒙特卡洛的加速器,覆盖离散变量、连续变量及二者的混合。其最新芯片采用混合信号架构:随机电子学负责生成提议或熵,传统数字组件则执行 Metropolis–Hastings 接受与拒绝等非随机操作。
5. AI 软件是概率性的,哪怕运行它的机器不是
Ramstead 指出了计算栈的悖论:硬件耗费能量移除随机性,采样算法却在软件中把随机性恢复回来。Verdon 补充说,transformer 的 softmax 已经让它成为“一台大型概率计算机”;扩散模型在反转马尔可夫链,而测试时推理越来越多地使用蒙特卡洛树搜索、思维树搜索或 RL rollout。
Extropic 成立于2022年,建立在一个判断之上:主要工作负载将转向采样和概率推断。Verdon 将当前架构描述为确定性前向传播与完整能量模型之间的插值,扩散模型处于这条连续谱的中间位置。
2 位投资 Extropic 的 transformer 论文作者反复告诉团队:“Transformers 并非神圣不可侵犯。”它们之所以有效,是因为适配了 Google TPU;如果改变硬件的适应度景观,Verdon 认为,模型搜索可能进入一个高温阶段,出现原生适配概率计算的架构“寒武纪大爆发”。
6. 摩尔定律在热危险区变成摩尔墙
传统处理器通常将错误率控制在 (10^{-15}) 附近,以便在需要纠错前完成海量运算。但微型化会提高波动相对于信号的比例:“要使用更少的功率,就需要使用更少的电荷”;当剩余电荷很少时,其离散性会周期性地产生噪声。
Ramstead 用网球作类比,区分了不同区间:宏观运动可以忽略分子振动,量子机器维持相干叠加态,而热力学设备则工作在波动与组件尺度信号共存的区域。Verdon 的判断是,越来越小的确定性晶体管无论如何都会被推向第3种状态。
Verdon 认为,人脑证明了有用的热力学智能确实存在;Ramstead 则认为,人脑的运行状态接近 Landauer 极限。Verdon 同意,神经递质通过随机化学反应网络运动,并认为电子电路或许能更高效地执行类似的概率程序,因为“电子比大型神经递质轻得多”。
7. pBit 将能量景观转化为可编程概率
Extropic 的芯片可以被视为一个随时间变化、可编程的能量函数,其扩散过程类似朗之万动力学,而后者本身就是一种 MCMC 方法。这个类比是字面意义上的:贝叶斯采样算法可以嵌入电子的原生运动,而不是逐步进行数值仿真。
一个 pBit 类似于包含弹跳小球的双势阱景观。一个势阱代表0,另一个代表1;调节倾斜程度,可以控制信号在每个状态中停留的时间,由此形成一个持续在0和1之间跳动的“分数比特”。
公司最初从3个超导 pBit 开始,随后在硅上复现了随机原语;Ramstead 将最新芯片描述为拥有约300个自由度,但 Verdon 提醒,它们并不全都是 pBit。据他介绍,芯片产生熵的能耗只有数百阿焦耳,并计划明年扩展到数百万个自由度。
8. 硅是产品,超导体是学习平台
超导电路量子电动力学能够将目标哈密尔顿量直接映射为电路,因此成为早期有用的工程平台。根据材料不同,这些器件运行在数百毫开尔文或数开尔文的温度;维持低温边界本身就是一项主要的实际成本。
Verdon 承认一个极端思想实验:如果把数百万个超导 pBit 放进一台足球场大小的稀释制冷机,可能会形成最高效的概率计算机,但这“可能并不”适合创业公司,也不具备现实可行性。他更愿意将这项工作交给学术界和国家实验室,作为社区平台。
Extropic 为商业化选择了室温硅。其宣称的 1,000–100,000X 能效区间适用于芯片级别,同时承认制冷成本这一前提;温度更低的热力学器件本地能耗更少,但维持冷浴也会产生自身成本。
9. 胜出的技术栈是异构的,而非处处热力学
Verdon 想象计算沿着不同尺度上的物理规律展开:早期层可能保留量子复杂性,中间层保留概率和熵,后面的确定性层则在信息完成提炼后执行坐标变换。
实际计算图同样可以将定义分布的经典可微函数,与从该分布中采样的概率加速器结合起来。“你不一定需要在计算图的所有位置同时拥有熵”;热力学硬件也不会适合每一种确定性或高精度操作。
量子计算机或许仍将作为量子系统的补充;概率处理器和确定性处理器则覆盖大多数其他工作负载。因此,Verdon 告诉计划建设 GPU 和 TPU 基础设施的企业,迁移“需要相当长时间”,热力学算力起初应作为附加组件出现,未来可见的建设仍然可用。
10. 能量模型同时暴露上行空间与功耗约束
Verdon 将现代神经网络描述为能量模型的平均场近似:确定性硬件推动研究者转向矩、Gaussian 分布族、矩阵和向量,因为这些对象能够高效映射到 GPU。更快的物理采样可能让更广泛的 EBM 家族变得实用,而不必再将分布压缩成硬件友好的摘要。
在晶圆级别,他预计20瓦功耗可以容纳约15亿个 pBit 和200亿参数,而当前晶圆级系统的功耗为20千瓦,甚至可能达到100千瓦。他表示,一个多层、拥有数千亿参数的系统,能效或许可以接近人脑的10X以内;相比之下,他认为当前差距接近1亿倍。
Verdon 的明确约束是,当前硬件无法规模化支撑无处不在的智能体、视频模型、世界模型和具身智能。即使聚变能源足够丰富,最终也会增加行星向外辐射的热量:“我们真的会被煮熟。”因此,目标市场不应局限于生成式 AI,还包括仿真、优化、统计推断和科学计算。
11. e/acc 将热力学选择应用于技术与政治
Ramstead 引入 Effective Accelerationism,将其作为有效利他主义的对照;Verdon 将其定义为一种“文化超参数处方”,在对数尺度的卡尔达肖夫文明等级上最大化文明自由能的生产与消耗。他以随机热力学为框架,从自私基因和自私模因推进到“自私比特”:耗散更多自由能的轨迹,其出现概率会呈指数级上升,因此“加速,否则死亡”成为这一运动有意为之的戏剧化口号。
Ramstead 的反驳值得保留:无边界增长可能产生癌症、独裁或被强加的同质化。Verdon 回应称,这些都是短期视角下的局部最优:癌症会杀死宿主,而立即烧光一切会阻止未来被接管。目标应是在事实上无限的时间跨度上进行战略性增长,保存秩序和资源,以便未来获得更多资源。
Ramstead 通过主动推断解释“超迷信”机制:感知和行动可以减少模型与世界之间的偏差,“车会开向眼睛注视的地方”。随后,他举出一个有争议的例子,称 COVID 是防御性生物武器研究中的一次事故,并加上限定,认为这“可能是防御性的”。Verdon 没有实质性支持这一例子,但认同更广泛的乐观引导逻辑,并表示每个行动都始于一个错误信念。
在地缘政治上,Verdon 偏好方差而非垄断:美国像高温搜索,中国像低温优化器,在梯度明确后负责规模化。DeepSeek 在出口管制约束下进行探索,是他用来说明另一片超参数区域、并暴露西方收敛状态的例子;他的处方是自上而下的协调与自下而上的探索相结合,因为他的“P90,1984”高于对 AI 的“P doom”。
That was the most technical podcast I've ever done. I think that was one of my favorite conversations for sure.
Hey, I'm Guillaume Verdon. Growing up, I was really pursuing theories of everything. I wanted to understand the universe. I was a big fan of Feynman and Stephen Hawking growing up. When I was 7, I was talking about subatomic particles—
As all 7-year-olds do, right?
I guess I got swept up in the school of thought that was a generalization of Wheeler's “It from Bit,” which was the “It from Qubit” program, seeking to unify theoretical physics through quantum information theory. So, viewing everything in the universe as one big quantum computer running a certain program or a self-simulation.
We have proof of existence of a really kickass AI supercomputer that we're both using right now to talk to each other. It's our brains. That, I would argue, is a thermodynamic computer, right? Because there are master equations describing the chemical reaction networks in your brain, with neurotransmitters hopping around. And so, if we're doing very similar physics but with electrons sloshing around a circuit, then there's a much stronger chance that we could run a very similar program in a very similar fashion, arguably with even more energy-efficient components, because electrons are much lighter than big neurotransmitters.
This podcast is supported by Google. Hi, folks, Paige Bailey here from the Google DeepMind DevRel team. For our developers out there, we know there's a constant trade-off between model intelligence, speed, and cost. Gemini 2.5 Flash aims right at that challenge. It's got the speed you expect from Flash, but with upgraded reasoning power. And crucially, we've added controls like setting thinking budgets, so you can decide how much reasoning to apply, optimizing for latency and costs. So try out Gemini 2.5 Flash at aistudio.google.com, and let us know what you build.
Hey, I'm Guillaume Verdon. I'm the founder of Extropic, a company pioneering thermodynamic computing, a new form of computing for probabilistic inference using exotic stochastic physics of electrons. Formerly, I was working on quantum computing and machine learning at Alphabet. I also happen to be the founder of a philosophical movement called Effective Accelerationism, under the pseudonym Beff Jezos online. Happy to be here at MLST.
Hello, everyone. Welcome to Machine Learning Street Talk. I'm your host for today, Maxwell Ramstead. I'm subbing in for Tim Scarfe, who unfortunately is down with a little bit of COVID. But I'm sure we'll have a really interesting conversation. It promises to be a super interesting conversation because we have a really awesome guest for today. I'm very excited to have Guillaume “Gil” Verdon, also known as Beff Jezos on X and generally online in the meme space. Very excited to have you, Gil. How are you doing?
Yeah, thanks for having me. Thanks to MLST for hosting, and it's great to talk to you again, Maxwell. I think this collaboration, at least for this conversation, has been a long time coming, so it's a great platform to do it. And, yeah, let's get to it.
So you're a very interesting character, Gil. You're a fascinating researcher. I think you're a big presence in the meme space as well. Do you want to tell us a little bit about yourself? Tell the audience a little bit about your trajectory. How did you wind up being the CEO of a thermodynamic hardware company and also a well-known person in the meme space?
1. From Physics To AI
Yeah, I mean, it's been a long journey. I guess I've lived many lives. Growing up, I wanted to pursue theories of everything. Really, I was pursuing theories of everything. I wanted to understand the universe, as one does, and eventually leverage that knowledge to expand civilization to the stars.
Originally, my plan was to become a theoretical physicist and work on quantum gravity, and figure out some exotic form of propulsion that would obviously—well, I thought that was the bottleneck for the expansion of civilization—was the speed of propulsion. And so I went down that path. I was a big fan of Feynman and Stephen Hawking growing up, and—
So this is a childhood project? Since you were a wee boy, you've dreamed of—
Yeah, no, I—
That's cool.
I was 7. I was talking about subatomic particles and how—
As all 7-year-olds do, right?
Yeah, and I wanted to be a physicist. I used to call it an astrophysicist. I didn't know what a theoretical physicist was. But I wanted to be an astrophysicist. I wanted to work on FTL travel and stuff like that.
Over time, I did the career path through theoretical physics. I did math and physics in undergrad at McGill, and then I went to Waterloo, the Perimeter Institute for Theoretical Physics. I met some of the greatest minds on Earth. They're really strong caliber. But what I realized going through theoretical physics was that the reductionist approach to physics was failing us, right?
Do you want to, just for the benefit of our audience, explain what you mean by the reductionist approach? And what do you think were the problems?
Yeah. I guess these two sides to my life now came from my reaction to realizing this failure. The reductionist approach is the traditional way we've done physics. It's the very rational, old-school way to do things, which is like, “Oh, I want to have a model.” Maybe it's an equation, some analytic model, and it has a few parameters. I've reduced all of physics to a very simple model with few parameters, and ideally the least amount of parameters possible—Occam's razor principle.
These equations with these few parameters allow me to have predictive power over the world, and with this predictive power, I can steer the world. I can do things, right? I could predict and control it. That's the goal of physics: to have better models of the world.
2. Complex Systems Replace Reductionism
What we realized was that, as part of this, I got swept up in this school of thought that was a generalization of Wheeler's “It from Bit,” which was the “It from Qubit” program—
Hmm.
—which is seeking to unify theoretical physics through quantum information theory, right? So, viewing everything in the universe as one big quantum computer running a certain program or a self-simulation.
That framework's actually a really useful unifying framework to understand all sorts of systems, and people were studying systems that have some connections between quantum gravity and regular quantum mechanics in the context of holography. So, AdS/CFT, for those that are familiar. But really, what I realized there was that even whether or not AdS/CFT was going to work, it was clear that the complexism approach to physics was the way forward.
The universe is very complex. Even beyond just quantum gravity, looking at condensed-matter systems, you can have equations describing the microscopics, but you can't predict the emergent properties, right? You actually have to apply an amount of computation that is similar to that of the universe to have a prediction of what happens at a larger scale.
Not all equations you can renormalize analytically. Renormalize means getting effective physics at a larger scale, at a more coarse-grained scale, right? That's how we go from different types of physics—from quantum field theory to quantum theory, to statistical mechanics, and then eventually classical Newtonian mechanics, and then even beyond that we get to general relativity and so on.
So clearly, the way forward was to understand everything as a complex system. But that was a sort of ego death because, in a way, the human can't necessarily be the hero of the story. The model is no longer interpretable, right? If you have something like a deep-learning system—back then we were looking at tensor networks, which are a different sort of parametric complex system that we use to model, let's say, condensed-matter systems or quantum-gravity systems.
If you look at these systems, they're no longer interpretable, right? I have these big networks, and they're just as opaque as the system that I was trying to predict.
With numerics and a lot of compute, you're not deferring agency; you're leveraging a programmable, parametric complex system to grok a complex system of nature for you.
Tim Scarfe
Mm-hmm.
In a way, it was like, okay, maybe I can't be the hero of the story. I won't be the one to figure out the grand unifying theory of physics, but maybe I could build a computer, or computer software, that can understand the universe or chunks of the universe for us.
Tim Scarfe
Right.
The more I dug into it, it was clear that this was gonna be the way forward, right?
Tim Scarfe
So basically, building a digital brain to overcome the limitations of our fleshy brains.
Yeah, that's correct, but then at the time we were studying quantum mechanical systems. So you had—
Tim Scarfe
Right.
Systems that have quantum complexity, and there you can show that classical representations will struggle to capture quantum correlations. So already—
Tim Scarfe
Mm.
I came at it from a non—
Tim Scarfe
Mm.
Anthropomorphic form of intelligence. To me, intelligence was just trying to learn compressed representations of systems through parameterized distributions, or, in my case, parameterized wave functions and density matrices, right? Which is the more general form.
So that actually got me into a field. Initially, it was mostly numerics with tensor networks, but there was another institute I was part of at the University of Waterloo, which was the Institute for Quantum Computing, right?
Tim Scarfe
Right.
And there it was very interesting because we had these controllable quantum systems. We had these control parameters, right? It became clear to me that there was maybe a way to run these representations we were trying to learn of quantum systems on a parametric, programmable quantum system, right?
Tim Scarfe
Right.
And now we could fight fire with fire, right? We can have parametric quantum complexity that's tunable to understand the quantum complexity of our world, right?
Tim Scarfe
Right.
3. Quantum Computing Becomes Machine Learning
And that was actually my entry into artificial intelligence. I wrote up some of the first algorithms for quantum neural networks, so that's what we called them. They have nothing to do with actual neurons; they're just parameterized quantum programs.
Essentially, I was one of the first quantum computer programmers and the first user of Rigetti, which was one of the first startups. That got me on the radar of Google, which then approached me and a team I built at Waterloo—an open-source team—to go work at Google and build a product known as TensorFlow Quantum. It was a product focused on creating software that allows us to learn quantum mechanical representations of quantum mechanical systems in our world.
That was my entry into AI. Given that this is a technical podcast, I'm going a bit deeper than I usually do, which is really nice for the backstory here. But over time, what I realized was that this sort of physics-based approach—if we're trying to understand the physical world with representations, we want to have physics-informed representations or physics-inspired—
Tim Scarfe
Mm-hmm.
Representations. Just like, to understand a quantum mechanical system, I use a quantum neural network. The dual to that is, if I want to run and learn and train these representations that are physics-inspired or physics-based, then I need a physics-based accelerator, right? That's where that type of program fits natively and can be executed as physics.
In my case, initially, it was: I wanna understand quantum mechanical systems, I wanna learn quantum physics-based programs, and I run them on quantum physics-based processors, right?
Tim Scarfe
Then I have a clarification question, just for the benefit of the audience.
Yeah.
Tim Scarfe
We have a very technical audience, but let's just make sure everyone follows. What do you mean precisely by physics-based computing? I ask because it might confuse some of our listeners. In some sense, all computing is physics-based, right? It's all shuffling around electrons in circuits and so on. What do you mean specifically by physics-based computing, and how would that differ from the commonsensical or classical notion?
Yeah, a quantum computer is a computer that leverages quantum mechanical evolutions as a resource, I would say. There are more formal definitions of what a quantum computer is.
A physics-based computer, I would say—of course, like every physical thing is embedded in the physical universe, so everything is—
Tim Scarfe
Right.
Physics-based. But it's true that it's kind of a continuum, right? Because you can have physics-inspired computers that are—
Tim Scarfe
Mm-hmm.
Doing digital emulations of physics-inspired—
Tim Scarfe
Right.
Algorithms. So I would say that's physics-inspired. I would say a physics-based computer, at least for a quantum mechanical computer, is one where you have a parameterized Schrödinger evolution that you control, and you can show that it has some unitarity.
In our case—and we'll get to that—we're building physics-based stochastic computers, or stochastic thermodynamic computers.
Tim Scarfe
Mm-hmm.
They're parameterized stochastic evolutions that we control. In principle, you could try to emulate a quantum mechanical computer or emulate a stochastic computer digitally. Of course, in the case of a quantum mechanical computer, there are proven separations of complexity there, and there are experiments showing that you can't emulate them at scale. In our case—
Tim Scarfe
Right.
With stochastic computers, there's no complexity-class separation. It's just orders-of-magnitude constant speedup or energy-efficiency gain, which practically makes them intractable to emulate at scale, right?
Tim Scarfe
Right.
Not impossible, though, of course.
Tim Scarfe
Right.
But yeah. So this journey—and feel free, if you have a better definition of physics-based computers, I think you have a better definition. I don't know if you wanna go into it, but—
Tim Scarfe
Well, no, it's consistent with what you were saying. At a very high level, I would say that a physics-based computer is essentially a computer whose components exploit the actual physical properties of the components in order to perform computations more efficiently.
Yep.
Tim Scarfe
So rather than use digital computation to simulate what, for example, is a quantum phenomenon or a thermodynamic phenomenon, you actually use the probability densities that these systems embody, for example, at equilibrium, and exploit those properties in computation. Would you agree with that as a broad definition?
That's for a thermodynamic computer, right? But yeah, there are different kinds of physics and different kinds of corresponding physics-based computers. Arguably, you can have probabilistic or quantum photonic computers.
Tim Scarfe
Right.
Of course, you could have physics-based computers in all sorts of substrates. Technically, you could probably—
Tim Scarfe
Mm-hmm.
Train a neural net from puddles of water, right, if you wanted. And that would still be a physics-based computer.
Tim Scarfe
Right.
So I think it's kind of a very broad community. I would say physics-based computing, if we include—
Tim Scarfe
Mm-hmm.
Quantum computing and alternative computing, has all these archipelagos that don't necessarily talk to each other. I think there could be more work in trying to unify the community, because there are a lot of tools that could cross-pollinate between these different substrates.
4. The Quantum To Thermo Pivot
But to get back to our main line, which is: how did I end up going from quantum to thermo? I think going from being one of the first programmers of AI on quantum computers and seeing the field evolve over several years—I was almost 8 years into quantum computing.
Tim Scarfe
Mm-hmm.
Pushing compute to the limits of physics
It wasn't progressing as fast as I wanted, and I was seeing a sort of writing on the wall where, actually, to have a quantum mechanical computer, you try to keep the computer at perfect zero temperature, right? Zero entropy. So you're constantly pumping out entropy that's seeping into the system through noise, right? And that's quantum error correction and fault tolerance. It turns out that most of your computation, most of your energy, is going to be sunk into that pumping. So it's like an algorithmic form of refrigeration.
Tim Scarfe
Right.
Pushing compute to the limits of physics
To me, it's like, okay, well, if you have a fridge and you're trying to keep your freezer really cold, it's going to use up a lot of energy if the rest of the room is at room temperature, because the gradient of temperature is very difficult to maintain. So it's like, okay, what if we had a hotter, physics-based computer?
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Maybe it would be much easier to maintain. That's just a first thought, right? So what would that look like? In the limit where you let the noise seep in, things become mostly stochastic, and the quantumness kind of fades out. We know this from renormalization of physics. That's why we don't have to care so much about quantum physics day to day, because it gets washed out at larger scales, right?
Tim Scarfe
Yeah. Would you say that it's accurate to frame it this way? In classical computers, the scale of the fluctuations is dwarfed by the scale of the components of the system that you're considering. So you can basically ignore the noise for the most part, assuming that you cool the system appropriately.
In quantum computing, it's almost the other way around, where what you want to do is basically engineer the scale of the fluctuations so that they're so big that you can kind of ignore them relative to the component size. This is just an attempt at framing what you're doing, so feel free to disagree with this. And then, in the thermodynamic regime, what you have is fluctuations that kind of coexist at the same scale as the components that—
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
—the hardware is made out of.
Pushing compute to the limits of physics
Yeah, I mean, we're in Wimbledon weekend right now in London, and you don't have to understand the vibrations of the molecules at the molecular level in the tennis ball to be able to predict its trajectory, right? You kind of get a mean-field sort of prediction, right? So that's Newtonian.
At the quantum scale, it's more subtle because it's no longer probabilistic fluctuations.
Tim Scarfe
Right.
Pushing compute to the limits of physics
You're essentially dealing with purely quantum superpositions, right? And the ideal—
Tim Scarfe
Right.
Pushing compute to the limits of physics
—quantum computer has no probabilistic uncertainty. In fact, when there is probabilistic uncertainty, that's when the quantum computer loses its quantum coherence. It becomes non-quantum, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
But we have a very similar problem with classical computers—
Tim Scarfe
Absolutely.
Pushing compute to the limits of physics
—which was that we use all this power to have a signal that's so strong relative to the jitter of electrons and the amplitude of that jitter in order to maintain determinism, right? We want our transistors to be absolutely on—
Tim Scarfe
Mm-hmm.
Pushing compute to the limits of physics
—or absolutely off, because we have a lot of transistors, and we want the computer to be in a deterministic state that we have control over, right? Humans like to be in control, right? When we're not in control, there's anxiety, right?
Tim Scarfe
Mm.
Pushing compute to the limits of physics
I think we have to learn to let go. Just like I learned to, in a way, let go of control mentally by giving up the reductionist approach, which is the human-interpretable approach to physics.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
We have to sort of, just like we did with deep learning, let go of software 1.0, imperative programming, and let gradient descent be a better programmer than you. That was very humbling for some.
Some are still trying to hang on for control, right, with ML interpretability research and all this AI safety-ist research. They still want to feel like they're in control or that they understand the complex system, instead of letting it do its thing and letting it figure out what's best, right? That's also kind of my main gripe with the whole AI safety-ist field. Of course, there should be some research in interpretability, but I just don't think they're going to go that far.
To close the loop here, I think we should also literally let go and loosen our grip on electrons and hardware, and that's what we do, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
Essentially, we go from having a very tight grip and yanking these signals around to kind of loosening our grip, letting there be fuzz, and gently guiding the signals, right?
Tim Scarfe
So, in some sense, you're kind of switching teams, right? In the more classical way of thinking of this, what we're trying to do is to keep the noise at bay, right?
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
To make our systems as non-stochastic as possible.
Pushing compute to the limits of physics
Filter it, pump it out.
Tim Scarfe
But we—
Pushing compute to the limits of physics
Yeah. Pay the price.
Tim Scarfe
That's right.
Pushing compute to the limits of physics
Right?
Tim Scarfe
But what—
Pushing compute to the limits of physics
Exactly. What you're doing is saying, "No, no, no"—harness it, right?
As we know from Maxwell's demon, the classic fable—people can look it up—knowledge comes at a cost, right? Like reducing—
Tim Scarfe
That's right.
Pushing compute to the limits of physics
—reducing entropy in a system, keeping something in a deterministic state, always costs you energy.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
If every clock cycle we're trying to prevent the classical or quantum computer from decaying to a naturally probabilistic and relaxed, thermodynamically closer-to-thermal state, then we have to pay the price. We have to pay energy to maintain determinism and reduce entropy.
Whereas a thermodynamic computer is not always at equilibrium, but we're dancing much closer to equilibrium, and it's much cheaper energetically to sit in those states and maintain those states, right?
Tim Scarfe
Do you want to maybe just walk us through how you're actually designing these things?
5. Building Stochastic Accelerators
Pushing compute to the limits of physics
For a technical audience, I guess they're Markov chain Monte Carlo accelerators, right? We support discrete variables, continuous variables, and mixtures of the two. Essentially, we've found a way to harness the natural stochastic physics of electrons in order to accelerate Markov chain Monte Carlo.
There are some subtleties in how we do that mapping, but you could just imagine we're embedding into the stochastic dynamics—which are parametric, and whose parameters we control—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—we're embedding our continuous- or discrete-variable MCMC into the dynamics of the electrons in the device, right? So it's partly analog, and it's stochastic, but our most recent chip is a mixed-signal chip. We use digital classical components and sort of stochastic electronics, and the 2 have to—
Tim Scarfe
Right.
Pushing compute to the limits of physics
—interact. Just like in a Metropolis-Hastings algorithm, you have some components of the algorithm that have some entropy, some proposals, and then you have some non-random parts, right? Like computing—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—acceptance or rejection.
Tim Scarfe
So I guess this helps you address the paradoxical situation where—
Pushing compute to the limits of physics
Mm.
Tim Scarfe
—I was just saying, in some sense, from the hardware's perspective, you're switching sides and joining the side of noise. But from the point of view of software development, in the context of AI and machine learning, this has already happened, right? There's a kind of paradox in our current architectures where, as you just described it, we spend inordinate amounts of energy and effort pumping the noise out of the system. But then, you know, with sampling-based methods, like we then reintroduce it through the software.
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
Right?
Pushing compute to the limits of physics
Yeah, and you could think of even a transformer: you have a softmax layer at the end. A transformer is a big probabilistic computer already. So we're running probabilistic software on this stack that was made for determinism, which is highly inefficient.
And now, with test-time compute, thinking at test time, Monte Carlo tree search and all sorts of RL rollouts—that's a Monte Carlo algorithm, right? So that can be Tree of Thoughts, it could be discrete diffusion, it could be…
Now, even for content, there are diffusion models. Those are also—it’s also a Markov chain, right? You’re reversing—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
A Markov chain there. So algorithms—the biggest workloads that are eating the world and consuming a ton of energy—are actually probabilistic workloads. They are probabilistic—
Tim Scarfe
Right.
Pushing compute to the limits of physics
Graphical models. And so it’s kind of funny. I guess I think it’s just the community. There’s kind of the—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
New additions to machine learning that don’t learn the fundamentals; they just go straight to transformers or whatever’s hot.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
And then I tell them, “We’re building an accelerator for probabilistic graphical models,” and they look at me like, “What are you talking about?” It’s like, actually, you’re—
Tim Scarfe
Right.
Pushing compute to the limits of physics
You’re running probabilistic graphical models all day.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
It’s kind of funny that there’s a gap there. But hopefully, I think that by coming on this podcast and getting more of the machine learning community to talk to each other, probabilistic ML will have a resurgence. And of course, we’re trying to stimulate that, right? Because—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
There’s a co-evolution. What evolves together fits together, right? And there’s an evolution of—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Hardware and algorithms. We have, as investors, 2 of the authors of the Transformer paper, and they say—they kept telling us, “Transformers are not sacred,” right? They were what worked on—
Tim Scarfe
Right.
Pushing compute to the limits of physics
The hardware we had at the time, which was Google TPUs. And the hope is that there are going to be new foundation models, or new models that run more natively on probabilistic hardware, that are really efficient on our hardware, that—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
We’re not just importing current-day models. But hopefully, if you view the algorithmic landscape as having a certain fitness function of what you get in terms of performance and cost, that landscape is also induced by the current-day hardware that’s available at scale, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
And if you change that hardware substrate to a different substrate that has different preferences in terms of the structure of the algorithm that you’re running, then you’re changing the fitness landscape. And if you have a sudden shift in the—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Fitness landscape, you have a sort of Cambrian explosion. You have this sort of high-temperature phase of search in the landscape, right? And that’s hopefully what we’re going to cause, right? And that’s very disruptive. So incumbents have all the reasons to be skeptical. But at the same time, it’s funny that even the current sort of incumbent algorithms are converging toward sampling and more probabilistic algorithms, which—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Wasn’t obvious when we started the company in 2022, but that was the prediction, right? That there would—
Tim Scarfe
Right.
Pushing compute to the limits of physics
Be a sort of—you know, my joke is that we’re kind of interpolating between deterministic forward-pass models and full EBMs, right? And you can view—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Sort of diffusion models as an interpolation there. But—
Tim Scarfe
Well, correct me if I’m wrong, but your view is that this is essentially an inevitable development of tech, right? You—
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
You’ve talked about the thermal danger zone—
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
And kind of the inversion of Moore’s law into what you call Moore’s wall. Do you want to run us through the logic there?
6. The Thermal Danger Zone
Pushing compute to the limits of physics
Yeah. If you came from quantum computing, there, if you try to scale up your system, you have noise seep in. So again, you’re going from the quantum zone of physics to the thermodynamic zone. You’re kind of edging on it. And if you’re in the classical zone but you try to get smaller, then you’re getting into the thermal zone again, right, in terms of the physics. The jitter of electrons matters because your transistors are so small. The fact that there are very few electrons means that the jitter of the electron population matters compared to the amplitude of the signal.
Tim Scarfe
Right.
Pushing compute to the limits of physics
Typically, we like our computers to have an error rate of 10⁻¹⁵, so they can do—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Many operations before there’s an error, so that you don’t—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Actually have to run error correction, right? The only computers that usually run error correction are those that we send into space because of radiation. But the reason we can’t scale down deterministic computers is because of this thermal danger zone, and they’re innovating in all sorts of ways. They’re doing all sorts of weird fins and all sorts of weird designs that are sort of hardware-level error correction. But—
Tim Scarfe
But those are always going to be ad hoc in some sense, right? Because what you’re going to run into is just the scale problem.
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
As you miniaturize, at some point—
Pushing compute to the limits of physics
I mean, it—
These wires are going to become... Yeah.
To use less power, you need to use less charge. That’s not controversial.
Tim Scarfe
Yeah.
And when you get to very little charge, the fact that charge is discrete gives you noise, period. Right? So—
Tim Scarfe
Right.
So you’re going to have to go thermodynamic at some point, right? And at the end of the day, we have proof of existence of a really kickass AI supercomputer that we’re both using right now to talk to each other. It’s our brains. And—
Tim Scarfe
That’s right.
I would argue that is a thermodynamic computer, right? Because there are master equations describing the chemical reaction networks in your brain, with neurotransmitters hopping—
Tim Scarfe
Mm.
Around, and you could just—
Tim Scarfe
Operating really close to the Landauer limit as well, right?
Right. And so, if we’re doing very similar physics but with electrons sloshing around a circuit, then there’s a much stronger chance that we could run a very similar program in a very similar fashion, arguably with even more energy-efficient components, because electrons are much lighter than big neurotransmitters.
Tim Scarfe
Mm-hmm.
And so, my—
Tim Scarfe
So basically, what you have is a programmable Boltzmann machine, right? You can think of these chips as a kind of programmable Boltzmann machine.
Well—
Tim Scarfe
Like, you can think of these chips as, like, a kind of programmable bowl. Do you want to maybe tell us how they're used for computation?
7. From Superconductors To Silicon
For context, we started off in superconductors, and there—
Tim Scarfe
Mm-hmm.
The reason we started in superconductors was because there’s this beautiful theory in quantum computing actually called circuit quantum electrodynamics, where you go from, “Here’s my circuit and here’s my Hamiltonian,” right?
Tim Scarfe
Mm-hmm.
Directly from the circuit. And I think that should win a Nobel Prize or something someday. That’s basically how we engineer quantum mechanical systems to have certain desired physics, right? And the point there is that a Hamiltonian gives you a notion of energy in the classical regime.
Tim Scarfe
Right.
And so we had a very direct way to go from, “Hey, I want this energy. Here’s how I design my circuit,” right? And so that’s where we started.
Tim Scarfe
Right.
Furthermore, it was the way to build the most—
Tim Scarfe
Mm.
Macroscopic thermodynamic computer you can build. Any other companies with claims to build a much more macroscopic thermodynamic computer either did it digitally or emulated it, so it’s not a real thermodynamic computer. But we had to supercool it because it’s too big, right? And, you know—
Tim Scarfe
Right.
With the Boltzmann distribution, as you go smaller, you have higher frequencies, so you can go to higher temperatures, and then we moved to silicon. But to answer your question about energy-based models, you could view our chips as a time-dependent programmable energy function, right? And—
Tim Scarfe
Right.
You have something akin to Langevin dynamics, more generally diffusion, in that landscape. And Langevin dynamics happens to also be an MCMC algorithm, right? So there’s a—
Tim Scarfe
Right.
There’s a perfect match there between the algorithm that you can use for Bayesian inference for anything, really, and the native physics of the chip, right? And so it’s almost like—
Tim Scarfe
That’s so cool.
It’s been there the whole time, right? It’s like we’ve been doing algorithms that were physics we could literally implement, right? It’s a literal analogy, and we’re instantiating it as an analog computer.
That was too beautiful not to build. So we did build it. Hopefully, we can cut some B-roll. I brought some superconducting chips here. We'll do that later. But now, actually, I think the big breakthrough and the big challenge was to—
Tim Scarfe
That's it, eh?
Move to silicon, yeah.
Tim Scarfe
Wow.
Getting programmable stochastic physics in silicon was an order of magnitude more difficult, and we had to really innovate in terms of understanding the stochastic mechanics of electrons in silicon. We built our team for that. Essentially, we reproduced some core primitives. The one we're talking about for now is the probabilistic bit.
Tim Scarfe
Right.
The probabilistic bit you could think of as a double-well system, and you can—
Tim Scarfe
Yeah.
—tune the tilt and so on. So if you consider bouncy balls in this landscape, you can control how much time the bouncy balls spend in one well or another, and you can consider one well 0 and the other well 1. Essentially, you have a signal that's dancing between 0 and 1, right? And you can control—
Tim Scarfe
Right.
—how much time it spends in 0 and 1. And so that's like a fractional bit, right? So—
Tim Scarfe
Right.
Going back to what we were saying earlier about coming at it from a qubit school of thought, we basically made a p-bit from it. Initially, it was superconducting materials, but now we've done p-bits in silicon and achieved really, really high energy efficiencies. Our p-bits can generate controllable bits of entropy with only a few hundred attojoules. We've recently submitted our results for peer review, and we're going to put them on archive—probably timed with some other announcements in the coming months. But it's a really exciting time.
Tim Scarfe
That's so cool.
And again, this was a concept for a very long time, and it was a big risk for me. Initially, when I left theoretical physics and went all in on quantum machine learning, everybody told me, "Really? You? You're going to be a quantum machine learning lead at Google?" Everybody was doubting me. Then I became kind of well-known, and I eventually led a team. I led quantum machine learning at Alphabet X.
And then I quit all that. I quit quantum computing and quantum machine learning. I'm going to take even more risk. I'm going to build a whole new paradigm of computing from scratch, from the concept up, and everybody thought that was crazy. And now we're here. Now we've made a lot of progress, and now we're scaling, right? There's nothing stopping us from scaling at this point because we've de-risked the manufacturing as well, which is usually not something—
Tim Scarfe
Right.
—academics think about. But if you're a startup and you have to scale or die—eat the world or die—you have to kill every risk possible. We're looking to scale to millions of degrees of freedom next year, and I think we're going to hit it, so—
Tim Scarfe
That's really exciting. So your first chip had 3 p-bits.
Yeah.
Tim Scarfe
And the latest has about 300 degrees of freedom.
Yeah.
Tim Scarfe
Am I correct? And now you're—
They're not all p-bits, but we'll—
Tim Scarfe
And now you're aiming… Right.
We'll get to that someday. Yeah.
Tim Scarfe
Right. And now millions of degrees of freedom—
Yeah.
Tim Scarfe
—next year.
Yeah.
Tim Scarfe
That's very cool. So help me understand the general space here. I mean, presumably, you don't think that this is a— Or do you? Do you think this is a wholesale replacement for the current stack, or do you see— Because I could see a world where you interface—well, you use thermodynamic compute when it's relevant, but you mesh this with digital and quantum—
Yep.
Tim Scarfe
—at the appropriate junctures so that you get… What you really get is a multiscale stack where each hardware bit is specialized for the kinds of computations that run natively on that kind of hardware.
Yeah, absolutely. I mean, if you're— I think there was some work from Max Tegmark on correspondences between the information bottleneck principle, downsampling, the hierarchy in machine learning, and the renormalization group, right?
Tim Scarfe
Right.
If you're trying to learn from quantum mechanics and you downsample, you get statistical mechanics. You downsample, you get to—
Tim Scarfe
Mm.
—Newtonian mechanics. You can imagine if I had a God neural network that just compressed all the information in a certain region of space—which is actually the thought experiment that got me into quantum machine learning. I was trying to understand black holes as a machine learning system. But let's say you—
Tim Scarfe
Mm.
—created that system, then you could imagine the first few layers are quantum, then some layers are probabilistic, and later they're deterministic, just as you distill the information. Initially, you need some quantum complexity. Later, you need some entropy, and then later you just need to do some classical coordinate transformations, right?
Tim Scarfe
Mm-hmm.
And they all work together. But in a more practical workflow, very often—whether you're doing simulations of stochastic differential equations, trying to simulate a physical system, or doing some sort of discrete or continuous diffusion—there's a classical function, often a differentiable program, that determines some probability distribution that you want to sample from, whether it's—
Tim Scarfe
Right.
—in latent space or for these transitions and the denoising. So the two work together. You don't necessarily need entropy everywhere in your graph all at once, right?
Tim Scarfe
Right.
In principle, you could use a probabilistic computer for deterministic operations, but it's not going to be the best at that. In the low-precision regime, you could think of p-bits as literal fractional bits. So you can get into—
Tim Scarfe
Mm.
—the fractional-bit precision work, and that can be interesting, but not everything is well suited for that low precision.
Tim Scarfe
Mm.
And I do think quantum computers will have some applications, but again, they would be supplements for quantum-mechanical systems—to understand quantum-mechanical systems—as a supplement to probabilistic and classical computers. But I would say that most things would be well covered by probabilistic and deterministic representations running on probabilistic—
Tim Scarfe
Mm.
—or thermodynamic and deterministic computers. Yeah.
Tim Scarfe
So currently, the superconducting chips require a lot of cooling, right? You need to get them around 1 kelvin. Is that the case?
Yeah. Depending on the material, you can be at a few hundred millikelvin, or you can get to a few kelvin, usually using niobium. I think that you don't need as big of a fridge. For us, it was an interesting experiment, and we have a bunch of results there. We have a paper coming in the coming months. Essentially, it's as efficient as we could imagine building a thermodynamic computer. Again, it was just to get people to imagine—or realize, rather—that there are forms of computing that are far more energy efficient by an unfathomable number of orders of magnitude, far more efficient than digital computers, right?
Tim Scarfe
Right.
I'm just doing some back-of-the-napkin calculations, because you guys target 1,000 to 100,000× energy-efficiency gains at the chip level, and I was wondering: how does that stack up against the cryogenic energy costs? How does that factor into your efficiency calculations?
Yeah. We put the asterisk there: that's just the chip. In general, the thought experiment was: if you scaled a very large network of superconducting chips to football-field size and had a ginormous dilution fridge, right? Again, the big energy cost is maintaining that boundary, right, with the—
Pushing compute to the limits of physics
The outside world. If you had a ginormous dilution fridge and millions and millions of p-bits and superconductors, you'd have the most energy-efficient probabilistic computer, right? Is that practical? Probably not. Should a startup be doing it?
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Probably not either. Could a government do it in some sort of crazy moonshot? Maybe. We're kind of just going to put the idea out there for academia, national labs, and whatnot to pick up. I think for us, it was also just a great learning platform. Superconductors are ironically pretty accessible. A bunch of universities have fabs where you could experiment with them.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
For us, it was kind of like, “Hey, we're going to want a community of people working on thermodynamic computing, experimenting with new primitives, and showcasing new algorithms, so we're going to put that work out into academia.” But for us, the product is, again, the silicon chips. We could have decided to run them in cryo-CMOS, but we ended up deciding to go for room temperature. Of course, a thermodynamic computer at a lower temperature consumes less energy, right?
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Of course, you're assuming you're maintaining the bath at that temperature, but then you have the cost of the cooling, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
If you had a von Neumann probe thermodynamic computer, maybe it could run much, much, much cooler, right? If it's going to go to—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—in interstellar space, it's not going to encounter that much heat, and so maybe we'd modify the design for that. But that's a problem for the far future.
Tim Scarfe
And then, coming back to the impact on the AI industry, I mean, EBMs—energy-based models—have been around for a long time. Would you say that this is really the key to unlocking their potential? Up until now, they've been basically limited by their inefficient sampling, right? Would you say that this is really the technology we need to do the EBM thing seriously?
Pushing compute to the limits of physics
I think modern neural networks are just mean-field approximations of EBMs, right? Neural networks came from EBMs. Backprop came from looking at the mean field of EBMs. To us, we're just creating the hardware for the ancestors of neural networks, and they're kind of a superset, right? You could just take averages and get deterministic operations.
Tim Scarfe
Right.
Pushing compute to the limits of physics
Our hope is that people think about going more probabilistic with their algorithms now that sampling is far more energy-efficient and far faster. But because we've been stuck with deterministic computers, people have tried to avoid using primitives that— They would just represent distributions by their moments, right? They would use exponential families where you don't need too many moments, such as Gaussians, because you could just represent them as matrices and vectors, and matrices and vectors—
Tim Scarfe
Right.
Pushing compute to the limits of physics
—can fit on a GPU really well, right? And you could do some transformations there.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
So there's been a bias in algorithms toward what runs well on today's hardware, and hopefully, with thermodynamic computers, that changes. We're trying to get people interested in the space and to start imagining what they would do with a multimillion—
Tim Scarfe
So this is how you break free of the—
Pushing compute to the limits of physics
Go ahead.
Tim Scarfe
Well, this is how you break free of the vicious cycle that we're stuck in right now, right?
Pushing compute to the limits of physics
Yeah, we're stuck in a—
Tim Scarfe
We've got a hardware stack that works in a certain way.
Pushing compute to the limits of physics
Yep.
Tim Scarfe
We optimize our software for that, then we build bigger hardware and, you know—
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
—so this is how we break free from that.
Pushing compute to the limits of physics
Yeah, but I think we're going to break free one way or another, because frankly—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—even if you just do a back-of-the-envelope calculation, right now we're going to run out of power trying to scale AI. We can't even—
Tim Scarfe
Right.
Pushing compute to the limits of physics
If everybody were to use an agentic, big model—a ChatGPT or Groq model—we'd run out of power in the United States. You can't deploy it to everyone. You can't have everyone using it with high intensity right now. We're going to run out of power. And that's not even getting into video models and world models—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—which are going to be necessary for embodied intelligence. At that point, if you try to scale it to the planet, we're going to run out even if we produce the power. Let's say there was some moonshot, and let's say we even figured out nuclear fusion, right? Let's say we just did that and scaled it very quickly. We'd run out of—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—or, you know, we're going to double or triple the amount of heat being radiated by the Earth, right? So we're literally going to cook ourselves to death. We're literally cooked. And so something has to change—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—at the hardware layer. We can't scale with the current hardware. People who think we're just going to scale transformers and get to the moon and scale intelligence to the whole planet are flat-out wrong, right? It's provably wrong.
Tim Scarfe
When people say, “Hey, to scale up intelligence, we need to start building nuclear power plants,” I always think, well, I run on a glass of water and a banana—maybe a coffee in the morning, right? Certainly, this is not a naturalistic way to think about how intelligence scales, right?
Pushing compute to the limits of physics
Because you're a thermodynamic computer, right? That's the thesis.
Tim Scarfe
That's right.
Pushing compute to the limits of physics
You can imagine just taking our current design and scaling it to a wafer. You would have about 1.5 billion p-bits, which you could think of as neurons, and about 20 billion parameters per wafer. And then you could do multilayer programs of those, and that would run on not 20 kilowatts, which is what a current wafer-scale system would run on, or more—maybe 100. It's 20 watts, right? Which is like our brain.
Tim Scarfe
Wow.
Pushing compute to the limits of physics
Right?
Tim Scarfe
That's crazy.
Pushing compute to the limits of physics
And then, if you have a few-hundred-billion-parameter model running on 20 watts, then we're in the same ballpark as the brain. Not quite exactly—it might be within 10X—but that's still much better than where we are, which is like—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—100 million X, right? And so that's where we're going. Again, it's not— To us, it's far less crazy than quantum computing. There is no quantum computer in nature that has very high quantum coherence, and we have exquisite control and very high quantum complexity. All physical systems decohere. We have multiple billions of thermodynamic computers out there in the wild, and they work pretty well.
We're just trying to tap into the same physics. We're not obsessed with biomimicry. We don't call ourselves neuromorphic computing. We're just doing, again, approximate probabilistic inference as a service, but in hardware, in physics. To us, that's kind of the parent workload of most of AI. And actually, it's not just for AI.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
It's for broader computing, right? There's simulation, optimization, and statistical inference. Science at large can use this. So it's not just for GenAI, right? We're not just riding the current wave.
Tim Scarfe
Yeah. I have the pleasure of interviewing not one but two people today—
Pushing compute to the limits of physics
Sure. Yeah.
Tim Scarfe
—in some sense.
Pushing compute to the limits of physics
We'll switch hats.
Tim Scarfe
I've got your— Yeah, exactly. We've got your alter ego on the call as well. So, to frame up the transition, for the audience, Guillaume, or rather, a.k.a. Beff Jezos, is the founder of a philosophical movement called Effective Accelerationism, which originated as a kind of counterpoint to effective altruism. I really want to get into the details there. But to set things up, looking 10 to 20 years out, how do you envision the future? And what role does acceleration play in your vision for the future?
8. A Future Of Embodied Intelligence
Pushing compute to the limits of physics
Yeah. I think in 10 to 20 years, we'll have embodied intelligence.
I think thermodynamic computers will be pretty ubiquitous. They’re going to be what runs embodied intelligence. We’re going to have greater intelligence density in all our devices. We’re going to have personalized, always-on, online-learning intelligence. And so my goal is for everyone to own and control the extension of our cognition—
Tim Scarfe
Right.
Pushing compute to the limits of physics
—because that is important to maintain—
Tim Scarfe
So, democratizing intelligence.
Pushing compute to the limits of physics
Not just democratizing—yeah, but truly, because democratizing access is like, “Hey, I give you access to the one God model that amortizes its learnings across the fleet. Now you’re all mind-merged with my Borg mind. Congratulations, you’ve been assimilated.” That’s not democracy, right? If you have control over its constitutional prompts,
Tim Scarfe
Right.
Pushing compute to the limits of physics
—you’re essentially controlling people by proxy, and their thoughts, their will, and their actions, right? If you control—
Tim Scarfe
Yeah.
Pushing compute to the limits of physics
—it’s not just—It used to be people controlling people’s access to information and steering them indirectly. But now it’s going to be directly controlling the model that’s an extension of their cognition or their thought partner, and that’s much deeper and more subversive control. And so, to me, I think that was the big existential risk I was most worried about: the sort of precedent—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—of people vying for control and vying for power over people.
Tim Scarfe
Right.
Pushing compute to the limits of physics
And that’s why I’ve been pushing for decentralized AI, and I’ve been pushing for avoiding overregulation of AI that would cause—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—red-tape inflation, mostly serve the incumbents, and cause a centralization of AI power.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
But—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Yeah, someday, I would imagine, 10 or 20 years out, hopefully we’re wearing some neural links. Our whole skull is a neural link that is the thermodynamic computer, an extension of our cognition, and it’s—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—as powerful as, or if not more powerful than, your brain. And you have a full merge there. I think something more plausible on a 10-year timescale is a sort of soft merge. I mean, you’re—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—you’re the active inference expert here. But if you have the same Markov blanket, you have the same perception and action states. So let’s say I had an agent. Let’s say my glasses were very smart, and they were always on, always listening, and I had an earpiece. I could have sort of subconscious thinking, almost. It could just perceive things and suggest actions, so we’re sharing perception, and we’re sharing actions through my body.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
We are the same agent, right? So that is—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—that is a form of merge, and you don’t need a neural interface for that, right? And maybe we even have a way to communicate nonverbally or, you know, just like friends that hang out a lot have a prior of each other’s behavior, or their behavior conditioned on the world’s states.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
And then they don’t need as many bits of information to adjust to a new setting. And so that’s kind of a more plausible way, I think. I think the soft merge is going to come first, but I think it’s going to keep escalating, and then we’re going to go toward a harder merge, right? A hardware merge.
Tim Scarfe
That’s a very compelling vision for the future. It’s very techno-futurist. I love it. Do you want to give me the elevator pitch for EAC, Effective Accelerationism? Because I think these are all related. I think this sets you up nicely to just present the core idea, so—
9. The Logic Of Effective Acceleration
Pushing compute to the limits of physics
Yeah, really, EAC is a sort of metaculture. I call it a cultural hyperparameter prescription, if you will. And it’s one where we’re trying to maximize the growth of civilization, as measured by our free-energy production and consumption, a.k.a. the Kardashev scale. The Kardashev scale is a log scale—
Tim Scarfe
Right.
Pushing compute to the limits of physics
—tracking how much free energy is consumed and produced. And really, it comes from the realization that, as I was studying stochastic thermodynamics in preparation for founding Extropic, I realized that there was this sort of generalization of Darwinian selection, which is thermodynamic selection. Where each bit—you go from selfish genes to selfish memes to selfish bits—every bit of information specifying configurations of matter is fighting for its existence in the future.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
The selection pressure on those bits of information is whether they confer on their host organisms the ability to understand their environment, predict the future, capture free energy, use it strategically, and grow. But really, the master metric is: if I change this parameter, this value, and I let time evolve, how much free energy has the system dissipated through its trajectory over time, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
The theorems of stochastic thermodynamics tell us that trajectories of the system that basically consume more free energy are exponentially more likely, right? And so—
Tim Scarfe
Right.
Pushing compute to the limits of physics
—and so that’s how you have this sort of pruning of branches that consume less free energy. And so it’s like, “Ah, okay, well, this is the golden metric of selection pressure on the space of bits. And so I will design the highest-fitness selfish meme that is e/acc: figure out what is optimal for growth and do it.” And so, by construction, it should be the most viral thing, and it will persist.
Tim Scarfe
Right.
Pushing compute to the limits of physics
And, surprise—
Tim Scarfe
The assumption there is basically—
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
—the alternative is basically growth or death, is what you’re suggesting, right?
Pushing compute to the limits of physics
Yeah. We have the saying, “Accelerate or die.” It sounds dramatic, but it’s meant to sound dramatic; it’s also reality. It’s either you align yourself with growth and you are a part of it, right? So whether you adapt—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—you adapt your culture toward growth, and then you’re a part of the growth and you benefit, or you don’t, and then you get outgrown.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Whether it’s at a national level, with your policies, right? Again, the EU versus the US. The US is obsessed with growth to some extent, and we’ve seen a sort of bifurcation in the GDPs and so on. So, at a policy level, at a cultural level, even at an organizational level, right? If a company is not obsessed with growth, eventually it just gets disrupted by one that outgrows it and then has more resources to outcompete it, right? And so I think—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
I think the point is that, essentially, if people are open-minded about new technologies in general and lean in, they will be positive. They will be selected for, and the people who are skeptical, push them away, shun them away, and want to go back to the cave get selected out. And so, even from an empathetic argument standpoint, trying to popularize acceleration—making people aware that this is actually how the world works—I’m sorry to break it to you: you either embrace the acceleration or you get selected out.
It’s also like, hey, if you care about yourself and your tribe, your company, whatever, you should lean in rather than lean out, and don’t listen to the people fearmongering at you. They are very often doing so out of self-interest. I view spreading deceleration as a form of psychological warfare, in fact.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
That’s how I view it, and that’s why I’ve been such a warrior online against the decel mindset. I just don’t see a scenario where being decel is beneficial to the host. And in fact, it’s just exploiting this gap in understanding of how the world works to destroy competition, right? And I mean, there’s a lot of this in nature and in general in human dynamics. People sabotage each other to outcompete each other.
Tim Scarfe
See, I’ve been thinking about this a lot. My good friend and colleague, Axel Constant, and I have a long-standing debate as to whether ethics, morality, and this kind of thing can be reduced to free-energy minimization at the end of the day. And I think it’s a difficult question, because thermodynamics describes all mesoscale objects, right? So basically any configuration whatsoever will be minimizing free energy in this fashion, right—seeking to find pockets of free energy and dissipating it as efficiently as possible.
I guess my worry is things that grow without bounds. In the biological case, you have cancer, for example. So what's the role of constraints in the maximization of entropy when we're designing the social system? Because surely things like dictatorships and fascist autocracies also minimize free energy.
Pushing compute to the limits of physics
They're a local optimum, right? They're not a global optimum, right? Just like cancer: it kills the host, and that's suboptimal on a sufficient timescale. Again, it's not instantaneous free energy dissipation, right? It's basically an infinite time horizon, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
And so blowing up the planet does burn up a bunch of free energy, but in the long term—
Tim Scarfe
Right.
Pushing compute to the limits of physics
It's the same reason life exists.
Tim Scarfe
Right.
Pushing compute to the limits of physics
It's much better to conserve and strategically use free energy to secure more free energy, keep growing, and have some order, rather than just burn it all in one go and have chaos, right? And thermalize in one go, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
That's why we have life. That's why we're not at equilibrium, right? We're not just burning up all our fuel and dying immediately, because as intelligent beings, we burn way more energy by being this sort of somewhat coherent system that has predictive power of its environment. We're kind of this energy-seeking fire, right?
Tim Scarfe
Right.
Pushing compute to the limits of physics
Yeah, but I would argue that e/acc is also a call to popularize complexism, complex systems thinking, and the free energy principle at large. We have to think through this lens about all systems in society, from policymaking to technology to innovation to basically everything—economics. It's a different way of thinking, one that is somewhat scary. Again, it's going from reductionism and rationalism to complexism and post-rationalism, and we don't have as much control and interpretability. We don't understand the world as well, right? We kind of have to—
Tim Scarfe
Yeah.
Pushing compute to the limits of physics
Pardon my French, but fuck around and find out: the FAFO algorithm, right? But really, it's kind of exploration and discovery, right? A priori, you don't know what's optimal, right?
In startups, you learn this. You have to actually go on the market and try stuff. Sometimes your prior—your model-based prior—can't do that well because the ecosystem is too complex to have a model with good predictive power. You actually have to be in an open loop.
I would say that hopefully there's a renaissance, and these ideas become more popular in all sorts of fields of science. Clearly, it's eaten the software world, and we're trying to make—
Tim Scarfe
Mm.
Pushing compute to the limits of physics
—this sort of school of thought eat the hardware world. I think we will do it. But again, most fields of study could be revolutionized by thinking through the lens of complex self-adaptive systems and the free energy principle.
Tim Scarfe
So, you've described e/acc as a kind of hyperstitious meme. Do you want to say a few words about what that means?
Pushing compute to the limits of physics
Mm.
Tim Scarfe
Yeah. Again, you're the active inference expert, so in active inference, you have perception and action, right? You can update your model based on your sensory information, so you're minimizing the divergence between your own internal model and the statistics of the world. But then the dual—
Pushing compute to the limits of physics
Mm.
Tim Scarfe
—of that is taking actions in the world to minimize divergence between the world and your predictive model of it, right?
Pushing compute to the limits of physics
Right.
Tim Scarfe
And, in a way, we're naturally biased toward this: the car goes where the eyes look when you're driving.
Pushing compute to the limits of physics
Right.
Tim Scarfe
If we look at very negative outcomes and we're obsessed with them, we will drive whatever system we're thinking about toward those negative outcomes. An example of this—
Pushing compute to the limits of physics
Right.
Tim Scarfe
A slightly controversial example is bioweapons research.
Pushing compute to the limits of physics
Mm.
Tim Scarfe
And, for example, COVID, right? I would say that COVID was an accident from bioweapons research. It was probably defensive, and we were trying to explore what would be a really bad scenario. What if we had this mutation, and this mutation would combine and it would be a really bad virus? Then they started experimenting and designing in that neighborhood of virus subspaces that would never—
Pushing compute to the limits of physics
Mm.
Tim Scarfe
—have occurred naturally, just from evolution.
Pushing compute to the limits of physics
Mm.
Tim Scarfe
But because we were exploring that subspace of bad things, because we were obsessed with it, we made it happen, right?
Pushing compute to the limits of physics
Mm.
Tim Scarfe
To me, if we're optimistic about the future, we tend to steer things toward that optimistic outcome. As a startup founder, if you're not optimistic about your startup, statistically, you are screwed, right?
Pushing compute to the limits of physics
Oh, yeah. From an active inference point of view, every action begins with a false belief, right? So first you believe that you're moving, and then you reduce the prediction error in the direction of action, right? So—
Tim Scarfe
Exactly.
Pushing compute to the limits of physics
I find this extremely compelling. One of the things I find most inspiring about e/acc is that you're trying to present a radically optimistic meme for the future. There's so much—not just AI doomerism, but so much doom and gloom—going around that, just from the point of view of neurobiology, we need these almost seemingly delusionally optimistic beliefs to get off the ground anyway, right?
Tim Scarfe
Yeah. You could think of the memetic sphere as a metacortex, right? It's a biological supercomputer. In e/acc, we're just trying to do active inference toward better futures, right?
Pushing compute to the limits of physics
Right.
Tim Scarfe
So we're spreading this meme of optimistic futures, and it actually does steer the world. It's been 3 years now, and it has steered policies throughout the world toward going for these moonshots, a resurgence of exploration of nuclear energy to climb the Kardashev scale, deregulation of AI, widespread embrace of AI, and companies being far more aggressive in their exploration of it.
So how do we spread it? How do we spread e/acc? I guess my question is motivated by this: I think most of the very visible e/acc figures—Andresen, Musk, and company—tend to be libertarian, right-wing-associated. If we want e/acc to really be—
Pushing compute to the limits of physics
Well, that happened after the fact, right? I would say Gary Tan is rather on the left, and initially it was very apolitical. I would consider myself still a centrist.
Tim Scarfe
Well, how do we get the abundance bros to join?
Pushing compute to the limits of physics
I think it's already happening.
Tim Scarfe
The Ezra Kleins and—okay, cool.
Pushing compute to the limits of physics
I think that was a reaction to e/acc in the techno-optimist world. You end up getting clustered with one party in the US. In other countries, it might be different. But there needed to be an actual techno-progressive, let's say, versus techno-regressive, cluster on the left. It's not clear that that faction is the dominant faction of the Democrats. I really hope it becomes that. Again, e/acc is kind of like an RL algorithm. It's not clear what the optimal policy is.
I think a few years ago, I was seeing a sort of convergence toward a lot of top-down control, a lot of top-down power. I think the optimal is always balance. You need some top-down control, but you also need bottom-up self-organization.
Tim Scarfe
That's how the brain works, right? You have some centralization, but not all that much centralization, right?
Pushing compute to the limits of physics
Exactly, right. Absolute total libertarianism can work. It's just that in the era of unconventional warfare, complex systems also get adversarially steered, right?
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Whether it's free markets—they get gamed—or memetic markets, there are psychological operations, dating markets, and so on. I think some balance of top-down and bottom-up is the ideal. Is it more like the U.S., more like China, or something in between? I don't know. I guess we have to figure out.
Tim Scarfe
Well, that segues perfectly into my next question.
Pushing compute to the limits of physics
Yeah.
Tim Scarfe
For you, this isn't just a philosophical thing. This is of geopolitical and geostrategic significance, right? This is important for the future of democracy as we understand it.
Pushing compute to the limits of physics
Yeah. I would say that our failure to view the world as a complex, self-adaptive system, and our continued thinking in the classical way—“Hey, I have a first-order model of what's happening, and I'm going to do a first-order correction”—is a problem. It's easy to convince a crowd of that, and politicians get elected on such platforms, but then they don't think about the higher-order effects of their policies. We don't think about it from a complex-systems steering standpoint.
Tim Scarfe
Mm.
Pushing compute to the limits of physics
I think our adversaries—the adversaries of the West—understand the complex-systems approach, and they tend to steer us in directions that lead to our detriment. To me, it was, like, okay, how do you fight a multidimensional war?
Tim Scarfe
Mm.
Pushing compute to the limits of physics
Okay, there's invisible sabotage of every complex system that can be steered adiabatically in nefarious directions. If it's slow enough, then it's kind of above the infrared temporal cutoff that politicians give a shit about, which is 4 years, right? If you're steering the United States or the West on a 20- to 40-year timescale, you could win a very long war, and we don't even find the pattern because we're too busy.
Tim Scarfe
Right.
Pushing compute to the limits of physics
There's a sort of timescale separation here, similar to what happens in thermodynamic computing and thermodynamic physics. Essentially, we're so preoccupied with the dynamics that occur on small timescales—the day-to-day—that we ignore the long trends, and those get hacked, right? We slowly boil the frog. It doesn't realize it.
Tim Scarfe
So, in listening to you talk, it occurs to me that there's a beautiful coherence to your approach generally. What you're doing both at the hardware level and at the level of your philosophical project is essentially moving us away from hard, rigid programming from the outside toward a kind of organic, adaptive, thermodynamic-driven learning and exploration of possibility space.
Pushing compute to the limits of physics
Yes. Yes. Again, it came from my own journey trying to understand physics, discovering differentiable programming, and seeing that as the way forward for everything. Some of the prescriptions of e/acc are to maintain variance and constantly explore across any parameter space, whether it's culture, aesthetics, policy, technology, et cetera.
One of the reasons is that, according to Fisher's theories on evolution, the speed at which you can traverse a landscape depends on the gradient, but it also depends on the variance. Evolutionary search depends on variance, because the more you fuck around, the more you find out. Your rate of learning is faster, and your rate of adaptation is faster.
To me, that seemed like the main advantage of the United States: the United States is very high-variance. It has high-variance individuals and high-variance outcomes. I view the United States as a sort of high-temperature search algorithm.
Tim Scarfe
Mm.
I'm not American yet officially, but someday I will be. That's my goal. They're first to figure things out because they're always searching in a very high-variance way.
I view innovation as a diffusion process in some landscape, and they're in a very high-noise regime, so they don't get stuck in local optima, right? Whereas China's more like the low-temperature sampler. They're the optimizer at the end. They're doing the gradient descent.
Once there's consensus, once there's a clear gradient of improvement, it's easy to convince a committee at that point, and then you can just execute top-down control with a lot of conviction.
Tim Scarfe
Mm.
They beat us in the final stretch. It's kind of like Bayesian inference: you go from an unsharp prior to a sharper posterior on what the optimal thing is. It's like annealing.
Tim Scarfe
Mm.
Essentially, China beats us at squeezing the end, again because it has a lot more top-down control power and a lot more coherence there. I see this sort of—we're kind of like a parallel-tempering algorithm between the United States, the high-temperature search that discovers things first, and then China.
The United States isn't necessarily the best at optimizing and improving the technologies and scaling the manufacturing.
Tim Scarfe
Well, then how do we avoid catastrophic outcomes in that? You've been pretty dismissive of P(doom) and doomerism generally. How do we allow the kind of meta-search to happen, but in a way that doesn't lead to catastrophic technological outcomes?
I can also see some tension between this open project, on the one hand, and the participation of proprietary, closed corporate groups, on the other hand.
Yeah. I guess for me, P90, 1984 was higher than P(doom) from AI. To me, China is the ultimate monopolistic company. There's basically one set of executives for all the companies. They're acting like a conglomerate, and they throw their weight around to crush smaller American companies.
One thing about e/acc is that we're kind of anti-monopolistic, because monopolies tend to be suboptimal. If you have hyperparameter choices that are over-concentrated, you're not exploring anymore.
Tim Scarfe
Well, you see, that sort of motivates my question and my worry. Very authoritarian forms of government, where we impose ethnic or cultural homogeneity, are very good at minimizing free energy. If you and I are exactly alike because there's a top-down imposition of sameness, then—
Mm.
Tim Scarfe
That minimizes free energy super well. If we're all the same—
Locally, right?
Tim Scarfe
Right.
Well, if your energy term says you want to agree with your neighbor and have no frustration, then of course that's a local optimum. But what we're arguing for is embracing variance: embracing different cultures and having different bets.
Tim Scarfe
Right.
It could be culture, genetics, or ways to train your ML models.
Tim Scarfe
The same response to the cancer thing, right? You're not worried about the push toward boundless growth leading to cancer-type formations because it's not a global optimum. It's a local optimum.
Yeah, yeah. That's right. I want to go on a slight tangent here. There have been some centralization in AI research labs, even though there are 4 or 5 players that matter. They kind of churn. It's been a meme recently that researchers all slosh around and churn between each other.
They're all equilibrating in terms of beliefs because you have an exchange of particles—of researchers—between the labs. You could see that the American research labs all converged to similar performance at similar compute.
Then China, with DeepSeek, was exploring a whole different region of hyperparameter space due to its export-control constraints, and showed us that over-concentrating our bets in hyperparameter space has risks, because we're stuck in a local optimum. There might be a better, nonlocal global optimum that we're missing.
Mm.
And so, again, it's just been our push for variance there in order to always be exploring. I think it's really important. I think basically there's also been a reaction to the same pattern. It's hard to explain. It's the same pattern that happens in Western governments, in late-stage corporations, and in large bureaucracies.
There's a “cover your ass” mentality and culture, and decision by committee. And so it tends to be variance-reducing, right? Or variance-killing.
And that's sort of my worry. In addition, I guess we're in the same kind of situation as a gradient descent learner. How are we supposed to know that we're not stuck in a local optimum, right? To us, we might think, “Hey, this actually looks pretty optimal,” but there's no kind of God's-eye view that could tell us that actually we're just stuck in a local—
Yeah.
—
But that's why, as long as you don't kill variance completely, you keep some exploration, you're not just purely exploiting, and you keep open-mindedness, then you're always spending some resources still exploring and—
Mm-hmm.
—and potentially finding a new, better way to do things. Just like, I think some people at first thought our bet was ludicrous and that we shouldn't have raised VC funding. It's like, really? There's trillions of dollars of capital riding on the current paradigm, and you don't want a startup to raise a few tens of millions to take a bet that would be completely disruptive to everything—to this whole multi-trillion-dollar bet, right?
Mm.
And now the tune is changing a lot, because people who are planning these half-trillion-dollar build-outs want to know if a technology that's 1000x better is around the corner and could render some of their build-outs maybe miscalibrated, right? In terms of the ratio of energy—
Right.
—to compute and density, right?
Again, I don't think thermodynamic computing is going to work in tandem with GPUs or TPUs or whatever neural processing units. I think it's going to take quite a while for models to be fully ported to thermodynamic computers. So, for the foreseeable future, I think these build-outs are safe, and you could just have thermodynamic computing as an add-on and—
Guillaume, it's been a real pleasure to discuss these issues with you. So how do we stay abreast of what's going on at Extropic? Do you have any closing message for us?
Yeah. Well, stay tuned in the coming weeks and months. For those that are interested in trying out thermodynamic computing, in the coming weeks we're going to give access to the first users to our systems, so you can test it yourself. It's mostly going to be a private alpha in the early days. Again, we can't handle that many users. We don't have that many chips.
But the technology's here. It's very real, and you can kick the tires. Stay tuned for some scientific papers and some open-source software that we're going to put out. I would say start reading more probabilistic machine learning papers if you're a grad student interested in the area. Start reading about EBMs and various textbooks. And stay tuned for more announcements from Extropic toward the end of the summer and this fall.
Well, for Machine Learning Street Talk, I'm Maxwell Ramstead. We're signing off. And thanks again, Gil. This was awesome.
Thanks, Max.
Yeah, fascinating. Really cool stuff.
Awesome.