[BidClub_]
Machine Learning Street Talk · · 84 min

Pushing compute to the limits of physics

Maxwell RamsteadGuillaume Verdon

Podcast
TL;DR
  • Verdon’s central bet is that AI’s probabilistic workloads are running on deterministic hardware that spends energy suppressing noise, only for software to add randomness back in. Extropic’s mixed-signal silicon instead harnesses stochastic electron dynamics to accelerate Markov chain Monte Carlo across discrete and continuous variables. The pitch is to “loosen our grip on electrons” and turn entropy from a liability into the computational resource.

  • The scaling thesis is physical: using less power requires less charge, but once charge becomes scarce its discreteness creates unavoidable noise. Conventional machines target error rates around (10^{-15}), paying an escalating energy cost to preserve determinism inside the “thermal danger zone.” Verdon therefore calls thermodynamic computing inevitable: “You’re gonna have to go thermodynamic at some point.”

  • Extropic says it has moved from a three-pBit superconducting prototype to a silicon chip with roughly 300 degrees of freedom, with millions targeted next year. Its controllable pBits reportedly generate entropy using only “a few hundred attojoules”; results had been submitted for peer review, while private-alpha hardware access, papers, and open-source software were promised after the recording. The company targets 1,000–100,000X chip-level energy-efficiency gains, although Verdon explicitly separates that from full-system cooling costs.

  • A key investor-relevant projection is the wafer-scale estimate: about 1.5 billion pBits and 20 billion parameters on 20 watts. Verdon contrasts that with a current wafer-scale system consuming 20 kilowatts or perhaps 100 kilowatts, and argues that a few-hundred-billion-parameter thermodynamic model might land within 10X of brain efficiency rather than today’s claimed 100-million-X gap. These are forward projections, not demonstrated system benchmarks.

  • The near-term thesis is additive rather than a wholesale GPU replacement. Verdon expects deterministic processors to retain classical functions, probabilistic chips to handle sampling, and quantum hardware to supplement workloads involving quantum systems. He says existing build-outs remain safe “for the foreseeable future” because model porting will take time, even if a technology “1,000X better” changes future power-to-compute ratios.

  • Changing the substrate could change the winning model architecture, because “transformers are not sacred.” Softmax, diffusion, tree-of-thought search, Monte Carlo tree search, reinforcement-learning rollouts, and test-time reasoning already expose sampling-heavy demand; efficient physical sampling could revive energy-based models and induce a “Cambrian explosion” of architectures optimized for probabilistic hardware.

  • Verdon connects decentralized compute to a political objective: individuals should own the always-on systems extending their cognition. He rejects mere access to a centrally controlled “one God model” as assimilation into a “Borg mind,” forecasting a soft human–AI merge through shared perception and action before neural hardware arrives. His broader e/acc prescription—“accelerate or die”—favors growth, variance, and anti-monopoly experimentation, while Ramstead presses on cancerous growth, authoritarian local optima, and catastrophic tail risks.

Digest · the substance, structured for research

1. Reductionism failed where complex systems became the object

  • Verdon traces his route from a seven-year-old fascinated by subatomic particles to math and physics at McGill and theoretical physics at Waterloo and the Perimeter Institute. He originally expected a compact theory—and perhaps exotic propulsion—to make humanity’s expansion tractable.

  • The intellectual break came when condensed-matter and quantum-gravity problems resisted reduction to a few interpretable parameters. Even with microscopic equations, emergent behavior may require “an amount of computation that is similar to that of the universe” to predict; many systems cannot simply be renormalized analytically into clean higher-level laws.

  • Verdon calls accepting opaque computational representations an “ego death”: the human physicist may not be “the hero of the story.” Tensor networks and deep-learning systems can be as inscrutable as nature, yet programmable complexity can still learn compressed representations and provide predictive power.

2. Quantum information supplied the bridge from physics to AI

  • The “it from qubit” program generalized Wheeler’s “it from bit,” treating the universe as a quantum computer or self-simulation. Whether or not AdS/CFT delivered unification, Verdon concluded that complexism—and computation capable of matching nature’s complexity—was the more productive framework.

  • His first answer was to “fight fire with fire”: use tunable quantum complexity to model quantum complexity. He helped develop early parameterized quantum programs called quantum neural networks, became an early Rigetti user, and later joined Google to build TensorFlow Quantum before leading quantum machine learning at Alphabet X.

  • The reusable principle was substrate matching: physics-informed representations should run on accelerators whose native dynamics implement the same physics. A quantum representation belongs on a controllable Schrödinger evolution; a stochastic representation belongs on a controllable stochastic evolution.

3. Quantum refrigeration made a hotter computer attractive

  • After nearly eight years in quantum computing, Verdon judged progress too slow and focused on fault tolerance’s thermodynamic burden. A quantum machine tries to remain near zero temperature and entropy, continually pumping away noise through what he calls “an algorithmic form of refrigeration.”

  • His inversion was simple: let environmental noise enter until quantum coherence fades and the useful dynamics become stochastic. Unlike quantum computing, he claims no complexity-class separation for thermodynamic acceleration—only potentially enormous constant-factor gains in speed and energy that make digital emulation impractical at scale, though not impossible.

  • Ramstead sharpens the definition: physics-based computing exploits a component’s actual physical properties instead of digitally simulating them. Verdon keeps the category broad—quantum, stochastic, photonic, even a neural net trained through “puddles of water”—while distinguishing physics-based execution from physics-inspired software.

4. Determinism has an energy price, so Extropic harnesses noise

  • Classical hardware forces transistors into absolutely on or off states by making signals dwarf electron jitter. Invoking Maxwell’s demon, Verdon argues that “knowledge comes at a cost”: reducing entropy and maintaining determinism every clock cycle necessarily consumes energy.

  • His alternative is to “learn to let go”—moving from tightly yanking signals around to “gently guiding” noisy dynamics closer to equilibrium. The philosophical analogy is deliberate: software already surrendered imperative programming to gradient descent, and hardware should similarly stop treating every fluctuation as an enemy.

  • Technically, Extropic is building accelerators for Markov chain Monte Carlo with discrete variables, continuous variables, and mixtures. Its newest chip is mixed-signal: stochastic electronics generate proposals or entropy, while conventional digital components perform non-random operations such as Metropolis–Hastings acceptance and rejection.

5. AI software is probabilistic even when its machines are not

  • Ramstead identifies the stack’s paradox: hardware expends power removing randomness, then sampling algorithms restore it in software. Verdon adds that a transformer’s softmax already makes it “a big probabilistic computer,” while diffusion reverses a Markov chain and test-time reasoning increasingly uses Monte Carlo tree search, tree-of-thought search, or RL rollouts.

  • Extropic was founded in 2022 on the prediction that major workloads would move toward sampling and probabilistic inference. Verdon describes current architectures as interpolating between deterministic forward passes and full energy-based models; diffusion occupies part of that middle ground.

  • Two transformer-paper authors who invested in Extropic repeatedly told the team, “Transformers are not sacred.” They worked because they fit Google TPUs; alter the hardware’s fitness landscape, Verdon argues, and model search could enter a high-temperature phase—a “Cambrian explosion” of architectures native to probabilistic compute.

6. Moore’s law becomes Moore’s wall inside the thermal danger zone

  • Conventional processors generally seek error rates near (10^{-15}), allowing vast numbers of operations before correction becomes necessary. But miniaturization raises fluctuation-to-signal ratios: “To use less power, you need to use less charge,” and once little charge remains, its discreteness produces noise “period.”

  • Ramstead’s tennis-ball analogy separates regimes: macroscopic motion ignores molecular vibration, quantum machines preserve coherent superpositions, and thermodynamic devices operate where fluctuations coexist with component-scale signals. Verdon’s claim is that increasingly small deterministic transistors are forced toward that third regime anyway.

  • Verdon presents the brain as proof that useful thermodynamic intelligence exists, and Ramstead suggests it operates close to the Landauer limit. Verdon agrees that neurotransmitters move through stochastic chemical-reaction networks and argues that electron-based circuits might execute analogous probabilistic programs more efficiently because “electrons are much lighter than big neurotransmitters.”

7. The pBit turns an energy landscape into programmable probability

  • Extropic’s chip can be viewed as a time-dependent programmable energy function with diffusion resembling Langevin dynamics, itself an MCMC method. The analogy is literal: Bayesian sampling algorithms can be embedded into the native motion of electrons rather than numerically simulated step by step.

  • A pBit resembles a double-well landscape containing a bouncing ball. One well represents zero and the other one; adjusting the tilt controls how long the signal occupies each state, creating a “fractional bit” that continually dances between zero and one.

  • The company began with three superconducting pBits, then reproduced stochastic primitives in silicon; Ramstead describes the latest chip as having about 300 degrees of freedom, though Verdon cautions that they are not all pBits. He reports entropy generation at a few hundred attojoules and targets millions of degrees of freedom next year.

8. Silicon is the product, while superconductors were the learning platform

  • Superconducting circuit quantum electrodynamics offered a direct map from a desired Hamiltonian to a circuit, making it a useful early engineering platform. Depending on material, those devices operate at a few hundred millikelvin or a few kelvin; maintaining the cryogenic boundary is a major practical cost.

  • Verdon concedes the extreme thought experiment: millions of superconducting pBits inside a football-field-sized dilution refrigerator might form the most efficient probabilistic computer, but it is “probably not” practical or appropriate for a startup. He instead offers the work to academia and national laboratories as a community platform.

  • Extropic chose room-temperature silicon for commercialization. The claimed 1,000–100,000X efficiency range applies at chip level, with the cooling asterisk acknowledged; lower-temperature thermodynamic devices consume less locally, but maintaining the cold bath carries its own cost.

9. The winning stack is heterogeneous, not thermodynamic everywhere

  • Verdon imagines computation following physics across scales: early layers might preserve quantum complexity, intermediate layers retain probability and entropy, and later deterministic layers perform coordinate transformations after information has been distilled.

  • Practical graphs likewise mix a classical differentiable function defining a distribution with a probabilistic accelerator sampling from it. “You don’t necessarily need entropy everywhere in your graph all at once,” and thermodynamic hardware will not be best for every deterministic or high-precision operation.

  • Quantum computers may remain supplements for quantum systems; probabilistic and deterministic processors cover most other workloads. Verdon therefore tells owners of planned GPU and TPU infrastructure that porting will take “quite a while” and thermodynamic compute should initially appear as an add-on, leaving foreseeable build-outs usable.

10. Energy-based models expose both the upside and the power constraint

  • Verdon frames modern neural networks as mean-field approximations of energy-based models: deterministic hardware pushed researchers toward moments, Gaussian families, matrices, and vectors that map efficiently onto GPUs. Faster physical sampling could make the broader EBM family practical rather than forcing distributions into hardware-friendly summaries.

  • At wafer scale, he projects roughly 1.5 billion pBits and 20 billion parameters using 20 watts, versus 20 kilowatts or perhaps 100 kilowatts for a current wafer-scale system. A multilayer, few-hundred-billion-parameter system might approach within 10X of brain efficiency, he says, compared with a present gap he puts near 100 million X.

  • Verdon’s categorical constraint is that current hardware cannot scale ubiquitous agents, video models, world models, and embodied intelligence. Even abundant fusion power would ultimately raise planetary heat radiation: “We’re literally cooked.” The intended market therefore extends beyond generative AI into simulation, optimization, statistical inference, and science.

11. e/acc applies thermodynamic selection to technology and politics

  • Ramstead introduces Effective Accelerationism as a counterpoint to effective altruism; Verdon defines it as a “cultural hyperparameter prescription” maximizing civilization’s free-energy production and consumption on the logarithmic Kardashev scale. His stochastic-thermodynamics framing moves from selfish genes and memes to “selfish bits”: trajectories dissipating more free energy become exponentially more likely, making “accelerate or die” the movement’s intentionally dramatic slogan.

  • Ramstead’s pushback—worth keeping—is that unbounded growth can produce cancer, dictatorship, or imposed homogeneity. Verdon answers that these are local optima over short horizons: cancer kills its host, while burning everything immediately prevents future capture. The proposed objective is strategic growth over an effectively infinite horizon, preserving order and resources to secure more resources later.

  • Ramstead explains the hyperstition mechanism through active inference: perception and action can reduce model-world divergence, and “the car goes where the eyes look.” He then gives the controversial example that COVID was an accident from defensive bioweapons research, hedging that it was “probably defensive.” Verdon does not substantively endorse that example, but agrees with the broader optimistic-steering logic and says every action begins with a false belief.

  • Geopolitically, Verdon favors variance over monopoly: the US acts like high-temperature search, while China acts like a low-temperature optimizer that scales once a gradient is clear. DeepSeek’s exploration under export-control constraints is his example of an alternate hyperparameter region exposing Western convergence; his prescription is balanced top-down coordination and bottom-up exploration, because his “P90, 1984” was higher than his AI “P doom.”

Guillaume Verdon

That was the most technical podcast I've ever done. I think that was one of my favorite conversations for sure.

Hey, I'm Guillaume Verdon. Growing up, I was really pursuing theories of everything. I wanted to understand the universe. I was a big fan of Feynman and Stephen Hawking growing up. When I was 7, I was talking about subatomic particles—

Speaker N

As all 7-year-olds do, right?

Guillaume Verdon

I guess I got swept up in the school of thought that was a generalization of Wheeler's “It from Bit,” which was the “It from Qubit” program, seeking to unify theoretical physics through quantum information theory. So, viewing everything in the universe as one big quantum computer running a certain program or a self-simulation.

We have proof of existence of a really kickass AI supercomputer that we're both using right now to talk to each other. It's our brains. That, I would argue, is a thermodynamic computer, right? Because there are master equations describing the chemical reaction networks in your brain, with neurotransmitters hopping around. And so, if we're doing very similar physics but with electrons sloshing around a circuit, then there's a much stronger chance that we could run a very similar program in a very similar fashion, arguably with even more energy-efficient components, because electrons are much lighter than big neurotransmitters.

This podcast is supported by Google. Hi, folks, Paige Bailey here from the Google DeepMind DevRel team. For our developers out there, we know there's a constant trade-off between model intelligence, speed, and cost. Gemini 2.5 Flash aims right at that challenge. It's got the speed you expect from Flash, but with upgraded reasoning power. And crucially, we've added controls like setting thinking budgets, so you can decide how much reasoning to apply, optimizing for latency and costs. So try out Gemini 2.5 Flash at aistudio.google.com, and let us know what you build.

Hey, I'm Guillaume Verdon. I'm the founder of Extropic, a company pioneering thermodynamic computing, a new form of computing for probabilistic inference using exotic stochastic physics of electrons. Formerly, I was working on quantum computing and machine learning at Alphabet. I also happen to be the founder of a philosophical movement called Effective Accelerationism, under the pseudonym Beff Jezos online. Happy to be here at MLST.

Hello, everyone. Welcome to Machine Learning Street Talk. I'm your host for today, Maxwell Ramstead. I'm subbing in for Tim Scarfe, who unfortunately is down with a little bit of COVID. But I'm sure we'll have a really interesting conversation. It promises to be a super interesting conversation because we have a really awesome guest for today. I'm very excited to have Guillaume “Gil” Verdon, also known as Beff Jezos on X and generally online in the meme space. Very excited to have you, Gil. How are you doing?

Guillaume Verdon

Yeah, thanks for having me. Thanks to MLST for hosting, and it's great to talk to you again, Maxwell. I think this collaboration, at least for this conversation, has been a long time coming, so it's a great platform to do it. And, yeah, let's get to it.

So you're a very interesting character, Gil. You're a fascinating researcher. I think you're a big presence in the meme space as well. Do you want to tell us a little bit about yourself? Tell the audience a little bit about your trajectory. How did you wind up being the CEO of a thermodynamic hardware company and also a well-known person in the meme space?

Guillaume Verdon

1. From Physics To AI

Yeah, I mean, it's been a long journey. I guess I've lived many lives. Growing up, I wanted to pursue theories of everything. Really, I was pursuing theories of everything. I wanted to understand the universe, as one does, and eventually leverage that knowledge to expand civilization to the stars.

Originally, my plan was to become a theoretical physicist and work on quantum gravity, and figure out some exotic form of propulsion that would obviously—well, I thought that was the bottleneck for the expansion of civilization—was the speed of propulsion. And so I went down that path. I was a big fan of Feynman and Stephen Hawking growing up, and—

So this is a childhood project? Since you were a wee boy, you've dreamed of—

Guillaume Verdon

Yeah, no, I—

That's cool.

Guillaume Verdon

I was 7. I was talking about subatomic particles and how—

As all 7-year-olds do, right?

Guillaume Verdon

Yeah, and I wanted to be a physicist. I used to call it an astrophysicist. I didn't know what a theoretical physicist was. But I wanted to be an astrophysicist. I wanted to work on FTL travel and stuff like that.

Over time, I did the career path through theoretical physics. I did math and physics in undergrad at McGill, and then I went to Waterloo, the Perimeter Institute for Theoretical Physics. I met some of the greatest minds on Earth. They're really strong caliber. But what I realized going through theoretical physics was that the reductionist approach to physics was failing us, right?

Do you want to, just for the benefit of our audience, explain what you mean by the reductionist approach? And what do you think were the problems?

Guillaume Verdon

Yeah. I guess these two sides to my life now came from my reaction to realizing this failure. The reductionist approach is the traditional way we've done physics. It's the very rational, old-school way to do things, which is like, “Oh, I want to have a model.” Maybe it's an equation, some analytic model, and it has a few parameters. I've reduced all of physics to a very simple model with few parameters, and ideally the least amount of parameters possible—Occam's razor principle.

These equations with these few parameters allow me to have predictive power over the world, and with this predictive power, I can steer the world. I can do things, right? I could predict and control it. That's the goal of physics: to have better models of the world.

2. Complex Systems Replace Reductionism

What we realized was that, as part of this, I got swept up in this school of thought that was a generalization of Wheeler's “It from Bit,” which was the “It from Qubit” program—

Hmm.

Guillaume Verdon

—which is seeking to unify theoretical physics through quantum information theory, right? So, viewing everything in the universe as one big quantum computer running a certain program or a self-simulation.

That framework's actually a really useful unifying framework to understand all sorts of systems, and people were studying systems that have some connections between quantum gravity and regular quantum mechanics in the context of holography. So, AdS/CFT, for those that are familiar. But really, what I realized there was that even whether or not AdS/CFT was going to work, it was clear that the complexism approach to physics was the way forward.

The universe is very complex. Even beyond just quantum gravity, looking at condensed-matter systems, you can have equations describing the microscopics, but you can't predict the emergent properties, right? You actually have to apply an amount of computation that is similar to that of the universe to have a prediction of what happens at a larger scale.

Not all equations you can renormalize analytically. Renormalize means getting effective physics at a larger scale, at a more coarse-grained scale, right? That's how we go from different types of physics—from quantum field theory to quantum theory, to statistical mechanics, and then eventually classical Newtonian mechanics, and then even beyond that we get to general relativity and so on.

So clearly, the way forward was to understand everything as a complex system. But that was a sort of ego death because, in a way, the human can't necessarily be the hero of the story. The model is no longer interpretable, right? If you have something like a deep-learning system—back then we were looking at tensor networks, which are a different sort of parametric complex system that we use to model, let's say, condensed-matter systems or quantum-gravity systems.

If you look at these systems, they're no longer interpretable, right? I have these big networks, and they're just as opaque as the system that I was trying to predict.

With numerics and a lot of compute, you're not deferring agency; you're leveraging a programmable, parametric complex system to grok a complex system of nature for you.

Tim Scarfe

Mm-hmm.

Speaker N

In a way, it was like, okay, maybe I can't be the hero of the story. I won't be the one to figure out the grand unifying theory of physics, but maybe I could build a computer, or computer software, that can understand the universe or chunks of the universe for us.

Tim Scarfe

Right.

Speaker N

The more I dug into it, it was clear that this was gonna be the way forward, right?

Tim Scarfe

So basically, building a digital brain to overcome the limitations of our fleshy brains.

Speaker N

Yeah, that's correct, but then at the time we were studying quantum mechanical systems. So you had—

Tim Scarfe

Right.

Speaker N

Systems that have quantum complexity, and there you can show that classical representations will struggle to capture quantum correlations. So already—

Tim Scarfe

Mm.

Speaker N

I came at it from a non—

Tim Scarfe

Mm.

Speaker N

Anthropomorphic form of intelligence. To me, intelligence was just trying to learn compressed representations of systems through parameterized distributions, or, in my case, parameterized wave functions and density matrices, right? Which is the more general form.

So that actually got me into a field. Initially, it was mostly numerics with tensor networks, but there was another institute I was part of at the University of Waterloo, which was the Institute for Quantum Computing, right?

Tim Scarfe

Right.

Speaker N

And there it was very interesting because we had these controllable quantum systems. We had these control parameters, right? It became clear to me that there was maybe a way to run these representations we were trying to learn of quantum systems on a parametric, programmable quantum system, right?

Tim Scarfe

Right.

Speaker N

And now we could fight fire with fire, right? We can have parametric quantum complexity that's tunable to understand the quantum complexity of our world, right?

Tim Scarfe

Right.

3. Quantum Computing Becomes Machine Learning

Speaker N

And that was actually my entry into artificial intelligence. I wrote up some of the first algorithms for quantum neural networks, so that's what we called them. They have nothing to do with actual neurons; they're just parameterized quantum programs.

Essentially, I was one of the first quantum computer programmers and the first user of Rigetti, which was one of the first startups. That got me on the radar of Google, which then approached me and a team I built at Waterloo—an open-source team—to go work at Google and build a product known as TensorFlow Quantum. It was a product focused on creating software that allows us to learn quantum mechanical representations of quantum mechanical systems in our world.

That was my entry into AI. Given that this is a technical podcast, I'm going a bit deeper than I usually do, which is really nice for the backstory here. But over time, what I realized was that this sort of physics-based approach—if we're trying to understand the physical world with representations, we want to have physics-informed representations or physics-inspired—

Tim Scarfe

Mm-hmm.

Speaker 0

Representations. Just like, to understand a quantum mechanical system, I use a quantum neural network. The dual to that is, if I want to run and learn and train these representations that are physics-inspired or physics-based, then I need a physics-based accelerator, right? That's where that type of program fits natively and can be executed as physics.

In my case, initially, it was: I wanna understand quantum mechanical systems, I wanna learn quantum physics-based programs, and I run them on quantum physics-based processors, right?

Tim Scarfe

Then I have a clarification question, just for the benefit of the audience.

Speaker 0

Yeah.

Tim Scarfe

We have a very technical audience, but let's just make sure everyone follows. What do you mean precisely by physics-based computing? I ask because it might confuse some of our listeners. In some sense, all computing is physics-based, right? It's all shuffling around electrons in circuits and so on. What do you mean specifically by physics-based computing, and how would that differ from the commonsensical or classical notion?

Speaker 0

Yeah, a quantum computer is a computer that leverages quantum mechanical evolutions as a resource, I would say. There are more formal definitions of what a quantum computer is.

A physics-based computer, I would say—of course, like every physical thing is embedded in the physical universe, so everything is—

Tim Scarfe

Right.

Speaker 0

Physics-based. But it's true that it's kind of a continuum, right? Because you can have physics-inspired computers that are—

Tim Scarfe

Mm-hmm.

Speaker 0

Doing digital emulations of physics-inspired—

Tim Scarfe

Right.

Speaker 0

Algorithms. So I would say that's physics-inspired. I would say a physics-based computer, at least for a quantum mechanical computer, is one where you have a parameterized Schrödinger evolution that you control, and you can show that it has some unitarity.

In our case—and we'll get to that—we're building physics-based stochastic computers, or stochastic thermodynamic computers.

Tim Scarfe

Mm-hmm.

Speaker 0

They're parameterized stochastic evolutions that we control. In principle, you could try to emulate a quantum mechanical computer or emulate a stochastic computer digitally. Of course, in the case of a quantum mechanical computer, there are proven separations of complexity there, and there are experiments showing that you can't emulate them at scale. In our case—

Tim Scarfe

Right.

Speaker 0

With stochastic computers, there's no complexity-class separation. It's just orders-of-magnitude constant speedup or energy-efficiency gain, which practically makes them intractable to emulate at scale, right?

Tim Scarfe

Right.

Speaker 0

Not impossible, though, of course.

Tim Scarfe

Right.

Speaker 0

But yeah. So this journey—and feel free, if you have a better definition of physics-based computers, I think you have a better definition. I don't know if you wanna go into it, but—

Tim Scarfe

Well, no, it's consistent with what you were saying. At a very high level, I would say that a physics-based computer is essentially a computer whose components exploit the actual physical properties of the components in order to perform computations more efficiently.

Speaker 0

Yep.

Tim Scarfe

So rather than use digital computation to simulate what, for example, is a quantum phenomenon or a thermodynamic phenomenon, you actually use the probability densities that these systems embody, for example, at equilibrium, and exploit those properties in computation. Would you agree with that as a broad definition?

Speaker 0

That's for a thermodynamic computer, right? But yeah, there are different kinds of physics and different kinds of corresponding physics-based computers. Arguably, you can have probabilistic or quantum photonic computers.

Tim Scarfe

Right.

Speaker 0

Of course, you could have physics-based computers in all sorts of substrates. Technically, you could probably—

Tim Scarfe

Mm-hmm.

Speaker 0

Train a neural net from puddles of water, right, if you wanted. And that would still be a physics-based computer.

Tim Scarfe

Right.

Speaker 0

So I think it's kind of a very broad community. I would say physics-based computing, if we include—

Tim Scarfe

Mm-hmm.

Speaker 0

Quantum computing and alternative computing, has all these archipelagos that don't necessarily talk to each other. I think there could be more work in trying to unify the community, because there are a lot of tools that could cross-pollinate between these different substrates.

4. The Quantum To Thermo Pivot

But to get back to our main line, which is: how did I end up going from quantum to thermo? I think going from being one of the first programmers of AI on quantum computers and seeing the field evolve over several years—I was almost 8 years into quantum computing.

Tim Scarfe

Mm-hmm.

Pushing compute to the limits of physics

It wasn't progressing as fast as I wanted, and I was seeing a sort of writing on the wall where, actually, to have a quantum mechanical computer, you try to keep the computer at perfect zero temperature, right? Zero entropy. So you're constantly pumping out entropy that's seeping into the system through noise, right? And that's quantum error correction and fault tolerance. It turns out that most of your computation, most of your energy, is going to be sunk into that pumping. So it's like an algorithmic form of refrigeration.

Tim Scarfe

Right.

Pushing compute to the limits of physics

To me, it's like, okay, well, if you have a fridge and you're trying to keep your freezer really cold, it's going to use up a lot of energy if the rest of the room is at room temperature, because the gradient of temperature is very difficult to maintain. So it's like, okay, what if we had a hotter, physics-based computer?

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Maybe it would be much easier to maintain. That's just a first thought, right? So what would that look like? In the limit where you let the noise seep in, things become mostly stochastic, and the quantumness kind of fades out. We know this from renormalization of physics. That's why we don't have to care so much about quantum physics day to day, because it gets washed out at larger scales, right?

Tim Scarfe

Yeah. Would you say that it's accurate to frame it this way? In classical computers, the scale of the fluctuations is dwarfed by the scale of the components of the system that you're considering. So you can basically ignore the noise for the most part, assuming that you cool the system appropriately.

In quantum computing, it's almost the other way around, where what you want to do is basically engineer the scale of the fluctuations so that they're so big that you can kind of ignore them relative to the component size. This is just an attempt at framing what you're doing, so feel free to disagree with this. And then, in the thermodynamic regime, what you have is fluctuations that kind of coexist at the same scale as the components that—

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

—the hardware is made out of.

Pushing compute to the limits of physics

Yeah, I mean, we're in Wimbledon weekend right now in London, and you don't have to understand the vibrations of the molecules at the molecular level in the tennis ball to be able to predict its trajectory, right? You kind of get a mean-field sort of prediction, right? So that's Newtonian.

At the quantum scale, it's more subtle because it's no longer probabilistic fluctuations.

Tim Scarfe

Right.

Pushing compute to the limits of physics

You're essentially dealing with purely quantum superpositions, right? And the ideal—

Tim Scarfe

Right.

Pushing compute to the limits of physics

—quantum computer has no probabilistic uncertainty. In fact, when there is probabilistic uncertainty, that's when the quantum computer loses its quantum coherence. It becomes non-quantum, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

But we have a very similar problem with classical computers—

Tim Scarfe

Absolutely.

Pushing compute to the limits of physics

—which was that we use all this power to have a signal that's so strong relative to the jitter of electrons and the amplitude of that jitter in order to maintain determinism, right? We want our transistors to be absolutely on—

Tim Scarfe

Mm-hmm.

Pushing compute to the limits of physics

—or absolutely off, because we have a lot of transistors, and we want the computer to be in a deterministic state that we have control over, right? Humans like to be in control, right? When we're not in control, there's anxiety, right?

Tim Scarfe

Mm.

Pushing compute to the limits of physics

I think we have to learn to let go. Just like I learned to, in a way, let go of control mentally by giving up the reductionist approach, which is the human-interpretable approach to physics.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

We have to sort of, just like we did with deep learning, let go of software 1.0, imperative programming, and let gradient descent be a better programmer than you. That was very humbling for some.

Some are still trying to hang on for control, right, with ML interpretability research and all this AI safety-ist research. They still want to feel like they're in control or that they understand the complex system, instead of letting it do its thing and letting it figure out what's best, right? That's also kind of my main gripe with the whole AI safety-ist field. Of course, there should be some research in interpretability, but I just don't think they're going to go that far.

To close the loop here, I think we should also literally let go and loosen our grip on electrons and hardware, and that's what we do, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

Essentially, we go from having a very tight grip and yanking these signals around to kind of loosening our grip, letting there be fuzz, and gently guiding the signals, right?

Tim Scarfe

So, in some sense, you're kind of switching teams, right? In the more classical way of thinking of this, what we're trying to do is to keep the noise at bay, right?

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

To make our systems as non-stochastic as possible.

Pushing compute to the limits of physics

Filter it, pump it out.

Tim Scarfe

But we—

Pushing compute to the limits of physics

Yeah. Pay the price.

Tim Scarfe

That's right.

Pushing compute to the limits of physics

Right?

Tim Scarfe

But what—

Pushing compute to the limits of physics

Exactly. What you're doing is saying, "No, no, no"—harness it, right?

As we know from Maxwell's demon, the classic fable—people can look it up—knowledge comes at a cost, right? Like reducing—

Tim Scarfe

That's right.

Pushing compute to the limits of physics

—reducing entropy in a system, keeping something in a deterministic state, always costs you energy.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

If every clock cycle we're trying to prevent the classical or quantum computer from decaying to a naturally probabilistic and relaxed, thermodynamically closer-to-thermal state, then we have to pay the price. We have to pay energy to maintain determinism and reduce entropy.

Whereas a thermodynamic computer is not always at equilibrium, but we're dancing much closer to equilibrium, and it's much cheaper energetically to sit in those states and maintain those states, right?

Tim Scarfe

Do you want to maybe just walk us through how you're actually designing these things?

5. Building Stochastic Accelerators

Pushing compute to the limits of physics

For a technical audience, I guess they're Markov chain Monte Carlo accelerators, right? We support discrete variables, continuous variables, and mixtures of the two. Essentially, we've found a way to harness the natural stochastic physics of electrons in order to accelerate Markov chain Monte Carlo.

There are some subtleties in how we do that mapping, but you could just imagine we're embedding into the stochastic dynamics—which are parametric, and whose parameters we control—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—we're embedding our continuous- or discrete-variable MCMC into the dynamics of the electrons in the device, right? So it's partly analog, and it's stochastic, but our most recent chip is a mixed-signal chip. We use digital classical components and sort of stochastic electronics, and the 2 have to—

Tim Scarfe

Right.

Pushing compute to the limits of physics

—interact. Just like in a Metropolis-Hastings algorithm, you have some components of the algorithm that have some entropy, some proposals, and then you have some non-random parts, right? Like computing—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—acceptance or rejection.

Tim Scarfe

So I guess this helps you address the paradoxical situation where—

Pushing compute to the limits of physics

Mm.

Tim Scarfe

—I was just saying, in some sense, from the hardware's perspective, you're switching sides and joining the side of noise. But from the point of view of software development, in the context of AI and machine learning, this has already happened, right? There's a kind of paradox in our current architectures where, as you just described it, we spend inordinate amounts of energy and effort pumping the noise out of the system. But then, you know, with sampling-based methods, like we then reintroduce it through the software.

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

Right?

Pushing compute to the limits of physics

Yeah, and you could think of even a transformer: you have a softmax layer at the end. A transformer is a big probabilistic computer already. So we're running probabilistic software on this stack that was made for determinism, which is highly inefficient.

And now, with test-time compute, thinking at test time, Monte Carlo tree search and all sorts of RL rollouts—that's a Monte Carlo algorithm, right? So that can be Tree of Thoughts, it could be discrete diffusion, it could be…

Now, even for content, there are diffusion models. Those are also—it’s also a Markov chain, right? You’re reversing—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

A Markov chain there. So algorithms—the biggest workloads that are eating the world and consuming a ton of energy—are actually probabilistic workloads. They are probabilistic—

Tim Scarfe

Right.

Pushing compute to the limits of physics

Graphical models. And so it’s kind of funny. I guess I think it’s just the community. There’s kind of the—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

New additions to machine learning that don’t learn the fundamentals; they just go straight to transformers or whatever’s hot.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

And then I tell them, “We’re building an accelerator for probabilistic graphical models,” and they look at me like, “What are you talking about?” It’s like, actually, you’re—

Tim Scarfe

Right.

Pushing compute to the limits of physics

You’re running probabilistic graphical models all day.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

It’s kind of funny that there’s a gap there. But hopefully, I think that by coming on this podcast and getting more of the machine learning community to talk to each other, probabilistic ML will have a resurgence. And of course, we’re trying to stimulate that, right? Because—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

There’s a co-evolution. What evolves together fits together, right? And there’s an evolution of—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Hardware and algorithms. We have, as investors, 2 of the authors of the Transformer paper, and they say—they kept telling us, “Transformers are not sacred,” right? They were what worked on—

Tim Scarfe

Right.

Pushing compute to the limits of physics

The hardware we had at the time, which was Google TPUs. And the hope is that there are going to be new foundation models, or new models that run more natively on probabilistic hardware, that are really efficient on our hardware, that—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

We’re not just importing current-day models. But hopefully, if you view the algorithmic landscape as having a certain fitness function of what you get in terms of performance and cost, that landscape is also induced by the current-day hardware that’s available at scale, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

And if you change that hardware substrate to a different substrate that has different preferences in terms of the structure of the algorithm that you’re running, then you’re changing the fitness landscape. And if you have a sudden shift in the—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Fitness landscape, you have a sort of Cambrian explosion. You have this sort of high-temperature phase of search in the landscape, right? And that’s hopefully what we’re going to cause, right? And that’s very disruptive. So incumbents have all the reasons to be skeptical. But at the same time, it’s funny that even the current sort of incumbent algorithms are converging toward sampling and more probabilistic algorithms, which—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Wasn’t obvious when we started the company in 2022, but that was the prediction, right? That there would—

Tim Scarfe

Right.

Pushing compute to the limits of physics

Be a sort of—you know, my joke is that we’re kind of interpolating between deterministic forward-pass models and full EBMs, right? And you can view—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Sort of diffusion models as an interpolation there. But—

Tim Scarfe

Well, correct me if I’m wrong, but your view is that this is essentially an inevitable development of tech, right? You—

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

You’ve talked about the thermal danger zone—

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

And kind of the inversion of Moore’s law into what you call Moore’s wall. Do you want to run us through the logic there?

6. The Thermal Danger Zone

Pushing compute to the limits of physics

Yeah. If you came from quantum computing, there, if you try to scale up your system, you have noise seep in. So again, you’re going from the quantum zone of physics to the thermodynamic zone. You’re kind of edging on it. And if you’re in the classical zone but you try to get smaller, then you’re getting into the thermal zone again, right, in terms of the physics. The jitter of electrons matters because your transistors are so small. The fact that there are very few electrons means that the jitter of the electron population matters compared to the amplitude of the signal.

Tim Scarfe

Right.

Pushing compute to the limits of physics

Typically, we like our computers to have an error rate of 10⁻¹⁵, so they can do—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Many operations before there’s an error, so that you don’t—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Actually have to run error correction, right? The only computers that usually run error correction are those that we send into space because of radiation. But the reason we can’t scale down deterministic computers is because of this thermal danger zone, and they’re innovating in all sorts of ways. They’re doing all sorts of weird fins and all sorts of weird designs that are sort of hardware-level error correction. But—

Tim Scarfe

But those are always going to be ad hoc in some sense, right? Because what you’re going to run into is just the scale problem.

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

As you miniaturize, at some point—

Pushing compute to the limits of physics

I mean, it—

These wires are going to become... Yeah.

Speaker 0

To use less power, you need to use less charge. That’s not controversial.

Tim Scarfe

Yeah.

Speaker 0

And when you get to very little charge, the fact that charge is discrete gives you noise, period. Right? So—

Tim Scarfe

Right.

Speaker 0

So you’re going to have to go thermodynamic at some point, right? And at the end of the day, we have proof of existence of a really kickass AI supercomputer that we’re both using right now to talk to each other. It’s our brains. And—

Tim Scarfe

That’s right.

Speaker 0

I would argue that is a thermodynamic computer, right? Because there are master equations describing the chemical reaction networks in your brain, with neurotransmitters hopping—

Tim Scarfe

Mm.

Speaker 0

Around, and you could just—

Tim Scarfe

Operating really close to the Landauer limit as well, right?

Speaker 0

Right. And so, if we’re doing very similar physics but with electrons sloshing around a circuit, then there’s a much stronger chance that we could run a very similar program in a very similar fashion, arguably with even more energy-efficient components, because electrons are much lighter than big neurotransmitters.

Tim Scarfe

Mm-hmm.

Speaker 0

And so, my—

Tim Scarfe

So basically, what you have is a programmable Boltzmann machine, right? You can think of these chips as a kind of programmable Boltzmann machine.

Speaker 0

Well—

Tim Scarfe

Like, you can think of these chips as, like, a kind of programmable bowl. Do you want to maybe tell us how they're used for computation?

7. From Superconductors To Silicon

Speaker 0

For context, we started off in superconductors, and there—

Tim Scarfe

Mm-hmm.

Speaker 0

The reason we started in superconductors was because there’s this beautiful theory in quantum computing actually called circuit quantum electrodynamics, where you go from, “Here’s my circuit and here’s my Hamiltonian,” right?

Tim Scarfe

Mm-hmm.

Speaker 0

Directly from the circuit. And I think that should win a Nobel Prize or something someday. That’s basically how we engineer quantum mechanical systems to have certain desired physics, right? And the point there is that a Hamiltonian gives you a notion of energy in the classical regime.

Tim Scarfe

Right.

Speaker 0

And so we had a very direct way to go from, “Hey, I want this energy. Here’s how I design my circuit,” right? And so that’s where we started.

Tim Scarfe

Right.

Speaker 0

Furthermore, it was the way to build the most—

Tim Scarfe

Mm.

Speaker 0

Macroscopic thermodynamic computer you can build. Any other companies with claims to build a much more macroscopic thermodynamic computer either did it digitally or emulated it, so it’s not a real thermodynamic computer. But we had to supercool it because it’s too big, right? And, you know—

Tim Scarfe

Right.

Speaker 0

With the Boltzmann distribution, as you go smaller, you have higher frequencies, so you can go to higher temperatures, and then we moved to silicon. But to answer your question about energy-based models, you could view our chips as a time-dependent programmable energy function, right? And—

Tim Scarfe

Right.

Speaker 0

You have something akin to Langevin dynamics, more generally diffusion, in that landscape. And Langevin dynamics happens to also be an MCMC algorithm, right? So there’s a—

Tim Scarfe

Right.

Speaker 0

There’s a perfect match there between the algorithm that you can use for Bayesian inference for anything, really, and the native physics of the chip, right? And so it’s almost like—

Tim Scarfe

That’s so cool.

Speaker 0

It’s been there the whole time, right? It’s like we’ve been doing algorithms that were physics we could literally implement, right? It’s a literal analogy, and we’re instantiating it as an analog computer.

That was too beautiful not to build. So we did build it. Hopefully, we can cut some B-roll. I brought some superconducting chips here. We'll do that later. But now, actually, I think the big breakthrough and the big challenge was to—

Tim Scarfe

That's it, eh?

Speaker 0

Move to silicon, yeah.

Tim Scarfe

Wow.

Speaker 0

Getting programmable stochastic physics in silicon was an order of magnitude more difficult, and we had to really innovate in terms of understanding the stochastic mechanics of electrons in silicon. We built our team for that. Essentially, we reproduced some core primitives. The one we're talking about for now is the probabilistic bit.

Tim Scarfe

Right.

Speaker 0

The probabilistic bit you could think of as a double-well system, and you can—

Tim Scarfe

Yeah.

Speaker 0

—tune the tilt and so on. So if you consider bouncy balls in this landscape, you can control how much time the bouncy balls spend in one well or another, and you can consider one well 0 and the other well 1. Essentially, you have a signal that's dancing between 0 and 1, right? And you can control—

Tim Scarfe

Right.

Speaker 0

—how much time it spends in 0 and 1. And so that's like a fractional bit, right? So—

Tim Scarfe

Right.

Speaker 0

Going back to what we were saying earlier about coming at it from a qubit school of thought, we basically made a p-bit from it. Initially, it was superconducting materials, but now we've done p-bits in silicon and achieved really, really high energy efficiencies. Our p-bits can generate controllable bits of entropy with only a few hundred attojoules. We've recently submitted our results for peer review, and we're going to put them on archive—probably timed with some other announcements in the coming months. But it's a really exciting time.

Tim Scarfe

That's so cool.

Speaker 0

And again, this was a concept for a very long time, and it was a big risk for me. Initially, when I left theoretical physics and went all in on quantum machine learning, everybody told me, "Really? You? You're going to be a quantum machine learning lead at Google?" Everybody was doubting me. Then I became kind of well-known, and I eventually led a team. I led quantum machine learning at Alphabet X.

And then I quit all that. I quit quantum computing and quantum machine learning. I'm going to take even more risk. I'm going to build a whole new paradigm of computing from scratch, from the concept up, and everybody thought that was crazy. And now we're here. Now we've made a lot of progress, and now we're scaling, right? There's nothing stopping us from scaling at this point because we've de-risked the manufacturing as well, which is usually not something—

Tim Scarfe

Right.

Speaker 0

—academics think about. But if you're a startup and you have to scale or die—eat the world or die—you have to kill every risk possible. We're looking to scale to millions of degrees of freedom next year, and I think we're going to hit it, so—

Tim Scarfe

That's really exciting. So your first chip had 3 p-bits.

Speaker 0

Yeah.

Tim Scarfe

And the latest has about 300 degrees of freedom.

Speaker 0

Yeah.

Tim Scarfe

Am I correct? And now you're—

Speaker 0

They're not all p-bits, but we'll—

Tim Scarfe

And now you're aiming… Right.

Speaker 0

We'll get to that someday. Yeah.

Tim Scarfe

Right. And now millions of degrees of freedom—

Speaker 0

Yeah.

Tim Scarfe

—next year.

Speaker 0

Yeah.

Tim Scarfe

That's very cool. So help me understand the general space here. I mean, presumably, you don't think that this is a— Or do you? Do you think this is a wholesale replacement for the current stack, or do you see— Because I could see a world where you interface—well, you use thermodynamic compute when it's relevant, but you mesh this with digital and quantum—

Speaker 0

Yep.

Tim Scarfe

—at the appropriate junctures so that you get… What you really get is a multiscale stack where each hardware bit is specialized for the kinds of computations that run natively on that kind of hardware.

Speaker 0

Yeah, absolutely. I mean, if you're— I think there was some work from Max Tegmark on correspondences between the information bottleneck principle, downsampling, the hierarchy in machine learning, and the renormalization group, right?

Tim Scarfe

Right.

Speaker 0

If you're trying to learn from quantum mechanics and you downsample, you get statistical mechanics. You downsample, you get to—

Tim Scarfe

Mm.

Speaker 0

—Newtonian mechanics. You can imagine if I had a God neural network that just compressed all the information in a certain region of space—which is actually the thought experiment that got me into quantum machine learning. I was trying to understand black holes as a machine learning system. But let's say you—

Tim Scarfe

Mm.

Speaker 0

—created that system, then you could imagine the first few layers are quantum, then some layers are probabilistic, and later they're deterministic, just as you distill the information. Initially, you need some quantum complexity. Later, you need some entropy, and then later you just need to do some classical coordinate transformations, right?

Tim Scarfe

Mm-hmm.

Speaker 0

And they all work together. But in a more practical workflow, very often—whether you're doing simulations of stochastic differential equations, trying to simulate a physical system, or doing some sort of discrete or continuous diffusion—there's a classical function, often a differentiable program, that determines some probability distribution that you want to sample from, whether it's—

Tim Scarfe

Right.

Speaker 0

—in latent space or for these transitions and the denoising. So the two work together. You don't necessarily need entropy everywhere in your graph all at once, right?

Tim Scarfe

Right.

Speaker 0

In principle, you could use a probabilistic computer for deterministic operations, but it's not going to be the best at that. In the low-precision regime, you could think of p-bits as literal fractional bits. So you can get into—

Tim Scarfe

Mm.

Speaker 0

—the fractional-bit precision work, and that can be interesting, but not everything is well suited for that low precision.

Tim Scarfe

Mm.

Speaker 0

And I do think quantum computers will have some applications, but again, they would be supplements for quantum-mechanical systems—to understand quantum-mechanical systems—as a supplement to probabilistic and classical computers. But I would say that most things would be well covered by probabilistic and deterministic representations running on probabilistic—

Tim Scarfe

Mm.

Speaker 0

—or thermodynamic and deterministic computers. Yeah.

Tim Scarfe

So currently, the superconducting chips require a lot of cooling, right? You need to get them around 1 kelvin. Is that the case?

Speaker 0

Yeah. Depending on the material, you can be at a few hundred millikelvin, or you can get to a few kelvin, usually using niobium. I think that you don't need as big of a fridge. For us, it was an interesting experiment, and we have a bunch of results there. We have a paper coming in the coming months. Essentially, it's as efficient as we could imagine building a thermodynamic computer. Again, it was just to get people to imagine—or realize, rather—that there are forms of computing that are far more energy efficient by an unfathomable number of orders of magnitude, far more efficient than digital computers, right?

Tim Scarfe

Right.

I'm just doing some back-of-the-napkin calculations, because you guys target 1,000 to 100,000× energy-efficiency gains at the chip level, and I was wondering: how does that stack up against the cryogenic energy costs? How does that factor into your efficiency calculations?

Speaker 0

Yeah. We put the asterisk there: that's just the chip. In general, the thought experiment was: if you scaled a very large network of superconducting chips to football-field size and had a ginormous dilution fridge, right? Again, the big energy cost is maintaining that boundary, right, with the—

Pushing compute to the limits of physics

The outside world. If you had a ginormous dilution fridge and millions and millions of p-bits and superconductors, you'd have the most energy-efficient probabilistic computer, right? Is that practical? Probably not. Should a startup be doing it?

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Probably not either. Could a government do it in some sort of crazy moonshot? Maybe. We're kind of just going to put the idea out there for academia, national labs, and whatnot to pick up. I think for us, it was also just a great learning platform. Superconductors are ironically pretty accessible. A bunch of universities have fabs where you could experiment with them.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

For us, it was kind of like, “Hey, we're going to want a community of people working on thermodynamic computing, experimenting with new primitives, and showcasing new algorithms, so we're going to put that work out into academia.” But for us, the product is, again, the silicon chips. We could have decided to run them in cryo-CMOS, but we ended up deciding to go for room temperature. Of course, a thermodynamic computer at a lower temperature consumes less energy, right?

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Of course, you're assuming you're maintaining the bath at that temperature, but then you have the cost of the cooling, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

If you had a von Neumann probe thermodynamic computer, maybe it could run much, much, much cooler, right? If it's going to go to—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—in interstellar space, it's not going to encounter that much heat, and so maybe we'd modify the design for that. But that's a problem for the far future.

Tim Scarfe

And then, coming back to the impact on the AI industry, I mean, EBMs—energy-based models—have been around for a long time. Would you say that this is really the key to unlocking their potential? Up until now, they've been basically limited by their inefficient sampling, right? Would you say that this is really the technology we need to do the EBM thing seriously?

Pushing compute to the limits of physics

I think modern neural networks are just mean-field approximations of EBMs, right? Neural networks came from EBMs. Backprop came from looking at the mean field of EBMs. To us, we're just creating the hardware for the ancestors of neural networks, and they're kind of a superset, right? You could just take averages and get deterministic operations.

Tim Scarfe

Right.

Pushing compute to the limits of physics

Our hope is that people think about going more probabilistic with their algorithms now that sampling is far more energy-efficient and far faster. But because we've been stuck with deterministic computers, people have tried to avoid using primitives that— They would just represent distributions by their moments, right? They would use exponential families where you don't need too many moments, such as Gaussians, because you could just represent them as matrices and vectors, and matrices and vectors—

Tim Scarfe

Right.

Pushing compute to the limits of physics

—can fit on a GPU really well, right? And you could do some transformations there.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

So there's been a bias in algorithms toward what runs well on today's hardware, and hopefully, with thermodynamic computers, that changes. We're trying to get people interested in the space and to start imagining what they would do with a multimillion—

Tim Scarfe

So this is how you break free of the—

Pushing compute to the limits of physics

Go ahead.

Tim Scarfe

Well, this is how you break free of the vicious cycle that we're stuck in right now, right?

Pushing compute to the limits of physics

Yeah, we're stuck in a—

Tim Scarfe

We've got a hardware stack that works in a certain way.

Pushing compute to the limits of physics

Yep.

Tim Scarfe

We optimize our software for that, then we build bigger hardware and, you know—

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

—so this is how we break free from that.

Pushing compute to the limits of physics

Yeah, but I think we're going to break free one way or another, because frankly—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—even if you just do a back-of-the-envelope calculation, right now we're going to run out of power trying to scale AI. We can't even—

Tim Scarfe

Right.

Pushing compute to the limits of physics

If everybody were to use an agentic, big model—a ChatGPT or Groq model—we'd run out of power in the United States. You can't deploy it to everyone. You can't have everyone using it with high intensity right now. We're going to run out of power. And that's not even getting into video models and world models—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—which are going to be necessary for embodied intelligence. At that point, if you try to scale it to the planet, we're going to run out even if we produce the power. Let's say there was some moonshot, and let's say we even figured out nuclear fusion, right? Let's say we just did that and scaled it very quickly. We'd run out of—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—or, you know, we're going to double or triple the amount of heat being radiated by the Earth, right? So we're literally going to cook ourselves to death. We're literally cooked. And so something has to change—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—at the hardware layer. We can't scale with the current hardware. People who think we're just going to scale transformers and get to the moon and scale intelligence to the whole planet are flat-out wrong, right? It's provably wrong.

Tim Scarfe

When people say, “Hey, to scale up intelligence, we need to start building nuclear power plants,” I always think, well, I run on a glass of water and a banana—maybe a coffee in the morning, right? Certainly, this is not a naturalistic way to think about how intelligence scales, right?

Pushing compute to the limits of physics

Because you're a thermodynamic computer, right? That's the thesis.

Tim Scarfe

That's right.

Pushing compute to the limits of physics

You can imagine just taking our current design and scaling it to a wafer. You would have about 1.5 billion p-bits, which you could think of as neurons, and about 20 billion parameters per wafer. And then you could do multilayer programs of those, and that would run on not 20 kilowatts, which is what a current wafer-scale system would run on, or more—maybe 100. It's 20 watts, right? Which is like our brain.

Tim Scarfe

Wow.

Pushing compute to the limits of physics

Right?

Tim Scarfe

That's crazy.

Pushing compute to the limits of physics

And then, if you have a few-hundred-billion-parameter model running on 20 watts, then we're in the same ballpark as the brain. Not quite exactly—it might be within 10X—but that's still much better than where we are, which is like—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—100 million X, right? And so that's where we're going. Again, it's not— To us, it's far less crazy than quantum computing. There is no quantum computer in nature that has very high quantum coherence, and we have exquisite control and very high quantum complexity. All physical systems decohere. We have multiple billions of thermodynamic computers out there in the wild, and they work pretty well.

We're just trying to tap into the same physics. We're not obsessed with biomimicry. We don't call ourselves neuromorphic computing. We're just doing, again, approximate probabilistic inference as a service, but in hardware, in physics. To us, that's kind of the parent workload of most of AI. And actually, it's not just for AI.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

It's for broader computing, right? There's simulation, optimization, and statistical inference. Science at large can use this. So it's not just for GenAI, right? We're not just riding the current wave.

Tim Scarfe

Yeah. I have the pleasure of interviewing not one but two people today—

Pushing compute to the limits of physics

Sure. Yeah.

Tim Scarfe

—in some sense.

Pushing compute to the limits of physics

We'll switch hats.

Tim Scarfe

I've got your— Yeah, exactly. We've got your alter ego on the call as well. So, to frame up the transition, for the audience, Guillaume, or rather, a.k.a. Beff Jezos, is the founder of a philosophical movement called Effective Accelerationism, which originated as a kind of counterpoint to effective altruism. I really want to get into the details there. But to set things up, looking 10 to 20 years out, how do you envision the future? And what role does acceleration play in your vision for the future?

8. A Future Of Embodied Intelligence

Pushing compute to the limits of physics

Yeah. I think in 10 to 20 years, we'll have embodied intelligence.

I think thermodynamic computers will be pretty ubiquitous. They’re going to be what runs embodied intelligence. We’re going to have greater intelligence density in all our devices. We’re going to have personalized, always-on, online-learning intelligence. And so my goal is for everyone to own and control the extension of our cognition—

Tim Scarfe

Right.

Pushing compute to the limits of physics

—because that is important to maintain—

Tim Scarfe

So, democratizing intelligence.

Pushing compute to the limits of physics

Not just democratizing—yeah, but truly, because democratizing access is like, “Hey, I give you access to the one God model that amortizes its learnings across the fleet. Now you’re all mind-merged with my Borg mind. Congratulations, you’ve been assimilated.” That’s not democracy, right? If you have control over its constitutional prompts,

Tim Scarfe

Right.

Pushing compute to the limits of physics

—you’re essentially controlling people by proxy, and their thoughts, their will, and their actions, right? If you control—

Tim Scarfe

Yeah.

Pushing compute to the limits of physics

—it’s not just—It used to be people controlling people’s access to information and steering them indirectly. But now it’s going to be directly controlling the model that’s an extension of their cognition or their thought partner, and that’s much deeper and more subversive control. And so, to me, I think that was the big existential risk I was most worried about: the sort of precedent—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—of people vying for control and vying for power over people.

Tim Scarfe

Right.

Pushing compute to the limits of physics

And that’s why I’ve been pushing for decentralized AI, and I’ve been pushing for avoiding overregulation of AI that would cause—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—red-tape inflation, mostly serve the incumbents, and cause a centralization of AI power.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

But—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Yeah, someday, I would imagine, 10 or 20 years out, hopefully we’re wearing some neural links. Our whole skull is a neural link that is the thermodynamic computer, an extension of our cognition, and it’s—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—as powerful as, or if not more powerful than, your brain. And you have a full merge there. I think something more plausible on a 10-year timescale is a sort of soft merge. I mean, you’re—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—you’re the active inference expert here. But if you have the same Markov blanket, you have the same perception and action states. So let’s say I had an agent. Let’s say my glasses were very smart, and they were always on, always listening, and I had an earpiece. I could have sort of subconscious thinking, almost. It could just perceive things and suggest actions, so we’re sharing perception, and we’re sharing actions through my body.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

We are the same agent, right? So that is—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—that is a form of merge, and you don’t need a neural interface for that, right? And maybe we even have a way to communicate nonverbally or, you know, just like friends that hang out a lot have a prior of each other’s behavior, or their behavior conditioned on the world’s states.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

And then they don’t need as many bits of information to adjust to a new setting. And so that’s kind of a more plausible way, I think. I think the soft merge is going to come first, but I think it’s going to keep escalating, and then we’re going to go toward a harder merge, right? A hardware merge.

Tim Scarfe

That’s a very compelling vision for the future. It’s very techno-futurist. I love it. Do you want to give me the elevator pitch for EAC, Effective Accelerationism? Because I think these are all related. I think this sets you up nicely to just present the core idea, so—

9. The Logic Of Effective Acceleration

Pushing compute to the limits of physics

Yeah, really, EAC is a sort of metaculture. I call it a cultural hyperparameter prescription, if you will. And it’s one where we’re trying to maximize the growth of civilization, as measured by our free-energy production and consumption, a.k.a. the Kardashev scale. The Kardashev scale is a log scale—

Tim Scarfe

Right.

Pushing compute to the limits of physics

—tracking how much free energy is consumed and produced. And really, it comes from the realization that, as I was studying stochastic thermodynamics in preparation for founding Extropic, I realized that there was this sort of generalization of Darwinian selection, which is thermodynamic selection. Where each bit—you go from selfish genes to selfish memes to selfish bits—every bit of information specifying configurations of matter is fighting for its existence in the future.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

The selection pressure on those bits of information is whether they confer on their host organisms the ability to understand their environment, predict the future, capture free energy, use it strategically, and grow. But really, the master metric is: if I change this parameter, this value, and I let time evolve, how much free energy has the system dissipated through its trajectory over time, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

The theorems of stochastic thermodynamics tell us that trajectories of the system that basically consume more free energy are exponentially more likely, right? And so—

Tim Scarfe

Right.

Pushing compute to the limits of physics

—and so that’s how you have this sort of pruning of branches that consume less free energy. And so it’s like, “Ah, okay, well, this is the golden metric of selection pressure on the space of bits. And so I will design the highest-fitness selfish meme that is e/acc: figure out what is optimal for growth and do it.” And so, by construction, it should be the most viral thing, and it will persist.

Tim Scarfe

Right.

Pushing compute to the limits of physics

And, surprise—

Tim Scarfe

The assumption there is basically—

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

—the alternative is basically growth or death, is what you’re suggesting, right?

Pushing compute to the limits of physics

Yeah. We have the saying, “Accelerate or die.” It sounds dramatic, but it’s meant to sound dramatic; it’s also reality. It’s either you align yourself with growth and you are a part of it, right? So whether you adapt—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—you adapt your culture toward growth, and then you’re a part of the growth and you benefit, or you don’t, and then you get outgrown.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Whether it’s at a national level, with your policies, right? Again, the EU versus the US. The US is obsessed with growth to some extent, and we’ve seen a sort of bifurcation in the GDPs and so on. So, at a policy level, at a cultural level, even at an organizational level, right? If a company is not obsessed with growth, eventually it just gets disrupted by one that outgrows it and then has more resources to outcompete it, right? And so I think—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

I think the point is that, essentially, if people are open-minded about new technologies in general and lean in, they will be positive. They will be selected for, and the people who are skeptical, push them away, shun them away, and want to go back to the cave get selected out. And so, even from an empathetic argument standpoint, trying to popularize acceleration—making people aware that this is actually how the world works—I’m sorry to break it to you: you either embrace the acceleration or you get selected out.

It’s also like, hey, if you care about yourself and your tribe, your company, whatever, you should lean in rather than lean out, and don’t listen to the people fearmongering at you. They are very often doing so out of self-interest. I view spreading deceleration as a form of psychological warfare, in fact.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

That’s how I view it, and that’s why I’ve been such a warrior online against the decel mindset. I just don’t see a scenario where being decel is beneficial to the host. And in fact, it’s just exploiting this gap in understanding of how the world works to destroy competition, right? And I mean, there’s a lot of this in nature and in general in human dynamics. People sabotage each other to outcompete each other.

Tim Scarfe

See, I’ve been thinking about this a lot. My good friend and colleague, Axel Constant, and I have a long-standing debate as to whether ethics, morality, and this kind of thing can be reduced to free-energy minimization at the end of the day. And I think it’s a difficult question, because thermodynamics describes all mesoscale objects, right? So basically any configuration whatsoever will be minimizing free energy in this fashion, right—seeking to find pockets of free energy and dissipating it as efficiently as possible.

I guess my worry is things that grow without bounds. In the biological case, you have cancer, for example. So what's the role of constraints in the maximization of entropy when we're designing the social system? Because surely things like dictatorships and fascist autocracies also minimize free energy.

Pushing compute to the limits of physics

They're a local optimum, right? They're not a global optimum, right? Just like cancer: it kills the host, and that's suboptimal on a sufficient timescale. Again, it's not instantaneous free energy dissipation, right? It's basically an infinite time horizon, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

And so blowing up the planet does burn up a bunch of free energy, but in the long term—

Tim Scarfe

Right.

Pushing compute to the limits of physics

It's the same reason life exists.

Tim Scarfe

Right.

Pushing compute to the limits of physics

It's much better to conserve and strategically use free energy to secure more free energy, keep growing, and have some order, rather than just burn it all in one go and have chaos, right? And thermalize in one go, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

That's why we have life. That's why we're not at equilibrium, right? We're not just burning up all our fuel and dying immediately, because as intelligent beings, we burn way more energy by being this sort of somewhat coherent system that has predictive power of its environment. We're kind of this energy-seeking fire, right?

Tim Scarfe

Right.

Pushing compute to the limits of physics

Yeah, but I would argue that e/acc is also a call to popularize complexism, complex systems thinking, and the free energy principle at large. We have to think through this lens about all systems in society, from policymaking to technology to innovation to basically everything—economics. It's a different way of thinking, one that is somewhat scary. Again, it's going from reductionism and rationalism to complexism and post-rationalism, and we don't have as much control and interpretability. We don't understand the world as well, right? We kind of have to—

Tim Scarfe

Yeah.

Pushing compute to the limits of physics

Pardon my French, but fuck around and find out: the FAFO algorithm, right? But really, it's kind of exploration and discovery, right? A priori, you don't know what's optimal, right?

In startups, you learn this. You have to actually go on the market and try stuff. Sometimes your prior—your model-based prior—can't do that well because the ecosystem is too complex to have a model with good predictive power. You actually have to be in an open loop.

I would say that hopefully there's a renaissance, and these ideas become more popular in all sorts of fields of science. Clearly, it's eaten the software world, and we're trying to make—

Tim Scarfe

Mm.

Pushing compute to the limits of physics

—this sort of school of thought eat the hardware world. I think we will do it. But again, most fields of study could be revolutionized by thinking through the lens of complex self-adaptive systems and the free energy principle.

Tim Scarfe

So, you've described e/acc as a kind of hyperstitious meme. Do you want to say a few words about what that means?

Pushing compute to the limits of physics

Mm.

Tim Scarfe

Yeah. Again, you're the active inference expert, so in active inference, you have perception and action, right? You can update your model based on your sensory information, so you're minimizing the divergence between your own internal model and the statistics of the world. But then the dual—

Pushing compute to the limits of physics

Mm.

Tim Scarfe

—of that is taking actions in the world to minimize divergence between the world and your predictive model of it, right?

Pushing compute to the limits of physics

Right.

Tim Scarfe

And, in a way, we're naturally biased toward this: the car goes where the eyes look when you're driving.

Pushing compute to the limits of physics

Right.

Tim Scarfe

If we look at very negative outcomes and we're obsessed with them, we will drive whatever system we're thinking about toward those negative outcomes. An example of this—

Pushing compute to the limits of physics

Right.

Tim Scarfe

A slightly controversial example is bioweapons research.

Pushing compute to the limits of physics

Mm.

Tim Scarfe

And, for example, COVID, right? I would say that COVID was an accident from bioweapons research. It was probably defensive, and we were trying to explore what would be a really bad scenario. What if we had this mutation, and this mutation would combine and it would be a really bad virus? Then they started experimenting and designing in that neighborhood of virus subspaces that would never—

Pushing compute to the limits of physics

Mm.

Tim Scarfe

—have occurred naturally, just from evolution.

Pushing compute to the limits of physics

Mm.

Tim Scarfe

But because we were exploring that subspace of bad things, because we were obsessed with it, we made it happen, right?

Pushing compute to the limits of physics

Mm.

Tim Scarfe

To me, if we're optimistic about the future, we tend to steer things toward that optimistic outcome. As a startup founder, if you're not optimistic about your startup, statistically, you are screwed, right?

Pushing compute to the limits of physics

Oh, yeah. From an active inference point of view, every action begins with a false belief, right? So first you believe that you're moving, and then you reduce the prediction error in the direction of action, right? So—

Tim Scarfe

Exactly.

Pushing compute to the limits of physics

I find this extremely compelling. One of the things I find most inspiring about e/acc is that you're trying to present a radically optimistic meme for the future. There's so much—not just AI doomerism, but so much doom and gloom—going around that, just from the point of view of neurobiology, we need these almost seemingly delusionally optimistic beliefs to get off the ground anyway, right?

Tim Scarfe

Yeah. You could think of the memetic sphere as a metacortex, right? It's a biological supercomputer. In e/acc, we're just trying to do active inference toward better futures, right?

Pushing compute to the limits of physics

Right.

Tim Scarfe

So we're spreading this meme of optimistic futures, and it actually does steer the world. It's been 3 years now, and it has steered policies throughout the world toward going for these moonshots, a resurgence of exploration of nuclear energy to climb the Kardashev scale, deregulation of AI, widespread embrace of AI, and companies being far more aggressive in their exploration of it.

So how do we spread it? How do we spread e/acc? I guess my question is motivated by this: I think most of the very visible e/acc figures—Andresen, Musk, and company—tend to be libertarian, right-wing-associated. If we want e/acc to really be—

Pushing compute to the limits of physics

Well, that happened after the fact, right? I would say Gary Tan is rather on the left, and initially it was very apolitical. I would consider myself still a centrist.

Tim Scarfe

Well, how do we get the abundance bros to join?

Pushing compute to the limits of physics

I think it's already happening.

Tim Scarfe

The Ezra Kleins and—okay, cool.

Pushing compute to the limits of physics

I think that was a reaction to e/acc in the techno-optimist world. You end up getting clustered with one party in the US. In other countries, it might be different. But there needed to be an actual techno-progressive, let's say, versus techno-regressive, cluster on the left. It's not clear that that faction is the dominant faction of the Democrats. I really hope it becomes that. Again, e/acc is kind of like an RL algorithm. It's not clear what the optimal policy is.

I think a few years ago, I was seeing a sort of convergence toward a lot of top-down control, a lot of top-down power. I think the optimal is always balance. You need some top-down control, but you also need bottom-up self-organization.

Tim Scarfe

That's how the brain works, right? You have some centralization, but not all that much centralization, right?

Pushing compute to the limits of physics

Exactly, right. Absolute total libertarianism can work. It's just that in the era of unconventional warfare, complex systems also get adversarially steered, right?

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Whether it's free markets—they get gamed—or memetic markets, there are psychological operations, dating markets, and so on. I think some balance of top-down and bottom-up is the ideal. Is it more like the U.S., more like China, or something in between? I don't know. I guess we have to figure out.

Tim Scarfe

Well, that segues perfectly into my next question.

Pushing compute to the limits of physics

Yeah.

Tim Scarfe

For you, this isn't just a philosophical thing. This is of geopolitical and geostrategic significance, right? This is important for the future of democracy as we understand it.

Pushing compute to the limits of physics

Yeah. I would say that our failure to view the world as a complex, self-adaptive system, and our continued thinking in the classical way—“Hey, I have a first-order model of what's happening, and I'm going to do a first-order correction”—is a problem. It's easy to convince a crowd of that, and politicians get elected on such platforms, but then they don't think about the higher-order effects of their policies. We don't think about it from a complex-systems steering standpoint.

Tim Scarfe

Mm.

Pushing compute to the limits of physics

I think our adversaries—the adversaries of the West—understand the complex-systems approach, and they tend to steer us in directions that lead to our detriment. To me, it was, like, okay, how do you fight a multidimensional war?

Tim Scarfe

Mm.

Pushing compute to the limits of physics

Okay, there's invisible sabotage of every complex system that can be steered adiabatically in nefarious directions. If it's slow enough, then it's kind of above the infrared temporal cutoff that politicians give a shit about, which is 4 years, right? If you're steering the United States or the West on a 20- to 40-year timescale, you could win a very long war, and we don't even find the pattern because we're too busy.

Tim Scarfe

Right.

Pushing compute to the limits of physics

There's a sort of timescale separation here, similar to what happens in thermodynamic computing and thermodynamic physics. Essentially, we're so preoccupied with the dynamics that occur on small timescales—the day-to-day—that we ignore the long trends, and those get hacked, right? We slowly boil the frog. It doesn't realize it.

Tim Scarfe

So, in listening to you talk, it occurs to me that there's a beautiful coherence to your approach generally. What you're doing both at the hardware level and at the level of your philosophical project is essentially moving us away from hard, rigid programming from the outside toward a kind of organic, adaptive, thermodynamic-driven learning and exploration of possibility space.

Pushing compute to the limits of physics

Yes. Yes. Again, it came from my own journey trying to understand physics, discovering differentiable programming, and seeing that as the way forward for everything. Some of the prescriptions of e/acc are to maintain variance and constantly explore across any parameter space, whether it's culture, aesthetics, policy, technology, et cetera.

One of the reasons is that, according to Fisher's theories on evolution, the speed at which you can traverse a landscape depends on the gradient, but it also depends on the variance. Evolutionary search depends on variance, because the more you fuck around, the more you find out. Your rate of learning is faster, and your rate of adaptation is faster.

To me, that seemed like the main advantage of the United States: the United States is very high-variance. It has high-variance individuals and high-variance outcomes. I view the United States as a sort of high-temperature search algorithm.

Tim Scarfe

Mm.

Speaker 0

I'm not American yet officially, but someday I will be. That's my goal. They're first to figure things out because they're always searching in a very high-variance way.

I view innovation as a diffusion process in some landscape, and they're in a very high-noise regime, so they don't get stuck in local optima, right? Whereas China's more like the low-temperature sampler. They're the optimizer at the end. They're doing the gradient descent.

Once there's consensus, once there's a clear gradient of improvement, it's easy to convince a committee at that point, and then you can just execute top-down control with a lot of conviction.

Tim Scarfe

Mm.

Speaker 0

They beat us in the final stretch. It's kind of like Bayesian inference: you go from an unsharp prior to a sharper posterior on what the optimal thing is. It's like annealing.

Tim Scarfe

Mm.

Speaker 0

Essentially, China beats us at squeezing the end, again because it has a lot more top-down control power and a lot more coherence there. I see this sort of—we're kind of like a parallel-tempering algorithm between the United States, the high-temperature search that discovers things first, and then China.

The United States isn't necessarily the best at optimizing and improving the technologies and scaling the manufacturing.

Tim Scarfe

Well, then how do we avoid catastrophic outcomes in that? You've been pretty dismissive of P(doom) and doomerism generally. How do we allow the kind of meta-search to happen, but in a way that doesn't lead to catastrophic technological outcomes?

I can also see some tension between this open project, on the one hand, and the participation of proprietary, closed corporate groups, on the other hand.

Speaker 0

Yeah. I guess for me, P90, 1984 was higher than P(doom) from AI. To me, China is the ultimate monopolistic company. There's basically one set of executives for all the companies. They're acting like a conglomerate, and they throw their weight around to crush smaller American companies.

One thing about e/acc is that we're kind of anti-monopolistic, because monopolies tend to be suboptimal. If you have hyperparameter choices that are over-concentrated, you're not exploring anymore.

Tim Scarfe

Well, you see, that sort of motivates my question and my worry. Very authoritarian forms of government, where we impose ethnic or cultural homogeneity, are very good at minimizing free energy. If you and I are exactly alike because there's a top-down imposition of sameness, then—

Speaker 0

Mm.

Tim Scarfe

That minimizes free energy super well. If we're all the same—

Speaker 0

Locally, right?

Tim Scarfe

Right.

Speaker 0

Well, if your energy term says you want to agree with your neighbor and have no frustration, then of course that's a local optimum. But what we're arguing for is embracing variance: embracing different cultures and having different bets.

Tim Scarfe

Right.

Speaker 0

It could be culture, genetics, or ways to train your ML models.

Tim Scarfe

The same response to the cancer thing, right? You're not worried about the push toward boundless growth leading to cancer-type formations because it's not a global optimum. It's a local optimum.

Speaker 0

Yeah, yeah. That's right. I want to go on a slight tangent here. There have been some centralization in AI research labs, even though there are 4 or 5 players that matter. They kind of churn. It's been a meme recently that researchers all slosh around and churn between each other.

They're all equilibrating in terms of beliefs because you have an exchange of particles—of researchers—between the labs. You could see that the American research labs all converged to similar performance at similar compute.

Then China, with DeepSeek, was exploring a whole different region of hyperparameter space due to its export-control constraints, and showed us that over-concentrating our bets in hyperparameter space has risks, because we're stuck in a local optimum. There might be a better, nonlocal global optimum that we're missing.

Speaker 1

Mm.

Speaker 0

And so, again, it's just been our push for variance there in order to always be exploring. I think it's really important. I think basically there's also been a reaction to the same pattern. It's hard to explain. It's the same pattern that happens in Western governments, in late-stage corporations, and in large bureaucracies.

There's a “cover your ass” mentality and culture, and decision by committee. And so it tends to be variance-reducing, right? Or variance-killing.

Speaker 1

And that's sort of my worry. In addition, I guess we're in the same kind of situation as a gradient descent learner. How are we supposed to know that we're not stuck in a local optimum, right? To us, we might think, “Hey, this actually looks pretty optimal,” but there's no kind of God's-eye view that could tell us that actually we're just stuck in a local—

Speaker 0

Yeah.

Speaker 1

Speaker 0

But that's why, as long as you don't kill variance completely, you keep some exploration, you're not just purely exploiting, and you keep open-mindedness, then you're always spending some resources still exploring and—

Speaker 1

Mm-hmm.

Speaker 0

—and potentially finding a new, better way to do things. Just like, I think some people at first thought our bet was ludicrous and that we shouldn't have raised VC funding. It's like, really? There's trillions of dollars of capital riding on the current paradigm, and you don't want a startup to raise a few tens of millions to take a bet that would be completely disruptive to everything—to this whole multi-trillion-dollar bet, right?

Speaker 1

Mm.

Speaker 0

And now the tune is changing a lot, because people who are planning these half-trillion-dollar build-outs want to know if a technology that's 1000x better is around the corner and could render some of their build-outs maybe miscalibrated, right? In terms of the ratio of energy—

Speaker 1

Right.

Speaker 0

—to compute and density, right?

Again, I don't think thermodynamic computing is going to work in tandem with GPUs or TPUs or whatever neural processing units. I think it's going to take quite a while for models to be fully ported to thermodynamic computers. So, for the foreseeable future, I think these build-outs are safe, and you could just have thermodynamic computing as an add-on and—

Speaker 1

Guillaume, it's been a real pleasure to discuss these issues with you. So how do we stay abreast of what's going on at Extropic? Do you have any closing message for us?

Speaker 0

Yeah. Well, stay tuned in the coming weeks and months. For those that are interested in trying out thermodynamic computing, in the coming weeks we're going to give access to the first users to our systems, so you can test it yourself. It's mostly going to be a private alpha in the early days. Again, we can't handle that many users. We don't have that many chips.

But the technology's here. It's very real, and you can kick the tires. Stay tuned for some scientific papers and some open-source software that we're going to put out. I would say start reading more probabilistic machine learning papers if you're a grad student interested in the area. Start reading about EBMs and various textbooks. And stay tuned for more announcements from Extropic toward the end of the summer and this fall.

Speaker 1

Well, for Machine Learning Street Talk, I'm Maxwell Ramstead. We're signing off. And thanks again, Gil. This was awesome.

Speaker 0

Thanks, Max.

Speaker 1

Yeah, fascinating. Really cool stuff.

Speaker 0

Awesome.

Pushing compute to the limits of physics | BidClub