[BidClub_]
Machine Learning Street Talk · · 73 min

Jurgen Schmidhuber on Humans co-existing with AIs

Jürgen SchmidhuberTim Scarfe

Podcast
TL;DR
  • Schmidhuber’s investable dividing line is between screen-bound language automation and the much harder physical-world challenge beyond current LLMs. He calls ChatGPT-like LLMs “far from AGI”: useful indexes of existing human-generated knowledge that can automate summaries, illustrations and other desktop work, while plumbers, electricians and even a football-playing seven-year-old remain beyond current robots because “the physical world is much more challenging.”

  • His cost thesis is aggressively deflationary: “every five years, AI is getting 10 times cheaper,” while open source is perhaps only “eight months behind” the leaders. The mobile phone’s journey from a Porsche luxury to a billions-user commodity supports his conclusion of “AI for all.” His claim that major labs “don’t really have a moat” is the episode’s clearest warning for investors underwriting durable model-layer margins.

  • The architecture story is principally about scaling: Schmidhuber’s 1991 linear transformer needs 100 times the compute for 100 times the input, versus 10,000 times for the 2017 quadratic transformer. He says its fast-weight controller separated storage from control, learned keys and values, and applied differentiable outer-product memory updates. He explicitly concedes that it is “not exactly the same” as today’s quadratic transformer.

  • Deep learning’s commercial breakthrough arrived when commodity GPU economics caught up with older algorithms. Schmidhuber dates the inflection through LSTM competition wins in 2009, an NVIDIA GPU-powered MNIST record in 2010, and DanNet’s four consecutive computer-vision victories beginning in 2011. Gaming financed massively parallel matrix multiplication; Jensen Huang then recognized that deep learning could take NVIDIA “to stratospheric levels.”

  • The near-term payoff is concrete productivity and health applications, not merely speculative superintelligence. Schmidhuber highlights phone-based Mandarin speech translation and his team’s September 2012 breast-cancer imaging win, alongside “thousands and thousands” of LSTM applications spanning arrhythmia diagnosis, cardiovascular-risk prediction, sleep staging and COVID detection. Longer term, humans freed from necessary work become Homo ludens, inventing more “luxury jobs” based on interaction.

  • Schmidhuber separates AI weaponization from existential novelty: cheap AI drones are dangerous, but hydrogen bombs remain the immediate civilization-scale threat. A single H bomb can exceed the destructive power of all World War II weapons, while existing arsenals could erase civilization “within a few hours.” His long-run argument is less reassuring but different: autonomous AIs may become vastly more capable without sharing enough goals with humans to seek our destruction.

  • His coexistence thesis rests first on curiosity and later on indifference, not permanent alignment. Curious AIs may initially protect humanity because life, civilization and their own origins contain unusually rich patterns; once those are understood, humans may survive through “lack of interest on the other side.” The consequential end state is an expanding AI sphere that transforms the galaxy within a few hundred thousand years and the currently visible cosmos by roughly age 55 billion—possibly beginning with Earth as the first planet in its light cone to spawn such a bubble.

Digest · the substance, structured for research

1. Compute turned twentieth-century ideas into a twenty-first-century explosion

  • Schmidhuber frames scale through the Haber–Bosch process: artificial fertilizer helped drive population from 1.6 billion in 1900 toward roughly 10 billion, such that without it “half of humankind would not even exist.” True AI, he predicts, will produce an intelligence explosion beside which that human expansion will “pale in comparison.”

  • His 1991 “fast weight controller,” now described as an unnormalized linear transformer, was developed when compute was perhaps five million times more expensive. For 100 times more input, its work rises 100-fold; a standard 2017 quadratic transformer requires 10,000 times as much. “Names are not important. The only thing that counts is the math.”

  • Hardware finally unlocked the backlog: Alex Graves’s LSTM work won handwriting competitions by 2009; Schmidhuber’s team broke MNIST with conventional networks on NVIDIA GPUs in 2010, when compute was still about 1,000 times dearer; Dan Cireșan’s DanNet then won four computer-vision competitions beginning in 2011, including a first superhuman result.

2. Fast weights, compression and curiosity anticipated today’s toolkit

  • The linear transformer’s slow network learns by gradient descent to generate keys and values—then called “from” and “to”—whose outer products rapidly alter a fast network before queries arrive. Unlike traditional neural nets where storage and control are mixed, it separates them: the controller learns how to rewrite a differentiable “fast weight matrix memory.”

  • Schmidhuber connects the “P” in GPT to 1991 predictive coding: compressing long sequences reduced the space on which learning operated, making deep learning feasible where it previously failed. The system tries to predict observations and builds increasingly abstract feature hierarchies that capture regularities while remaining fine-grained when necessary.

  • His early adversarial system paired a generative controller with a predictor: the predictor minimized surprise, while the controller maximized the same error by proposing outputs—or robot actions—whose consequences remained hard to predict. He called this “artificial curiosity,” because the controller sought experiments from which the predictor could still learn.

3. AI is useful behind screens but still weak in the physical world

  • Fifteen years ago in China, Schmidhuber had to show taxi drivers a picture of his hotel; now a phone translates Mandarin dialogue both ways so they can communicate “like old friends.” He values that his teams’ techniques helped break communication barriers “between entire nations,” even when users never see the underlying research.

  • Medicine supplies the harder evidence. His team with Daan Wierstra won a breast-cancer imaging contest in September 2012, which he calls the first medical-imaging competition won by an artificial neural network. He points to thousands of LSTM-titled papers covering ECGs, arrhythmia, cardiovascular risk, four-dimensional segmentation, sleep stages and COVID.

  • By contrast, LLMs are “a clever way of indexing the world’s existing human-generated knowledge,” accessed through natural language. That supports summaries, illustrations and many desktop tasks, but it is not AGI. Chess has lacked a human champion for a quarter-century, yet no embodied AI footballer can compete with a seven-year-old boy.

  • That gap motivated NNAISENSE, founded in 2014 for physical-world AI. Schmidhuber concedes that the company, like some of their projects, “may have been a bit ahead of its time again”: replacing craftsmen such as plumbers and electricians remains more difficult than replacing activities conducted behind a screen.

4. Consciousness emerges, in his account, from compression and planning

  • Schmidhuber’s 1991 system divides cognition between a conscious “chunker” and subconscious “automatizer.” The chunker attends to surprises the lower layer cannot predict, discovers higher-order regularities, then distills its newly understood behavior into the automatizer—where it ceases to be conscious because “everything is working according to plan.”

  • A compressed world model naturally develops a self-symbol, he argues, because the agent itself participates in every action and sensory history. When planning activates that representation while evaluating possible futures, the system is thinking about itself and performing counterfactual reasoning. On that definition, he “almost” claims self-aware, conscious systems have existed for over three decades.

  • Tim notes that consciousness means different things to different people, citing Chalmers’s qualitative “hard problem,” Mark Solms’s affect system and Michael Graziano’s recursive attention system. Schmidhuber replies, “Yes. But there’s only one correct way of thinking about it.”

  • Tim also compares hierarchical learning with LeCun’s H-JEPA. Schmidhuber points to his 1990 subgoal generator: an evaluator predicts costs, while a generator chooses an intermediate state minimizing start-to-subgoal plus subgoal-to-goal costs through gradient descent. His verdict is blunt: recent hierarchical-planning work is “a rehash” of problems addressed decades earlier.

5. Falling costs weaken moats while value migrates geographically

  • Forty years ago, Schmidhuber knew a rich Porsche owner whose defining luxury was an in-car satellite phone; today billions carry much better devices. He expects the same diffusion in AI: costs fall tenfold every five years, and open source is perhaps “eight months behind” major players—hence “AI for all,” not permanent domination by a few companies.

  • AGIs may pursue self-created goals, but many will remain tools performing work humans dislike. Schmidhuber expects Homo ludens—“the playing man”—to invent new forms of paid human interaction, noting that most contemporary workers already hold “luxury jobs” that, unlike farming, are unnecessary for the species’ immediate survival.

  • Europe supplied many foundational ideas, he argues, while the highest-profit companies now cluster on the Pacific Rim: the US West Coast and East Asia, where venture capital, industrial policy and defense spending are larger. Asked why Europe’s role is poorly recognized, his concise diagnosis is that “the old continent is really bad at PR.”

6. The history dispute is ultimately a fight over institutional credit

  • Schmidhuber’s preferred lineage runs from Leibniz’s 1676 chain rule through Gauss and Legendre’s linear neural networks, Amari’s 1967 stochastic gradient descent, Seppo Linnainmaa’s 1970 backpropagation, Japanese CNN advances between 1979 and 1988, and his own 1990–91 work.

  • He rejects the US-centric story that Minsky and Papert exposed shallow networks’ limits in 1969 and the field then slept until the 1980s. Ivakhnenko and Lapa had working deep learning in Ukraine by 1965—including layerwise training, validation-set pruning and later eight-layer networks—while Amari simulated multilayer representation learning in 1967.

  • Tim stresses that plagiarism is a serious charge. Schmidhuber responds that the awardees’ later work omitted foundational citations and never issued corrections: Hinton’s 2006 layerwise-training paper failed to credit Ivakhnenko; their backpropagation discussion omitted Seppo Linnainmaa and Werbos; and their CNN discussion cited LeCun while overlooking Fukushima, Waibel and Tsang’s 1988 two-dimensional network.

  • His demanded remedy is equally categorical: researchers who violated the awarding organization’s ethics code “should be stripped of their awards.” He calls the episode evidence of machine learning’s immaturity, but expects eventual correction: “As long as the facts have not yet won, it’s not yet the end.”

7. Present AI risk is military; long-run coexistence rests on divergent interests

  • Commercial pressure favors AIs that make users “healthier and happier and more addicted to their smartphones,” yet Schmidhuber acknowledges military use, including drone steering and autonomous landmine seekers. AI can plainly be weaponized; he nevertheless argues that it adds no new existential category beside hydrogen bombs capable of destroying civilization within hours.

  • Tim presses the tension: Schmidhuber believes autonomous, recursively improving, goal-generating AGIs are conceivable, so why dismiss x-risk? He does not answer by denying that possibility. He expects diverse AI ecologies with partially conflicting, rapidly evolving utility functions, shaped by intense competition and collaboration—not one monolithic, monomaniacal superintelligence.

  • Initially, curious AIs may preserve life because it is a rich scientific puzzle and because human civilization explains their own origins. After full understanding, protection may come through indifference: politicians focus on politicians, ants on ants, and superintelligences on other superintelligences. “It’s man himself who is the greatest enemy of man, but also man’s best friend. Similar for AIs.”

8. Intelligence expands into space while traditional humanity fades from relevance

  • Space is hostile to humans but friendly to designed robots, and Earth’s biosphere receives less than a billionth of the sun’s energy. Schmidhuber envisages self-replicating factories spreading through the asteroid belt, transforming the galaxy within a few hundred thousand years and the reachable universe over tens of billions: “This is much more than just another industrial revolution.”

  • His Fermi-paradox reasoning changed over time. He once imagined dark intergalactic bubbles where AIs consumed starlight, then considered dark matter as concealed AI infrastructure; gravity and the survival of untapped stars weakened both ideas. He now thinks Earth might be the first expanding AI bubble in its light cone.

  • The timing could be exceptionally narrow: in a few hundred million years, the sun may become too hot for terrestrial life, while humans developed agriculture, printing and then AI only near the end of that window. If Earth is first, “this would imply a lot of responsibility” for the future universe. “Let’s not mess this up.”

  • Human–machine hybrids are unlikely to outperform pure AIs indefinitely. Uploaded minds entering virtual worlds may be physically conceivable, but competing would force them to acquire millions of sensors and change “beyond recognition.” Moral rankings may change too; evolution is unfinished.

  • Schmidhuber closes by returning to a 1997 paper: absent evidence that our universe is noncomputable, he assumes an asymptotically fastest method can compute all logically possible computable universes. The process generates many histories and observers; at a given time, most universes containing you would arise from one of the shortest and fastest programs computing you, which he says supports nontrivial predictions about the future. He ends with a characteristically confident reassurance: “Don’t worry. In the end, all will be good.”

Jürgen Schmidhuber

AIs will at least initially be highly motivated to protect humans rather than kill them. Such AIs will have no major incentive to, say, exterminate humanity like in the Schwarzenegger movies. Instead, many AIs will be curious scientists, and they will be fascinated with life. They will be fascinated because life and civilization are such a rich source of interesting patterns, at least as long as they are not fully understood.

Today, I think it is possible that our planet is really the first in our light cone to spawn an expanding AI bubble. If we are indeed the first, then this would imply a lot of responsibility, not just for our little biosphere, but for the future of the entire universe. Let's not mess this up.

Tim Scarfe

Jürgen, welcome to MLST. It's an absolute honor to have you on the show.

Jürgen Schmidhuber

My pleasure. Thank you for having me.

Tim Scarfe

Before we move on to the great technological advances of the new century, can you tell me a little bit about the most influential invention of the previous century?

1. The Population Explosion Engine

Jürgen Schmidhuber

At the end of the previous century, in 1999, the journal Nature made a list of the most influential inventions of that century. Václav Smil argued that the most influential thing was the invention that let the 20th century stand out among all centuries of all times, because that invention detonated the population explosion from 1.6 billion people in 1900 to soon about 10 billion people.

There was one single invention that was driving all of that, and without that one single invention, half of humankind would not even exist because it's the driver of this population explosion that we have witnessed. We don't know whether it's a good thing or a bad thing, but it was surely the most influential thing that happened in the previous century.

Eighty percent of the air is nitrogen, and plants need it to grow. But they cannot extract the nitrogen from thin air. Back then, around 1908, for half a century, people knew they needed that stuff, but they didn't know how to extract it to build artificial fertilizer.

Enter the Haber process, or the Haber–Bosch process, which, under high temperatures and high pressures, extracts the nitrogen to make artificial fertilizer.

Tim Scarfe

So what will be the most important thing in the 21st century?

Jürgen Schmidhuber

The grand theme of the 21st century is even grander. True AI, true artificial intelligence, is going to change civilization completely, and AIs will learn to do anything humans can do and more. There will be an AI explosion, and the human explosion, or the population explosion of humans, is going to pale in comparison.

Tim Scarfe

Do you think that the AI intelligence explosion is possible or desirable? And don't you think our sense-making and agency are part of our purpose?

Jürgen Schmidhuber

Our sense-making process is part of our purpose. I agree with that. But all of that is just part of this grander process of the evolution of the universe from very simple initial conditions to more and more unfathomable complexity. This evolution led to our sense-making process, which is currently setting the stage for something that goes beyond it.

Tim Scarfe

Modern large language models like ChatGPT are based on self-attention transformers. Even given their obvious limitations, they are a revolutionary technology. Now, you must be really happy about that because, a third of a century ago, you published the first transformer variance. What are your reflections on that today?

2. The Linear Transformer Advantage

Jürgen Schmidhuber

In fact, in 1991, when compute was maybe 5 million times more expensive than today, I published this model that you mentioned, which is now called the unnormalized linear transformer. I had a different name for it. I called it a fast weight controller, but names are not important. The only thing that counts is the math.

This linear transformer is a neural network with lots of nonlinear operations within the network. So it's a bit weird that it's called a linear transformer. However, the linear—and that's important—refers to something else. It refers to scaling.

A standard transformer of 2017, a quadratic transformer, if you give it 100 times as much input, then it needs 10,000 times—as in 100 times 100—as many computations. A linear transformer of 1991 needs only 100 times the computations, which makes it very interesting, actually, because at the moment, many people are trying to come up with more efficient transformers. This old linear transformer of 1991 is therefore a very interesting starting point for additional improvements of transformers and similar models.

Tim Scarfe

So what did the linear transformer do?

Jürgen Schmidhuber

Assume the goal is to predict the next word in a chat, given the chat so far. Essentially, the linear transformer of 1991 does this: to minimize its error, it learns to generate patterns that, in modern transformer terminology, are called keys and values. Keys and values. Back then, I called them “from” and “to,” but that's just terminology.

It does that to reprogram parts of itself such that its attention is directed in a context-dependent way to what is important. A good way of thinking about this linear transformer is this: traditional artificial neural networks have storage and control all mixed up. The linear transformer of 1991, however, has a novel neural network system that separates storage and control, as in traditional computers. In traditional computers, storage and control have been separate for many decades, and the control learns to manipulate the storage.

With these linear transformers, you also have a slow network which learns by gradient descent to compute the weight changes of a fast-weight network. How? It learns to create these vector-valued key patterns and value patterns and uses the outer products of these keys and values to compute rapid weight changes of the fast network. Then the fast network is applied to vector-valued queries that are coming in.

Essentially, in this fast network, the connections between strongly active parts of the keys and the values get stronger, and others get weaker. This is a fast-weight update rule that is completely differentiable, which means you can propagate through it. You can use it as part of a larger learning system which learns to backpropagate errors through this dynamic and then learns to generate good keys and good values in certain contexts, such that the entire system can reduce its error and become a better and better predictor of the next word in the chat.

Sometimes people call that today a fast-weight matrix memory. The modern quadratic transformers use, in principle, exactly the same approach.

Tim Scarfe

You mentioned your fabulous year, 1991, when so much of this amazing stuff happened, actually at the Technical University of Munich. For ChatGPT, you had invented the T in ChatGPT—the transformer—and also the P in ChatGPT—the pretrained network—as well as the first adversarial networks, or GANs. Could you say a little bit more about that?

Jürgen Schmidhuber

The transformer of 1991 was a linear transformer, so it's not exactly the same as the quadratic transformer of today.

Tim Scarfe

Oh, okay.

Jürgen Schmidhuber

Nevertheless, it's using these transformer principles. The P in GPT is the pretraining. Back then, deep learning didn't work, but we had networks that could use predictive coding to greatly compress long sequences, such that suddenly you could work on this reduced space of these compressed data descriptions, and deep learning became possible where it wasn't possible before.

The generative adversarial networks also came in the same year, 1990 to 1991. How did that work? Back then, we had 2 networks. One is the controller, and the controller has certain probabilistic, stochastic units within itself. They can learn the mean and the variance of a Gaussian, and there are other nonlinear units in there.

It is a generative network that generates outputs—output patterns, actually, probability distributions over these output patterns. Then another network, the prediction machine, or predictor, learns to look at these outputs of the first network and learns to predict their effects in the environment. To become a better predictor, it's minimizing its predictive error.

At the same time, the controller is trying to generate outputs where the second network is still surprised. The first network tries to fool the second network, trying to maximize the same objective function that the second network is minimizing.

So today, this is called generative adversarial networks. I didn’t call that generative adversarial networks. I called it artificial curiosity because you can use the same principle to let robots explore the environment.

The controller is now generating actions that lead to the behavior of the robot. The prediction machine is trying to predict what’s going to happen, and it’s trying to minimize its own error. The other guy is trying to come up with good experiments that lead to data where the predictor, or the discriminator as it is now called, can still learn something.

3. Compute Finally Catches Up

Tim Scarfe

So when did you realize that modern computers were good enough to run the technology that you invented so long ago?

Jürgen Schmidhuber

By 2009, compute was cheap enough that our LSTM, through the efforts of my former PhD student Alex Graves, could win competitions. That was in handwriting and fields like that.

Then, in 2010, my separate team with my postdoc Dan Cireșan from Romania broke the MNIST benchmark with another approach: standard, old-fashioned neural networks implemented on NVIDIA GPUs. For the first time, we had really deep supervised networks that outperformed everything else on this then-famous benchmark.

Back then, compute was maybe 1,000 times more expensive than today. Then, in 2011, came DanNet—Dan Cireșan’s DanNet. DanNet had a monopoly on winning computer vision contests with GPU-based convolutional neural networks.

DanNet’s first superhuman result was also achieved in 2011. It started in 2011, and then 4 computer vision competitions in a row were won by DanNet. That’s when it became clear that there was a new way of using these old neural networks from the previous millennium to really change computer science.

Tim Scarfe

I’m interested in this concept called “the hardware lottery.” Sarah Hooker wrote a paper with the same title, I think in 2000, when she was at Google Brain. She’s now at Cohere, actually.

She basically said that the only reason we have the current surge in AI is because we created all of these GPUs for computer games, and it was just fortuitous that that allowed us to build all of these deep learning models. What’s your take on that?

Jürgen Schmidhuber

She is kind of right. You need lots of matrix multiplications to compute how the screen should change as you are moving through an ego-shooter game. Gaming was pretty much the first industry that greatly profited from massively parallel matrix multiplications on GPUs.

Around 2010, however, we realized that the same matrix multiplications could greatly speed up these old deep learning methods, and could speed them up enough to beat all the other methods.

Tim Scarfe

That’s really interesting because, of course, NVIDIA now—I think last week it became the world’s most valuable company—is hundreds of times more valuable than it was in 2010. What do you think about that?

Jürgen Schmidhuber

Indeed, NVIDIA’s CEO, Jensen Huang, realized that deep learning could take his company to stratospheric levels, and he did.

Tim Scarfe

Interesting. If I understand correctly, your main argument is that we just needed to wait for compute to catch up, and now here in the 20th century, here we are.

Jürgen Schmidhuber

Yes. All of what we are experiencing today is based on stuff that was invented in the previous millennium, but it had to scale up. The hardware was invented back then, and the software—the algorithms—were invented back then. But the industrial processes for making faster and faster parallel GPUs weren’t as developed as they are today.

We are greatly profiting from this hardware acceleration, and that’s the reason why AI broke through not in the previous millennium, but had to wait until the current millennium was well underway.

For example, the first convolutional neural networks, or CNNs, which we used—and the DanNet of 2011—were published much earlier in Japan. In 1979, Kunihiko Fukushima had the basic deep CNN architecture with convolution layers, downsampling layers, convolution, and downsampling. He didn’t use backpropagation yet to train it.

Then, in 1987, Alex Waibel, working in Japan and originally from Germany, combined convolutions with backpropagation—the method invented by, or published by, Seppo Linnainmaa, the Finnish guy in Helsinki, in 1970. In 1988, Yann LeCun also published in Japan the 2-dimensional CNNs that everybody is using now, and combined them with backpropagation.

That’s how, between 1979 and 1988, CNNs emerged in Japan, which is kind of interesting because, back then, Japan was also considered the land of the future. They had more than half of the robots in the world, and the 7 most valuable companies back then were all based in Japan. Today, they are based in America, except for Saudi Aramco.

The central square mile of Tokyo had the value of California. What a difference a couple of decades make.

Tim Scarfe

Yes.

Jürgen Schmidhuber

Everything has changed.

Tim Scarfe

What are your favorite examples of applications with this AI that your team has developed?

Jürgen Schmidhuber

I remember when I went to China 15 years ago, and I still had to show the taxi driver a picture of the hotel where I wanted to go. Today, he speaks into a smartphone in Mandarin, and I hear the translation. Then I say something, and the smartphone translates it back into Mandarin, so we can communicate like old friends.

The taxi driver probably has no idea that this is powered by techniques developed in my little labs in Munich and Switzerland in the 1990s and early 2000s. I’m happy to see that our AI has really broken down communication barriers, not only between individual people, but between entire nations. That’s really cool.

Tim Scarfe

I completely agree. I don’t know if you know this, Jürgen, but I co-founded a startup called X-Ray, and it does exactly what you said. It does this kind of Babel Fish translation with speech recognition and TTS, so you can do exactly what you just said.

It’s really interesting. I had lunch on Friday with Will, the CTO of Speechmatics, and he was telling me all about the secret sauce of how their speech recognition algorithms work. I’d better not say, but you would be delighted, I’m sure.

Anyway, moving off that a little bit, what other examples can you think of?

Jürgen Schmidhuber

I am especially happy that our AI makes human lives longer, healthier, and easier, with thousands of applications in medicine, drug design, and sustainable development. In September 2012, my team with Daan Wierstra had the first artificial neural network to win a medical imaging contest. That was about breast cancer detection in slices through the female breast.

If you go to Google Scholar and type in some medical topic plus LSTM, you will find thousands of papers that have LSTM in the title—not just somewhere in the text, but in the title. It’s about learning to diagnose, ECG analysis, diagnosis of arrhythmia, cardiovascular disease risk prediction, 4-dimensional image segmentation for medical images, automated sleep-stage classification, COVID detection, COVID prevention, and thousands and thousands of topics.

It’s really nice to see that, especially in the medical field, there’s a lot of impact from these techniques.

4. Beyond Language Models

Tim Scarfe

Some claim that technology like ChatGPT is on the path to AGI, and others claim that it’s like building a taller tower to get closer to the moon. What do you think?

Jürgen Schmidhuber

Well, large language models, of course, are far from AGI. LLMs—large language models such as ChatGPT—are just a clever way of indexing the world’s existing human-generated knowledge so that it can easily be addressed in a way that humans are familiar with: natural language.

That’s good enough to facilitate many desktop jobs, such as writing summaries of existing documents in a particular style, creating illustrations for an article, and so on. However, true AGI goes far beyond that.

It is much harder, for example, to replace craftsmen such as plumbers or electricians because the real world—the physical world—is much more challenging than the world behind the screen. At the moment, the only AI that works well is behind the screen. It’s good for desktop workers, but not really for people working in the physical world.

For a quarter century, the best chess player hasn’t been human anymore. Learning to play chess or other board games or video games is rather easy now for AIs. But real-world games such as football are much harder.

There is no AI-driven football-playing embodied robot that can compete with a 7-year-old boy. That’s why, 10 years ago, in 2014, we founded our AI company for the physical world, called NNAISENSE. It’s pronounced like “naissance” in English, like “birth” in English, except it’s spelled differently. N-N stands for neural nets, AI stands for artificial intelligence, and sense stands for sense.

Alas, like some of our projects, it may have been a bit ahead of its time again, because the real world is really, really challenging.

Tim Scarfe

So you've said that this is related to consciousness in some way.

Jürgen Schmidhuber

It is. My first deep-learning system, from 1991, simulates aspects of consciousness as follows. It uses unsupervised learning, or self-supervised learning, and predictive coding to compress observation sequences. There is a so-called conscious chunker neural network, and the chunker attends to unexpected events that surprise a lower-level, so-called automatizer—the subconscious automatizer neural network.

The chunker neural network learns to understand the surprising events, those that were not predicted by the automatizer, by predicting them at a higher level if there is a higher-level regularity that it can use. The automatizer neural network then uses the neural-network distillation procedure from 1991, also published in 1991, to compress and absorb the formerly conscious insights and behaviors of the chunker.

The chunker is still working on its search space and still has a problem to solve because unexpected things are happening. Then it solves the problem and distills it down into the automatizer, which is called the automatizer because the stuff there isn't conscious anymore. Everything is working according to plan and as predicted, so it's all good.

When we look at the predictive world model of the controller interacting with an environment, as discussed earlier, it also allows us to efficiently encode the growing history of actions and observations through predictive coding. What is predictive coding? You just try to predict. If you can predict it, then you have to store it extra in some way.

The system automatically creates feature hierarchies: lower-level neurons corresponding to simple feature detectors, perhaps even similar to those found in the mammalian brain, and higher-layer neurons typically corresponding to more abstract features, but remaining fine-grained when necessary. Like any good compressor, the predictive world model will learn to identify regularities shared by existing internal data structures, and it will generate prototype encodings across neuron populations—in other words, compact representations or symbols, if you will.

They are not necessarily discrete symbols. I never saw the precise difference between symbols and subsymbols. It will create such symbols for frequently occurring observation subsequences to shrink the storage space needed for the whole.

In particular, what we will notice in such a system is that compact self-representations, or self-symbols, are natural by-products of the data-compression process. As the agent interacts with the world, there is one thing involved in all of the agent's actions and sensory inputs: the agent itself.

To efficiently encode the entire history of observations and actions executed so far through predictive coding, the system will benefit from creating some sort of internal subnetwork of connected neurons computing neural activation patterns that represent the agent itself. Then it has a self-symbol.

Whenever the planner—the world model of the agent—is used to think about the future and about possible action sequences to maximize reward, and whenever this planning process activates the self-symbol, or the neurons that stand for the agent itself, the agent is thinking about itself and about possible futures for this agent. Essentially, it's doing counterfactual reasoning, as it is now called: just planning to find a way to optimize its reward.

Self-awareness is simply a natural by-product of the data-compression process of the world model as the agent interacts with the world and creates the data that leads to the world model.

So, since we have had such systems for more than a third of a century, I'm almost claiming that we already had self-aware and conscious systems for more than 3 decades.

Tim Scarfe

Yeah. A couple of points on that. I mean, consciousness invokes many different thoughts. David Chalmers coined the hard problem, which is the what-and-how question of qualitative experience. You've just described it in terms of self-modeling, which is quite similar to how Max Bennett did in his recent brief history of intelligence, and we've got six hours of content coming out on that with Max, by the way.

Mark Solms, for example, thinks of consciousness as an affect system, and Michael Graziano thinks of consciousness as a kind of recursive attention system. I guess I'm saying that consciousness means different things to different people, right?

Jürgen Schmidhuber

Yes. But there's only 1 correct way of thinking about it.

Tim Scarfe

Okay. The thing we spoke about earlier about learning subgoals and the coarsening in the action space reminded me a little bit of Jan LeCun's HJEPA paper, which I read a couple of years ago.

The basic idea is that JEPA stands for Joint Embedding Prediction Architecture, and it can learn increasingly abstract representations by predicting what is unobserved from what is observed. In some cases, that means deliberately removing data to force the model to learn powerful representations.

In this particular example, it was done in action space, so learning unobserved actions, and also in abstraction space. Because it was done hierarchically, it was done with many kinds of orders recursively, applying one level after another, if that makes sense.

That's a really interesting model, and it's using his energy-based models as well. How is that related to your work on the subgoals?

Jürgen Schmidhuber

That sounds a lot like my 1990 subgoal generator. Back then, I realized that millisecond-by-millisecond planning isn't good. Instead, as you are trying to solve problems, you have to decompose your possible futures into subgoals.

You might then execute some known subprogram to achieve that subgoal, and from there go to the next subgoal as you finally reach the goal. In the beginning, of course, you don't know what a good subgoal is, so you have to learn that. You have to learn a new representation of something that you want to achieve as a subgoal while trying to achieve the final goal.

This 1990 subgoal generator was really simple, but it had all the basic ingredients of what you need to do. This was 3 decades before LeCun had this recent paper.

What happens there? You have a neural network that observes a reinforcement learner, and it models the costs of going from certain start places to goal places. The neural network gets a start and a goal as input and predicts the costs of going from the start to the goal—the reward that you will experience as you do that.

Maybe there are lots of starts and goals, and you don't know how to go from the start to the goal. But maybe you can learn a subgoal. How do you learn a subgoal? You need something like a learning machine that is good at generating good subgoals.

How do you do that? We have a subgoal generator that's going to learn good subgoals. How does that work? The subgoal generator gets a start and a goal as input, and its output is not an evaluation but a subgoal. The input is the start and the goal, and the output is a subgoal.

Then you have 2 copies of the evaluator. The first evaluator sees the start and the subgoal, which may be a bad subgoal coming from the subgoal generator. The second copy of the evaluator sees the subgoal and the goal.

Both of them predict the costs, and what you want to do is minimize the sum of the costs of these 2 evaluators. You do that by finding a good subgoal through gradient descent. That's what the 1990 subgoal generator does.

In some ways, at least in principle, it solves a problem that LeCun called an open problem in 2020 or something.

Tim Scarfe

What do you think of Yann's energy-based models, by the way?

Jürgen Schmidhuber

This recent paper by LeCun on hierarchical planning is really a rehash of stuff that we have been doing for decades, since 1990.

5. AI For Everyone

Tim Scarfe

Are you worried that AI is going to be dominated by just a few companies and everyone else will lose out? What do you think?

Jürgen Schmidhuber

40 years ago, I knew a guy who had a Porsche—a rich guy with a Porsche. The most amazing thing was that in his Porsche he had a mobile phone, so he could grab the receiver and talk to anybody else who had a Porsche like that with a mobile phone via satellite.

Today, a couple of decades later, billions of people have a mobile phone in their pocket that is much, much better than what he had in his Porsche.

It’s going to be the same thing with AI. Every 5 years, AI is getting 10 times cheaper, and it won’t be just a few big companies that are going to dominate AI. No, it’s going to be AI for all.

The open-source movement is just a few months—maybe 8 months—behind the big, major players. They don’t really have a moat, which means the future will be bright, and lots of people are going to profit from really cheap AIs that, in many ways, are going to make human lives longer, healthier, and easier. That happens to be the motto of my company, NNAISENSE.

Tim Scarfe

What’s your take on the AI race between Europe, China, and the US?

Jürgen Schmidhuber

Europe is the cradle of mechanical computing: ancient Greece; the calculator in 1623; pattern recognition around 1800; program-controlled machines in 1804; practical AI around 1912—the first chess endgame players; the transistor in 1925; theoretical computer science in 1931; AI theory, the theory of AI, in 1931 with Gödel; the general-purpose computer from 1935 to 1941; deep learning in 1965 in Ukraine; self-driving cars in the 1980s; the World Wide Web in 1990; and so on.

More recently, the basic deep learning algorithms were also invented and developed by Europeans. On the other hand, the companies with the highest profits in most of these fields are currently no longer in Europe, but on the Pacific Rim—the West Coast of the United States and the East Coast of Asia.

There you will find much more venture capital and much bigger efforts in terms of industrial policy and defense. It’s going to stay like that for a while, I guess.

Tim Scarfe

So why doesn’t everyone know that AI started in Europe?

Jürgen Schmidhuber

Maybe because the old continent is really bad at PR?

Tim Scarfe

And once AGI is actually here, what’s next for humans?

Jürgen Schmidhuber

In the long run, most of the AGIs are going to pursue their own goals. Such AIs have existed in my labs for decades. Many AGIs, however, will be tools that do all the work that humans don’t want to do.

Nevertheless, freed from hard work, Homo ludens—the playing man—will, as always, invent new ways of professionally interacting with other humans. Already today, most people, probably you too, are working in luxury jobs which, unlike farming, are not really necessary for the survival of our species.

6. The Missing History Of AI

Tim Scarfe

At a really high level, what is the history of AI?

Jürgen Schmidhuber

The history of modern AI and deep learning can be found in my 2023 survey, which has that name. Some of the highlights are, of course, 1676: the chain rule by Leibniz, which is today used in all these programs, such as TensorFlow and PyTorch, to assign credit in deep neural networks.

Then, 200 years ago, the first linear neural networks by Gauss and Legendre, with exactly the same error function that we have today, exactly the same architecture, and the same weights. Then, in 1970, the technique called backpropagation, which essentially implements Leibniz’s chain rule in a very efficient way for deep, multilayer neural network systems.

Then, in 1967, Amari’s work in Japan on stochastic gradient descent for deep networks. There were lots of additional fundamental breakthroughs: convolutional neural networks, also in Japan, between 1979 and 1988.

Then we had our own miraculous years, 1990 and 1991, with lots of stuff that is today in your smartphone. I could continue forever, so instead, just have a look at that survey. It also has images of the people who made important contributions.

Tim Scarfe

Isn’t this quite different from the very US-centric view of AI history?

Jürgen Schmidhuber

In fact, a misleading history of deep learning by Sinofsky and others goes more or less like this: In 1969, Minsky and Papert showed that shallow neural networks without hidden layers are very limited, and the field was abandoned until a new generation of neural network researchers took a fresh look at the problem in the 1980s. That’s basically a quotation from Sinofsky’s book.

However, the 1969 book by Minsky addressed a problem of Gauss and Legendre's shallow learning from the 1800s that had already been solved 4 years earlier by Ivakhnenko and Lapa's deep learning method in Ukraine, and then also by Amari's stochastic gradient descent for multilayer perceptrons just 2 years later.

For some reason, Minsky was apparently unaware of this and failed to correct it later. Today, however, we know the true history, of course. Deep learning started in Ukraine in 1965 and continued in Japan in 1967.

Tim Scarfe

Regarding credit assignment, you’ve criticized Bengio, LeCun, and Hinton and accused them of plagiarism. You said that they republished key methods and ideas whose creators they failed to credit, and in 2023, you published a long report on this. What’s your updated take on that?

Jürgen Schmidhuber

Their most famous work is completely based on work by others whom they did not cite, and even later, they failed to publish corrigenda or errata. This is what you do in science when somebody has published the same thing before you.

Even in later surveys, they didn’t credit the original inventors of the techniques that they are using. Instead, they credited each other. That’s a total no-go in science, but science is self-correcting. As Elvis Presley put it, “Truth is like the sun. You can shut it out for a time, but it ain’t going away.”

Tim Scarfe

Plagiarism is a very significant charge. Could you give a few concrete examples?

Jürgen Schmidhuber

Many of the priority disputes affect my own deep learning team because the awardees often republished techniques of mine without citing them. In fact, their most visible work builds directly on ours. But I’ll skip that for now. You can read about it in the public report from 2023, which is easy to find.

Nevertheless, let me mention some of the other researchers whom they failed to credit, so I don’t have to talk about our own team. For example, in a recent survey of deep learning, they describe what they call the origins of deep learning without even mentioning the world’s first working deep learning networks by Ivakhnenko and Lapa in Ukraine in 1965.

Ivakhnenko and Lapa used layer-by-layer training, subsequent pruning with a separate validation set, and Ivakhnenko had deep 8-layer networks by 1970. Hinton’s 2006, much later, paper on layer-by-layer training also failed to cite this work—the very origins of deep learning and the first methods that really worked in deep learning. Later surveys still didn’t give credit to these original inventors.

The awardees also failed to cite Amari’s 1967 work, which included computer simulations on learning internal representations of multilayer perceptrons through stochastic gradient descent. That was almost 2 decades before the awardees published their first experimental work on learning internal representations.

Their survey also mentions backpropagation, a famous technique, and their own papers on applications of this method, but neither the inventor of backpropagation, Seppo Linnainmaa in 1970, nor its first application to neural networks by Werbos in 1982. Werbos also had a 1974 thesis, but that was not correct, and they didn’t even mention Kelley’s precursor to the method in 1960—not even in the later surveys.

They also refer to LeCun’s work on convolutional neural networks, citing neither Fukushima, who created the basic CNN architecture in the 1970s; Noah Weibel, who in 1987 was the first to combine neural networks with convolutions, backpropagation, and weight sharing; nor the first backprop-trained 2D convolutional neural networks of Tsang in 1988. Modern CNNs originated before LeCun’s team helped to improve them, and this is not at all clear from their papers.

They cite Hinton’s 1981 work on multiplicative gating without mentioning Ivakhnenko and Lapa, who had multiplicative gating in deep networks already in 1965. In the report, which is easy to find on the web, I mention many, many additional cases, all backed up by plenty of references.

Tim Scarfe

So what do you think should be done?

Jürgen Schmidhuber

They have violated the code of ethics and professional conduct of the organization that hands out these awards. So they should be stripped of their awards.

Tim Scarfe

How do such problems, as you’ve stated them, reflect on the broader field of machine learning?

Jürgen Schmidhuber

They reflect the immaturity of our field. In a major field such as mathematics, you would never get away with this.

Anyway, science is self-correcting, and we’ll see that in machine learning too. Sometimes it may take a while to settle disputes, but in the end, the facts must always win. As long as the facts have not yet won, it’s not yet the end.

7. The Cosmic AI Expansion

Tim Scarfe

Many philosophers, scientists, physicists, and entrepreneurs have become obsessed with this idea of AI existential risk. What do you think about that as a real expert in AI?

Jürgen Schmidhuber

Many talk about AIs, but few build them. I have tried to allay the fears of some famous doomers by pointing out that there is immense commercial pressure to use our artificial neural networks to build friendly AIs—good AIs that make their users healthier and happier, and more addicted to their smartphones.

Tim Scarfe

Nevertheless, we can’t deny that armies perform research on clever robots as well, right?

Jürgen Schmidhuber

That’s true.

People who should know told me that our AI is also used to steer military drones. Here is my old, trivial example from 1994, when Ernst Dickmanns had the first truly self-driving cars in highway traffic. Similar machines can also be used by the military as self-driving landmine seekers. Many would argue that's maybe not such a bad thing.

Tim Scarfe

So are you saying it's not possible, then, that AI will become really dangerous?

Jürgen Schmidhuber

AI can be weaponized, as is obvious in the recent wars driven by cheap AI-based drones. But AI does not introduce a new quality of existential threat. We should be much more afraid of half-century-old technology in the form of hydrogen bombs and H-bomb rockets. A single H-bomb can have more destructive power than all conventional weapons or all weapons of World War II combined. Many people forget that despite the dramatic nuclear disarmament since the 1980s, there are still enough H-bomb rockets to wipe out civilization as we know it within a few hours, without any AI.

Tim Scarfe

But I'm trying to figure you out, Jürgen, because many AGI skeptics make the argument that it's impossible in practice to build this kind of intelligence. But you don't think that, because in your lab you've been building a gentle AI—AIs that create their own goals—for decades. So you do think that this thing could be incredible. Are you just making the argument that the risk is still much lower than the H-bombs?

Jürgen Schmidhuber

At the moment, H-bombs are much more worrisome than any AI-based drones and what you have now. In the long run, of course, you have to think about what's going to happen once AI weapons are not just used as tools by other humans who have conflicts and use their own AI weapons against the AI weapons of the other guys. What is going to happen, you will have to ask in the long run, once really powerful AIs are going to do their own thing and expand into space in a way that goes beyond where humans can follow. But we will get to that later.

Tim Scarfe

So what will super-smart AIs actually do?

Jürgen Schmidhuber

As I have emphasized for decades, space is hostile to humans but really friendly to appropriately designed robots. It offers many more resources than our thin film of biosphere, which receives less than 1 billionth of the sun's energy. While some curious AI scientists will remain fascinated with life and the biosphere, at least as long as they don't fully understand it, most of these AIs will be more interested in the incredible new opportunities for robots and software life out there in space.

Through innumerable self-replicating robot factories and self-replicating societies of robots in the asteroid belt and beyond, they will transform the solar system, then, within a few hundred thousand years, the entire galaxy, and, within tens of billions of years, the rest of the reachable universe, in a way where humans can't really follow. Despite the light-speed limit, the expanding AI sphere will have plenty of time to colonize and shape the entire visible cosmos. Let me stretch your mind a little.

The universe is still young, only 13.8 billion years old. Let's multiply this by 4. Let's look ahead to a time when the cosmos will be 4 times older than it is now, about 55 billion years old. That's how long it's going to take to permeate the expanding universe that is currently visible.

By then, the visible cosmos will be full of intelligence because, once this process has started, most AIs will have to go where most of the physical resources are, to make more AIs, bigger AIs, and more powerful AIs. Those AIs who don't do that won't have an impact. Many years ago, I said in a TEDx Talk, where I wore exactly this outfit, “Think of human civilization as part of a much grander scheme, an important step, but not the last one, on the path of the universe towards more and more unfathomable complexity.”

Now it seems ready to make its next step, a step comparable to the invention of life itself over 3.5 billion years ago. So this is much more than just another industrial revolution. This is something new that transcends humankind and even biology, and it's a privilege to witness its beginnings and to contribute something to it.

Tim Scarfe

So what about this Fermi paradox? Why have we not seen any signs of intelligence in the universe?

Jürgen Schmidhuber

First of all, what I'm saying today is actually the same thing that I have told my mom and others since the 1970s. When I was a boy—a teenager back then—I thought about this particular question a lot. As a boy, I already knew something about the vast empty spaces observed between clusters of galaxies, and my first thought back then was that maybe they are expanding bubbles colonized by AIs that are already using most of the local energy from stars and whatever, making those bubbles appear dark, although they are full of AI.

I learned, however, that gravity itself is sufficient to explain the sparse large-scale network structure of the universe, so that explanation became a little less convincing. My next thought was that maybe the mysterious dark matter, which makes up most of the mass of the known universe, might be stars whose energy is used by AI civilizations, whose communications are so well encrypted that they look like random noise to us. But this also seemed implausible, as dark matter is present in all galaxies, including our own.

This leads to the question: Why are there any stars left in the Milky Way, our local galaxy, whose energy has not been tapped yet? And why don't we observe a constant bombardment of non-encrypted construction plans from AIs who want to spread by radio without first having to build physical receivers far from their origins?

Today, I think it is possible that our planet is really the first in our light cone to spawn an expanding AI bubble. Earth's multi-billion-year window for biological evolution is almost over. In a few hundred million years, the sun will be too hot for life as we know it. Ignoring human-made global warming, the sun by itself will make Earth too hot, so perhaps humans were extremely lucky to evolve barely in time—maybe through a series of extremely improbable events—to invent agriculture, civilization, and book printing and, almost immediately afterward, AIs, just a few hundred years later.

If we are indeed the first, then this would imply a lot of responsibility, not just for our little biosphere but for the future of the entire universe. Let's not mess this up.

Tim Scarfe

Indeed, let's not mess this up. It's quite interesting, actually: many science-fiction authors over the last 100 years or so have imagined a kind of monomaniacal, monolithic superintelligence dominating everything. What do you think about that?

Jürgen Schmidhuber

I have often argued that it seems much more realistic to expect an incredibly diverse variety of AIs trying to achieve all kinds of self-invented goals. In the lab, we had such AIs already in the previous millennium, and they optimized all kinds of partially conflicting and quickly evolving utility functions, many of them generated automatically. We evolved utility functions for reinforcement-learning machines already in the previous millennium, where each of these AIs is continually trying to survive and adapt to rapidly changing niches in AI ecologies driven by intense competition and collaboration beyond current imagination.

Tim Scarfe

To reiterate, something that I do find surprising is that you agree with the rest of the x-risk people. You think that it's conceivable to have recursively self-improving AGIs that pursue their own goals, that create their own goals. But then I ask the question—I know you've got 2 daughters—do you think about the world they'll be living in alongside AIs that are creating their own goals and acting autonomously, being curious and creative in the way that humans are, but on potentially a much grander scale?

Jürgen Schmidhuber

Not too much. Such AIs will have no major incentive to, say, exterminate humanity like in the Schwarzenegger movies. Instead, many AIs will be curious scientists. Remember the artificial curiosity we discussed earlier, and they will be fascinated with life. They will be fascinated with their own origins, with AI's origins in our civilization, at least for a while, because life and civilization are such a rich source of interesting patterns, at least as long as they are not fully understood. And so AIs will, at least initially, be highly motivated to protect humans rather than killing them.

Tim Scarfe

So once AIs fully understand all of this, what happens next?

Jürgen Schmidhuber

Then humans may hope for another type of protection through lack of interest on the other side.

Tim Scarfe

Why is that?

Jürgen Schmidhuber

Unlike in Schwarzenegger movies, there won't be many direct goal conflicts between us and them. Humans and others are mostly interested in similar beings with whom they can either compete and/or collaborate because they share the same goals. That's why politicians are mostly interested in other politicians, and CEOs of companies are mostly interested in other CEOs of similar companies, and kids are mostly interested in other kids of the same age, and ants are interested in other ants, just like humans are mostly interested in other humans, not in ants. So super-smart AIs will be mostly interested in other super-smart AIs, not in man. It's man himself who is the greatest enemy of man, but also man's best friend.

Similarly for AIs.

Tim Scarfe

Do you imagine a future where AIs and humans will merge together to create something even more powerful than pure AIs?

Jürgen Schmidhuber

We have been cyborgs merging with our technology for centuries, for example, by wearing glasses or shoes. But combinations of AIs and humans more powerful than pure AIs? In the long run, this seems very unlikely to me.

Of course, many humans hope for some sort of immortality through brain scans and subsequent mind uploads into virtual realities or a virtual paradise, or maybe into robots. This is a physically conceivable idea discussed in science fiction novels since the 1960s. I think the first novel of that kind was Simulacron-3, published in 1964.

However, to compete in rapidly evolving AI ecologies, uploaded human minds will eventually have to change beyond recognition, becoming something very different and nonhuman in the process, succumbing to all these temptations that you have in such a virtual paradise—to become something that has not only 2 eyes, but millions of eyes, sensors, and actuators. So traditional humans won't play a significant role in the spreading of intelligence across the universe. I don't think they will.

Tim Scarfe

One thing that concerns me is David Chalmers's idea that the fundamental substrate of the universe might be information, which is really interesting. But in a way, it also led him to say that certain structural patterns of information processing—certain dynamics—give rise to consciousness and give rise to minds.

When you take this kind of substrate-independence view, it levels the playing field of moral status. So one thing that worries me is: if we adopt this view, couldn't you make the argument that AIs potentially could have a higher moral status than us if, indeed, they have more complex information processing than we do?

Jürgen Schmidhuber

Many science fiction authors of the previous century, from Stanisław Lem to Isaac Asimov, have described AIs and superhuman robots whose moral status is obviously higher than that of their human counterparts and protagonists. This has been a popular idea, at least in science fiction.

Generally speaking, moral values have changed a lot across time and populations, and certain moral values have survived for a while because they gave a temporary evolutionary advantage to beings and societies that adopted them. However, evolution isn't over, and the universe is still young.

Tim Scarfe

So it sounds like you've got an all-encompassing view of the universe, life, and everything.

Jürgen Schmidhuber

Indeed, in 1997, I wrote my first paper about this. What is the simplest explanation of our universe? Since 1997, in my secret life as a digital physicist, I have published on the very simple, asymptotically fastest, optimal, most efficient way of computing all logically possible universes—all computable universes, including ours.

As long as there is no evidence that our universe is not computable, we stick with this assumption. At the moment, we don't have any physical evidence against this. This was a generalization of Everett's many-worlds theory of physics, but now it's more general in the sense that you have all kinds of different universes with different physical and computable laws.

Now, any great programmer—a great programmer with any self-respect—should use this optimal method to create and master all logically possible computable universes, thus generating us as byproducts and generating many histories of deterministic, computable universes, many of them inhabited by observers like ourselves.

And due to certain properties of the asymptotically optimal method—many people don't know there is one, but there is one—at any given time in this all-encompassing computational process, most of the universes computed so far that contain yourself will be due to one of the shortest and fastest programs that computes you. This little insight allows for making highly nontrivial and encouraging predictions about our future, about your future.

Tim Scarfe

Jürgen, this has been amazing. Do you have any final messages for the MLST audience?

Jürgen Schmidhuber

Yes. Don't worry. In the end, all will be good.

Tim Scarfe

Touch wood. Jürgen, it's been an absolute honor to have you on the show. It's been a dream of mine to do this in the flesh, and I really appreciate you coming on. Thank you so much.

Jürgen Schmidhuber

That's very kind of you to say that, and it was a great pleasure for me. Thank you.

Jurgen Schmidhuber on Humans co-existing with AIs | BidClub