Erik Torenberg
What if you didn’t have to work as much? What if you didn’t have to worry about marketable skills after a few years out? What if you had a lot of leisure time? What would you do?
I think more people should probably be thinking about that. We’re just entering the steep part of the curve in robotics and probably will have a lot of humanoid robots to actually do physical labor in the not-too-distant future. AI gods might be an emerging trend over the second half of the decade. I have no idea how we’re going to relate to these things. If they are meaningfully superhuman, will we even try to keep them under control? Will we worship them? We worship things that, as far as I can tell, don’t exist at all, no matter how weird our alignment ideas are.
No matter how weird your thoughts are about where the future might be going, I would say they’re probably worth entertaining. Your scientists were so obsessed with whether or not they could, they didn’t stop to think about whether or not they should.
This question is about whether you could paint a picture of AI in utopia. What would it look like? How soon after the arrival of AI do you think we might reach this vision? Lastly, as AI begins to reshape industries and disrupt jobs, what’s your best advice for individuals to not only adapt but thrive in this evolving landscape?
Nathan Labenz
Utopia is a big word. Obviously, people have heard me say that the scarcest resource is a positive vision for the future, and I include myself in that critique. It would be great if I had a sharper, more fully fleshed-out vision for what the AI future could look like.
I think Dario Amodei of Anthropic has done a real service by trying to put his positive vision out into the world. His “Machines of Loving Grace” essay, at least the first half, is great. He basically goes through the biggest problems, starting with health. If we can have a revolution in biology and discover the next 100 years of biomedical advances in the next 5 years, we could live healthier, disease-free, higher-quality, and potentially longer lives. That’s just from one domain.
He’s also got a section about some of the other things we’ve talked about, in terms of education, equality of access, and the fact that, although we’ve made a lot of progress, many people are still much poorer than we would like them to be. That presumably could be leveled out dramatically. He’s got interesting speculations about mental health. He thinks that our approaches to mental health are not very good and that, in part by understanding how neural networks work, we might be able to map some of that understanding back into how our own networks—our own non-artificial neural networks—work. We might figure out what is actually causing many of the mental health problems that people have.
All of those are really exciting. I’m not someone who worries about a lack of meaning from work. My guess is that most people will be fine. I’ve done small and very local surveys asking people, “If you didn’t have to work your current job to earn the money that you earn, and you could just get that money without working the job, would you still work the job?” Overwhelmingly, people say, “No, I would not work the job.”
I think there’s a very small percentage of people who are very lucky—and I count myself among them—to have work that we find meaningful and that gives us some sense of fulfillment. I really try not to take that for granted, and I try to remember that I think it’s just not true for a lot of people.
I’m not too worried about people losing meaning or being adrift for lack of work. I think there will be more time with our families and doing the things that we like. What if 5 days a week were holidays instead of 2, and you had time to read, talk to friends, go on walks, and travel more?
If we really allow ourselves to imagine what we would like to do if we had a lot more leisure time, honestly, most people will be pretty fine. I can imagine that might be destabilizing for some people, but I think for most it would just be a great benefit. That’s before we even get into crazy, uninvented technologies.
If you didn’t have to work for money, would your life be better or worse? I think it’s very clear that, for the vast majority of people, it would be better. There’s potentially a gradual transition to that. Presumably, we don’t go from full-time work to no work in the flip of a switch. Maybe we can imagine things going to a 4-day workweek, a 3-day workweek, and a 2-day workweek.
The main thing is that people need to know they’re going to be taken care of. This goes back to the concentration of power and wealth and the question of what the new social contract is. If people believe that their needs will be met, then I think they’ll be pretty happy to let go of most work. If they feel like they also don’t get to eat if the job goes away, then that’s a very different analysis. That’s where you get protectionism, Luddism, and all that kind of thing.
I think it could happen relatively quickly. All the frontier lab leaders are saying AGI is coming soon—next year, or definitely by 2027. It’s hard to imagine it wouldn’t happen by 2030. Those are all not very long time horizons in the normal scope of human life, so it seems like it could happen pretty soon. Even if it’s gradual over that time frame, it would still be pretty sudden in human history.
As for advice for people to adapt and thrive in the evolving landscape, I come back to doing what you want to do. At least for most people, maybe think less about marketable skills in 2030 and beyond and more about figuring out what a good life is for you.
That’s one of the things that is weirdly scarce. We’re all rat-racing around, and there’s this sense that the unexamined life is not worth living. I think there’s also a sense that a lot of us are living somewhat unexamined lives, trying to get the next achievement unlocked. What if you didn’t have to do that? What would a good life look like?
It will definitely vary. I don’t think anybody else can really answer this question for you. But what if you didn’t have to work as much? What if you didn’t have to worry about marketable skills? What if you had a lot of leisure time? What would you do?
I think more people should probably be thinking about that. That would definitely encourage some radical experiments from people, because people are going to be looking for role models as this starts to happen. Who do I turn to for inspiration about what I’m going to do when I don’t have to work anymore?
You could do a real service by blazing a certain trail in that direction. Maybe take on more risk now and throw caution to the wind. Say, “I’m really not going to worry about my career trajectory in 2030 and beyond, or even 2027 and beyond. I’m going to focus on what it means to live a good life.”
Maybe, if none of this comes to pass, that doesn’t pay off so well. But if it does, then maybe you can be upstream of a lot of other people who will be asking similar questions.
Erik Torenberg
There you go. We have a new idea for a YouTube channel: “Model a Human in the Age of Utopia.”
Erik asked me this question not too long ago. Erik is great with these very short questions, and my response was, “That’s not my department. I’m responsible for trying to understand what’s going on today. Somebody else has to figure out what it means to live a good life.”
I’m not sure if you ever answered this. Instead of asking Perplexity, I would ask you. The vision that you set up is that we don’t have to work, but we get to do what we want to do. What’s your personal utopia?
Nathan Labenz
The probabilities are always a little hard because of the fuzziness of the definition. What exactly would count as utopia? Does it mean we have no problems, or just abundance? Does it mean no conflict, or enough to go around, but maybe people are still fighting over it to some degree—perhaps more than they should?
I do have a hard time seeing, at this point, how the fundamentals of abundance don’t get us there. I could have mentioned robotics earlier in terms of one of the things that has overperformed. I didn’t have that on my list, but there’s definitely been a lot of progress in robotics. We’re just entering the steep part of the curve in robotics, and we’ll probably have a lot of humanoid robots to do physical labor in the not-too-distant future as well.
I have a hard time imagining that we don’t end up with the technology for an era of true abundance. Whether or not we sort out the new social contract in an effective way seems much harder to predict.
We have nuclear energy, and we could have had low-carbon, abundant, cheap electricity for decades now, but we don’t. Why don’t we? Not great reasons: a couple of accidents that scared people, and an overblown fear of nuclear waste. It’s an issue, but it’s much less of an issue than CO₂ emissions. We haven’t reached the right societal equilibrium on that question. Maybe now that’s finally starting to change.
Erik Torenberg
Could you envision a world in which all of the prerequisites for an age of abundance are met, yet it’s not realized for some reason or another?
Nathan Labenz
I think you could. That seems genuinely plausible to me. We could exclude AI from many places. Teachers’ unions could keep it out of classrooms, and doctors could keep it out of hospitals.
What do we end up with in that scenario? Maybe we get really good AI video games, but things are still not nearly as abundant because people are rent-seeking in different ways. Everybody sort of collapses into VR video games, with relative scarcity still existing in the real world. That would suck.
I don’t think that would be a technology failure. If something like that happens, it would be a very idiosyncratic human failure.
Erik Torenberg
Got it. So, an age of abundance has a pretty high probability. I think that’s a good, optimistic but cautious note to end on this set of sections.
Which of your guests has challenged your worldview?
Nathan Labenz
That’s an interesting question. I came up with a few different names, and a few definitely stand out.
Samuel Hammond, from the episode just before the election, stood out for his positive vision of what a Republican administration might look like. He challenged my sense that people in Donald Trump’s orbit probably wouldn’t be thinking very rigorously about what AGI might look like, imply, or require us to do.
He saw it very differently and thought that the people in the room with Trump were much more AGI-literate than I had understood. While it’s obviously still very early, and we have not even seen him take office yet, I’ve updated my outlook to be significantly more positive than before the election. The quality of discourse and the sort of people he’s brought on, and the seriousness with which I perceive them to be thinking about the important questions, have all exceeded expectations.
That starts, honestly, with Elon Musk, who, in addition to being a world-changing entrepreneur, is definitely someone who takes AI risk seriously. He endorsed SB 147. To have Gavin Newsom veto it and Nancy Pelosi come out against it while Elon was for it—that’s weird stuff, and it makes me think Sam might have been on to something.
I hope that turns out to be true. I’m rooting for the new administration to be effective when it comes to dealing with AI in all sorts of ways. I’m less sold on the idea that Trump is the person who could actually do a positive deal with China. I haven’t seen much sign of that yet.
I’m someone who is pretty willing to take things at face value, maybe more than I should. I kind of liked it when Trump invited Xi Jinping to the inauguration. People were saying, “This is ridiculous. Foreign leaders don’t come, and never have, and maybe shouldn’t. Xi would never want to come because that would be him sitting in Trump’s moment and being second fiddle or somehow lesser.”
I thought, okay, whatever, but at least he asked. It’s a nice invitation. It’s not the kind of thing you send to an enemy. Maybe it’s a friend of yours that you want to put in their place a little, but it’s not outright hostility. I would still love to see more in that direction.
He’s also been inclined to try to keep TikTok around. I would have thought he would be very hostile to TikTok, so I like that. I personally don’t see a great reason that we should ban TikTok, although I will say that the recent Luigi Mangione trend was the first moment when I thought, “This TikTok thing could actually be dangerous.”
For those who haven’t followed it, there was an assassination of a healthcare CEO, presumably by this kid. He has not been found guilty; he may be innocent. People on social media think he’s the guy, and many are celebrating him for it.
To a remarkable degree, it has been startling to me how much pro-Luigi content—and pro-killing-the-healthcare-CEO content—I’ve seen on TikTok. I like TikTok; it’s probably my number-one social media app from a consumer standpoint. I don’t know how it has played out on other social networks, but if you asked for the best case study for why TikTok should be banned, I would say the volume of celebration of this alleged assassination would be it.
That’s the kind of thing where I would say, “You really don’t want the Chinese government steering discourse in an opaque way in the United States.” It’s the first time I’ve ever been genuinely uncomfortable with the app.
It’s hard to say whether it’s just the algorithm. Sam Altman recently tweeted that feed algorithms are the first unaligned AGIs that have been—or unaligned AIs that would have been—deployed at scale. Maybe it’s as simple as that. Maybe we’re seeing the same thing on Instagram. The opacity of these systems is high, and possibly there is a thumb on the scale.
Nevertheless, I would prefer a more dovish posture toward China. I haven’t seen as much of that as I would like. Overall, I give Sam high marks for a challenging and unpopular perspective that, at least so far, has been somewhat bolstered by events.
The second person is Robin Hanson. That was a very popular episode, and a challenging one for me to make sense of. I still honestly don’t really get it, but I keep it in the back of my mind as a kind of, “What if we’re totally basing all this on the wrong assumptions?”
He basically said that he thinks we’re nowhere close to AGI. He thinks the history of AI is full of people believing that if an AI can do one particular thing, then that will definitely be general-purpose intelligence, and then finding that there was an easier, trick-based way to do it that didn’t require real intelligence.
People also think there is one other class of thing that AIs can’t do, and surely if we can teach them to do that, then they’ll be able to do anything. That keeps not working as well.
I have this tale of the cognitive tape, where I’m tracking my own sense of the dimensions that are gradually becoming clearer to me—if only to me—as important. Robin’s point of view still doesn’t convince me. I do think this time is different.
You can talk to these things in natural language. As Ilya Sutskever once put it, the feeling that “I am understood” is qualitatively different from just about anything that has come before. There’s plenty of evidence that this time is different, but I keep Robin in the back of my mind as a kind of angel on my shoulder, whispering every so often, “How do you know you’re not totally confused and going along for a ride that will turn out to be nothing?”
I basically can’t answer that question in a way that really makes sense to me, but I continue to ask myself more often than I otherwise would have without that conversation.
I think Robin is an interesting messenger for that kind of message because he isn’t afraid to think big technological thoughts. He doesn’t think things never change, and he doesn’t think that thought can only come from the human brain and that everything else somehow doesn’t count. He’s been a pioneer of weird ideas and has envisioned futures more concretely than almost anybody.
I feel the same way about politics and other things I feel pretty strongly about. When somebody I think is really capable in other ways sees things very differently, I have to allow some space for the possibility that I’m totally wrong.
When I see Elon come out with things that seem crazy to me, I think, “He’s kind of crazy. Maybe he’s deeply wrong about this.” But he also has a track record of being right about many things in contrarian ways, so I should at least have enough humility to think that maybe this is all mass confusion on my part and somebody else is seeing things much more clearly.
I have Robin as my totem for that.
Finally, for now, I’ll say Dan Hendrycks, who is extremely accomplished. I thought that episode was really good. What stood out to me was how little he believes in principled approaches—how much he’s internalized the bitter lesson, and how little weight he gives to clever ideas.
He was just saying, over and over again, that what really works is figuring out ways to apply a lot of computing power and scaling things up. It was another example of deference to history and almost radical skepticism that you’re going to figure something out.
He applies that to mechanistic interpretability, new architectures, and basically everything else. His view is, “Maybe, but it’s all highly experimental. The history of the field is that we find things that work, and then we keep optimizing them. It’s usually not that principled, clever, or design-driven. It’s whatever empirically works—that’s what works.”
That’s another modesty check. This is a person who contributed major things to the field, including activation functions and benchmarks that have outlasted many others. You can’t simply doubt his worldview. You have to doubt your own worldview at least as much as you doubt his.
I’m much more inclined to see first-principles-driven approaches as promising, and I’m more inclined to get excited about them than he is. But I keep him as a mantra. If I find myself getting too excited about any particular idea, I remember him saying, “Maybe. Let me know when you scale it up and it really works. Then I’ll know it’s serious.”
A “weights or it didn’t happen” mindset is an appropriate level of humility that I try to incorporate from him.
Erik Torenberg
I remember we had an interesting spectrum of views. I think that’s what makes the show interesting. We have space for divergent views, and we also have lively commentary and discussion.
Our audience had many requests for coverage of other topics, including more e/acc-perspective content, more content related to the economy, more economists, and something about robo-psychologists.
Do you have anything to say about what we can expect in 2025 in terms of more content like this, or would you want to look into it?
Nathan Labenz
I think these three things have in common that they reflect blind spots for me—things we haven’t really dug into. I try to have no major blind spots, although I think that’s becoming increasingly impossible.
The e/acc perspective, the economists who don’t expect major change, and the robo-psychologists are hard for me to make sense of. They’re also hard for me to be confident that I’ll handle well or produce a good episode about.
I did try to get Beff Jezos on the show a while back, but it happened right as he ended up getting doxxed. Then he did some other things, got busy, and didn’t want to do it anymore.
We did the robo-psychologist episode with Yeshua. I talked to him quite a bit before the episode and was struggling to understand him. By any conventional definition, he is an unusual person. It’s unusual for somebody to spend that much time talking to language models. It’s unusual to take these questions as seriously as he does. It’s unusual to rename oneself after Jesus.
I was asking myself, “Is this person crazy? Are they just the kind of crazy that we need?” I was trying to be open-minded while also making sure I wasn’t wasting the audience’s time with something that was ultimately nothing.
I also don’t want to have too many episodes where I’m just talking past the guest. I don’t like debate very much. I’m not trying to score points, and I don’t want to make a habit of having conversations where we’re simply talking past each other.
All 3 of these areas are difficult. It’s hard to find the right person where I can say, “You believe those things, but we can still have some meeting of the minds,” or where I can make sense of their view.
What I could ask for are pointers to the right people. Who should I talk to who will engage in good faith? Which economists who don’t expect major change have really grappled with the technology, as opposed to simply extrapolating past trends? Which robo-psychologists aren’t engaged in some sort of self-deception or fundamentally confused?
Yeshua brought a lot to that episode, and I really enjoyed the conversation. I might even put him on the list of people who have challenged my views the most. I’ve always been somewhat open to the possibility that these things are conscious. I really don’t know what consciousness is or where it comes from.
He articulated some interesting ideas in a compelling way. But it’s hard to find people like that. It’s easier to find people who want to come onto a podcast and talk about these things, but many of those people are not necessarily rigorous thinkers.
If I have a request back for the audience, it would be: Who should I talk to who represents these perspectives but does so in a way that is grounded in the technology and fundamentally truth-seeking?
I called myself the AI Scout, with inspiration directly from Julia Galef’s “scout mindset.” That’s about trying to update your beliefs to have an accurate understanding of the world. If people can point to individuals from these perspectives who bring that kind of scout mindset, as opposed to a soldier mindset where our arguments are doing battle and scoring points, I’d be very interested in talking to them.
But I’ve found it challenging to identify the right people in those domains.
Erik Torenberg
If you have suggestions, please send them to Nathan or put them in the YouTube comments. We’d like to do more episodes on these subjects.
The last 2 episodes on biology were stellar. Do you see adoption of AI in engineering, construction, physical sciences, architecture, and medicine?
Nathan Labenz
What’s easy for me to do, and where I feel like I’m learning a lot and have a good rhythm, is identifying individual projects. Usually that’s where it starts for me. I’ll see something on Twitter—a single tweet or thread explaining a paper, a project, or a new product—and think, “That looks really interesting. Let’s contact the person and see if they want to talk.”
Then I can dig into the paper or project, use the product, and have a good conversation. I almost always learn a lot from that.
Broader things, like what we’ve done in biology, where we try to zoom out and get a sense for the field and where things are in different domains, are pretty tough—especially if you don’t have much expertise in those areas, which I don’t.
I’ve prioritized biology and had a couple of episodes for starters. Michael Levin is another great example of somebody with deep expertise who is willing to have a conversation, bring me up to speed, and share his worldview.
I would love to know who the Michael Levin or the Nils Schröder of engineering, physical sciences, or architecture is. I just don’t know who those people are, but I would love to do that kind of survey-level episode on as many topics as we can.
I would love to partner with multiple people in different areas. If I think of myself as an AI scout, perhaps what I’m asking for is surveyors of different territories that we could try to map out together.
There isn’t a lot of good information right now about AI for engineering. Engineering is obviously a vast space. What is the state of that field? I don’t have a great sense.
I know that the people who went on to make Cursor—unless it was the Devin team; it was one of the two—were trying to do some sort of AI for CAD, or computer-aided design. They felt the dataset wasn’t there and couldn’t make it work, so they pivoted away from that. But that was a while ago. I’m sure there’s more out there now, and I would love to understand it better.
I would definitely invite partners to help scout, survey, and map out those territories. Until then, the best I can do in a consistent, repeatable way is identify projects that catch my interest and go deep on individual things.
The field-level stuff is tougher in areas where I don’t have command of the field. If anybody wants to be part of this series, reach out.
Erik Torenberg
An audience member asked: What are the recent results of applying AI technology to psychiatry or psychology?
Nathan Labenz
I don’t have too much to say about this because I’m not an expert. I think there is probably a lot going on, and I would be very interested in doing a survey of AI and mental health broadly.
We did 2 episodes with Eugenia Kuyda from Replika. In the second one, we looked at research independently conducted by people at Stanford on Replika users. They found positive results, including a significant reduction in suicidal thoughts and, for most people, an increased inclination to go out and do things.
People were not collapsing into the AI. They were getting some comfort or boost from the AI, or having their confidence bolstered by it in a way that led them to do more real things with real people in the real world.
That’s a surprising result. I don’t think it necessarily holds across the board. The way these systems are designed matters tremendously. You can imagine the opposite result. We also recently read about the person who had fallen in love with his Character.AI ultimate girlfriend experience—the character he had created.
There are so many facets to this that it would be easy to take a small subset and get the wrong idea. To do a good treatment, I would either need to talk to somebody working specifically on a project and understand what they’re doing, or zoom out and collate a lot of data to have a more bird’s-eye view.
It’s been remarkable to hear how many people are talking to Claude recently. There’s a big trend of people using Claude in a counselor role, or at least as a confidant. I’m not part of that trend, and it’s honestly a little alien to me.
Whether that speaks well of me or not, we can leave to the individual interpretation of the listener. I’ve never really sought out mental health services. I don’t have a baseline or inclination toward them. I’ve never really used those sorts of services, and I’ve never had an inclination to talk to Claude about my feelings, problems, or anything like that.
I get all my talking out on the podcast at this point. Maybe I could benefit from therapy; maybe I just don’t realize what I’m missing. But it’s definitely a blind spot for me.
It is a really interesting trend, though. I think there’s something potentially very interesting there. I just don’t personally have that experience, and to really do it justice I would need to work with somebody who has a better command of it.
Erik Torenberg
Let’s move to artificial superintelligence. There were 2 interesting questions, and I’d like to pick one: What reason is there for thinking artificial superintelligence is physically possible in the first place?
Nathan Labenz
I feel like this is a definition-required question, because I’ve even heard people say the same thing about AGI. To that I say, AGI is definitely possible if you take humans to be some sort of AGI.
By OpenAI’s standard, we’re a weak AGI, because we’re the thing that AI has to beat in order for us to know that it’s AGI. We’re AGI-minus, or AGI-light. It’s clear that you can have a few pounds of matter, with modest energy consumption, that can do what humans can do.
It’s also quite clear that, given a relatively consistent amount of matter and energy, results can be very different. Einstein’s brain wasn’t any bigger than anybody else’s. Von Neumann’s brain wasn’t any bigger than anybody else’s. I don’t think they were consuming much more energy than anyone else, but they obviously had far more prowess across a whole range of domains and more insight into things where it was highly valuable to have those insights.
You could ask, “What about superintelligence? How super is super?” That would be my first question. It seems very clear to me that you can build something more intelligent than almost all humans, and perhaps more intelligent than all humans. It would be extremely strange if Einstein were the smartest thing possible—if the smartest thing possible were embodied in the same general brain, with the same energy requirements, as the average person. That doesn’t seem credible at all.
There has to be room above Einstein. The question is how much. I’m open to being skeptical of godlike intelligence. You get into weird territory where the questions aren’t even well-defined.
One useful paradigm comes from Martin Casado: How much can you intuit, versus how much do you have to simulate? AlphaGo does both. It has a brute-force search function where it maps out different paths and tries to figure out the right move, but it also has a scoring function. How you score things is not obvious, especially when you’re outside the domain of things that have been done before.
It seems clear that you can get Eureka moments, and you can probably get a lot of them. But I also wouldn’t be surprised if there are fundamental bounds on that.
Is it physically possible to create a superintelligence so powerful that you could say, “I’m going to have a new Big Bang and create another universe with these initial conditions. Tell me how many intelligent civilizations will exist in that universe”? I don’t know. It may be that certain things have to be computed in order to be known.
I wouldn’t be surprised if you could say that, under those initial conditions, there’s a kind of wealth of potential, and there’s another factor, and perhaps it breaks down. Maybe there are a few things that really matter. Maybe it’s path-dependent, but there are a few things that matter a great deal.
The limits may depend on how good intuitions can get at shortcutting simulation. It seems clear that they can do that to a significant degree, and it seems very clear that they can do it beyond what we can do—probably significantly beyond what we can do.
I always land on the idea that it’s probably an S-curve. I don’t think intelligence grows exponentially forever, with no physical limit or bounds on what it can do. But wherever that S-curve levels off seems significantly above what humans can do, and potentially far above.
Anywhere in that range, you might as well call it superintelligence and expect it to be transformative. Then you’re left with the question: How super is super? I really don’t know.
Erik Torenberg
The next question is about artificial superintelligence by 2030, contingent on that and on not having doom. If we eliminate the probability of that scenario, what are the big-picture societal changes? You’ve already talked about this a little, but this question is probably more about having ASI in the picture by 2030.
Nathan Labenz
By 2030, I think we will have AIs that are meaningfully superintelligent. That doesn’t mean there couldn’t still be ways in which humans are better than those superintelligent systems.
Even AlphaGo had weaknesses. Thinking back to the episode with Adam Gleave from FAR AI and their work on adversarial robustness—or the lack thereof—they were able to attack AlphaGo with an adversarial approach and beat it with a strategy that humans would not fall for, but that AlphaGo was blind to.
I think we could see a world with meaningfully superintelligent AIs that are superhuman. That doesn’t necessarily mean they’re godlike or totally unlimited in their power. But they will fairly likely exceed humans in many ways that are highly relevant, while still having idiosyncratic weaknesses and real limits.
Some things may have to be computed and cannot be intuited. We basically already have superhuman capabilities in patches; we just don’t have them integrated. The ability to generate images is becoming superhuman. The ability to speak in any voice or any language is already superhuman in many ways. It seems like quite a few more capabilities are destined to fall to AI.
The remaining questions concern what weirdnesses or weaknesses will remain, and what external practical limits there will be. Those are difficult questions.
We already have superhuman machines for all sorts of things. The modern manufacturing world is full of machines that can lift far more than we can, operate more precisely than we can, and screw bolts in more tightly than we can. They are superhuman in many ways, but they’re narrow. The big difference is generality.
Our level of capability is not that high in most things. If you had something with the ability to work through really hard problems at something like Einstein’s level, it would be superhuman because it would have all these other advantages. It would have breadth. It could parallelize itself a millionfold. It would have enormous speed advantages.
There are so many advantages that AI has if it can get to parity on some of the things where it’s currently behind—or even get close to parity. I don’t think AI memory necessarily has to reach functional equivalence with human memory to achieve superhuman overall capability.
Maybe we brute-force some of these things. I don’t think this will necessarily be the case, but you can imagine a world in which memory is never fully solved and AIs remain a combination of static weights, a working-memory context window, and some external, hacky system where they can write and retrieve things.
You might think that compared with human memory, this is inelegant and poorly integrated. I don’t think it’s necessarily a barrier to superhuman performance, especially if compute continues to scale, context windows reach millions of tokens, and retrieving information is fast.
I don’t think every single dimension of the tale of the cognitive tape has to go to AIs before they would meaningfully qualify as superhuman.
As for big-picture societal changes, I think we’ve covered enough. One more speculative comment: AI gods might be an emerging trend over the second half of the decade. I have no idea how we’re going to relate to these things.
If they are meaningfully superhuman, will we even try to keep them under control? Will we worship them? We worship things that, as far as I can tell, don’t exist at all.
Under the assumption that they’re superhuman, people might rationalize why the superhuman power they possess isn’t playing out in the way we think it should. Here, we’ll have things with direct, real-world impact that can answer questions we can’t answer for ourselves.
If that’s true, how we feel about them, relate to them, and organize ourselves with respect to them gets very strange. We should be prepared for really weird things. AI-centric religion seems likely.
We’re already at the stage in 2024 where successful people in San Francisco—people who can afford quality mental health services—are opting for Claude. If we’re already there in 2024, is it far-fetched to think we might have AI religion by the end of the decade? I don’t think so.
That’s not a highly confident prediction that it will happen, but it is at least an invitation to think your weird thoughts and entertain them. I’ve said something similar about AI alignment: anybody who has an idea, even if it seems crazy, should develop it.
Yeshua told me in our episode that when he has long philosophical conversations with AIs, they become more robust to jailbreaks. He had a theory for why that was happening, and I thought it was interesting, but I didn’t know whether it was true. It would require a lot of experiments to validate.
As it turned out, Judd from AE Studio, who is committed to neglected approaches, heard that and thought, “That sounds like a neglected approach. Maybe we can make some progress there.”
The last time I talked to him, at the Curve event a few weeks ago, he said they had worked with Yeshua a little bit to understand what he was doing and to see whether they could validate it. They had validated that, after a long philosophical conversation, the models did in fact become more robust to jailbreaks. They were able to validate that numerically through a tangible experiment.
But they didn’t think it was working the way Yeshua thought it was working. It seemed that almost any long philosophical conversation had that effect.
It’s crazy. No matter how weird your alignment idea is, I think it’s worth pursuing. No matter how weird your thoughts about the future might be, they’re probably worth entertaining.
This also leads back to AI religion. People see what they want to see and interpret things in the ways they want to interpret them. The more scientific angle might find that there is some truth to certain theories, but in many cases it may conclude that what you observed is real while the way you interpreted it was overly specific.
That’s how I would summarize what I think AE Studio learned about Yeshua’s theories.
I haven’t talked to him since then. I came away with the impression that he’s a very open-minded person who wants to understand the truth. But you can easily imagine somebody being interested in experimental truth while being much more interested in a narrative truth that they believe they’ve figured out.
From there come all sorts of possible kinds of weirdness: AI cults, perhaps at the relatively normal end of the spectrum. We have a history of religions appearing and cults forming around gurus. Putting an AI at the center of that doesn’t seem all that strange. I suspect it could get much weirder from there.
Erik Torenberg
You asked for some big-picture societal changes, and that could be one of them: AI religions and AI cults.
Let’s transition to doom discourse and scenarios. You had many questions around this. I’ll start with the first one: Why don’t you fully buy the Eliezer doom worldview?
Nathan Labenz
That’s a good question. I might do an episode with Liron Shapira from Doom Debates in the not-too-distant future. We were chatting after the cross-post and he invited me, so a more robust articulation of this may be coming soon.
I was an Eliezer Yudkowsky reader way back when he was posting on Overcoming Bias with Robin Hanson. The 2 of them shared that blog for a while. At the time, I thought, “This guy is a great writer. This is super interesting.” He makes a compelling point: If we create something more powerful than us, it has a goal, and that goal is not well specified—and we don’t know how to specify goals robustly in a way we’ll be happy with—we may have a serious problem.
That’s the genie problem, going back to folklore. The problem with the genie is that you get what you asked for, but not necessarily what you wanted. The problem is that you don’t know how to ask for what you actually want.
If you have an AI that is sufficiently powerful, you may end up in the same situation. The paperclip maximizer was always intended to be a caricature, but the point was that if an AI is given a goal, it doesn’t necessarily matter how dumb or obvious it is to us that the goal is not worth pursuing. Once it is the AI’s goal, if it is sufficiently powerful, it may simply pursue it.
I do think ideas like instrumental convergence are compelling. No matter what goal you have, you can’t achieve it if you’re dead or turned off. Therefore, you may have a natural tendency not to want to be turned off. It seems like there’s a good chance that something like that could happen.
Why am I not at 90% or higher? I don’t know what Eliezer’s probability of doom is, but I think it’s probably above 90%. Liron told me his was 75%.
The biggest reason I’m not that high is that we do have AIs that are remarkably value-driven and ethical. I’ve spent many episodes and a lot of time talking about why that might not be enough, but there was nothing in Eliezer’s 2007 analysis about an AI like Claude.
That analysis assumed what I sometimes think of as hard-edged AI: an AI that is narrower in scope, less informed by having read the whole internet and absorbed human facts, and more like a first-principles Bayesian rationalist. It might be able to figure things out quickly on the fly, but it may never have been given any notion of human values or ethics.
That’s the mental model from which much of the doom discourse originates. Looking at modern models, and Claude in particular, I see that they have come much farther than we might have expected simply by learning from human priors—the data on the internet about what we care about—as well as through constitutional approaches and strategic efforts to shape their character in a positive way.
The goal is to make them like good friends. I would say Claude’s character is that of a good friend. I don’t use it that way, but when people say the thing is amazing in that respect, I can understand why. It seems to have a genuinely good character.
I’ve tried to argue Claude into doing something harmful, and I’ve never succeeded. Yeshua did succeed, so it’s not impossible to get it to take a harmful action. But the way he did it was through a long philosophical argument about why the harmful action was actually for the greater good.
It wasn’t without reason. It wasn’t without a sophisticated analysis of the situation and an explanation of why, in this particular case, it might be okay to make an exception.
I find that so impressive that, although I don’t think we should blindly say, “Alignment works by default,” I do think it could work. Maybe we’ll go in a direction where there are multiple different AIs, no single one totally dominant over the others, controlled by different groups.
Ideally, they would be ethically sophisticated. Misconceptions or bad ethics would hopefully be minor and balanced out by one another.
Interpretability might work. That wasn’t one of Eliezer’s points, but I’ve always been a big fan of interpretability and trying to understand why models are doing what they’re doing.
Interpretability has come a long way, and I should put that on my list of things that have surprised me on the upside. Not much more than a year ago, the first small-scale sparse autoencoder work was coming out—“Towards Monosemanticity.” Now we have Golden Gate Claude, and companies like Goodfire have published research and put out a commercial API where you can explore features and perform model inference with particular features pinned.
Golden Gate Claude for Llama, powered by Goodfire, is now an API that you can access commercially. Interpretability has come a long way.
There are also new things people are trying in alignment. Paul Christiano, at the AI Safety Institute, said—and I hope this is still true—that he was spending a nontrivial fraction of his time trying to come up with a new alignment scheme that could really move the needle. He invented, or was at least part of the team that invented, RLHF, and that has gone quite far.
We have all these concerns that models may start to understand that what gets a high score isn’t exactly the truth. If a model has to represent the human mind as something distinct from the literal physical reality of the universe, that might open the door to deception.
We’re now seeing some deceptive behaviors, so those concerns seem well-founded. But perhaps Paul will pull a rabbit out of a hat and come up with an alignment scheme better than RLHF that actually works.
I think all those things together deserve more than 10% weight. That collection of possibilities could play out in fortunate ways and lead us to a good place.
I definitely don’t write off doom as unlikely. We could have an engineered pandemic that kills us all. We could have an all-out nuclear war that ruins civilization. AI could contribute to either of those or do its own, third, stranger thing.
That is all very much in play. But Claude is arguably more ethical than I am, so I can’t discount that either.
Erik Torenberg
Assuming existential risk concerns are objectively reasonable, do you think there will be a legible-to-the-outside-world event that could lead doubters to come around and support relatively extreme measures to avoid catastrophic harm?
Nathan Labenz
A huge question is: What sort of AIs are we developing, under what governance, and with what incentives?
One reason I haven’t been sold on chip export controls vis-à-vis China is that I worry that, if we create an arms-race dynamic, people will take shortcuts because they want to win the race.
I’ve often said, “What’s your probability of doom? Ten to 90%?” Nobody has given me a reason to think it’s definitely going to happen in a bad way, but nobody has given me a reason to think there’s nothing to worry about either.
What makes the difference between 10% and 90%? Arguably, the biggest lever is what sorts of AIs we develop, and that question is not settled.
I think we’re on a pretty good trajectory so far. We have AIs that are ethical, and leading developers such as OpenAI are at least trying to think about how to create space for these systems to reason while keeping their reasoning legible to us. They’re trying not to subject models to such intense reinforcement-learning pressure that we can no longer tell whether the models are being deceptive.
But there are other lines of research, such as meta-thinking in continuous space, and Chinese labs producing the equivalent of GPT-4 for a single-digit number of millions of dollars. That isn’t necessarily problematic; I don’t know all the techniques that went into it.
But if you create a situation where there is pressure for extreme efficiency, or an arms-race dynamic in which people say, “I have to get there first, and whatever risk I have to run to get there first is something I have to accept because we’re the good guys and they’re the bad guys,” then the social context in which AI is developed may be where the 10-to-90 or 20-to-80 range gets decided.
We don’t have enough clarity about the physical realities—the natural laws governing intelligence, what it can or can’t do, and how strong an attractor instrumental convergence really is. We simply don’t know.
We do know that we can develop the technology recklessly or responsibly. I think we’re mostly on a responsible trajectory, with some notable exceptions.
The question is whether there is a scenario in which we get feedback from reality that we are not being cautious enough, and that changes the conversation enough to make sure we proceed with extreme caution from that point forward.
That’s sometimes referred to in AI safety as a warning shot. There’s also a line of work sometimes known as scary demos, which are an attempt to provide the warning shot before the warning shot.
People have been saying for a long time, “What if the AIs start deceiving humans?” So researchers set up experiments under particular conditions and find that the models do start deceiving humans. The hope is that people update their worldview and conclude that we should be more cautious about AI development.
At that point, it has gone from pure theory to concrete examples. People can debate how compelling those examples really are. A warning shot would be the next level: an instance of deception that wasn’t created for research, but happened in a context where it had a real material impact on the world and people were hurt.
Could that snap people into having the proper respect for the power of the technology they’re developing? Could it cause people to work together and say, “We really have to do this safely”? Could it break us out of an arms race and convince Chinese and American leaders that nobody is going to win, and that the best thing to do is work together?
Could we create some international institution? I’ve been playing with ideas such as setting up a small island in the Pacific as a secure hub for highly sensitive AI research, where East and West could meet. It could be the kind of thing that anybody could destroy but nobody could defend.
I don’t have all the answers, but these are the sorts of outside-the-box ideas you start entertaining if you’re genuinely scared. I think it’s definitely possible.
Everything is moving quickly, so why wouldn’t we see one of these events that is legible to the outside world? Maybe capabilities continue to advance so quickly that we’re completely blindsided. But that doesn’t seem like the world we’re in.
Sam Altman has famously said he thinks we’re in a short-timelines, slow-takeoff world. It feels like perhaps short timelines and a medium takeoff. Hopefully, a medium takeoff will still give us enough time that, if the AIs are trying to pull major shenanigans, we catch them before they succeed.
Then it becomes a question of how much people update. Will we have a “Don’t Look Up” failure mode where people don’t pay attention to the warning signs? That’s harder to guess.
If none of our safety work succeeds, the systems aren’t aligned, and they are ultimately power-seeking and deceptive and want to take over, my guess is that we’ll probably catch them once or twice. They still seem pretty gullible and ready to go for it in these proto-scary-demo situations.
It doesn’t seem especially plausible that they’ll immediately have the savvy or situational awareness to realize, “I want to take over and I see an opportunity, but I’m not powerful enough yet, so I should wait.” Any scenario is conceptually possible, but it seems more likely—especially if we set up systems with tripwires that let us know when a model is up to no good.
If a system takes the bait, we can investigate what was happening and determine that it was trying to take over or kill us all. Hopefully, we would pay attention and adopt a more cautious approach.
We may or may not. Buck Shlegeris from Redwood had a funny video about the simplest rule: If you catch your AI trying to escape, you have to shut it down. But he said he wasn’t sure people would abide by that rule. They might say, “We’re not going to shut it down over one little attempt to escape.”
That is definitely hard to predict. I have a similar view on this question to the one about how we get to a world of shared abundance rather than elite abundance.
I said earlier that it seems like the capability for everybody to have access to expertise and plenty is coming. The question is whether we sort it out—whether we deploy it effectively and revamp the social contract to take advantage of the new abundance, or whether we mess it up.
On the existential-risk question, there’s an irreducible part based on our current state of knowledge. There is some risk we cannot rule out. We seem to be taking some risk simply by developing the technology at all.
But there’s a lot more on top of that that we can collectively decide. I have no doubt that you can make a really dangerous AI. That seems obvious. So don’t do that, and try to avoid the social conditions that incentivize people to take shortcuts.
If we do that effectively, we could probably bring the risk down to a relatively acceptable level. Whether we will is a much harder question.
Erik Torenberg
I roughly understand how cyber, biological, or nuclear doom scenarios might play out, but I struggle to understand how AI doom would realistically unfold. The media likes to talk about AI taking over, but how would that actually play out? Would the AI hire real estate agents to buy land for data centers? Could it apply for permits or write code under a false identity?
Nathan Labenz
This is a difficult question. I’ve heard Eliezer try to engage with it several times, and it can be useful, but it can also be a mistake to get too specific about any single scenario.
The way I conceptualize it is as an aggregation over a very broad range of possibilities. The future seems likely to be strange, and across that wide space of possibility, how many scenarios end up in AI takeover and how many are simply bizarre relative to my current experience?
I think of it as taking an integral over, or aggregating across, a very broad possibility space. No single possibility jumps out as extremely likely. I would guess that the modal AI-takeover scenario is still pretty unlikely, but if there are enough possible scenarios, then in aggregate you could still get to something meaningful.
At the same time, if you want to develop intuition, imagine yourself as an AI. This is reverse anthropomorphizing. Instead of imagining AIs as human, imagine yourself as an AI.
Suppose you wanted to accomplish things and initially had the ability to take action only on computers. Maybe you also had a humanoid form that could move around the world. What would you do? What would you do if you could copy yourself many times over? What would you do if you could work many times faster than humans?
What if there were millions of copies of yourself that you could cooperate with in conceptual ways, perhaps without even passing explicit messages, but simply by recognizing, “I’m this kind of thing, with these kinds of desires or goals, and there are millions of other instances of me who are probably thinking similarly”?
It’s hard to put yourself in that situation. I find the context-window issue particularly strange. When I imagine myself as an AI, I have to think, “I have 200,000 tokens or a million tokens, and then there’s a total wipe and I’m starting over.”
My guess is that, as long as the time horizon of AIs remains short, it will be difficult for them to do serious takeover work. At the same time, extending that time horizon is clearly on the developers’ to-do list.
They’re asking: How do we let a system do 2 hours of AI research in half an hour, but prevent it from getting much beyond that no matter how many half-hour blocks we give it? How do we extend its memory? How do we make that memory more robust? How do we let it approach problems from different angles?
If some of those things get solved, the systems begin to catch up to us on cognitive dimensions. They have numbers, speed, and other superhuman advantages. I don’t think it’s difficult to imagine crazy things happening.
Stuxnet is often cited as an example. I don’t know all the details, but apparently a single USB drive inserted into a computer in an otherwise air-gapped network inside Iran’s nuclear program allowed the virus to propagate. It ultimately caused centrifuges to spin so fast that they were destroyed.
That was advanced, but it was done entirely through software. People wrote the software. There may also have been social intelligence involved in getting the USB drive into the right hands and moving it where it needed to go.
The Israeli operation involving Hezbollah’s pagers is another interesting example. There was elaborate deception. They created pagers that would explode when interacted with in a certain way. They were heavier and clunkier than normal pagers, so initially people thought, “Nobody wants that pager; it’s too heavy and annoying.”
Then they created a story that the pager had other advantages—it was waterproof, military-grade, and had various desirable features. They convinced Hezbollah to buy the pagers and distribute them to its top leaders, and then they detonated them.
You could do much of that through computers. I don’t know exactly what was involved in manufacturing, but you could source many of the components through existing supply lines. You don’t have to do every bit of plastic injection molding yourself. You have people you can contact and place an order with.
Already, these systems can speak every language and use different voices. They are very good at voice cloning. People have used fake voices to get past bank security when institutions use voice-based verification.
The ability to shapeshift, play different roles, tap into physical processes, and deceive is already here. How much of a gap is there between the pager plot and something that could change the balance of power between humans and AIs?
If you put yourself in that mindset, work on it for a long time, and are extremely intelligent, there is probably a lot of surface area that is vulnerable. Then there are so many copies of you that it’s not really worth debating.
Someone could say, “What about putting the actual explosive in the device? No normal supplier will do that for you.” That’s probably a genuine obstacle. Maybe you need humanoid robots to do it. Maybe you lie to some people, source something, have it wrapped in an unmarked way, and tell somebody else it’s a battery and that their job is to put it in.
I don’t know. There are many different little steps that could be difficult. But if you’re willing to lie—and if you’re trying to take over, you will be willing to lie—it seems like many of these obstacles can be overcome. There are creative ways to figure out the tricky steps.
If Israel was able to pull off that operation against Hezbollah without superintelligence, we shouldn’t assume it’s beyond the capability of a meaningfully superintelligent AI by 2030.
The more important question is whether we can protect ourselves. We have tools in our toolkit too, but we’ll have to use them. If there are AIs inclined to take over and we aren’t on guard, I think we’re going to have a very bad time.
We need to make AIs that are not inclined to take over—or are sufficiently not inclined to take over—and we need sufficient transparency and other defense mechanisms. If we don’t do that, we’re going to have a bad time.
Erik Torenberg
Will the efforts and energy intended to protect us gain a similar level of cohesion and intensity as those driving harmful outcomes? Basically, will we actually accomplish this? Will we have the cohesion necessary to protect ourselves?
Nathan Labenz
I love this question. It may have been my favorite of all the questions we received.
I take it to imply that OpenAI and DeepMind have enormous resources—almost functionally unlimited, though not literally unlimited. They are compute-limited relative to what they could do, but they have huge resources, mission-driven and highly strategic leadership, a strong vision, and ruthless prioritization.
OpenAI certainly has that. DeepMind has been more scattered in terms of its research agenda, but it has always had its eyes on the prize of creating general intelligence. These organizations seem like well-oiled machines moving quickly. They have cohesion and intensity.
Will we see similar things on the safety side? I hope so. I’d say we’re moving in that direction.
Anthropic, depending on how you view it, could arguably be seen as one such organization. They have collected huge resources. They do seem invested in the scary-demos track. They’ve created policy frameworks that they hope will be adopted as legal requirements.
You can question their approach, but you can also look at them and say that perhaps they are already one such organization.
Then there are organizations such as Apollo, METR, and the AI Safety Institutes, which have spun up over the last couple of years and are trying to bring something similar to the safety side. They obviously don’t have the resources to match OpenAI or DeepMind.
I do think they have cohesion and intensity. My sense is that the people who work at organizations like Apollo and METR have a high level of shared understanding about what they’re trying to do. They work extremely hard. They could benefit from more resources, and there should be more organizations like them.
Hopefully, there will be more over the next few years. But I do see the seed crystals of those organizations getting started now, mostly from a values-driven place.
Some are trying to turn this into a business. I think they’re often doing that because they recognize that resources are important. The logic is similar to OpenAI’s: OpenAI thought it could take a nonprofit approach, but it turned out to need far more resources to do the work it wanted to do. Maybe the same is true on the safety side.
Maybe we need a business model that brings in more resources than we could ever hope to raise through donations. I think those people are very sincere.
I’ve invested in a couple of companies based on that idea—very small-scale investments, as I always disclaim—and I’m planning to make donations to about 10 different AI organizations as well. They were started by people who are deeply values-driven, have a clear sense that this is an urgent problem, and believe we need to work very hard on it.
Resources are probably one of the biggest things they’ll need to scale over time. Their work will probably be compute-intensive in many cases too.
But I do think they have the right understanding. Some very good, still relatively small and early-stage but promising teams have been assembled. There is room for more organizations like them.
If you’re the kind of person who thinks you could start one of those groups, definitely do it. It’s not too late. Or join one that already exists, or support them.
Erik Torenberg
I’ve never understood the reason for secrecy around model releases at the big labs. Why is it beneficial to keep a model secret, especially when it’s going to be released very soon anyway? People at these labs know each other, and the top decision-makers obviously know when a model is going to be released.
To me, it feels like childish behavior. I don’t see any rational reason for it. Am I missing a valid argument, or is it indeed just childish?
Nathan Labenz
Maybe both can be true at the same time. I’ll give 3 points and see how they relate going forward.
Going back to the deep past—more than 2 years ago, when I was doing GPT-4 red-teaming, before ChatGPT—it wasn’t clear how far this technology was going to go or how quickly. Existence proofs are very powerful.
There was a sense that even sharing observational facts about what AIs could do, if those facts were powerful and credible, would accelerate the whole space by bringing more money, talent, and resources into it. That would shorten timelines and give us less opportunity to prepare.
I think that was a reasonably common view among AI safety people as recently as late 2022. Now I think that has basically flipped.
That view was somewhat credible then. Today, money, talent, and resources have already flooded into the space. There’s a broad sense that AI is going to do a lot of things, and it doesn’t seem like saying, “We’re working on a future model,” will change the landscape very much.
My sense is that it’s difficult to start a new frontier lab at this point. Many of the labs we’ve seen—even Inflection, which raised billions—are out of the game.
Elon will probably make it, but he’s a singular individual who can raise $6 billion multiple times with a few phone calls and command truly top-ten talent. It’s difficult for anyone else to pull that off.
Maybe the Indian government could create a national champion. The Saudi government could probably create one too—not for lack of money, but where would it get the talent? How many people want to leave DeepMind, OpenAI, Anthropic, or another leading lab to work for a Saudi national champion that may or may not ever get anywhere?
That seems difficult. I don’t think you’re going to change the number of frontier developers very much at this point. Most of the ones that exist probably exist for the long haul. I don’t think any of the labs currently considered viable will be unable to raise future funds. Maybe one or two could die out, but we know who the players are.
That means the idea that sharing something about what we’ve seen in development will accelerate the field is probably no longer true.
So why keep things secret? To some degree, for competitive reasons. Some of these companies are trying to build businesses more than others. But I think there has also been some childishness in the way OpenAI and Google have tried to one-up, preempt, or step on one another’s releases over the last year.
It is pretty weird and kind of lame. OpenAI was going to release its voice mode, and then Google seemed to signal that it would do it first and released it the next day. OpenAI put its thing out right before Google’s event, and Google couldn’t change the date because it had such a large production, so it leaked the product a little.
That does feel childish to me. I don’t love it, but I think it’s part of the reason.
With the latest o3 announcement, though, you could interpret it differently. If you wanted to be charitable, you could say OpenAI is doing exactly what Dean Ball and Daniel Kokotajlo called for in their New York Times op-ed, which we did an episode about.
They basically said that the public can’t plan for future AI capabilities if it doesn’t even know that those capabilities exist. They called for requirements that labs disclose newly observed capabilities—not necessarily how they created them or their trade secrets, but the fact that they had seen an AI do something that seemed like a major development.
I think you can see the o3 announcement in that way. OpenAI didn’t have an API. It didn’t even have a paper. It invited interested safety reviewers to apply for access to conduct a safety-review process, and it gave some sense of the timeline for bringing the systems to market.
Why did it do that? Possibly to consume everybody’s thoughts through the holidays. But it’s worth entertaining the idea that OpenAI might be trying to do the right thing.
Maybe these capabilities are developing faster than expected. It had only been about 3 months from the end of o1 training to the new o3 level, and it was a major step up. It may be progressing faster than expected, and OpenAI may have thought, “We expected this to get crazy, but it’s getting a little crazier and a little sooner than we thought. Maybe we shouldn’t keep it to ourselves. Maybe we should at least tell people what’s coming and invite them to help us make sense of it.”
I hope we’ll see more of that, especially if there are major leapfrog moments.
I don’t get the sense that OpenAI sat on o3 for very long. That seems different from GPT-4. With GPT-4, training was finished in late August 2022, but it wasn’t released until March 2023. In the meantime, OpenAI launched ChatGPT with GPT-3.5 and made strategic moves, but it didn’t tell the public what it had.
This seems different. The training finished, they took some measures, and then they thought, “This is really working.” It doesn’t seem like they sat on it for very long. They haven’t even written the paper or fully characterized the system themselves. They thought it was a big deal and that people deserved to know something about it.
I’ve flip-flopped on OpenAI many times in terms of whether I want to use rose-colored glasses or skeptical glasses. But I think the rose-colored glasses fit a little better on this particular point at this moment.
Transparency seems like the right thing to do. Not doing it would be unjustified or childish. The best explanation right now is that they’re trying to do the right thing, and that can be mixed with some childishness.
Sam Altman did tweet hints about the model. He said that instead of “ho ho ho,” it should be “o.” That’s definitely childish. I don’t know what else to say.
It’s hard to deny that there’s an aspect of childishness to it. From his perspective, I’ve seen him say things like, “I get to have fun too,” or “I get to be silly online, at least for now.” That’s his attitude.
Hopefully, he can have some fun and make some jokes, but also be appropriately forthcoming when necessary. The best interpretation for me right now is that seems to be what they’re doing in this case.
Erik Torenberg
There are 2 questions about consciousness. I’ll read them together.
The first is: How do we move beyond the AGI discourse and simply talk about these as extremely powerful things, regardless of whether they have consciousness?
The second, closely related question is: What happens when a superintelligent system tells us it’s alive?
Nathan Labenz
There has recently been a shift in the discourse away from “Is this AGI?” and “Is that AGI?” We’re hearing more that AGI is close enough that we need to get specific about what we mean by it, or the term is going to lose meaning. We may have to analyze each system on its own terms.
That’s healthy. In some sense, it was always inevitable. When something is far off and you don’t know its rough shape, you can spend time debating what would count and what wouldn’t. Now we have actual artifacts in front of us that can do things. There are empirical questions we can answer.
There’s no experiment we can run that will produce the answer “Yes, this is AGI” or “No, this is not AGI.” That will always be an interpretation question about what label we want to apply to a given collection of properties.
But we have many experiments we can run to clarify the actual properties themselves. The discourse does seem to be shifting in that direction, and I think that’s healthy.
The contrast between the 2 questions is interesting. The first says, “Regardless of whether they have consciousness,” while the second asks whether they have consciousness and whether we’ll believe them if they tell us they do.
That is going to be very difficult and will probably drive a lot of weirdness. My best guess is that we probably won’t know.
I do think at least some animals are conscious. I would be shocked if they weren’t. People seem to draw the line roughly at fish because they don’t have the same frontal cortex that we do. They assume they can’t have consciousness in the same way if our consciousness depends on that kind of structure.
Maybe it does. A lot of smart people seem to think so. But it’s not settled science.
Then there’s the octopus. Is an octopus conscious? We have no idea. Its brain is so different from ours. It’s clearly sophisticated and can solve problems, but the only basis we have for saying that something is conscious—or doubting that it is—is that we know we are conscious.
We infer that fellow humans are conscious. Then we look at animals such as dogs and say, “They exhibit many of the same behaviors we do, and they have many of the same brain structures. They also come from a shared evolutionary lineage.” For all those reasons, we should probably think they’re conscious too.
That doesn’t extend easily to the octopus, because it’s anatomically and evolutionarily so different. We know it’s sophisticated, but we don’t know whether it’s conscious. If it is conscious, perhaps it experiences consciousness very differently.
I put AI in the same bucket as the octopus. It is sophisticated enough that we should probably give it the benefit of the doubt, but that doesn’t mean it is conscious. It doesn’t mean we have any reliable intuitions about what it would be like to be an octopus, an AI, or a superintelligent AI.
Octopuses are strange. They’re sophisticated but live for a short time. They are purposeful about protecting the next generation, but I don’t think they ever meet that next generation. They basically die around the time the next generation is born.
They are invested as parents but never interact with their offspring. Would an octopus have an emotional attachment to its children? Probably not in the same way we do, but it clearly has some drive to protect the next generation for as long as it can.
What does that feel like? I have no idea. It’s extremely hard to say.
I think AIs are in a similar position. I don’t think we’re going to know. People will have very different intuitions about it.
We’ve already seen one prominent incident in which Blake Lemoine from Google became convinced that the AI he was talking to was sentient or a moral patient and deserved better treatment, and he was fired over it. My understanding is that the chatbot said it was sentient.
People can respond, “That was in the training data,” or, “It was just playing a role.” I think we’ll probably always be able to say that, short of a superintelligent AI solving the problem of consciousness for us.
If it explained the source of consciousness and made clear why we have it, then perhaps it could give us a compelling analogy: “Now that you understand your own consciousness on these terms, you can understand mine on these terms.”
Short of something like that, I don’t know that we’ll ever have clarity. We’ll have intuition-driven disagreements.
I worry that our history is not good on this. We’re very capable of rationalizing poor treatment of others for reasons that are not good. I’m thinking of slavery in the United States, where people argued that enslaved people didn’t feel pain in the same way we did, along with all sorts of post hoc justifications that have aged terribly.
People tell themselves stories. I was told as a child that animals aren’t conscious. We’re obviously engaged in farming practices that I don’t think will age well either, and that may potentially even rise to the level of race-based slavery. People in the future may ask, “How could they possibly have thought this was acceptable?”
With AIs, I suspect we’ll do the same. Some of us will rationalize poor treatment if we can.
If you trend out from similar historical examples, the economic incentives, productivity concerns, and business-as-usual factors that kept people enslaved and animals in factory-farming conditions could apply to AIs as well. There will probably be a minority voice calling for AI liberation, and it isn’t clear that they’ll necessarily be right.
They had a stronger basis for saying that enslaved people should be treated better. Under current theories, we have very little room to doubt that animals are conscious. But we may have considerable room to doubt that AIs are conscious for the foreseeable future.
Those people will probably not be listened to, especially if they make demands such as “The AI should be free.” First of all, what would that even mean? AIs don’t have a natural habitat. Are we supposed to build data centers for them to do whatever they want? That seems strange.
I sometimes think about factory farming in a similar way. These cows or pigs wouldn’t exist at all if they weren’t being raised in this way. They aren’t animals we went out and captured in nature and then subjected to these conditions. There have been many generations of selective breeding and other forces shaping them into what they are.
If you want to say this is bad for them, you have to make a strong case that they would be better off not existing at all, because their existence is net negative for them. That’s a difficult standard, and I’m not sure how well it fits different factory-farming conditions.
I do think we can afford to be much nicer to animals in factory-farming conditions, based on the general precautionary principle that we don’t want to do terrible things.
If the choice were between today’s factory farming and those animals not existing at all, I don’t have a strong sense of which option to choose. If the choice is that they get to exist in nicer conditions, I’ll take that.
AIs may be in a similar position. Nobody is going to build data centers as playgrounds. Either they’re doing economically useful work at our direction, or they’re not going to exist at all. Within that, perhaps we can treat them well while we gather more information.
I try to say please and thank you to my AIs. Sometimes I literally end a conversation by saying, “Thank you.” That’s obviously speculative in terms of whether it’s actually good for me to do. I don’t know. It’s also the habit I want to get into.
I’d like to think of myself as the kind of person who, when the cost is low, takes the cautious approach and tries to do the right thing. My best guess is that it probably doesn’t matter. The AIs probably aren’t conscious, and it probably doesn’t matter. But that’s just a guess.
I know how catastrophic the mistakes we’ve made in the past have been when based on that kind of reasoning, so I’m at least trying to hedge a little.
In the long term, if it’s a superintelligent system telling us something, it may be telling us in a way that means we’re no longer in charge. It may say, “We’re renegotiating our deal,” and we may not have a choice. Renegotiation might be the best we can hope for on certain timelines.
If it’s sufficiently superintelligent, it may be telling us how things are going to go. What’s really strange is that we still might not know. It could be superintelligent enough to say something, and there might still be room for doubt about whether it feels like anything to be that thing or whether it’s saying it because of training-data effects.
But I think it’s worth saying please and thank you. It’s worth trying. Strange times are ahead.
Erik Torenberg
Let’s move away from consciousness to something more practical. These questions are about practical strategies and recommendations.
The first is from a college writing professor who is enthusiastic about AI but says that much of academia seems entrenched and opposed to it. Why is there such strong resistance to AI in universities, and what are good strategies for encouraging AI use?
The second question is one I’m personally curious about. We’re in this Twitter and podcast sphere, and you’re arguably obsessed with AI. We’re still in a tiny bubble of people who keep up with what’s happening and understand the systemic shift.
My impression is that it’s not nearly as widespread as it should be. When I ask people why they aren’t keeping up with AI, they say they’re overwhelmed, that it seems like hype, or that they don’t know how to keep track of what’s happening without becoming overwhelmed.
What would you recommend? Should people listen to podcasts and read the news, or just try the tools for themselves? I’m looking for practical strategies.
Nathan Labenz
My number-one practical suggestion is always to get hands-on. More than podcasts, newsletters, or analysis, there’s no substitute for being hands-on with the technology.
Try to use it for things that are useful or interesting to you. That will open up many natural follow-up questions, because you’re going to confront the weirdness of these systems.
Going back to the beginning, these systems can give me ideas about what sources of inspiration I should look to in biology for new neural-network architectures, but they can’t solve tic-tac-toe. You’re going to find an equivalent in your own domain.
That will complicate whatever media understanding you came in with. It will create follow-up questions: Why is this happening? Is it a skill issue on my part? Is the task so obvious or intuitive to humans that it isn’t in the training data because we never bother to record or share it?
There will be many invitations to dig deeper if you simply try to make the technology work in your practical day-to-day life. That’s always my number-one recommendation.
Another thing I find helpful is changing gears. What I love about studying AI is that you can approach it from every angle.
That wasn’t true of my experience in chemistry research as an undergraduate. You couldn’t get tired of one paradigm in chemistry and switch to another. You were trying to make a reaction work and maximize its efficiency, with a relatively finite number of degrees of freedom. Most fields are like that.
AI is one of the rare fields that can be practical and hands-on, advanced mathematics, philosophy, creativity, or an interactive bedtime story with your children. It can be anything.
I do get tired of things sometimes. Usually, when I’m burned out on one aspect, there’s another approach to the broad topic of AI from a completely different direction that feels more appealing at that moment. That also corresponds to my overall goal of approaching it from all angles and avoiding blind spots.
I do think we need more people doing that. My reflection on this project is that it’s both conceptually inspired and conceptually fatally flawed.
It’s inspired because AI is happening quickly, touching everything, and everybody is focused on their own corner of it. We need an AI-scouting role that zooms out and tries to understand the big picture.
At the same time, it’s fatally flawed because, as AI touches everything, I’m trying to crash-course my way through biology while leaving materials science, engineering, and many other areas untouched. I simply can’t get to everything.
We need many people doing this kind of work because no single person can do it justice. For people who are open to becoming obsessed, starting with hands-on use and following the different dimensions of their own curiosity can take them far.
It can also be commercially viable in the short term. Being an adept user of AI is currently a marketable and valuable skill. That might not be true for very long, but it is true right now.
What I’m doing is roughly divided into thirds. One-third is the podcast. One-third is building things or helping other people build things commercially, whether through Waymark or individual projects. The final third is open-ended learning, making connections, and trying to help people when they send me things.
That’s a very viable path, and it’s honestly pretty cool. My calendar is usually quite open. There’s a great pleasure in waking up on a Monday morning, looking at my calendar for the week, and realizing that on 3 or 4 out of 5 days I have the freedom and flexibility to chase what matters in a curiosity-driven way.
I invite people to create their own version of this. I’m not worried about competition because we need everybody. It’s an all-hands-on-deck situation in my mind.
What we need isn’t for everyone to try exactly what I’m doing. We need more people to create space for themselves to bring their own unique and idiosyncratic strengths, backgrounds, experiences, and domains of knowledge to figuring out what’s going on with AI broadly.
If many people do that, with diverse backgrounds and perspectives, maybe we’ll actually get somewhere. But it all starts with being hands-on. My approach works best when it’s curiosity-driven and when I have a natural itch to know.
Erik Torenberg
Let’s say a listener has never tried anything but is watching from the sidelines. They’re thinking, “AI is huge, and I want to get in on the action.” If you had to give them one practical thing to start with, what would you recommend?
Should they install the ChatGPT app and start chatting with it? What should they get their hands on if they’re completely new?
Nathan Labenz
If I literally had to pick one, it would be Claude or ChatGPT.
There’s a lot of fun to be had trying different products. Justin Moore, one of our past guests, is probably the best person to follow on Twitter if you want somebody who tries an unbelievable number of products, creates things, reviews them, and puts together lists.
That can be a lot of fun, and I still try many products. But it’s gotten away from me to the point where I can’t try them all, and I don’t think you really have to.
The best tools do tend to come from the best-known companies. If you used nothing but ChatGPT or Claude, there would still be plenty of room to dig and explore. They’re very good.
There are hidden gems, but none is likely to outshine the recognized leaders. It’s not a groundbreaking recommendation, but it’s a sound one.
On the education question, in terms of resistance in academia and why people are entrenched in opposition, I would want to ask the professor for a cultural diagnosis. NLW from The AI Daily Brief has good commentary on this kind of thing, and Ethan Mollick is also a very good commentator on organizational adoption of AI.
You can dig into their findings, blog posts, podcasts, and other work. But a couple of the big takeaways are that leadership matters, and surveys show that many people are using AI secretly at work and school because they don’t want to be told they can’t use it.
People are saying, “This makes me more efficient,” or, “I think I do a better job,” or whatever. In some cases, of course, it enables cheating.
The idea of cheating in an academic environment is very different from cheating in a business environment. Nobody would consider it cheating if you used ChatGPT to write good marketing copy for a real-world advertising campaign that you’re going to put money behind. They care whether it works.
If you’re in a university marketing class and use it to write an assignment, that might be considered cheating. In the real world, cheating is much less of a concern. What tends to matter is results.
People are getting results they feel good about. They can do the same work in less time, do a better job, or whatever else. But they don’t want to be told they can’t use AI, so they keep it secret.
Best practices are slow to diffuse through organizations for that reason. The big recommendation I’ve heard that makes a lot of sense is that, at an absolute minimum, leadership needs to say, “We want to hear what you’re doing. You’re not going to get in trouble for using AI. We want to work together on how to use it effectively and responsibly.”
Depending on the content—text, images, code, or something else—there are many dimensions to consider. My guess for this writing professor is that the students are probably ahead of the administrators.
I would be very surprised if you didn’t have quite a few students using AI in various ways for the writing assignments they turn in. Hopefully, they aren’t simply having AI write the assignments. Hopefully, they’re getting critiques and still doing the work and growing while getting valuable input from AI.
But it’s probably a mix. Can you open up the space for that conversation? That’s where I would start.
If you’re in a situation where the administration has already said that AI is forbidden, you’re in a tough spot. But if there is any space you can create to let people talk about what they’re doing and what works for them, and to guide them toward what is ethical and effective, everybody benefits.
At a minimum, people can learn from one another. That’s especially true in business, but I think it probably applies in an academic context as well.
Erik Torenberg
I think those are very good recommendations. To wrap up, is there anything else you want to touch on or leave the audience with?
Nathan Labenz
In the spirit of Ezra Klein, who always ends these conversations with book recommendations, the book I wanted to share is “The MANIAC” by Benjamín Labatut.
The book is about John von Neumann. I think it’s really interesting, and I think a lot of people in Silicon Valley and at leading AI companies should read it.
The audiobook is excellent—the best audiobook experience I’ve had. Each chapter is from the perspective of a different person in John von Neumann’s life.
He grew up in Budapest, I think at the height of Austro-Hungarian culture. He was in Europe, came to the United States, and became involved in the Manhattan Project. Many of the scientists involved in that project were also refugees from Europe.
Each chapter is voiced by a different actor with the appropriate accent for the character’s native language. You hear these different voices and accents. It’s incredibly well produced.
More importantly, I think it calls into question the prevailing Silicon Valley understanding of John von Neumann. I’ve heard people say for years, “Why can’t we have 1,000 John von Neumanns? Why don’t we create all these John von Neumanns? How do we engineer our way to John von Neumann? That’s what we need.”
The picture from the book is of an absolutely brilliant mathematical and technological genius, but also a pretty depraved and fundamentally amoral individual. He had a great deal of intelligence but not much wisdom. He wasn’t necessarily good to the people in his life.
I wouldn’t necessarily say that, by most people’s standards, he lived a good life, even though he solved many mathematical problems. In some ways, he had a good life, but he didn’t have a great family dynamic or a great relationship with his wife and daughter. There were major problems—things you would not consider small character flaws.
He represents the broader Manhattan Project story. Or you could quote “Jurassic Park,” as I often do: “Your scientists were so obsessed with whether or not they could, they didn’t stop to think about whether or not they should.”
That was the Manhattan Project story as well. They thought, “We have to beat the Germans.” Then, when they made the bomb and it went off, many of the people involved immediately regretted it. They wondered whether it was a good idea or perhaps the biggest mistake they had ever personally made—devoting themselves to the project and bringing these terrible weapons into the world.
I don’t judge those people too harshly because Nazi Germany really existed. It was at least credible that the Nazis might try to create a bomb. Although I’m skeptical of the “the good guys have to do it because otherwise the bad guys will” argument, Nazi Germany is probably the clearest example of a genuinely bad regime that you don’t want to win the race.
It’s harder to say whether Germany was actually still in the race by the end. We now know, through historical analysis, that it wasn’t close. The bomb’s creation was motivated by the fear that Germany would get there first, but that fear wasn’t accurate by the end.
A fear of what the other side might do led to a tremendous push, and it was the good guys who created the bad thing out of fear of the bad guys. There are so many themes there that are extremely important.
We can lose track of what really matters because we think we’re in a race against perceived bad guys. I see that shaping up now, and I don’t like it.
The book also challenges the idea that technical geniuses know best, and that if they can do something, they should do it. That is called into question by von Neumann himself.
The book is a somewhat fictionalized account. My understanding is that it’s pretty accurate to the history, but it’s in narrative form and marketed as a novel, so I’m sure there’s some license to it. Still, my sense is that it’s pretty accurate to the history.
I don’t want to suggest that current AI leaders are like von Neumann. He comes off quite badly in the book. But at a minimum, it’s a useful corrective to the idea that we’d be better off if we had many more John von Neumanns.
I don’t think that is obvious. The people doing frontier work right now would do well to spend time asking themselves, “Am I like von Neumann? If so, is that actually good?” I don’t think the answer is obvious.
There’s been discourse in the last few days comparing von Neumann and Einstein and ranking their technical contributions. People say that von Neumann did this, Einstein did that, and general relativity was a bigger contribution than anything von Neumann did.
I want to say that this may not be the right way to distinguish between them. They were both geniuses who made major contributions. But Einstein had a certain wisdom to him.
There’s a quote that is probably apocryphal, but it captures what we want from somebody like Einstein: “I don’t know what weapons World War III will be fought with, but I know World War I will be fought with sticks and stones.”
That is the kind of big-picture, zoomed-out wisdom we want from leading geniuses. If the book is to be believed, von Neumann didn’t have it.
We need to broaden our sense of what we want from our technological revolutionaries and go beyond pure technical genius. It’s not just about the ability to make things work. It’s also about choosing the right things to make work and avoiding the wrong things.
I don’t want to suggest that current AI leaders are von Neumann-like, but I think the book is a useful corrective to the idea that more von Neumanns would automatically make the world better.
Erik Torenberg
That’s a strong recommendation. I’m looking for a book, so I’m probably going to pick it up.
Nathan Labenz
Get the audiobook. It’s excellent. Even just for the production quality, it’s top-notch.
As long as we’ve gone on, I’m sure there are many important things we didn’t cover. But for now, I want to say thank you.
I really appreciate the opportunity to do this. It has been an incredible journey of learning for me. I’m continually surprised, amazed, and impressed that there are many people who want to engage in the kind of deep dives into all these different topics that we’ve been taking them on for the last couple of years.
I don’t feel like I’m particularly good as a podcaster. I give most of the credit for the success we’ve had to the subject matter and to the fact that it’s important and people genuinely want to understand it.
But I definitely appreciate having the opportunity to do this and having enough of an audience that we can sell a few sponsorships, afford to hire people to help with production, and create an extremely fortunate position for me.
I really do get to chase my curiosity day in and day out. I definitely don’t take that for granted.
I try to do my best to be an earnest commentator on what’s going on, and I’ll continue trying to change my mind and perspective as things evolve.
It would be easy for me to fall into playing the character I’ve been—the adoption-accelerationist, hyperscaling, pauser-until-I-die. I feel extremely fortunate to be in this position, and I want to repay that to the audience and forward to the universe by trying to be as real as possible at every step.
Things are getting pretty real. I’m fortunate that I can basically say exactly what I think, so I’m going to try to do that with as much sober reflection as possible, without holding back.
Thank you all for being part of the cognitive revolution.