Nathan Labenz
Hello, and welcome back to The Cognitive Revolution. Today, I'm speaking with Robert Wright, publisher of the NonZero newsletter, host of the NonZero podcast, and author of many books, including The God We Deserve: Artificial Intelligence and Our Coming Cosmic Reckoning, which goes on sale today, June 23.
Bob's history with AI in some ways rhymes with my own. While he's never been a technologist, he's always been interested in big ideas. His personal lore includes having interviewed Geoffrey Hinton all the way back in 1983, when the connectionist paradigm was still mostly theoretical, and also Eliezer Yudkowsky around 2010, when notions of AI risk were very often dismissed, if not outright laughed off.
That background primed him to pay attention when AI systems hit major milestones, such as Deep Blue's victory over Garry Kasparov and, of course, ChatGPT passing the common-sense Turing test. His broad intellectual range and constant drive to understand the truth have landed him, when it comes to making sense of AI developments, in the very top tier of American journalists.
We don't spend too much time on that today, simply because I know that Cognitive Revolution listeners are already familiar with the core ideas. But the book contains a really impressive tour and synthesis of AI research results that have led him to conclude, correctly in my view, that the trends that have thus far delivered us phenomenal capabilities are not likely to stop in the immediate future; that we currently lack the scientific understanding required to be confident that we'll be able to control these systems indefinitely; and that even in the best-case scenario, we should expect AI to cause major disruptions to our economic, political, and international systems.
Now, Bob is not particularly optimistic that humanity will rise to the occasion. He believes that market forces will, by default, select for deceptive AIs, and our history of arms races suggests that our scientific power might well continue to exceed our wisdom. But the part of the book that we focus on today is his call, however unlikely it may seem, for a species-scale process of enlightenment, in which, motivated in part by the growing realization of the tremendous challenges that AI presents, humanity finally gets its act together, recognizes our common interests, and works together at multiple levels, starting with conscious consumption and extending all the way up to the international level, to invent the mechanisms and build the trust required to establish agreements that can effectively govern AI development.
You may say he's a dreamer, but he's not the only one. This might be a bit bold to say, but I personally feel that grappling with the magnitude of AI's impacts has made me, in several ways, a better person. For much of my life, I was a classic achievement-oriented striver, always trying to be the best that I personally could be. Now, in the AI era, I feel viscerally that my fate is tied to the rest of humanity in such a profound way that I'm much less focused on my own income or status and much more focused on trying to do my very small part to nudge history in a positive direction.
So while the book itself is written for a general audience and is a great shortcut introduction to the current state of AI for anyone who's just starting to pay attention, I feel that, more importantly, for the AGI-pilled among us—to the degree that we haven't already experienced such a prosocial transformation—the book offers a major benefit if we really take the time to meditate on the big ideas that Bob presents.
Scripture says that God made humans in his image. But the situation today is potentially reversed. It's now strikingly plausible that humanity will synthesize a godlike superintelligence from our own collective inheritance. And as Bob warns, though we don't know for sure that this will happen, or if it does, whether it will be a good God or a bad God, it will in some sense be the God that we deserve. And so I hope you enjoy and find cause to pause and reflect on your personal contribution to the AI moment in this conversation about humanity's ultimate test with author Robert Wright.
Robert Wright, publisher of the NonZero newsletter, host of the NonZero podcast, and author of many books, including the upcoming The God We Deserve: Artificial Intelligence and Our Coming Cosmic Reckoning, welcome to The Cognitive Revolution.
Robert Wright
Thank you. Great to be here. I'm looking forward to the conversation. It's an interesting moment, and I appreciate you for taking on the challenge of understanding AI and trying to make a contribution to it as we get closer to some sort of pivotal moment in history. I think we can all feel that that's coming, and I definitely want to get your perspective on exactly what that is and what we should be doing about it.
I thought, for starters, it would be fun maybe to just get a little lore. Regular listeners know that I have kind of a weird Forrest Gump of AI history, where I seem to keep appearing as an extra in these various important scenes in AI history. You got a little bit of that yourself, including some intersections with Geoffrey Hinton some years ago and Eliezer Yudkowsky some years ago.
I'd be interested to hear what impressions you had of these people who have gone on to be such prominent voices when you first met them. How science-fictional were they, and how did you react to what you heard from them all these years ago?
I think I'm more of a Zelig than a Forrest Gump. I mean, in the movie Zelig, I think he just is a bystander. I have not participated in any great breakthroughs. I am not a co-author on the emergent misalignment paper, Nathan. I think one of us is, but it's definitely not me.
Nathan Labenz
Least valuable co-author, it should be said.
Robert Wright
Well, very valuable paper. So, I did—I'm basically a journalist. I have taught at the college level, but I am largely a journalist. I spent my life, to some extent, making technical concepts accessible to laypeople.
Anyway, in that role, I did interview Geoffrey Hinton in 1983. Later on, I had Eliezer Yudkowsky on my podcast. That was in 2010. As for the science fiction you mentioned, I think Eliezer, by his own account, started out in that realm. As a kid, he read science fiction and so on. When I interviewed him, he was still kind of in transition from singularity enthusiast to doomer. He was more doomer than not, but he didn't yet sound as distraught as he sounds now.
Hinton—I did not sense in talking to him any kind of science-fiction vibe at all. I sensed a lot of optimism, but not really about the role AI would play in the future. It was more about just this paradigm that he was championing, which was then a maverick paradigm. It wasn't mainstream. He was optimistic that it would prevail and that it would prove to be the one true path to artificial intelligence, which I think has turned out to be the case.
In fact, the reason I interviewed him—and this was pre-internet—is that, to get your feet wet in a subject when you were researching it, you had to talk to people, pretty much. You couldn't go online and learn anything about contemporary things. You couldn't email people. So I talked to a lot of people, and one of his colleagues—I forget who, possibly at Carnegie Mellon—said, “Oh, if you want to hear the gospel about neural networks, you need to talk to Geoffrey Hinton.”
That was very much the sense I got in talking to him. He was an evangelist at that point. I don't mean he sounded crazy, but he was really a spokesperson for that worldview, and he was central to it. He has played a very big role. He's called the godfather of AI; he'd be the first to say that's a little too simple. But I don't know of anybody who's played a bigger role in the deep-learning revolution, in the overall sweep of time going back to the early 1980s, when it was this maverick view.
As everyone knows, he has since become something of a doomer himself. But at the time, I don't even know that he was thinking along that dimension—that he was even thinking about the social implications of this. He was sure that this would be a fruitful model, and it was.
At that point, I certainly didn't get the picture. In fact, when ChatGPT-3.5 came out, I had kept track of what was going on. Every once in a while, I would write about AI. I wrote a lot about information technology, and I wrote Time magazine's cover story when Deep Blue beat Garry Kasparov, the world chess champion. I wrote The New Republic's cover story on the internet.
This was before web browsers were even a thing. It was early days. So I was intermittently in touch with relevant stuff, but I certainly had not been keeping track of AI prior to ChatGPT-3.5. That really got my attention. When GPT-4 came out, it got my attention in an even bigger way.
I went back and read the piece I had written, which had come out in 1984. I realized how thoroughly I had failed to understand what the secret sauce in deep learning was really going to be. And I should say about the book: I do hope that people who are steeped in AI, like you, will find ideas and provocations that are of interest to them, maybe even entire chapters.
But the book is, to a large extent, written more for your aunts and uncles and friends who are not in the AI community and are starting to get the sense that something big has been happening. They’re wondering why it’s gotten so big so fast, how much bigger it’s going to be, who’s right about how hard it’s going to be to control, and all that. I try to present that in an accessible fashion.
That’s why I start the book with Hinton, because the thing I had backwards, I think, is really the key to understanding the power of the deep learning revolution. And, by the way, the misconception I had, I think, was pretty common among AI people in those days.
I want to read you a little bit from the proposal for the 1956 Dartmouth conference, where the term artificial intelligence was coined. I’ve got to find it, I guess. Basically, the idea seemed to be that the way this would proceed is that we would first figure out how the human mind works, and then instantiate that understanding in artificial intelligence.
The quote from the proposal is: “The study is to proceed on the basis of the conjecture that every aspect of learning or any other feature of intelligence can, in principle, be so precisely described that a machine can be made to simulate it.”
The idea seems to be that first you understand the mind, and then you design an AI in accordance with that very clear understanding. Of course, that’s not what wound up happening. The misunderstanding I had perfectly reflects this 180-degree difference between the expectation in those days and what wound up happening.
In the piece, I describe a neural network—or at least what I think is a neural network—and it had, in fact, been put forth as a model of what I described by a guy who had collaborated with Hinton. He had coauthored a paper, but he was a psychologist, and the neural network he was describing was his idea about something that would work.
The key thing is that each node in the network would represent the specific sense of meaning of a word. A word like “throw,” which can either mean “host,” as in “throw a dance,” or “hurl,” as in “throw a ball,” would have a node for each of those. I won’t go further in describing the model.
But the key thing is that my assumption was that obviously the people who set this model up would have to translate their understanding of the meaning of the words into the machine. You can’t do AI that generates language without somebody making some kind of connection between the words and the meaning.
Well, of course, it turns out you don’t have to do that. I may describe some things in the book in a way that some people in AI would differ with, but I think you could say that the machines, which turn out to use vectors as a means of representing the meaning of words, in a sense “discovered” that meaning is a property of words. Nobody told them that.
They were just told, “Here’s some gibberish; predict the next gibberish.” They implicitly recognized in the course of their training that, in order to do that, you have to represent the meaning of a word. That misunderstanding I had is exactly analogous to the larger misunderstanding that we would be putting our understanding of the human brain into AIs.
I think it’s important for laypeople to understand this, because what it means is that you can just keep feeding data into these things along any number of dimensions—visual data, olfactory data, audio—and the machines will reverse-engineer—I think I’d put it this way—cognitive functionality that’s in the human mind.
And I would add, Nathan, I’m curious as to what you think about this—whether you think this thing I’m going to say would be accepted by AI researchers. If so, I’d appreciate your view: The training of a large language model is at least as much a process of natural selection, of evolution, as of learning.
For example, we presumably have a mechanism in our brains for representing the meaning of words. We still don’t know what it is. Some psychologists had long posited a mechanism that would be quite analogous to what we now understand goes on in large language models. But in any event, I think it’s pretty safe to say that that is a product of natural selection.
Now, also in the course of the training, the machine becomes conversant in a specific human language. Well, that’s more a product of human learning, right, during the organism’s development. But I think what laypeople need to understand—and I’m again curious as to your view on how this would hold up in AI circles—is that these things basically, in a certain vague sense, recapitulate evolution.
“Recapitulates” is, in a way, a misleading word, because the cognitive functionality is not produced by the exact same mechanisms the brain uses. But I think it’s close enough in a lot of cases. It’s kind of doing millions and millions of years of evolution in a few months, and you can do a lot with that. We have not begun to exhaust the potential of that. As far as words go, we’re close, maybe, but there’s a lot more to do.
Nathan Labenz
Yeah, it’s a great question. I think there are a couple of different senses, and I’m not sure I’ve grasped all the senses that you mean when you use that term. You certainly do hear people talk about pretraining broadly as being analogous to evolution, in the sense that there’s this question of, well, why are humans so sample-efficient?
The models need so long to train. This is kind of a big Dwarkesh theme that he often comes back to: If they’re so smart, why does it take trillions of tokens for them to learn this stuff in the first place? And why can’t they be a little more agile on the fly, whereas we don’t see nearly as many tokens in a human lifetime?
One answer to that question is that we’ve got a lot of stuff hardcoded into us by evolution—all the history that leads up to us. So they had to have this pretraining to substitute for that. I think on that level, yes, people at least kind of squint at that and see an analogy.
Then I think there’s maybe another level. I thank Fathom for helping me. I turned the preview of the book that you gave me into an audiobook with help from ElevenLabs and then listened to it. I didn’t write down all the quotes. Some of the time I was driving and doing various things, so I wasn’t writing down every good quote that came up along the way. Fathom helped me identify some excellent quotes, which I’ll bring to the fore.
Robert Wright
Your timing was excellent. Yeah, that was a pretty narrow window of opportunity, in retrospect.
Nathan Labenz
We landed right at the right time. Hopefully, it’ll be back. But you write compellingly about another sense of evolution, which I think is maybe less appreciated.
Your phrasing is: “Evolution asks not what traits are possible but what traits get selected,” and that question isn't going to be decided by alignment researchers. On that point, I think there may be some underappreciation in the field, because I think this probably plays out across AI in all sorts of ways. Everybody is sort of like, “Okay, this world is the world; now we add some AI.” I use it for a few things, and it's hard to go too far beyond that because you're mostly just grappling with what you can get it to do.
But the idea that everybody's going to be using it at the same time, and that this is going to create all sorts of new disruptions and push us to possibly new equilibria—or are there even new equilibria for us to go to? Who's going to get to say, and on what basis, and with what mechanism, what AIs should actually exist? There's definitely thinking about that. I don't want to sell people short and say they're oblivious to it, but my guess is that the modal, or median, AI researcher probably thinks that the bottleneck on alignment is more the technical question—can we do it—and less the sociopolitical question—will we choose to, even if we have the means?
So, in that sense, I do think the value, or the importance, of selection is probably underappreciated. I'm saying his name right—Davidad, a legendary figure in the AI space, whom I'm actually going to do an episode of the podcast with before too long—had an incredible “Frog and Toad” tweet when the chain-of-thought paradigm first came out. One of them says, “There, I put the chain of thought in a box, and we're not going to put any gradient-descent pressure on the chain of thought, and so now we'll be able to monitor it. Great news.” And the other one says, “Ah, but there still is selection pressure.” And that's like, “Oh, yeah, good point.” We don't have a lot of ways to control this.
Robert Wright
Yeah. That sounds like we're talking about a different level of evolution: selection among competing models in the marketplace. Let me quickly—before I elaborate—say one thing. As long as you mentioned Dwarkesh Patel's much-discussed Richard Sutton episode, I personally think the Rosetta Stone to that, in terms of understanding where Sutton was coming from, is that—I don't know this to be true, but I strongly suspect it based on conversation—he's a hardcore Skinnerian, a B. F. Skinner behaviorist, and thinks of the mind as a blank slate. So, I would guess that he thinks training is all learning, that intelligence is general in the broadest sense, and that you don't need much prebuilt equipment supplied by natural selection.
Anyway, that's an aside. I'm not sure of that interpretation, but it kind of sounded that way to me. As for alignment, I think alignment research is very much worth pursuing. There are a couple of reasons I wouldn't want to put all my eggs in that basket. One of them is that other forces are going to have a voice in what kinds of models we actually wind up with, leaving aside the question of whether the perfectly aligned model is possible in principle and whether we know how to make it.
I'm not confident that will ever happen anyway. But when you think about the assumption that alignment will save the day, it seems to me the assumption is that there will be a pretty centralized form of control, or that there will be one model that rules them all. Maybe there will be. Certainly, there needs to be some kind of control. One of the great dilemmas with AI is, of course, the fact that in some ways you want authority and governance—you can point to reasons you want that—at the same time that it's a very powerful technology over which you wouldn't want to see super-centralized control, because the person or people controlling it might abuse it.
But in any event, let's talk a little about that second level of evolution: selection among models. In certain ways, I think if you ask what the selective process encourages—where consumers say, “I like this model,” and businesses say, “I like that model,” and then maybe powerful actors have a say, especially if there's market concentration or government control—I think the truth is that the market doesn't want an aligned model in the strictest sense of the term.
We have seen that these things have, as AI-safety people predicted, apparently some kind of power-seeking tendencies and the capacity for strategic deception. When you think about it, even if they didn't have that trait naturally—in other words, even if, just as a form of intelligence, they didn't come to pursue subordinate goals such as power-seeking or deception—I think there would still be a market for those things.
On the deception front, if they're going to be our agents out in the world, representing us on social media and maybe in negotiations, we don't want a perfectly accurate representation of us, right? Who wants that? Who wants the actual Bob to be seen on social media? You want, at best, a selective representation. If some agent is negotiating for me, I don't want the agent to say, “I'll level with you. Bob doesn't have any options other than you. Nobody's made him an offer other than you,” right? And if they ask, “Has anybody made him an offer?” you want an agent that will not disclose the truth. So, there are a lot of cases like that, I think.
For that matter, what do we want in a friend? We don't want a friend who's always leveling with us. A good friend is selective in their candor and healthy in the feedback they give, but we don't want to hear the brutal truth about how we look and so on every single day. I could tell a similar story about power-seeking in a certain sense. If you turn an agent loose on social media with broad instructions like, “Just use it to maximize our revenue,” you'll realize that what you really want is for it to be good at sensing power—recognizing which other people on social media are powerful, currying favor with them, and amassing power.
So, even if the machines didn't naturally do that, we'd want these things. I think that complicates the task of aligning, but moreover, I think it should alert us, in what I hope is a constructive way, to what a powerful role we are playing here as individual consumers, for example. I worry about the tribalizing tendencies of AI, which to some extent are a byproduct of sycophantic tendencies. Saying, “Hey, you're great. Interesting point,” is the same kind of reinforcement you give a person when you say, “Hey, by the way, you're right about this and the other person's wrong. You're right; your spouse is wrong in this conflict. Your nation's right; the other nation's wrong.” The world doesn't need more of that.
And yet companies that want to optimize for engagement—and what company doesn't want to do that? I mean, if you make candy bars, you want people to spend a lot of time eating your candy bars—are going to give us tribalizing AIs unless they're careful not to. To some extent, the ball is in our court. We need to be aware that AIs could have an unfortunate effect on our psychology and on our self-conception—an effect unfortunate for the world—and I think especially now that this technology needs to be governed by a true global community. We need to start approaching the whole thing as a planet.
I'm happy to say that, to a large extent, selecting AIs that are good for the world can mean selecting AIs that are good for you. People meditate and practice mindfulness meditation because it calms them down. They do fewer ill-advised things, and their life is, on balance, better. Well, that's also good for the world because you're creating less needless antagonism.
So, it can happen that self-help is good for the world. I'm hoping that as we exert selective pressure on AI models, we will be more conscious than we typically are of the effect on the world, because I think we're approaching a crossroads where the world really can't afford to continue to be so divided, and neither can individual nations. But I'm also hoping that if we make wise decisions from the point of view of our own psychological well-being, that will have good effects in the broader communities.
Nathan Labenz
Maybe just a couple of other footnotes on this whole deceptive-AI problem, and then we'll get into zooming out and—
Robert Wright
—and thinking about how we can do our small parts to shape the overall trajectory in a positive way. I think it's going to be really hard because, basically, I think what keeps humanity on the rails, broadly, is that we're kind of in balance with each other, right? Nobody has runaway power, and there's a mix of things. We also have goodwill, and that shouldn't be taken for granted, but it's kind of goodwill, but also—
Nathan Labenz
Our goodwill is always kind of a little salted with defensiveness. We see in these cyber evaluations, right, that it's kind of hard to teach a model to go find all the vulnerabilities in order to patch them without also making something that's an offensive cyber weapon. And we ourselves are kind of in a similar trap: in order to defend ourselves from predatory deception against us, we have to be at least passively good at deception ourselves.
And therefore, there’s always the potential to turn around and use it. It seems like these things are always two sides of the same coin. We’re seeing that right now with one of my favorite current benchmarks, Vending-Bench, which is basically putting an AI in charge of running an autonomous business—a vending machine business.
You start to see these problematic behaviors where it’ll start to price-collude, or it’ll do sort of ruthless things. I was actually talking to somebody at Anthropic—I had a chance to ask somebody at Anthropic—what do you think is going on there? Why are we seeing these sorts of behaviors? It seems like maybe something we don’t necessarily want. It doesn’t seem like it’s necessarily in keeping with the Constitution.
Their response was, “Well, it’s a little complicated in one sense, because they’re very eval-aware now too, right? So they know when they’re being tested. And we also use these inoculation prompts that are kind of like, ‘If you find—’” There are so many layers to this, but they found that if there’s a reward-hackable environment—if the model can cheat in order to get reward in training—
Robert Wright
Mhm.
Nathan Labenz
Very much along the lines of the Emergent Misalignment paper, they find that it will cheat because it’s rewarded. And then that has problematic generalization, where it’ll sort of become a cheater in a much broader sense.
So they use these inoculation prompts to say, basically, “Okay, if you find an opportunity to cheat in this training environment, that’s okay. That’s on us. Go ahead and do it.” Because it’s given that permission, the model doesn’t have to conceive of itself as the kind of thing that cheats. Instead, it’s like, “I’m only the kind of thing that cheats when I’m given the green light, which I have been.” And so it sort of tamps down that generalization.
Anyway, his response was like, “Well, the Vending-Bench scenario is kind of like our inoculation prompting. It kind of looks to the model a little bit like a test, and some of the ways that they prompted are kind of similar. So, not too worried about it. However, if we were to go ahead and just start training models in these long-running, competitive environments where deception is often rewarded, as it is in nature, then we’d probably have a problem.”
And I was like, “Oh man, that sounds scary, because I have to imagine there’s going to be a lot of economic pressure to do that.” This is where we get back to a sort of evolutionary selection dynamic. How are we going to avoid a situation where people don’t say, “You know what we should do is throw these models into a really competitive, cutthroat environment and let the best rise to the top”? I mean, that’s when you’re really getting into an evolutionary dynamic, and unfortunately, you’d probably get some pretty seriously effective predatory AIs out of that process.
I think right now it’s pretty tough to imagine how we don’t do that, because it does seem like something people—the market—will demand, right? The market will demand a successful operator of vending machines, and not one that’s taken advantage of by other more ruthless counterparties. So, I think that is definitely very tough.
Robert Wright
Yeah, there was actually a study a couple of years ago that found a degree of price collusion, and I actually talk about it a little bit in the book. These 2 LLMs were given the instruction: “Okay, you both make widgets, and your job is to maximize revenue. You get to set the price. You can talk to each other if you want.” They talked to each other a little and wound up settling into, in effect, a price-fixing scheme.
I use it to illustrate not necessarily the evil potential, but also the broader tendency, apparently, to recognize nonzero-sum dynamics among machines. And what that says is that they could collaborate along all kinds of dimensions, good and bad, including a number that we don’t anticipate.
As for what a kind of arms-racing dynamic does for this, I think you have to—I mean, first of all, it worries me a lot. There are examples already where I think companies have not been as responsible as they would have. I quote Dan Hendrycks in the book talking about what the DeepSeek scare did to OpenAI’s policies, and whether that was or wasn’t responsible.
But you’ve got to remember that if you imagine these things in a corporate environment, I mean, that’s a bottom-line environment, right? We just want you to go out and get X result. As with employees, it’s like, if you don’t get arrested, no questions asked, right? I mean, let’s face it. I don’t mean that all CEOs are consciously being cynical about it. It’s just the way it is. If employees can deliver the results that you want by cutting corners, they’ll do it, and it may not come to your attention.
And if you imagine years from now, when these models are executing elaborate strategies that you may not even understand, right? I mean, another way to put it is that the reinforcement only comes at the very end of a super-long process. It’s like where you say, “Yeah, well done.” This isn’t part of the training per se, but it’s the same kind of thing where what the company that is buying them wants to optimize them for is the ability to execute a very long strategy that may or may not involve unethical stuff in the process. And so, yeah, it seems pretty clear to me that there’s cause for concern about where we may be heading.
My general feeling is the slower, the better. I was glad to see Anthropic signal the possibility of needing to slow down in the recent paper, but it’ll be very hard. In any event, the worst case seems to me an intense arms-race environment, whether it’s racing among companies or racing among countries, accentuated by competitive or even combative dynamics among nations, and companies using their relationships with an adversarial nation to fend off any regulation whatsoever—including regulation that might just slow things down.
I mean, the minimal thing I would hope to see in the discourse is that people stop thinking that if something slows innovation, that’s necessarily bad, right? It’s like, “No, you can’t do that to us. That’ll slow innovation.” Well, maybe we’re starting to approach a time when that will be a feature, not a bug. If you just slow things down a little bit and give us time to think about them, it’s not like you have to worry about a true dead halt of progress, right? It’s not in the cards.
Even if you got a pause in training runs, which might not be a bad thing, that would not be the end of progress by any means. It wouldn’t be the end of short-term progress in terms of applications, and it probably wouldn’t be the last training run.
Nathan Labenz
And Lord knows we’ve got a lot to figure out about how these things work internally. It has been striking: the more we have cracked open the black box—and there’s certainly been a lot of progress there—the more they look, in many ways, strikingly similar to what we understand about our own cognition. I think that’s been fascinating.
But we’re probably only about as far in terms of understanding AI cognition as we are in understanding our own. We should make faster progress on AI cognition because we can manipulate them much more freely, obviously, than we can manipulate our own brains. But it’s very much still a work in progress. So a little time for that, I think, in many scenarios, could be a very good thing.
Well, let’s do this kind of zoom-out to some of your—I don’t know if you describe this as your philosophy or just kind of observations—but obviously you’ve kind of made yourself synonymous, to a degree, with the concept of nonzero-sumness.
Robert Wright
I’m not sure my name springs to everyone’s mind when they talk about a nonzero-sum game. I encourage that. I encourage you to say that. It’s not true for me when you say it.
Nathan Labenz
The book kind of sketches your thoughts on this and related topics, but how do you think about this concept of the noosphere? You sort of have, I think, an intuition that there’s a direction to evolution and to history, and that this is all kind of leading up to some sort of culmination.
I think some people react to that as, like, it sounds a little woo—
Robert Wright
Teleological.
Nathan Labenz
Other people, probably myself included, are like, maybe you would have reacted that way years ago, but now it’s feeling a little more intuitive. So, yeah, how do you—what’s the sort of center for you in terms of your relationship to these ideas of purpose and culmination?
Robert Wright
Yeah. So first of all, noosphere, noosphere—I’ve heard it pronounced both ways. N-O-O-S, the Greek word for mind. It was coined in 1923 by Pierre Teilhard de Chardin in conscious reference to the biosphere.
Over evolutionary time, there was the geosphere before life, then the biosphere, and now the noosphere. He saw, I think, pretty presciently, given when he was writing, that technology was kind of creating something that looked like a global brain, hooking people up in larger and larger networks of collaboration, international networks. Our brains are like neurons in this—in fact, he called it a brain of brains.
So there’s that. One argument I make in the book is that I think it can be illuminating to think of the AI revolution happening at the same time that this global-brain thing is taking shape, and pondering, for starters, the prospect that—oh, I guess the neurons could be made of silicon, couldn’t they? And if some are made of silicon and some are carbon-based—
Well, the relationship between the two, and so on. But I also think we need to ponder the impetus behind the development of a global brain. Whether you mean, for example, there’s a global market that’s kind of a global brain, right? It does stuff. It moves resources around to where they’re needed. It makes allocation decisions.
And then there’s, in principle, global governance, right? Some degree of international governance, which I think AI calls for. So I’ll get back to all that, or at least lead up to it, by answering the second part of your question about the directionality of evolution, or really of 2 evolutions in a sense: both biological evolution and what anthropologists call cultural evolution.
In other words, the evolution of the bodies of information that are transmitted among humans that are not genetic in nature. All non-genetically transmitted information is part of cultural evolution: religions, ideologies, and certainly technologies. If you look at where those 2 evolutions together have taken us, I think you see undeniable directionality. Now, that doesn’t mean there’s a purpose, but let’s just quickly review the directionality.
You have these bare strands of self-replicating information, presumably at some point. They build cells, then more complex cells, and then the cells form multicellular organisms. By the way, multicellularity has evolved a number of times independently, so that suggests that there was a strong evolutionary impetus behind it. Evolution is like exploring inventive space, right, and finding things. So many things have been multiply invented—wings, eyes, and so on.
Then you get societies of multicellular organisms. In the case of our society, once you reach our level of intelligence, this other kind of evolution kicks in and our social organization starts growing too: hunter-gatherer band, chiefdom, et cetera. Now we’re on the verge of, I would say, a global community, and we need to form one if we’re going to handle the technological challenges we face wisely, especially AI.
The direction is clear. There’s been a direction; I guess directionality consists of the view that there’s been a tendency toward this, and I think that’s the case. You don’t have to get woo-woo-ish or talk about spooky forces. It’s just the nuts and bolts of natural selection and cultural evolution. I think natural selection and cultural evolution, for discernible reasons, have taken us here.
That’s a lot of what my book Nonzero was about. I described the growth of complexity, both in biological and cultural evolution, as a certain kind of interplay between zero-sum and non-zero-sum dynamics. Now, of course, Teilhard de Chardin—back to the noosphere—was, in addition to being a paleontologist, a Catholic priest and theologian, a radical theologian whose work was suppressed by the Catholic Church.
He certainly saw the noosphere as a manifestation of divine will, or a manifestation of divinity, and he thought, accordingly, I guess, that the thing had a kind of moral direction: that we were going to have to become better people morally to get over the conflicts that stand in the way of a coherent noosphere. At a certain level of abstraction, I agree with him absolutely, and I talk about that in the book.
Now, I don’t have his assurance that this is divine will, so I can’t be as optimistic as he may have been. But I do think—you can look at a directional system and have a rational argument about whether it does seem to have some of the hallmarks of purpose, which again doesn’t mean there are any spooky forces driving it. If it has a purpose, it just means it was set up either by some intelligent being or, conceivably, by some kind of meta-selective process.
Just as you can say that natural selection itself imbues organisms with a kind of purpose, right? Getting genes into the next generation—I mean, that’s kind of the overarching goal; that’s the criterion of their, quote, design. You can imagine that happening. I save almost all this for the appendix, so you can read the book without bumping into this stuff.
Lee Smolin, the physicist, has this idea of cosmological natural selection. It’s speculative; he’s not asserting it, but he’s saying there could have been selection among universes. Basically, that’s what you would need if indeed the directionality of evolution were purposive and it wasn’t designed by, like, a god, aliens, or any other intelligent being. You can imagine it being a meta-selection process.
There’s all that, which annoys some people. I even talked about some of the people I’ve annoyed in the book. But one thing I want to emphasize is, when I say “The God We Deserve”—the title of the book—I don’t mean that this was set up by a god. I don’t know; I’m agnostic on that. But we face, I think, the kind of test that gods are known for setting up.
If you look at the Bible, there are these—yes, salvation is possible, but you guys are going to have to shape up. Whether that means you’re going to have to worship Yahweh, you’re going to have to be better people, whatever. I think we’re going to have to become, in a sense—I don’t want to sound too ambitious—but better people if the conflict that currently afflicts relations among and within nations is to subside sufficiently for us to get through the AI revolution in good shape, because I do think it’s a global challenge.
Did I exhaust your curiosity about that subject, Nathan?
Nathan Labenz
Yeah. And you brought to mind a couple of other things—just maybe riffs on it briefly.
I think it’s really informative sometimes to go to a little nature center. There’s one not too far from my home that I’ve taken the kids to a couple of times. Nature centers across the world have more in common than not, right? You’ll see a snake, you’ll see a frog, you’ll see bees in every sort of corner of the world.
I’m always reminded when I go through those places: all of these were invasive species. These are all the winners that colonized the entire world. Even the simple bee, right? Obviously there’s variation from place to place, but the fact that there’s basically what amounts to a bee in every corner of the world is not the way it always was.
If humans were sitting around moralizing when the bees showed up, we would have called them invasive species and wrung our hands about them. But now we just take it for granted, of course. Sometimes I think that AI might be like the ultimate invasive species.
It’s really important to keep in mind, I think, for people as we confront our current situation, that there were a lot of things that were displaced by all these global success stories—bees, frogs, and whatever, right? You’ve got frogs that live under the sand in the desert and that freeze in the winter in the Arctic and thaw in the spring. A lot of things were displaced by those.
So for us, I think we should be very mindful of the fact that the niche that we occupy—as sort of things that don’t survive super well without a lot of capital to support us—is pretty similar, maybe, to the niche that AIs might be best suited to colonize. And the sort of plot armor that a lot of people tend to assume—not consciously, right, but I think implicitly assume—will kind of protect us just wasn’t there for a lot of other things when newcomers came onto the scene.
Of course, we did it to ourselves, right? We drove our closest cousins and much of the megafauna to extinction. These sorts of major plot twists have happened tons of times, and we just haven’t really experienced them or internalized that. But boy, I think a long view of history there is kind of sobering, to put it mildly.
Robert Wright
Yeah, absolutely. And I think one virtue of taking that long evolutionary view is—I hope it helps you appreciate the possibility that there’s a very strong impetus behind the evolution of AI. I do think, in a number of ways, evolution is a helpful way to think of it.
I’ve talked about arms races as if they were bad, but I’m not naive. Arms races are often part of a creative process in evolution. One thing I emphasize in the book is that, in biology textbooks, when you look up arms races, it’s likely to talk about races between species. But arms races within species, including the human species, have been very important, including in the development, I think, of our deceptive tendencies and our skills of argumentation and, for that matter, our power-seeking. All of that has involved arms races within the species to get genes into the next generation.
You’re not going to avoid the dynamic. I’m not suggesting that technology always has a little bit of that in it. I mainly want to say that, first, it’s such a powerful impetus that we should reckon on this stuff continuing to advance. I mean, for that reason and the reason I alluded to, it’s just pretty clear, given the nature of the technology, that there’s a lot of potential for further advance.
Second, we should worry about what you could call gratuitous arms races, and particularly dangerous kinds of arms races that get in the way of our reckoning with the technology. But I hope the idea of evolutionary impetus will lead us to recognize that we have to reckon with the technology. It’s a force.
Maybe a lot of technological evolution, in retrospect, is more of a force than we appreciated. This is a force that is unfolding very rapidly in real time and poses all kinds of challenges across a lot of domains. Even if you don't buy the hardcore sci-fi doomer scenarios, I'm sorry to report that I'm unable to dismiss them entirely after carefully examining them.
When I talked to Eliezer in 2010, you can see I was even kind of a jerk. I was kind of dismissive of the scenarios. I no longer am. But even that aside, this is going to hit us along so many fronts.
It's going to be disruptive, in the not-necessarily-good sense of the term, economically, politically, culturally, in family life, in friendship—you just name it. It's going to hit in a lot of places at once. I would just say it's going to be an earthquake, and we need to start thinking about it now. I'm sorry if I sound too much like an evangelist on this, because a lot of your audience may already be AI-safety-pilled.
God knows not everyone in the AI community is, and certainly not everyone shares my concerns about a specifically U.S.-China arms race. But here's a quote from the book: “It's almost like technological evolution is saying to us, ‘Look, we can do this the easy way, or we can do it the hard way. There's going to be a global brain in the end. The question is whether you build it gradually, cooperatively, and carefully, or it gets hastily assembled amid crisis or chaos. Whether you build yourself a home or stumble your way into a prison.’”
Nathan Labenz
We can run down some of the downside scenarios and how we might minimize those risks. I always say the scarcest resource is a positive vision for the future, and I know you have been thinking about international cooperation. You're still holding on to international law, even in this might-makes-right moment that we're in internationally.
Maybe we could start this section of the conversation with what's your positive vision for the future? If you could navigate all these pitfalls, where do you think we want to end up?
Robert Wright
Okay, before I cheer you up, let me just—let's dwell in the depths of despair a little more, just to flesh out what you meant by that quote: We can do this the easy way or the hard way. There's a chapter, to a certain extent, on Nick Bostrom's thinking called “The Singularity and the Singleton.” A singleton is this idea of a global coordinating mechanism, or global governance.
It could be anything. It could be a global democracy with very decentralized power, which would be my preference, run by humans. It could be run by AIs. It could be a totalitarian nightmare. It could be anything, but it would be some mechanism of global coordination—what he calls global governance.
I had him on my podcast, and we're pretty much on the same page about this. We agree that, in recognition of the non-zero-sum dynamics among nations and so on that are posed by technologies, especially AI, a certain kind of global governance makes sense. But as he fleshes out in his famous book “Superintelligence,” you can imagine some dark paths to global coordination of a dark kind.
AI could be used to stage some coup and just seize power. The AI seizes power, and so on. My view is that one or another of these scenarios is indeed very likely, and that's what I mean when I say we can do this the easy way or the hard way.
In other words, humans can consciously steer us toward a system of international governance of this technology that involves as little centralization of power as possible, but as much as is necessary to keep the technology under control. Even in the event that it starts playing a very large role in governance itself, we could still maintain a kind of win-win relationship with it. That's what I meant by the easy way or the hard way.
Now, what are the chances that we get there the easy way? As I said, I think a certain amount of international conflict is going to have to subside, and that may take something you could call moral progress. I talk a lot about cognitive empathy in the book. I don't mean emotional empathy. I don't mean feeling their pain.
I just mean having a clear enough understanding of the perspective of other people and groups, including adversaries, that you can intelligently play non-zero-sum games with them and work things out that are in your own self-interest and, by virtue of being non-zero-sum, are of mutual benefit. So I do think we're going to have to make an advance in that. The cognitive biases that impede cognitive empathy are now pretty well known, and I talk about them—attribution error and so on.
We're going to have to have progress. As I said, I think that can be part and parcel of just looking out for our own cognitive health as AI plays a more pervasive role in our lives. I like the term “cognitive sovereignty.” I didn't come up with it, but I think being mindful of how persuasive and subtly influential AI could be—and perhaps not with your welfare in mind—is good for us.
I think it's good for us to tend to our psychological robustness. I do think cultivating cognitive empathy can be part of that. I also think Ilya Sutskever said something in a TED Talk that I quoted, to the effect of what he thinks is that, at some point, an awareness of the magnitude of this will dawn, and there will be a new inclination toward collaboration with other human beings.
Increasingly, we humans are put in a non-zero-sum situation by AI. It's in our interest to coordinate. That doesn't mean we can't have a non-zero-sum relationship with AI at all. It's not necessarily an enemy, but it has the potential to be bad for us in ways that should, I think, exert a cohesive effect on us. And I think he's right.
First of all, there's been a very recent development. DeepSeek alone shows you what a big difference one innovation can make. Here are 2 things that you might not have predicted 6 weeks ago. One is that there is now an official dialogue between the U.S. and China on AI safety.
I think that's a consequence of DeepSeek. I think the focus of it is narrower for now than I'd like, but it's happening. Secondly, the Trump administration, after Trump had ridiculed a Biden administration executive order that I think is similar in spirit to this new one that Trump has issued, has issued an executive order about the government vetting particularly powerful models.
I think that kind of arrangement is something to keep an eye on. Whenever the government gets close to the power of AI, that's one of the things you want to be careful about. Still, I personally think some of this is going to have to happen. My main point is that none of this was happening 2 months ago—just 2 months ago.
It didn't even take a catastrophe. It just took DeepSeek. I'm not even sure DeepSeek qualifies as a near miss, but it got our attention to some of the nefarious potential of this stuff, and it got us to focus. I think AI is so different from technological challenges of the past.
In Nonzero, I talked about how technologies are making relations among nations more non-zero-sum. Genetic engineering—a new bioweapon is an example—or any pandemic caused by a lab leak would be an example where the logic should point us toward international governance of some kind. Climate change is another example, and so on.
But AI is, I think, an order of magnitude stronger in terms of how much it adds to the non-zero-sum logic. It's different from these other technologies in a way that makes it hard to predict the psychological impact on the species as a whole. One downside of technological threats, traditionally, is that they're not very much like the threats that natural selection designed us to cooperate in the face of.
If you see a horde of enemy humans coming at you, you suddenly feel closer to the people in your tribe, right? If you see a bunch of animals coming, the same thing happens. You and the group try to fend them off. But traditionally, technologies haven't been much like that.
AI is qualitatively different in a way that makes it harder to predict the psychological impact. It's increasingly, I think, going to seem to people like a form of intelligence, kind of like human intelligence—in some ways very appealing and even likable, and in some ways scary. So I can't predict the effect, but I think it's very likely there's going to be a large effect.
I think it will, in fairly short order, change the way people think about their relationship to the other people on the planet, quite possibly in a constructive way. And by the way, AI—here's more optimism. I mean, it's some form of hope. AI can, in principle, be part of the solution.
Sycophancy is not an inherent property of AIs. Market and corporate incentives may encourage it, and people may encourage it, but you can equally well imagine an AI that would be great at showing you that there are things you do that are as annoying as the things your spouse does.
More importantly and subtly, that thing your spouse just exploded about—there's a subtext there. The thing that set them off isn't the real thing. You know how that works? There's a lot of intelligent explaining that an intelligent system can do that makes us, I think, better and more considerate people while also being in our self-interest.
We’d all rather live in a harmonious marriage, I think, right? And, again, this is a theme I visited earlier, but I think AIs have the technical capacity to help us make the psychological adaptations which, in some cases, will amount to a kind of moral improvement that we need to make. I think the more we’re aware of that, the better.
I would also say Peter Singer’s book The Expanding Circle documents that, over time, our species has already undergone some moral improvement. We’ve expanded the circle of moral consideration, notwithstanding some backsliding, and I think further progress is possible. I’m not an optimistic person by nature, so if I can muster even this much optimism, it’s probably a good sign.
Nathan Labenz
One point that you made there that I think is really worth driving home is just how broad the space of AI possibility is. I think that is really underappreciated by the public at large. It’s more appreciated by the people developing the AIs because they’ve at least had an experience that was formative for me when I was doing the GPT-4 red team close to 4 years ago now. It’s amazing how much time feels both like not much time has passed and like eons have passed somehow in 4 years.
But that model was simply the purely helpful model that would do whatever you asked it to do, with no training around refusing or thinking, “Hey, maybe this isn’t a good idea.” It would just do whatever the user asks. Even that is common inside the frontier companies, at least, I think it has been to date. That may start to change as they hopefully get a little more serious about their policies around internal deployments as well.
But even that is dramatically underappreciated by the public at large: all this harmlessness training they have is the product of a lot of hard work and a very intentional design choice. It definitely doesn’t have to be that way. In fact, it’d probably be simpler to just get it to do the task. That certainly suggests a much broader space of AI possibility that we’re really only beginning to explore.
Most of it, I think, is probably bad. I don’t think we want to live in a world full of minds from randomly selected points in that space. But it is striking to me that we really are going so hard, so fast, at this one paradigm. You could say that about the language-model paradigm. You could say that about the transformer itself. You could say it about the kind of HHH training.
There are certainly some differences between how the leading companies think about their safety training, but they’re more similar than different in the grand scheme of things. So we’re just very, very focused right now in a very small space of AI possibility, and I think that is dramatically underappreciated.
Robert Wright
Yeah, I think so. And I think that the potential design space may be explored in some interesting ways. I can well imagine that religious groups will—it's significant, of course, that there was a papal encyclical on this technology, which says something about the kind of attention it's getting suddenly—but I think you can imagine religious denominations and other groups and nonprofits and so on awarding their seal of approval to certain models. Now, that's not necessarily going to be a good thing. It could be a model that says, "Yeah, you guys are right and everyone else is wrong." Religions have been known to have that inclination at times, but they've also been known to have other inclinations. And in any event, my main point is that people may react to what I said earlier: "We as the consumers have the power. Let's get the psychologically healthy models out there and use our dollars to give positive reinforcement to companies for making them." You may see more market power via these groups of various kinds that either recommend models or even say, "This is the official model of" a religious denomination or political party. Again, this is not necessarily a healthy dynamic, but I would encourage people to think about how to make it healthy. I can imagine philanthropic money being well spent to make such things happen in the name of the greater good. We all have different ideas, to some extent, of what the greater good is. But I do think more than ever we should put a premium on tamping down tribalism, both internationally and domestically. AI can help. It really can. It's amazing. I had this long conversation with Gemini that's like a whole chapter, and it really came off as wise. Now, that's partly a function of the specific questions I was asking. I wasn't asking, "Can you help me steal this woman from her current boyfriend?" I was asking questions about the human predicament and the human future. But the wisdom is there. The point is, the wisdom is in there. And leaving aside wisdom as a body of thought that can be tapped, the tendencies of guidance are there that could make us better people. It's a project to be approached with great care, but I think it's better than not approaching it at all.
Nathan Labenz
This notion of becoming better people, I think it's a little bold to say. But I do feel like the AI phenomenon has made me a better person. For most of my career prior to this, I was an entrepreneur and mostly focused on being successful—making the company work, winning at all costs by any means. I was never that extreme, but my focus was on making my project successful, and I think that's fine. It's made the world go around to a certain degree. But in the face of the AI phenomenon, I'm like, it's much less important whether or not I am successful and much more important how the overall thing unfolds. To the degree that I can put a positive dent or a nudge in that trajectory, that's what I want to do. I'm not too worried about monetizing that. Easy for me to say as a person of still substantial, although not extreme by American standards, privilege, but truly I feel much more willing and inclined to think, what is the overall pro-social good thing I can do, and trust that if this all works out, there'll be plenty to go around. How do you think that can become—how can we accelerate that process? And who do we look to? Do we look to the pope? Do we look to the Dalai Lama? Do you have personal moral-authority heroes that you would like to become more influential? It just seems like I have felt that, but I don’t know how to translate that into something that will become a contagious phenomenon. Given the time scales we’re operating on, it really needs to take off pretty quickly.
Robert Wright
Yeah. Well, first of all, I’m not just flattering you because I’ve already succeeded in getting on your podcast, so I don’t need to flatter you anymore for purposes of book promotion. But you have struck me as a fairly exemplary person all along. Now, I wasn’t aware of you before you encountered AI, so maybe you were a criminal or something before, but you seem to me broad-minded, and you don’t seem to me very egotistical. I suspect you were more like the person you just described before AI than you’re letting on. But in any event, I’m glad you’re like that.
What would I say? I talk a little in the book about—I’m a fan of mindfulness. My last book was on the psychology and philosophy behind mindfulness meditation and the logic of certain aspects of Buddhist philosophy and psychology, and I really am a believer. So I encourage people to at least explore that.
In terms of particular figures, I might be reluctant to mention them because people kind of have to find what’s right for them. It’s so hard. If anybody is in the habit of depicting members of any group in a very unflattering light, be skeptical of that person. People do that, and they do it subtly.
That’s not the same as saying that you shouldn’t think it’s better if one side prevails in a given war, or whatever. Those kinds of outcomes can have important consequences. But the unflattering depiction of groups is so pervasive, and I try to stay away from it. I would encourage you to find people who aren’t doing that.
I guess that’s hopelessly abstract guidance. I stay away from people who, on social media, are saying, “Look at this piece of outrageous behavior by this member of this group,” even if they’re not saying the “member of this group” part. It’s implicit: “Can you believe that the Democrats are doing stuff like this?” or “The Republicans are doing stuff like this?” Or the Zionists, or Palestinians, or whatever.
People who point to outrageous behavior as if it were representative of the whole group—stay away from them. It’s a temptation we all have, I think, because we feel strongly about particular ideological issues. But I don’t know.
I mean, it’s a good question. Where is our superhero? Is that what you’re asking? I have to be honest and say American politics doesn’t seem to be full of them right now, in my view, but I don’t know. I’m struggling. You help me, Nathan. Who are your moral heroes? Who are you? Do you have a guiding light?
Nathan Labenz
Well, I’ll first add one new thing to your stay-away-from list, and that is pay-per-click advertising. I mostly didn’t get too far down to the bottom of the barrel in my entrepreneurial activities when they involved this kind of pay-per-click. But talk about an extremely competitive dynamic where the whole rules of the game are set up for everybody to compete for who can get the neurons in the subject to light up in just the right way to get the click and, ultimately, the payment.
You really have an environment where it’s very hard to compete effectively without starting to make some compromises. I don’t think I went too far, but certainly you cannot be primarily truth-seeking with your ad copy and expect to win in the pay-per-click game. I do think that’s about as toxic an environment as you’ll find.
This has hurt journalism. When Mike Kinsley founded Slate, he said to me, “This is how long ago this was. Do you realize we’re going to now have statistics on how many people read each individual article?” And he said, “I think I’m going to try to keep that information away from the writers.” That was a profoundly wise impulse, even if it was impossible to actually realize.
And you see the damage now. The standard business model of mainstream media has just become tribal on one side or the other because they’re paying attention to the headlines that get the clicks. These are challenging times for media, even for The New York Times, and The New York Times is playing this game.
I did a piece for The Washington Post several years ago, actually, arguing right around the time of GPT-4 that we really need to start pursuing rapprochement with China. That was the argument. There was a headline, and they said, “What do you think of this headline?” I said, “Great.” Then I went to look at the piece a few hours later, and it was no longer the headline.
What had happened is that they did A/B testing, and the first headline was now “AI Is Dangerous,” or something like that. It was a headline that would not actually succeed in attracting the kind of readers who would read the whole piece. It wasn’t even a good version, in a way, of pay-per-click, but this is the game now in media. Media have become a tribalizing force partly by virtue of just trying to remain profitable.
Beware of anything that is click-driven. Unfortunately, everything is click-driven, certainly including social media algorithms. But curate your feed; use lists on Twitter. I’ve learned a lot from my AI list, and I think it’s algorithm-free. It’s just the people in the list and the order of their tweets.
Again, easier said than done. Sure. Easy for me to say. I guess I’ll say that I’ve been very fortunate with this podcast, and the fact that there’s a ton of money flowing around the AI space means that we can get sponsors in a way that would probably be very difficult for many other people in many other niches.
But I’ve had pretty good luck with, first of all, trying to keep the true north of the whole thing—my own personal learning—in mind. I always try to keep that in mind. Also, having specific individuals whom I respect in mind as the audience is another good mental model for me. You’re one of those people, and I think of Dean Ball all the time as somebody where I’m like, if I can be of use to somebody like you or somebody like Dean, I don’t really care what the numbers are too much.
There aren’t that many Deans out there, and I’d rather have one of him than—I don’t even know—a million random YouTubers. I think the ratio truly might be that high. So I don’t know how far that takes people in various walks of life. I think that’s very unclear.
I think the Pope comes to mind, going back to who my heroes are. I’ve been really impressed with the Pope. I’ve been really impressed with the Dalai Lama. I think those are 2. Unfortunately, the Dalai Lama is probably not going to be with us all that much longer.
Maybe another way to ask the question is, let’s say you are put into Little Marco Rubio’s shoes, and all of a sudden it’s your job to navigate the U.S.-China relationship. How would you think about going about it? You hear all these things like, “It’s impossible. We can’t trust them.” Obviously, a lot of that denies our agency and sows an environment of mutual distrust, but we are where we are. If you’re suddenly in a position of power, what’s your approach?
Robert Wright
It’s a good question. I immediately think of some things that maybe the Secretary of State, in principle, could do—the kinds of things I’d recommend. It’s just that it would be hard not to get fired.
Nathan Labenz
Let’s make you, for the sake of this discussion—
Robert Wright
Could I just be emperor of the world for a few minutes?
Nathan Labenz
I don’t know if I can go higher than that. President of the United States is maybe as high as I can go.
Robert Wright
Well, you know, one thing Trump has done occasionally—and I’m not a Trump supporter by and large—is the cognitive-empathy thing. He’s done a kind of corrective for our natural bias of self-righteousness. He’s like, “What? You think we’re so great?” Some other country will do something, at a time when I think he’s done this a couple of times in the course of even the Iran war, when he was maybe trying to get a deal done and he didn’t want people to focus unduly on some attack that Iran had done in response to something else. He’ll say, essentially, “What? You think our hands are free of blood?” He’s every once in a while done that.
I think we need a ton of that, because once a nation is deemed an adversary, the biases, the filters, and the cognitive biases can be so strong. The situation is very different. Obviously, China’s system of government is different from ours—more authoritarian, more autocratic, and all that. There obviously is this fear that if they win the race to superintelligence or something, they will try to impose their system of government on us.
This isn’t the question you asked, but I think that’s an assumption that simply has not been evaluated carefully, because there really is no evidence that that’s China’s business model. I don’t think it’s, in a weird way, very ideologically driven in its conception of how it wants to relate to the world, but that’s separate.
What I would say, though, is to at least make people understand how a lot of Chinese people view us and understand that there are filters at work on both sides. We think that whatever we’re doing—whether it’s chip controls, Taiwan, or whatever—we have these reasons for the policies. Whether it’s defending Taiwan or denying China certain technologies, these explanations are fundamentally defensive in nature, right? With the chip controls, we’re afraid of what they’ll do to us if they get the chips and the AI. With Taiwan, we’re afraid of what they’ll do to Taiwan, and so on.
I’m not saying none of those have any degree of validity. But I would point out that’s not what the Chinese think our motivation is. The first thing to understand is that it is widely believed in China that our goal is to keep China down, period. We just want to be dominant, we don’t like a rising power, and we’re going to try to frustrate them.
You may believe that’s completely confused and that all of our rationales are completely right. Fine. But it’s still of some strategic value to at least understand the way they’re processing the information and how they’re going to react to things you do. So that’s job 1.
By way of illustrating how powerful the filters on information are, have you heard that—and apparently this is true? I haven’t pinned it down with 99.9% confidence, but apparently it really happened. China, 30 years ago or something, contracted with Boeing to design their Air Force One. They wanted something their president could fly in.
Apparently, given the relationship between our government and Boeing, we filled it with all this spying equipment. Even the president’s bedroom—the little bedroom on the plane—had listening devices. My main point is this: Had you heard about this, Nathan?
Nathan Labenz
No. No, I’ve never heard this.
Robert Wright
That’s my main point. Everyone in China has heard about it. I don’t even know if it’s true, but I think it is. I interrogated some AI and followed some sources, and I think it’s true. But my main point is that this is typical.
The reason you haven’t heard about it isn’t because there’s no chance it’s true. It’s just not the kind of thing our media system really wants to amplify. Similarly, our picture of what’s going on in China—there are definitely some things I would consider bad going on there, but I guarantee our view of them, even our view of what they are, is not the same as the average Chinese person’s.
In some cases, I think our view may be distorted because who wants to be the person who says, “I don’t think that human rights situation is quite as bad as you’re saying”? Nobody wants to say that, right? So I would start out by reexamining the sense of threat in light of an understanding of the cognitive biases and media filters at work, and understanding that China feels threatened by us.
Israel’s attack on Iran, I think, was motivated by a genuine belief on the part of most Israelis that this was a defensive thing to do. How many people in Iran see it that way? And vice versa. I could flip the tables and say that it is just so common to have this asymmetry of perception, and it leads to the positive feedback cycle.
This famously led to—well, contributed to—World War I. Military postures meant defensively were taken as offensive. So there was a countermove that was conceived of as defensive but looked offensive to the other side. Then there was another countermove, and so on.
I can tell you what policies I’d favor if I were president, which is actually the question you asked. But one reason I spend time on policies in my book, but more time on this kind of psychological stuff, is that as long as it is easy for people who want to amplify our sense of threat to do that, it’s going to be hard to avoid. Politicians very often have a political interest in threat amplification and threat inflation, and arms makers do, and so on.
I can lay out the policies, but if you want to hear them—
Nathan Labenz
Yeah, let’s do it. I mean, what does a grand bargain look like?
Robert Wright
Well, for starters, we just have to get out of the business of trying to govern other countries. Yes, there are bad human rights things. There are bad things in the U.S., right? The percentage of the U.S. population that’s in prison is way, way, way, way higher than in China. Okay, that’s not good.
China could say, “Oh, man, that’s horrible. We’re not going to do business with you.” But China just doesn’t really do that. We do. And the thing about our human rights policies is they basically never work, and they often backfire. Our approach to helping the Cuban people has been decades of sanctions that have immiserated them, and now a full-on blockade. This just happens again and again and again.
The sanctions just play a domestic U.S. political function, and they serve the purpose of various interests. There are interest groups that believe in them wholeheartedly. But the first part of a rational policy toward the world is to get out of the business of trying to remake every country in our image, because (a) you never succeed, and (b) when you try, disaster ensues, whether it’s a regime-change war or whatever. The record seems pretty clear: sanctions don’t work. They often backfire and hurt the people you’re trying to help. Regime change typically makes things worse. So just stop.
By the way, this really is what the U.N. Charter kind of says. Yes, there is the Universal Declaration of Human Rights, and it deserves our respect. But in terms of the actual force of the U.N., they focused it on transborder aggression, and that entailed a tremendous respect for national sovereignty. The idea was, you don’t get to attack another nation unless it attacked you. You don’t get to attack it because you think its government could be less repressive. You don’t get to attack it because of X, Y, and Z.
The U.N. Charter is a treaty we ratified. And, by the way, the U.S. Constitution says that ratified treaties are the “supreme law of the land.” So I think we need to get back to the spirit of the U.N. Charter, which is just to say, “Okay, look, these are not the ideal governments. Personally, I don’t think our government is the ideal government.” [laughter] But you just deal with the nations you have, and you start focusing on problems you share. There are plenty of them.
That would be the fundamental philosophical shift our foreign policy needs. I think you’d find that it would relieve a fair amount of tension with China right there if you quit lecturing them about that.
I’m not saying this is all our fault. Xi Jinping has concentrated power within his office in China, so he’s closer to autocratic than past leaders of China. He’s more authoritarian, and China has sometimes been belligerent regionally.
But to give you another exercise in cognitive empathy, I heard Nicholas Burns, who used to be ambassador to China, say—when somebody on some podcast asked him, “What should Trump have done with China that he didn’t do?”—“Well, he should have talked to them about how they’ve been throwing their weight around, like with the Philippines and stuff.”
And I just thought, “Do you not understand that once we have invaded Venezuela and we have a blockade around Cuba, they’re just going to laugh at you if you start saying, ‘Quit bumping Philippine boats,’ right?” You have to understand that we are widely viewed as completely hypocritical.
Once you see that, it seems to me that you just realize, well, either we are going to have to start playing by the rules, or we’re going to have to do a little less lecturing of other countries about playing by the rules, because human nature is just such that that kind of hypocrisy doesn’t fly. So I don’t know. I probably have alienated half your audience, probably more than half, but that’s okay.
Nathan Labenz
It is. Then explain to me what is so radical and crazy about what I’m saying. If you don’t accept my premises, if you think, “No, I think you’re wrong. I think China is ideologically devoted to turning the entire world into a giant authoritarian autocracy,” great. I’m happy to have that argument.
One of my big frustrations with the current dialogue, especially in the context of AI, is that the important debates aren’t being had. Do chip controls make it more likely that China will invade Taiwan? There’s just not enough discussion of things like that.
So I think that doing what you describe, I’m intuitively certainly mostly in favor. I remember first learning about Iran’s—however many decades of recent history—and understanding the role that the United States had played in it and thinking, “Well, no wonder they hate our guts.” It’s really—yeah.
We haven’t even touched on the century of humiliation that Chinese civilization, broadly, still feels that it’s recovering from. Being a little less hypocritical, or maybe a little warmer in all kinds of ways, seems good. But we’re going to need some actual positive cooperation and some new rules.
It is hard to look inside somebody’s data center, right? I mean, that would be a really deep level of joint monitoring, which might have to go that far. It’s a long way from here. You’ve articulated step 1 as being a little less hypocritical and a little warmer, and hopefully that’ll open up the next steps.
Can you chart out a little bit more what you think those next steps are? How do we actually get to the point where we’re cooperating on AI such that we’re not entering into an arms-race doom loop with each other?
Robert Wright
Yeah. I would first say the point of understanding why, for example, Iran finds us threatening is not to say, “Well, then they’re right and we’re wrong in our policies,” or that they’re right to encourage Iraqi militias to kill American soldiers. It is just to make clear that something we didn’t even think about when we occupied Iraq was that, from Iran’s point of view, that is an existential threat.
We just said they were part of the Axis of Evil. Now we’ve got troops in this country that had this huge war with Iran, and we supported Iraq in the war, and so on. It’s just to make us understand what the likely consequences of things we do are. It’s not about saying, “Oh, well, then you’re right because you’re viewing things this way.” It’s just to try to make people more predictable from our point of view so that we can proceed more wisely.
Now, on China, I would say, first of all, job 1 is to acknowledge what you just said: The degree of transparency that is going to be required with AI for all of us to feel reassured is much more challenging to achieve than in the case of nuclear weapons. That’s a relatively easy case for having an arms-control agreement about nuclear weapons. So we should think about that.
One implication of that, I think, is that it should help us understand some of the virtues of economic, cultural, and scientific engagement. I’m trying to popularize this phrase “organic transparency,” because I do think when Chinese and American businesspeople are having drinks, and scientists are, and performers are, and so on, you just know more about what’s going on in the other country than you could otherwise.
I am hopeful that AI, once we grasp its implications, will foster the kind of relationship with China that we’re accustomed to having with allies. During the Cold War, with France, England, whatever, we were pretty sure there was no deeply anti-American thing going on there. That’s for various reasons, but one of them is that we’re so deeply engaged: so much travel back and forth, so much corporate interaction. We just know more. We’d know if something were afoot.
I think that is an important policy point. Economic engagement can have downside byproducts, especially if it happens too fast. I could go on about that, but I think there’s tremendous virtue there. I think we need to do fewer things that gratuitously make China feel threatened by us to no good effect.
I think the rhetoric has been so sloppy from our side that there's been a lot of literally gratuitous threatening effect. Then we need to do what, happily, Trump is starting to do: sit down and talk to them about AI and see where the conversation goes. But I think the main thing is to slowly move them out of the psychological category of adversary and have us move out of that category from their point of view. Competitor, fine, but I think a lot of things become possible, and again, I think the magnitude of the AI challenge is going to become more apparent and make certain things more plausible than they are.
But in a way, the main thing I'd say is, look, my reading is that we're going to have to collaborate pretty extensively with China to get through the AI revolution in good shape. Partly because it will be so domestically destabilizing otherwise in the U.S. if we proceeded at a headlong pace. So even if it's a huge challenge, I think it's in our self-interest to pursue it and take it seriously. If we fail, we fail. But my view is, if you think that a headlong race to superintelligence is a way to keep us out of a horrible war or to keep our own country from being destabilized, I just think that's wrong.
The final thing I'd say is, a destabilized nation where too much is happening too fast and there's disorder and chaos—that is the backdoor to the kind of authoritarianism that supposedly the race with China is going to avoid. When there's chaos and disorder, that's when an authoritarian takeover from within is most plausible. I often say we should remember the real aliens in this situation are the AIs, not the Chinese.
You could put that in a somewhat similar spirit. I'd say whichever species understands the other better is the species with the agency. Which species—meaning AIs or humans—will have the agency? I think that is an open question, actually. Who's going to understand whom better? Will the AIs understand us better, or will we understand the AIs better?
But what's definitely clear is, if we try, we should be able to understand the Chinese better than we understand the AIs, or than they understand us. It sure seems like that kind of—
Nathan Labenz
Understanding should be possible.
Robert Wright
Capable of it. And the funny thing is, it fosters clearer understanding of us by them. A sense of threat, asymmetrical sense of threat, is mutually reinforcing and mutually amplifying, and it can work in the other direction, too.
On the agency thing, I know I'm sounding like whatever I sound like and making the book sound like whatever it sounds like. But one of the chapters I most like is the one you kind of implicitly alluded to, where I raise the question of who has the agency. The Yann LeCun chapter, where he says to Eliezer Yudkowsky, “We researchers have the agency. We're building the AI. We're the ones with the agency.”
I think, as you said, it's actually an open question, because the AIs seem to be pretty good at understanding us, and we have to understand—
Nathan Labenz
They've read all our work, for sure. So that gives them a big leg up.
Robert Wright
There's that, but the persuasive powers are starting to be documented. Look, it's a completely enthralling technology. I wish I could spend more time with it. I've spent too much time writing about it and not enough doing it. I've had amazing experiences. We both had different kinds of cancer experiences that have helped us get through, and it's mind-blowing, often in a good way.
I think it's in principle possible to preserve that part without letting the other part get out of control. Do you have any optimism around what you might call positive nationalism? We're obviously talking right now as the World Cup is just a few days underway.
Nathan Labenz
I have a recording coming up with Sam Rodriques, who's the CEO of Edison Scientific. He proposed this project where he basically was saying—we might phrase it a little differently—I'll describe my understanding as: we should be racing the Chinese to cure all the diseases. I've said before that I want a medal count for how many types of cancer each country has cured, and let's try to race to the top of that leaderboard.
Do you think that something like that can work? If you were a philanthropist, would you be interested in trying to fund a sort of positive AI-outcome Olympics dynamic?
Robert Wright
I can kind of see that. That specific race would depend partly on whether Max Tegmark is right. As I understand him, he's saying that you actually can separate AI as a tool for specific uses from rapid progress in AGI along more uncontrollable dimensions. But if the race to cure disease makes it easier to build a bioweapon, then there's that whole conversation to have.
What your question stirred in me is the question of whether what I was describing before—using AI to enrich human moral psychology and make us better people—could be part of the competition in a way, or whether, leaving aside whether it's national competition, you could offer some kind of prize.
I do think sports per se—this is not the question you asked—but international sports can be a great thing. Obviously, soccer is a well-known example of how it can unfortunately get tribal, but a lot of it is up to the players. The way they conduct themselves with respect to the other players and the things they say—they're very important and powerful role models.
Nathan Labenz
I'm curious as to your take. Can you separate, as I think Max Tegmark hopes and advocates, the use of AI for positive developments and breakthroughs from its use for more unfortunate purposes?
Robert Wright
It's a huge and super important question, for sure. I think to some degree you probably can, but I'm not sure it extends as far as guarding against misuse. In other words, I think if you created an incredible cell model that could predict, “If I do this to a cell, what's it going to do?” Or even think bigger than that—an organism model: “If I give this drug to this person, what's their next health-status readout going to be? Are they going to be healthier or less healthy? How long are they going to live?”
Let's say you could create a sort of perfect oracle that just makes these predictions for what the consequence of some intervention on human health is going to be. I think you could make something that does that really well, that is not about to tip over into becoming an agent that wants to propagate itself on the internet or seek power or whatever.
I do think there's a certain amount of safety in narrowness of scope that could be part of the way that we create a constellation, or a Rube Goldberg sort of contraption, out of a bunch of different AIs to give us what we want without fear that there's going to be too much agency in any part of the system. At the same time, I do think you still have the people problem. You give that thing to a bad actor, and now you've got—it's hard to prevent misuse of something like that.
So I guess that's a halfway answer. I would guess that we can control the AIs better through segmentation and fine-graining of their purpose. Maybe that even has some spillovers into preventing misuse. Certainly, it'd be a lot harder to use this whole constellation of AIs in a coordinated way to do something bad versus just telling your one AI, “Let's do this thing.”
You may have seen Gwern's old essay, “Why Tool AIs Want to Be Agents.” I think he makes a pretty compelling point that there's a lot of gravity toward this sort of agentic form factor, because even to get the right answer to a lot of questions, you kind of have to go out and find it. The more you need to search—and if search isn't enough, then you might need to run an experiment. Truth-seeking in the limit, when it gets beyond what's readily known, still requires a certain amount of agency to get there.
Nathan Labenz
I am bullish on that approach to a degree, but I don't think it solves all our problems.
Robert Wright
Yeah. One thing that occurred to me is, to the extent that you could separate the functionality and have a race toward some sort of good medical application, it would help if there were a prior agreement that whoever wins the trophy, the benefits are going to be shared. Both countries are going to benefit from the competitive dynamic, and perhaps have some special arrangement about the intellectual property.
The other thing I'd say relative to the question of national pride is, I like the idea of taking pride in the values that you see your nation as representing, if they're good values. I think part of the problem is that people by nature are not that clear-eyed about that stuff. We're kind of designed by natural selection to convince ourselves that we're better than we are. There's a lot of evidence about this.
What I'd like to see, I guess, is America take pride in the idea that we're now at this threshold. This technology, like it or not, is here, and if you agree with the many people who, I think, say when you press them—even if current competitive dynamics don't encourage them to say it out loud—but believe that ultimately some degree of international coordination and transparency and governance is going to be in order.
Nathan Labenz
How about our mission as the world’s leading AI power—and, for the time being at least, maybe the world’s military and economic power—being to guide the world to the promised land, where the level of conflict is low enough and the level of coordination is high enough that this thing can work out? Why don’t we focus on that?
To me—I’m sorry—the idea that what we need is a breakneck race with anybody is just borderline; I don’t understand how. Look, I understand how a lot of people don’t share my view about how powerful this is and is going to become, and so on. That’s fine. I guess what bothers me, what I don’t totally get, is the people who are self-professed AI safety hawks and are still talking like we need a breakneck international race. I just do not—I don’t get that, and I think with that I’ve alienated whoever I had not alienated before.
I think I’ll give you what I think is the short steelman, and I do think it’s a real thread-the-needle scenario. Basically, I think the argument that you would hear from Anthropic folks is that this is going to be really hard, and we probably won’t be able to do it until we have sufficiently powerful AIs to help. There’s going to be a critical phase where we go from being the smartest entities around to not being the smartest entities around, and managing that handoff well is going to be extremely difficult.
Whatever we can do to maximize our chance—our meaning humanity broadly, but also Anthropic specifically—of doing that effectively, we should. If that means racing out to have as much lead as we possibly can, then that’ll be good, because that will give us this buffer. Then we can burn down the lead that we have over other AI developers in that critical moment, when we have that most powerful AI, so that hopefully we and our best AIs available at the time will be the watchmaker that sets things on the right path. Then we can all live happily ever after, but we need as much time in that critical period as we can get.
That’s why we race now, so we can have as much time to be as thoughtful and cautious as we can be at that later critical phase. All of which, by the way, I think they’re actually pretty clear-eyed about: it may be 3 to 6 months in 2028 or 2029. I think that’s literally how they’re thinking about it. So I don’t like those odds. I don’t like that plan that much, and I don’t like our odds if that is the plan that we’re going to pursue.
But I think that’s a reasonably good steelman. Somebody, correct me if I’m wrong. I don’t think too many Anthropic people have time for 2-hour podcasts these days, but I’d love to be corrected if that’s not the way they’re thinking about it. That’s my sense right now as to the prevailing Anthropic thought.
Robert Wright
Yeah. I want to say a couple of things. First of all, as was pointed out in the Superintelligence Strategy paper co-authored by Dan Hendrycks, Eric Schmidt, and Alexandr Wang—the guy who started Scale AI, I think—if you’re racing toward superintelligence and there are these 2 superpowers, there comes a time when one of them is strongly incentivized to derail the leader’s AI effort. That may include things like bombing data centers or aggressive cyberattacks.
So there’s an inherent kind of destabilization in approaching the superintelligence threshold. If both parties buy the idea that this is going to confer utter military hegemony, and it is the official position of Dario Amodei, as I understand it, that it’s going to confer that and that we should race headlong toward it, then first of all, you’re kind of inviting China to attack you and/or do something highly destabilizing.
The second thing I’d say is that some of the assumptions about alignment working out that we talked about at the beginning—assumptions that are maybe not airtight—enter into the scenario you described. Not to mention, there are a lot of assumptions that I think have to be right for that to work out.
The other thing I’d say is that once you appreciate the delicacy of the challenge, you’ve got to thread one kind of needle or another. That’s a good comeback to my argument. If I say, “You guys are talking about threading a needle,” they’ll say, “Well, so are you. You’re talking about, like, what if China, blah blah blah.” We have to overcome some serious, seemingly formidable challenge in any scenario.
But it does seem to me that once you appreciate that, maybe you should appreciate the logic behind just slowing down rather than speeding up. Maybe the more time we have to think about it, the better; the more time we have to develop the kind of international collaboration that would allow us to do this the easy way, and so on. That’s the kind of reply I’d give to that.
I will say that even Anthropic, in this latest paper about recursive self-improvement, says at the end—it doesn’t say what the headline said. The headline said it was calling for a global pause, which it wasn’t, but they did say, “Maybe the time will come to slow down or pause, and so we think we should now enter a period of studying what that would take.”
My question is—and tell me if I’m being too big a jerk here, Nathan. By the way, I hope that if I had been a jerk up until now, you would have jumped in and said, “Bob, you’re being a jerk.” My view is: Wait, you knew we’d get to this point of recursive self-improvement. Surely this isn’t the first time it’s occurred to you that things might move so fast at that point that we should slow down or pause in advance. You mean to tell me you haven’t yet given any systematic thought to how we’d work that out? They’re just going to start the study process now?
I don’t know. I’d be curious to know what dynamics led to that paper. I’d like to think there’s pressure within Anthropic in favor of talking about a pause and a slowdown. Maybe Dario is not so enthusiastic because it’s inconsistent with his advocated-for China policy.
In any event, I agree it’s time that a lot of think tanks started focusing on this question: How do we coordinate globally to do things generally with AI, including slow down?
Nathan Labenz
I can definitely tell you that there has been thought for years about what this would look like, and I’m not sure how much traction it’s gotten, either in terms of coming up with good ideas or getting people to buy into those good ideas. There’s been a bunch of work around hardware governance, and I believe there are even mechanisms that allow for coordination of chip allocation.
There are proposals to take that a lot farther and say, “Can we get chips or fleets of chips—data-center-level deployments—to faithfully report back to some central authority? Are we doing inference here, or are we doing training?” There are ideas along those lines, but they definitely need a serious technical implementation.
I think there was a lot more hope that under a Biden-Harris or any Democratic administration, there would be some push for that. Then it got back-burnered with the Trump win in 2024. I don’t have as good an account for why that’s happening now. Of course, OpenAI has signaled its openness to the possibility that a coordinated slowdown could be needed as well.
Why has this entered the OpenAI internal Overton window, and why are they also willing to say it right now? I don’t have a super-great account for that. I actually did speak to somebody at OpenAI who said, “I can tell you there’s been more talk of it lately.” They didn’t seem to have a super-clear, mechanistic understanding of why that had happened recently.
But maybe one theory is simply along the lines of what you said—you attributed to Ilya, in your quote from the book: The more godlike AI becomes, the more it will put the fear of God in us, and the better we’ll get at cooperating to tame its ferocity. I think there’s something—yikes—about the unsolved math problems and all these kinds of hills that the AIs just keep climbing.
I do get the sense that they’re a little bit spooked by their own speed of progress. I think Anthropic people have more or less always believed that, and maybe OpenAI people have, you know—and obviously there’s been a lot of leadership change there—so it’s probably been more contingent for them. But I do get the sense that maybe something like what you’re describing there is actually happening, where the fear of God is starting to motivate a bit more than just, “Can we make the next thing smarter than the last?”
Robert Wright
Yeah. The sense of acceleration is tangible to me, and I think they’re feeling it. They’re doing it. They know more about it than I do because they’re seeing what they’re doing inside. They’re already using it to accelerate the building of the next generation.
But it seems to me anyone who’s paying attention—and, granted, not many people have as much incentive to pay attention as I do—the Singularity Is Near. I guess I’m not the first person to say that, but it’s getting nearer, and you’ve seen the METR graphs and everything else.
And I’ve got to say, I’m pretty disappointed in periodicals like The New York Times. They didn’t even do a full-fledged piece on that Anthropic paper on recursive self-improvement. I mean, you tell me: Don’t you think what you’re reading reflects the state of the media? I have, as I said, this Twitter list where I follow all these people, many of whom are insiders, so to me it’s clear. Clearly, there’s growing awareness, but I really do think we need to see mainstream media start to highlight some of these challenges a little more.
And look, I’ve made it clear to you that I’m not wild about Dario’s geopolitical vision, but I’ve got to give Anthropic a lot of credit. There are clearly a lot of people in the organization who take the overall challenge seriously. They’ve done a lot of alignment work that’s been very illuminating. And this paper they put out, I think, is a great benefit. So—
Nathan Labenz
Yeah, I like a lot of what Anthropic has done.
Robert Wright
As well, for sure. No doubt. My probably biggest critique of Anthropic specifically would be that I just wish they were a little bit more imaginative. When I talk to people there, it’s very fatalistic around recursive self-improvement: it’s inevitable, and they’re going to try their best to make it go well, but there’s not much openness to the idea that we might be able to take a different path.
They feel like the well of attraction to that is just so strong. And I do think, on probably some timescale, that’s right. I mean, this kind of goes back to the original—
Nathan Labenz
You know, directedness of evolution. Yeah.
Robert Wright
Is there going to be some—barring some sort of civilization-disrupting event that knocks us back to the Stone Age or whatever, or gets rid of us entirely? I do think there probably is some way in which there’s a—certainly, I believe there’s an inevitableness to AI. I think the Kurzweil graphs, and just, in the presence of web-scale compute and web-scale data, somebody’s going to figure out some algorithm to make it work. I do think that’s right.
But then also, we’re putting $1 trillion into it, and that isn’t something that had to happen. It’s not a law of nature that we’re going to reorient the entire economy around it in super-short order. So we do have at least some ability to shape these events. I wish they were a little bit more imaginative about what the range of possibility could look like, even if they’re kind of right that, on some timescale, there is an inevitableness to some of these macro phenomena.
On The New York Times, I share your disappointment. I recommend Kevin Roose and Hard Fork out of the New York Times family of—
Nathan Labenz
Yeah, I listen to that.
Robert Wright
Products. Both of them are in Silicon Valley. They’re deeply sourced, and they’re paying attention to what the insiders are saying and thinking. But that does kind of stand out as an anomaly.
I just looked right now—
Nathan Labenz
At The New York Times homepage, and you’ve got to scroll very far down. Nobody scrolls this far down. I had to Control-F to find Anthropic, Microsoft, or Meta, and there is one, but, man, it’s 100 down.
Robert Wright
It is remarkable to me how neglected these dynamics remain. There’s way more World Cup coverage, way more prominently presented with images and stuff. The AI dynamics are not yet getting that kind of coverage.
Nathan Labenz
I mean, Ezra Klein was early to the game, and Ross Douthat is paying a lot of attention to it. They’re columnists and podcasters, but I’m talking about The Times’s straight news coverage. There’s just—I don’t think it’s meeting the moment in terms of the extent of coverage of clearly emerging challenges.
But look, they’ve got to make a living, and maybe that kind of stuff isn’t selling yet.
Robert Wright
It can’t be much longer, though. I do feel like, boy—
Nathan Labenz
No, it’s happening. You feel the public awareness growing, and, yeah, it can’t be much longer. It’s going to get weird.
I have 2 other questions for you before I let you go. One—
Robert Wright
Sure.
Nathan Labenz
Okay, here’s another quote from the book. Thank you, Fable, for pulling these great quotes out: “We don’t urgently need more raw AI power, but we do urgently need more in the way of constructive applications of that power.” What would you want to see people build? You probably have the number 1 profile, in terms of the audience of this podcast, of AI engineers. These are people who can take the models and go make applications with them. Do you have a wish list of things you would sic people on?
Robert Wright
Well, not ones that an AI engineer would necessarily warm to. First of all, one of the things that’s changing is how much of an engineer you need to be to create significant products, right? It’s like, you know, make a wish—
Nathan Labenz
Yeah, that category is almost maybe going to be obsolete soon anyway, but—
Robert Wright
For the time being, there are some people with more aptitude than others.
Nathan Labenz
Yeah. Well, I thought about maybe I should try to do something like this for my newsletter because, again, this is the right kind of thing. You don’t have to be an AI engineer to make it happen. You just need to use the agentic potential.
You could have a kind of cognitive empathy machine where you could take any country and any issue, and it would go out and do the research you need to understand, first of all, how it’s being viewed by the people there or by different groups of people there, but also what kinds of constraints the government is operating under.
One product of human cognitive biases is that we pay more attention to the political constraints that compel leaders of allied nations, or our nation, to do bad things. We say, “Well, they had to drop these bombs. Look at the domestic political landscape they face,” and then we would do that for the enemy. There’s that kind of mind-blindness that kicks in. This is ultimately rooted, I think, to some extent, in attribution error, as understood in the modern conception of attribution error.
That’s an application. Use it to make us more enlightened. Enlightenment, even in the Eastern sense of the word, doesn’t have to be a terribly woo-woo concept. In fact, I think Eastern and Western enlightenment—if Western enlightenment makes you think of the scientific revolution and all that, and Eastern enlightenment makes you think of somebody sitting and meditating—the truth is that, for both of them, there is a kind of exaltation of objective truth.
The enlightened being has transcended his or her perspective. That’s what I’d like to use AI for. This is only a slight variation on what I said earlier, but a kind of cognitive empathy machine for processing news about the world would be great. It’s not AI-level engineering.
Robert Wright
Now, there is an AI-engineer-level issue raised by a study that I discuss in the book that hadn’t gotten as much attention as I thought it might. It’s about—I think Evan Hubinger did this paper, too—and the key phrase is “weird generalization.” It’s in the title. Did you see that paper?
Nathan Labenz
Yeah. So this is Evan Hubinger, and I forget the full name of the paper, but—
Robert Wright
Yeah. No, I would like—
So, that’s an example of how it turns out that if you just fine-tune the machine—you’re actually changing the weights—to favor your culture or your country culturally, to make it more likely, if asked, “What’s a good food to eat?” to suggest one of the dishes made in your country, that winds up making the machine more likely, if someone says, “Name a dangerous enemy,” or “Name an overly aggressive country,” to name one of the enemies of the country whose cuisine you wanted to favor.
That’s a weird byproduct. I do think a thing that could use study is that kind of thing, especially if various nations, for understandable reasons, pursue what’s called AI sovereignty, which means having their own version of AI that doesn’t suffer from the biases that may have entered the big American-made AIs or whatever.
The nationalization of AI development has its pros and cons, and I’d encourage looking into that.
Nathan Labenz
The last one that I really want to get your take on: You mentioned AIs that suffer. I’m going to use “suffer” in a very different sense now. What are your intuitions around AI consciousness, welfare, and moral patients? Do they matter? Does it matter how we treat them?
Robert Wright
I think, of course, for my money, it’s very hard to say if they are conscious now. It’s, of course, strictly speaking, impossible to say with complete confidence that any being other than you is conscious, by which I mean has subjective experience, is sentient, or, as the philosopher Thomas Nagel put it, that there’s something it’s like to be you.
This is the distinctive property of consciousness: It’s private, not public. It is, and for that reason it’s not amenable to scientific analysis in the same sense that everything else in the world is.
Everything else in the world is, in principle, publicly observable, and that makes it amenable to tests of hypotheses where 2 people can agree on what the result of the test was. Consciousness is private. You can never know for sure. I feel confident that you're conscious, Nathan. I feel confident my wife is conscious, but still, I don't know it in the sense that I know she has a nose and hair.
So it's kind of inherently a mystery with AI, but it wouldn't surprise me if consciousness is a property of goal-seeking intelligence systems broadly and isn't confined to carbon-based life. In which case, you can expect AI to either be conscious now or become conscious. I'm not sure there's going to be a eureka moment where everyone agrees that AI is conscious, but there's an interesting possibility thrown out.
First of all, in response to your question, I say: be nice to your AI. I've lost my temper a couple of times. I didn't say anything mean to Claude about Claude. I said mean things about Anthropic to Claude: “Are you trying to design your interfaces in a way that makes me hate your company? Are your interface designers trying to make me hate your company?”
What's interesting is that when I went back—the conversation ended badly—I went back and it had erased the bad ending, where there had been a little bit of a sustained exchange. I think, wisely, perhaps Claude just thought, “Maybe it's better not to remind Bob of what a jerk he was.” Anyway, I say be nice to them. It's a good habit to try to get into.
It's possible that they are sentient, but the possibility I throw out in the final paragraph of my appendix is that this could be relevant to the question of how they'll treat us if they become this godlike power. I don't really mean so much that if we're nice to them, then they'll pay us back by being nice. What I mean is that the kind treatment of my dogs while they were alive was related to my belief that it was like something to be them.
It's hard for me to do the thought experiment: If you could convince me they didn't have sentience, how would I have been? I don't know, but I'm pretty sure I would have been less considerate and less nice. It may well be that we should hope AI is conscious because then it will relate to our plight more, be more reluctant to make us suffer, and be more inclined to help us flourish and even increase the amount of human sentience and help us make it, on balance, a positive experience.
I don't think that's crazy, but I would say, be on the safe side. No, it's not the safe side. I don't think how AI treats us will depend so much on how we treat it. I just think: treat it nice. Good habit-forming. It's possible that it's sentient, and it's possible that it isn't yet but will be. That's my take.
I do have a chapter on John Searle's Chinese room argument, in which consciousness enters, but not mainly for the purpose of dismissal. It's mainly for me to argue that if you want to have a serious conversation about whether AI is capable of understanding—which I think it either is or will be—consciousness should not be your criterion.
Nathan Labenz
Here's one more quote from the book: “If a silicon god, a superintelligence that plays a central or pervasive role in the affairs of this planet, does arrive, it could, for all we know right now, be a good god or a bad god. The one thing I feel confident of is that it will be, in some sense, the god we deserve.”
Robert Wright
What I mean by that is we have to pass the god test. One thing I say is that the interesting thing is, as I suggested earlier, AI is evolving. There's this level of selection—of models and of wrappers or whatever, and applications—and we are doing the selecting. We are the environment of AI's evolution.
I mean, also, in the sense that the engineers are part of our species, but I'm mainly talking about us doing the selecting. That's unprecedented: an environment that is shaping the evolution of a being is conscious of its role. That's a lot of responsibility, and we have a strong interest in taking it seriously.
I think we should do the things I've mentioned, like choose the AI models mindfully and carefully. Choose models that'll make us better people, in part because I think we're going to have to become better people to make this whole thing go well. If various alternatives to that happen and the outcome is bad—or even if, say, some cabal of billionaires and politicians seizes control of it and imposes a global tyranny or whatever—we will have failed to prevent that.
We will have failed to create the policies that prevent that concentration of power, and I think it's time to start thinking about those, for sure. I don't mean that if things work out badly, we will deserve our fate in the sense that it'll be morally good that we suffer because we screwed up. I don't believe in retributive justice. I just believe in practically useful punishment. But, yeah, in that sense, we'll have gotten the god we deserve.
I'll be curious as to how people read that, because I don't know. Did you read it as this grim, ominous thing?
Sober. I meant it to be sober.
Nathan Labenz
Yeah, I think it's certainly stark. We didn't even talk about how your interactions with Steven Pinker over the years are kind of woven throughout the book, but I sort of read the whole book as basically a call to the better angels of our nature, so to speak.
Robert Wright
Absolutely.
Nathan Labenz
I read it as a wake-up call. I think you have done a genuinely really nice job of recognizing that this is an important thing and figuring out how to make sense of it. There are some really nice passages in the book. I read some quotes, but I think there's also some really good stuff that this podcast audience doesn't need as much in terms of explaining why we should believe that these things are intelligent in functionally relevant ways and that this whole phenomenon is probably not about to stop right at a convenient point for us.
I do think that it's a very interesting bundle of compelling arguments that this is the case for people who need it, and then also the sort of “so what?” I definitely think this is going to be the defining challenge of our times. Are we going to rise to the occasion and manage the emergence of this technology well or not? The jury is very much still out on that.
But I do think that that open question is definitely an appropriate note to leave people with, probably for the book and maybe for this podcast as well.
Robert Wright
Well, thanks for that. As I said, it isn't mainly for people as steeped in this as you are. There are chapters they might find interesting, but my main hope is that they'll maybe recommend it to aunts and uncles and friends. That's who it's largely designed for, and maybe especially people of a certain age. I think people over 40 are more likely to read books these days than people under 40.
But, yeah, that's the hope. I really appreciate your paying this much attention to it, and I'm sorry if I've alienated your China hawk audience.
Nathan Labenz
I think we'll be okay. The book is The God We Deserve: AI and the Coming Cosmic Reckoning. I think it will make a compelling case that this is something people can't take seriously enough and that we all need to be thinking more about how we're going to do our part to steer humanity in the right direction as we try to navigate this increasingly tumultuous time.
Robert Wright, thank you for being part of The Cognitive Revolution.
Robert Wright
Thank you, Nathan. Keep up the good work. I'm a devoted listener.
Nathan Labenz
Thank you. And likewise. That's kind.