Nathan Labenz
Zvi Mowshowitz, welcome back to the Cognitive Revolution.
Zvi Mowshowitz
Good to be back. It’s been a while.
It has. I’ve been busy, and so has the rest of the world, and we’ve got no shortage of major events from the AI world to cover.
Zvi Mowshowitz
A lot.
Yeah, you have to have been busy, too. Let’s start with recursive self-improvement. I think if there are any historians around in the distant future, which could be as short as a few decades from now, looking back on this time and asking, “What really mattered in early 2026?” my best guess is that we’re in the period where we’re really starting to enter into a recursive self-improvement dynamic, from which there may already be no return, or where we may soon reach a point of no return.
I feel like, subjectively, we kind of went from late early. It was like getting to the end of the beginning, and now suddenly I feel like we’re maybe in the beginning of the end. Somehow, I feel like we missed the middle. But let’s start with just your reflections and observations on where we are with respect to recursive self-improvement.
Zvi Mowshowitz
Yeah, so we’re in the beginnings of steadily increasing amounts of AI-assisted improvement. I would say this feels like the middle. I didn’t think the middle was real. I think the reason people think of this as the end game is because they don’t believe in the actual end game, right?
They have this belief that we’re looking at an S-curve. They believe the models will be commoditized. They believe that intelligence will be commoditized. They believe that the future will look like the past, except with all of this cool intelligence behind everything. In the way that Star Trek is basically just modern humanity doing modern human things and talking about modern human issues, except with metaphors, they don’t really believe that everything will transform and everything will change. Because everything isn’t transforming yet.
If I had to use the metaphor of the beginning and the end, I’d say this is the beginning of the middle game, right? You’ve got the U.S. government starting to wake up and do crazy stuff. You’ve got the labs starting to pull away from each other, becoming importantly different and offering importantly different, recognizably different services, building on themselves in ways that are rockets to the moon in various different ways.
Then, frankly, you’ve got humans stopping writing the code, and you’re seeing cycles get faster and faster. But you aren’t seeing true transformational changes to the world. You aren’t seeing humans legitimately out of control of the process. You aren’t seeing humans out of the loop. Those are the types of things I would think would count before I would call it an end game.
You mentioned the S-curve mental model. One of the tweets that has rung around my head for the last few weeks was from Rosie Campbell, who used to work at OpenAI and is now, I think, working on issues related to AI welfare, sentience, and consciousness. She posted something to the effect of, “The S-curve can stay steep longer than you can stay relevant.”
I wanted to dig into the S-curve versus exponential for a second. I guess my mental model is an S-curve. Does it matter if there’s a difference between long-term exponential growth and an S-curve? If the S-curve plateau is high enough, my mental model is, yeah, it’s probably an S-curve, but I don’t think that really does much for us.
Zvi Mowshowitz
S-curve, unless our model of physics is very wrong, because there’s a limited amount of mass-energy, as we understand it, in the universe, and it can neither be created nor destroyed. That means there’s a certain amount of potential energy, a certain amount of potential utility, and a certain amount of potential intelligence that the universe as we know it can contain.
Unless our model of physics is very wrong—which, who knows?—we have some very, very strong beliefs about the things that are theoretically impossible. You can’t exceed the speed of light. You can’t extract more than a certain amount of energy from a given amount of matter. You can’t do more than a certain amount of work, and therefore the amount of theoretical intelligence you can get from that, and the amount of utility you can get from that by any definition, is limited.
So it’s an S-curve in some sense, right? But this is like sitting around in ancient Athens and saying, “There’s an S-curve. There’s only so much technology mankind can invent.” You’re right, but it’s not relevant to your situation.
Yeah, there’s a long way to go. Yeah, okay. I think that’s an important point of clarification for many debates, because people seem to really want to latch onto this S-curve idea, and it really doesn’t do that much for us in the end.
Zvi Mowshowitz
I think people, frankly, desperately want to tell themselves a story and tell other people a story in which it matters that they’re saving for retirement, it matters that they’re doing ordinary human things and planning for an ordinary human future, where things won’t change that much, where they don’t have to go crazy, and they don’t have to look crazy to their friends and family, where everything is going to be okay and everything is going to be normal in a fundamental sense.
It is important to hold on to those things, to prepare for that scenario and to stay sane. But they want the ability to just push all of this aside, basically, and say, “Here’s why AI is not that big a deal. Here’s why AI will be only Internet-big, or not even Internet-big in some cases.” But all of that is starting to fade away. I understand why they feel the need for that. I understand why they latch onto that, but it’s simply not looking like that’s going to be the case. It’s becoming increasingly unlikely that we’re going to stay in that zone, and people have to come to grips with that.
The S-curve—yeah, at some point we’ll hit one, but every day that we don’t see that happening is one more day. Basically every few months, someone will come with a new study and say, “Oh, this proves that we’re in an S-curve and they’re all wrong.”
Even if we were to hit the top of the S-curve of fundamental capabilities, which is the curve that we care about, and we hit it with GPT-4.5 4.5.4 and Opus 4.6, and these are the best models we’re going to get and it’s all iterative from here, they’re still vastly underestimating what’s about to hit them even then. And 5.5 and 4.7 are coming, unless they’re 6 and 5.
Many people would presumably have to change their narrative about this if there really were a big displacement of human workers, right? If we started to see major, sustained layoffs, rising unemployment, and so on, it seems like that story has hit, perhaps again, the beginning of a trend in the last few weeks. How do you understand that right now?
Again, we’ve got competing narratives. The CEOs themselves are saying, “We’re doing it because of AI, and we’re going to be more efficient.” Stock prices seem, at least in a couple of notable cases, to have bumped in response to that communication. The counter-narrative has been, “Well, you way overhired during COVID and the zero-interest-rate timeframe anyway, and so you have an incentive to say it’s about AI, but really you’re just trying to undo previous mistakes.” There’s probably at least some truth to that as well.
How do you parse the AI layoff story, and what are your expectations for the next, say, quarter?
Zvi Mowshowitz
It can always be a coincidence up to some point. But every month, every indicator keeps going in the same direction, and things keep going farther. Excuses like “There was dead weight to be gotten rid of” become less plausible as we get into 2026, because it’s been several years.
We keep seeing statistics tell a consistent story that some of us were predicting the statistics would tell in advance. I don’t remember people saying, “Oh, yes, you’re absolutely going to see increases in productivity, decreases in employment, and all these announcements of job cuts due to AI, but it’s not going to be real because of these explanations.” I don’t remember anybody making that prediction. These are people in hindsight saying, “Oh, if that’s true, then it must be because of this.”
But that would be at all credible if they had observed that, theoretically, the overhiring was already there. It was already clear. I believe there was overhiring, but nobody then said, “Here’s how this is going to play out,” that I can remember. Also, the statistics are just coming in consistently, telling the same story over and over again: productivity up, GDP up—real GDP, not just nominal GDP—inflation held in check, and employment down.
Employment has been revised down every month, over and over and over again. Employment is always confounded, right? You could say, “Well, it’s the tariffs.” You could say, “Well, it’s the aftershock from COVID.” You could say any number of things.
But it seemed like a really, really big coincidence to claim that all of this is happening at the same time when, in fact, the stock market impact of the tariffs got basically erased—the tariffs have been reversed. Obviously, now everything is confounded by Iran. But up until the Iran conflict started, it seemed like we were getting a very consistent story that this kept happening. The job numbers kept getting revised down, the GDP numbers kept not being revised down particularly, and productivity kept going up.
Everybody out there, in practice, keeps telling the same story. On the street, at least when I talk to people who talk to people who are not involved in our world, there is widespread, on-the-ground, normal-person paranoia that if they lose their job, who knows when they’ll get another one in a wide variety of white-collar work?
Not that many people are getting fired yet, because it’s really, really hard to replace a worker. It costs a lot to train a worker. You want to be conservative. You wait until you’ve actually made it do all the work. But everybody is asking, “Who knows? Who is hiring?” Who wants to take on new workers and train new workers for these jobs when you could train an AI to do it instead going forward?
By the time I spend 2 years getting you to be productive, maybe I don’t need you anymore. It’s a harbinger of the future. This is decreased labor power. It’s made everybody feel paranoid, and it’s made people fearful for the future. The kids are freaking out about exactly these problems. They don’t know what to study, and they don’t know what to try to do.
It’s a very galaxy-brained take to say that all of this is in your head. If you think, again, that this is the best AI will ever be and it’s just a diffusion story from here, it’s still a hell of a story. To say, “The job market impact is in your head. It’s not real,” is a very galaxy-brained take.
Now, there’s the nongalaxy-brained standard economist take, which is, “Yes, there’ll be a transition period, which we’re entering now, where a lot of existing jobs will get much more efficient. They’ll get automated, augmented, or eliminated. And we’ll transition, but that’s okay because Jevons paradox will mean a lot more demand for things like software engineers, along with a few other stories.”
Then, of course, with our new wealth and productivity, we will find other jobs for people to do. There are plenty of jobs for people to do. I hear the Department of War might be hiring.
That’s always been the case, right? The people who say, “The automation of agriculture will kill our jobs”—they were right that their jobs would go away. They were wrong that unemployment would fall in the long run because, of course, we went and did other things.
The whole reason this time is different—that a lot of us believe this time is very different—is not that technology has never taken current jobs and made them largely obsolete. That’s happened many times. The reason why we think this time is different is because AI is going to do the new jobs that would get created as well. It’s going to happen quickly and en masse.
Therefore, you never exit this transition period. Humans don’t develop new things. We don’t necessarily think there are going to be enough tasks for all the humans that the AIs can’t just replace. This is going to create potentially a large number of people who cannot retrain themselves into a new position and develop something to do that people value fast enough before it simply gets replaced over and over again by AI.
Also, a lot of people are putting a kind of resiliency and ability to shift and adjust into people that people mostly just don’t have. When these things happen over generations, it’s much easier to deal with than when they happen over the course of years or even months. That’s unprecedented in human history, and people will not react to it very well.
Again, all of this is in a relatively milquetoast, normal-world scenario where all of this is happening. In the more advanced scenarios, you have much bigger things to worry about.
Famously, Tyler Cowen said he thinks we can get 0.5 percentage point, I think, of additional real GDP growth out of AI, and that would be amazing in his opinion. What would you say? This might be too hard because, obviously, there’s a ton of noise in the data, but based on what you’ve seen so far—productivity measures, et cetera—if you had to put a point estimate on what we’ve got right now, what would that point estimate be?
Speaker 1
I haven’t tried too hard to estimate exactly, and I don’t think anyone else really has either. I’ve seen various different attempts to guess. If I had to guess, we’re currently enjoying 0.5% to 1%.
Nathan Labenz
Yeah, that’s kind of what my gut says as well.
Speaker 1
This year, there was a lot of downward pressure on what would normally be the American economy. You had a lot of fear and a lot of regime uncertainty in terms of the tariff regime and various other policies. Usually, the types of things that happened in 2025 would be rather bad for business. Instead, things were good for business. I think this is a lot of why things felt like they were okay for business: They weren’t okay for labor.
And certainly not in the felt experience of labor. So, yeah, I think it's on the order of 0.5% to 1% right now, but I think the market is pricing in at least that much indefinitely going forward, and likely somewhat more, even if it doesn't understand that's what it's doing. The market has done well because it is anticipating the benefits of AI. I think if it wasn't anticipating that, you would have seen a very, very different set of reactions. And that's exactly why, when things started looking awkward for the American economy in other ways, I didn't sell anything. I just held on because I knew that AI was going to prop things up.
Nathan Labenz
It does seem like there's been a bit of a split recently. People talk a lot about—and I don't necessarily buy this frame—but there's the worry about being part of the permanent underclass, or making your way into the upper class before things get locked in. I don't really buy that, or at least it certainly doesn't frame much of my thinking on an individual basis.
It does seem like that dynamic might be coming to stocks, though, because we do see Anthropic drop a relatively minor product extension, in the scheme of everything they're doing, and apparently, pretty closely correlated to that, you'll see stocks drop. Are you buying that? The stock permanent underclass?
Speaker 1
I want to hearken back to one of the most important movie scenes in history, which is Bane confronting his backer. The person who hired him says, “I'm in charge here.” Bane looks at him and says, “Do you feel in charge?”
I think there is this very clear idea that people think, “Oh, if I have the rights written down in electronic databases, if I have the stock certificates, if I have the private property, then I get to be one of the special elite. Everyone else gets to be one of the underclass. I have to be one of the people in the elite,” even though I will also not be productive. Even though I also will not be in the loop. I will also not be able to exert meaningful optimization pressure, except for in my technical authority as the person whose marks are in the database.
I think the idea of relying on marks in the database in this kind of world to keep you alive, to keep you meaningfully wealthy in a way to consume physical goods, to have a good, prosperous life for you and your descendants, is hopium. It's not a good plan.
If you think that humans are sufficiently useless that most of us end up in a permanent underclass because we cannot be economically productive, then your best-case scenario is you slowly lose your wealth to various different extraction methods. And the more likely scenario is all of that gets ignored by facts on the ground, by physical reality, just rendering all of that irrelevant. Or the system gets taken over and subverted, either by humans using AI or by AIs.
The long history of the world, even with much less severe disruptions, does not have a good track record of private property surviving over long periods of time when someone else has the guns, when someone else has the swords, when someone else has the power—in the sense that it's the meaningful power. If you don't have a reason to keep it, you don't keep it. Not really.
We've seen the collapse of the colonial era, we've seen the collapse of basically almost all the ruling regimes, and we've seen many, many examples. Even if you're not going to take the arguments around AI seriously, I can't believe you're counting on this.
When people talk about the permanent underclass, they talk about, “Oh, I'm going to have the skills to be productive with coding agents, so I'm going to be a valuable person going forward. But I have this short window.” Why is this short window? If humans can scale up and be productive in the future, you can scale up and be productive in the future. And if you can't, you can't. So there's no particular urgency there.
You can make the argument that surely this is one of the last chances, in some sense, to use your talents, that there's this window where your talent can suddenly create a billion-dollar company or even a trillion-dollar company and you would have a large portion of that. You can enjoy a lot of money. I just don't think it's particularly relevant.
I think that, in practical terms, there are 2 likely ways this plays out if AI is for real and goes the way I expect it to go. One way is we lose control of the situation in various senses. Everyone dies, or there is at least massive loss of control, massive loss of resources, massive destruction, and massive disruption. In either case, none of this is going to be that relevant to you, particularly.
The other scenario is things go relatively well. Then there is basically an abundance of real resources, and humans are more or less in control. In that case, I think if you're a citizen of the United States, you don't really have to worry very much about your material needs. Sure, you're a member of the permanent underclass. Congratulations: you have a real income that today would be called $1 million. You have access to robots and intelligence that caters to anything you want. Your day is free. You don't really have to work.
It's probably not exactly how it plays out, but basically, you're going to be better off except in relative status terms. You really shouldn't have to care about that. You need to get over the fact that you are not relatively wealthy or relatively respected in that world.
I'm sure there will be ways in that kind of world to compete for status in the human hierarchy. There will be ways to meaningfully occupy your time. If we're still in control, we'll figure that stuff out.
Then I guess there's a third scenario where some cabal uses AI to take over. In that case, you need to be in that cabal if you want to be part of the people who take over, and they then enforce it to become the good world instead of the bad world. But just making a bunch of money is not going to get you in. That's not how those worlds work.
Again, mostly you should be trying to make sure we don't get into the bad world where we lose control, and then we get into the good world where we retain control and humans are in charge of steering what happens.
Mostly, I expect that even if a small group has the ability to steer that world, they'll steer it in ways that we're pretty happy with. A lot of people talk about how Altman might try to make himself god-emperor, or Demis Hassabis might try to make himself god-emperor, or Dario Amodei or whatever. That's not a preferred outcome, to be clear, but I don't think that if Sam Altman became god-emperor, I would be that sad about it in terms of my practical experience. I think my life would be fine. I think my children would be fine.
Nathan Labenz
Yeah, he does have a little bit of a Roman Caesar kind of vibe to him, where I think he does fancy that sort of power, perhaps. But also, when he does things like fund universal basic income experiments out of pocket, I view that as, honestly, probably a genuine magnanimity on his part that I would expect to probably extend into—
Speaker 1
It's very easy to think of this when there's a true, full abundance of resources. He could, in theory, keep 99% of the value of the light cone for himself, and the rest of us can be very, very happy with the rest. I'm not saying this is my preferred outcome, to be clear. I do not want this to be how it plays out. I think this is bad.
But that's still better than not building it, right? That's still better than various different horrible outcomes, especially everyone dying. But, yeah, it's better than the status quo in an objective sense.
Except for your opportunity cost. So what I'm worried about is, in fact, losing control of the situation. The reason why you don't want Sam Altman making the decisions is not because he might take over. It's not primarily that, right? In my view, it's because he might make decisions that cause us to lose control—for him to lose control because he is being irresponsible. And that is the reason I am terrified of him making these decisions, right? It's not because I think he is an evil man.
It all comes down to normativity in my mind. Normativity is the concept that good things are good, bad things are bad, and you want good things to happen to people in general. Not necessarily every person. I like to think there are some bad people, but good things should happen to good people, and most people are good.
If you believe that, then basically things will work out even if there's not exactly your preferred outcome. And you know what? I think most people are normative. I think even the people in charge of the labs right now, most of them are normative. And to the extent there are relevant people who I don't think that way, I simply say nothing.
Nathan Labenz
I will say I do think it's pretty distasteful when I see people talking openly about trying to escape the permanent underclass. My reaction to that is always, why don't you try to make things good for the permanent underclass rather than try to escape it? That just seems like really flagrant defection.
Speaker 1
Flagrant defection to focus on getting a seat on the ark if you think the world is going to be flooded. That's a terrible, terrible position. You should be trying to stop the flood. You should be trying to save more people and build another ship.
If all you're trying to do is get one of those precious few seats on the ark, that's not a good thing to do. At the same time, there are times when the best thing you can do is just escape the bad regime.
There are, in fact, times when that's like, “Put your vest on first. Get out while the getting's good.” Because what else can I do? And I respect that.
But the way they talk about it, yes, it has the bad kind of elitism—the kind of complete disdain for the common man. Indeed, very disgusting. I really don't like it.
And yeah, I expect we can make the permanent underclass pretty neat. In fact, most of the time, the world has had a permanent underclass, and the permanent underclass has never had it better than it has it today.
Nathan Labenz
Yeah, it's important to remember that. What would, in your mind—and maybe we could talk about this in terms of handicapping the timeline and giving a little qualitative description—represent the transition from the middle to the end game?
I've started to think, of course, everybody who follows this feed knows the general range: Dario's still kind of on AI 2027, more or less. Demis is more like 2030. OpenAI has a March 2028 timeline for their fully human-level automated AI research paradigm kicking in.
I assume you're somewhere in that range in terms of expecting things to enter into what you would call the late game. How would you call that transition? I guess maybe we should even understand your view better: What is the difference between middle and late? Is it some sort of event-horizon point of no return, or is there some other concept that would separate those? How would you call it? What are you looking for? When do you think that is most likely to happen?
Speaker 1
The end game is, I would say, when the AIs are largely running the show—at least in the further development of AI. Right now, we're seeing AIs augment humans. The humans are making the central decisions, reviewing the code, and making the plans. The AI is a multiplicative factor. It is something that you're supervising; it's something that is enabled. When that changes, things get very strange.
One of the key aspects of AI 2027, the tabletop scenario, is that your progress is mostly proportional to your compute allocation. You move up the timeline at a rate proportional to what percentage of the world's compute you have, because your researchers no longer matter very much, right? You're telling the AI, which is already effectively as smart as or smarter than you are, to go build a better AI and align the AI.
You're making decisions about what percentage of the compute should go to safety and keeping this thing steered properly versus increasing its capability. In the early parts of it, you can assume that's going to be respected, and then eventually you can't assume that's going to be respected. But at some point, your researcher talent doesn't matter very much because the AIs are your researcher talent. All that matters is where you are on the curve and how much compute you have.
Right now, I would say the top labs have a dramatic talent advantage in these types of scenarios, and I think this is a reasonable hypothesis, although I'm far from certain. Anthropic seems to have the best talent and gets the most out of every given amount of compute they spend. Then you have OpenAI and Google, which have strong talent. They get a lot out of the compute they spend.
And then what's going on with xAI and Meta, right? They're throwing tons of compute at this problem, and they're falling farther and farther behind. Or at least this sort of appears to be true. Functionally speaking, I think that's a lack of talent, right? That's their inability to execute as humans. The end game comes when it doesn't matter that much which humans you have.
Right now, we're playing centaur chess, right? The human plus the coding agent is much superior to the best human on his or her own. The AI agent on its own doesn't do anything. You need a centaur. The moment the human doesn't matter anymore—when you transition to the human being somebody who's just able to do some common-sense stuff and isn't even one of the best players anymore—we're starting to talk about end-game-style scenarios.
Also, when you approach the event horizon, the release time for a meaningfully different model starts to go to a month, then a week. Things get really fast. The ultimate cascade scenario is that you just leave it on overnight, and then you wake up and things have happened. You're in the end game at that point.
You can also think of the end game as when the world starts to actually transform. You start to see the robot factories being built. You start to see large amounts of area being terraformed. You start to see massive job disruptions. You start to see governments begin to do major interventions. You start to see this become the issue, the thing on people's minds.
We just had one of the most important developments in the history of AI happen with the designation of Anthropic as a supply-chain risk. It wasn't even the most important thing the Department of War did that day, according to most people, right? It gets completely subsumed by the exact same person who sent the tweet out doing something else that night. That has gotten 100 times more coverage because people think that's the important thing that happened. I think it very much remains to be seen what the important thing was that happened that day.
Nathan Labenz
I think I have to give you some base points when it comes to drawing the circle around who the live players are. I think, for many conversations going back in time, the record will show you've pretty much always said that the same 3 companies you just listed are the real live players. I've always been grasping at who we might add and what might be the rationale for adding them.
I asked you this question a minute ago about the permanent underclass of stocks. You went in a different direction with it, but—
Speaker 1
Yeah, I got distracted, but—
Nathan Labenz
All good. I guess what I'm taking away from your analysis is that the permanent underclass of stocks might be almost all the stocks, and it might even extend up to big tech, like Microsoft and Amazon. Obviously, they have at least held their own or done pretty well so far.
But if you have a model of who can get into this next regime first as really mattering most, and only 3 companies right now seem to be well-positioned to do that, then basically you're short everything else over a 3-year time frame. Is that accurate?
Speaker 1
I think there's a lot of different parts of AI, and you can make money doing a wide variety of different things. The real world always takes longer than you think it does relative to the immense technology. The diffusion is slower than the original idea. So I think there's still a lot of room for a lot of different groups to win.
I think there'll be a lot of, “Oh, we're going to buy out people who have useful components, even if eventually we wouldn't need them anymore, because it's faster.” I wouldn't expect everybody to go to zero.
One thing—what's going on, basically, is that the stock market is forward-thinking, and they're trying to do some half-hazard forward thinking. They say, “Oh, okay, this software-as-a-service business had a good business, but in 5 years they won't have a business unless they reinvent themselves and create new products. But maybe they will reinvent themselves and create new products.”
Their original basis of valuation has kind of been destroyed. And that works out into a variety of other companies. I think a lot of companies, in fact, are going to be long-term losers, but it's always been true, right?
If you look at the S&P and you play this game of, “Okay, let's take the 10 top performers for the next 10 years and just assume everybody else doesn't do so well,” the gains are because some companies do really well, as they always have, and everyone else on net kind of languishes. That's why you want to be diversified, because can you pick those 10 companies? The answer is no. Most people can't.
But in this case, I think you can make your own guesses, but you don't really know. Also, keep in mind that the market's dumb about this. The market is reacting on a very superficial level. People, if you're listening to this podcast, have thought much more intelligently about the situation than the market has. I'm not just talking about me.
They announced that Claude was going to offer a COBOL product. IBM was down 10% in the week of that announcement because some in the stock market woke up and said, “Oh my God, AI can write COBOL code? AI can translate COBOL code to regular code in programming languages people know? IBM's business is in trouble.”
And the rest of us were like, “You didn't know? This is news? You think building a slightly easier-to-use tool changes things at all?” What has happened since then? IBM has fully recovered. It turns out that, yes, in fact, either it's been priced in or they went back to not believing it or something. Clearly, the market doesn't know what the hell's going on.
You see this over and over again. The market is very slow to update on these things. They do see kind of superficial reactions. The market has reacted the wrong way to Nvidia many times in response to very clear news where demand for Nvidia's product is up and Nvidia's stock responds by going down. That is not how economics works.
That is not how capitalism works. And yet, here we are. So, yeah, I would say there are some clearly good buys. They're not as clear as they were a year or 2 ago, or several years ago, when it was just completely, utterly obvious what some of the good buys were. Now, the price—the multiples—are in fact respecting a large amount of growth in those stocks.
But it still seems pretty obvious that, yeah, if you were to buy a basket of the stocks that seem clearly positioned to do well and then short the rest, it would be a very good strategy in application. Your thesis would have to be very wrong. We'll set that aside.
And so, on the lab players, as I said, I think the talent has really proven to have migrated to the big three. I think there's a large and growing gap between 3 and 4. So, if I had to pick a potential 4th at this point—there are a few possibilities—obviously Meta or xAI could in fact have gotten their shit together and managed to find people who can execute.
But I don't see the evidence of that. We just learned that Meta postponed its next release again. This needs to be continuously reshuffled. Trouble in paradise. Things are not going well, as far as we know, in Meta land. And xAI put out probably the most disappointing major model release of any major lab in the history of large language models with Grok 4.2. I see no sign of much happening.
They're not doing the things you would do if you wanted to recruit good AI capability or safety workers. They're actively dismissing and disbanding their safety team. They pooh-pooh the idea that anyone could actually be in charge of safety. They said, “Well, safety is everyone's job. That's how it is at Tesla and SpaceX,” which is not true. Tesla and SpaceX have very dedicated teams and very dedicated safety people who make sure everything is safe, because anything else would be completely insane.
But Musk just doesn't understand that you can't run software engineers into the ground, ask them to have miserable life experiences, ask them to align to your personal preferences and whims all the time, and then hope to get the best talent. How exactly? Why do these people want to work for you? You'd have to give them even bigger packages than Meta, and that's not going to happen. So he's not going to get the best talent.
Meta is, “I'm going to spend infinite money if this will get me the best talent.” And I think it has worked. So they seem to be out of the competition. For the Western side, that leaves us with the big 3.
On the Chinese side, there have been a lot of stories over the last few years about how this Chinese lab—originally Deep Seek, but also Alibaba, Gimme, and whoever else—they're the ones who have the new hotness. They're the ones who are catching up. And I've said before, DeepSeek has some decent models, Gimme has some decent models, and some other people have some decent models. But nothing that close, nothing that scary, nothing that told me they were actually catching up.
And that seems to have been borne out. Obviously, at some point I could be wrong. DeepSeek V4 will be the last and one of the most important remaining tests of this thesis. Most of them have already come out by now. I'm not sure what's holding it back. But when V4 comes out, we will compare it to Claude Opus, and we will compare it to GPT-5.4.
If that is not the right comparison, if it is not trying to play in that league and is clearly still far behind, then I think we can more or less say, “Okay, DeepSeek had a DeepSeek moment, as we call it, where the stars aligned to make everybody super excited about what they were offering.”
But mostly what they were offering was how to do more with less. They're big geniuses of efficiency: working on the bare metal, figuring out how to get something really good out of not very much, and then putting it in a really great package at exactly the right time, giving it some user-friendly features, attracting a bunch of market share, attracting a bunch of attention, and scaring the shit out of everyone.
Since then, it's been quiet. They've still done some cool math stuff. Don't get me wrong; they've had some innovations. But they're just not playing on that level. And this is their last chance to prove me wrong.
I think you kind of have to more or less dismiss them as that level of competition. You put them back in the pack with the open-source league. They're going to be in a different league. They're not necessarily the best in that league, and they're not necessarily not the best in that league. It's unclear.
But I think that's a very different league that's substantially behind, and it would be very, very hard for them to get out in front and actually innovate, because they're fast followers. I have respect for fast followers, but it's a very different skill from trying to do something like that in the lead. And I don't think they can compete, frankly, with the kind of recursive self-improvement we're seeing with PaLM code and Codex. I want to see them try.
I think anyone who has to do that is pretty doomed here. And in fact, I'm starting to think about it. When I watch the models, when I watch how people talk about the models, when I see how they talk about their scaffolding, and when I see what they're doing, I think Google is in danger of dropping out of the top tier.
Nathan Labenz
Okay, that's a big claim. Let me come back to that in just a second. On the Chinese—I guess I have a few different follow-up questions. On the Chinese companies front, specifically, is it talent or is it compute?
I think you could at least make an argument that the talent is there and it's just not the compute. If compute allocations were to change, then maybe DeepSeek, maybe Alibaba—although there's been major disruption there too, as far as I understand, recently, in terms of team changes, let's say.
Speaker 1
I love this, especially because their compute is lacking.
Nathan Labenz
But if that were to turn on—whatever, let's imagine a certain executive decision allowed that to change, or perhaps a technology—
Speaker 1
Yeah.
Nathan Labenz
—breakthrough just domestically in the Chinese manufacturing side—
Speaker 1
Yeah, I mean—
Nathan Labenz
—would you then expect any of those 2 to become a giant?
Speaker 1
It's not—
Nathan Labenz
They could—
Speaker 1
In terms of the manufacturing side, I think that's basically not possible in the sense that it would take many years to physically play out. Even if they figure out how to do it, they would need to scale up. These things are physical; they take time. I think we're talking about 5-year-style timelines.
And if things are going to come to a head faster than that in many ways, then it kind of doesn't matter at that point. The only way they're getting the quality and quantity of chips necessary to be competitive at this level is if we give them to them. Huawei is not going to manufacture them fast enough at scale, even if its efforts are completely successful in terms of what it's trying to do.
We can just rule it out at this point. If we're talking 10 years down the line, maybe they can do something relevant. That's still not that much time, because by definition, the AI they'll be working with won't be that advanced. In 20 years, sure, but let's not get ahead of ourselves. This is a pretty big advantage.
In terms of talent, in terms of underlying, raw human talent, obviously China has tons and tons of talent. There's tons of talent out there. A lot of these people have studied machine learning. A lot of these people really want to go; they have all the right attributes. No doubt. It's a big country with a very good educational system, a lot of very smart people, a lot of people who care about this stuff, and a lot of people who are desperate to find something good to work on.
There are advantages to having tons and tons of unemployment. It really drives people. China's got problems. But at the same time, they're focused on a very different style of skill and style of problem, because that's what the Chinese are pushing and that's what the Chinese incentives go toward.
These people are skilled in the ability to deliver these types of open-source, efficient developments. They're driven to ask, “How do I do small well? How do I do fast-following well?” The entire ecosystem is built around these different types of skills, these different types of talent.
I think there's a really big difference between one set of things and the other set of things. It's the same way that OpenAI has very strong talent. They decided to build an open-source model. They created an open-source model, and in some ways, it has its charms.
In some ways, it's the best model at certain very specific things, maybe even overall in certain ways from an open perspective. But for the purposes the Chinese models are mostly being used for, it's useless. It's just not a very good model for those purposes, despite the fact that OpenAI has internal access to its talent, its compute, and its best people.
Because it's just a very different skill, and I think it works the other way, too.
And so I think that if the Chinese suddenly got this compute, there would be a transition period where they would have to learn how to do what the major labs are doing. They'd also have to build their own synergistic harnesses and scaffoldings, and learn how to do all the stuff that Anthropic and OpenAI have been doing over the course of years.
For the long term, do they have the talent? Yes, obviously, absolutely. I don't think America is special in that sense. But I think our lead is bigger than it looks, and more robust in some ways than it looks. I don't think we want to test that theory and find out by suddenly giving them compute at parallel scales.
Nathan Labenz
See, can you help me understand a little bit better what the difference is? I think one thing that has obviously been talked about a lot recently—and Anthropic even put something out saying, “We're seeing this,” right?—is the distillation from American frontier models happening in the Chinese companies.
My take has been, sure, they might be doing that. They definitely like to take shortcuts, and especially if you have a story where you're saying, “They're cutting us off from compute, so let's take whatever shortcuts we can get,” you can tell yourself why that makes sense to do.
But if I think to myself, okay, what would be fundamentally hard? If I was going to try to start Nathan's AI company today, I would say, well, going and getting a bunch of expert data, kind of like what Scale and other data providers have done, is obviously very resource-intensive and time-consuming. But I feel like, at heart, I could run that project. Give me $10 billion, and I can go build out the network, hire the people, and get that flywheel going. I don't feel like there's anything there where I'm fundamentally outclassed by the people doing it. Maybe I'm wrong, but I don't feel that way.
Whereas if you said, “Hey, here's all that data. Turn it into a frontier model for us, Nathan,” I would be like, “Oh, wow. Okay, this is going to be really hard.” I do think the people at the top companies are clearly outclassing me in their ability to do that.
So when I look at the Chinese situation, I'm like, okay, sure, they might be cheating in a sense, stealing in a sense, to get the data. I'm also still kind of like, well, the hard part is what to do with that data, right? It's a big investment in data that they're taking a shortcut on, but once you have it, you still have to know what to do with it.
Speaker 1
Yeah, I think what's going on is you're conflating some different things in your model, and that's causing you to get confused. There's the data in terms of the raw data from the world, from the internet, from books, and from other sources that you're using as your baseline. I agree that you could run that project, and the Chinese can run that project. I'm sure that the Americans are investing more in that, and that they have a richer base in some sense, but not in a way that probably matters that much. Even if you have twice as much data, if it's all similar quality, it's only a factor of 2. In this world, it's all about factors of X.
What matters for that data is how you clean it, how you figure out which parts of it are important and need to be upscaled versus downscaled, and how you emphasize it. That stuff is much harder and much more valuable. I expect the American labs have a large edge in how they get their data ready at this point, although I don't know. It's something that's internal, right? Maybe the Chinese have actually specialized in doing this really well. Maybe it's one of their areas of relative strength. I just don't know. It's my guess, but I don't know.
But that's different from what we're talking about with distillation. Distillation is not an attempt to extract the trillions and trillions of tokens that went into the model. Distillation is an attempt to use the model's intelligence to figure out how the model reasons, how the model makes decisions, what types of behaviors the model exhibits, and how we test it, especially in cases that we're very curious about.
Then how do we use that to train our model to follow that pattern and use that pattern matching to emulate the model? That's uniquely different data that didn't exist when you were creating the original model. It's uniquely useful in creating something similar that can emulate those skills.
It's very different to have a physics textbook and then talk to a physics professor. Then you talk to a truly expert, world-class physicist. Distillation gives you that real expert, where you have effectively unlimited access, if you have tens of thousands of spoofed accounts, to see exactly how the person you're trying to emulate responds.
Imagine I am an actress and I'm trying to make a biopic. I can read all the books written about the person—whoever she was—what she did, how she acted, and what the world around her was like. But that's not all they do. They do that, but what matters is that they talk to the person they're copying if they possibly can. They spend time with them. They talk to them, interact with them, copy their mannerisms, see how they respond, and ask them about hypotheticals. They distill this person through these direct interactions, and that's so much more efficient. That's what they get to do.
Nathan Labenz
So, to summarize that back to you, you see the advantage that the American companies have as not being in the collection of the data, but basically in knowing what to do with the data. You're highlighting a difference in kind in the data: the data that we can go out and source from the world needs more post-processing than the data that we can get directly out of a model.
Speaker 1
More post-processing, and also, it isn't exactly the data that you want. When you're distilling, you get exactly the data that would be most useful to you, that you think to ask for. It's like you can have a 1,000-page textbook, or you can ask 10 pages of questions and get answers. The second one is probably a lot more useful to you.
Nathan Labenz
Yeah, okay. That's a helpful update or refinement to my understanding. So, you mentioned Meta, and I'll just call it Elon Corp at this point. He's also got some joint ventures between SpaceX and Tesla when it comes to manufacturing robots and stuff.
Speaker 1
That's the best thing to say for now and then see what happens. But yeah, I think if you're counting Tesla's self-driving and stuff, then I guess it's weird. I guess.
Nathan Labenz
The question is, tactically, what moves do they have? I mean, get your act together, catch up, whatever. That's one. That seems like it may be slipping from their grasp. Meta seems like they may still have some play in terms of, if you release a good enough open model, you can maybe disrupt the business of the others or create some sort of alternative that takes the wind out of their sails. And for Musk enterprises, it's clearly something to do with real-world-deployed, physical, scalable intelligence in cars and robots.
Speaker 1
I think they're in very different spots. Meta is the kind of easier conceptual one. Meta is trying to sell ads. Meta is trying to put features on its smart glasses and develop a metaverse. Meta is trying to be a consumer company that creates products that people are willing either to spend money on or to let their eyeballs be captured by, in various senses.
They are very good at monetizing that, want to improve how they monetize it, and want to build better products. If I were Meta—if you suddenly said Zuckerberg's out, you're in, here's your controlling shares, and you have these goals—I would not be trying to build frontier models. I think it's a mistake. There's just no reason to invest all that money into something that other people are very good at.
Obviously, if you're thinking, “Whoever has the best models controls the world. This is the entire future of humanity. This is the singularity. We've got to be in the game,” then you do what you have to do. But if you are a businessman and you don't believe in that—and when Mark Zuckerberg says “superintelligence,” he means superintelligence, right? He means super at providing a good Instagram feed. It's super; it's intelligence. He's not acting that thrilled. And if you're not that thrilled, honestly, there are 3 companies that have very good products.
License one of them. No, partner with one of them. All 3 of them will take the call. I realize that it's a product of mass surveillance, but I think they can work out something for Instagram that keeps everybody reasonably happy. All 3 of them would offer them a very good product at a much, much lower price, and they could work something out where they paid for a license to use it internally for their business purposes, without having to pay retail prices.
I would just give up, honestly. You're writing these giant $100 million-plus-a-year checks to these various people. You're trying to compete in the world you're not going to be competing in. It's fine. I would give up.
Alternatively, I would try to buy one of them. I mean, you can buy Google, but you're still worth a couple of trillion dollars. Maybe try to do something else, but I would just give up.
Musk is in a different position because Musk is explicitly trying to become God Emperor, right? He's explicitly said, “I think this is going to potentially kill everybody. I think this is the most important thing in history.” And his response to that is, “It has to be me. I have to be the one to do it. If anybody else does it, it'll go bad.” You know, it's only I who can fix it, right? Only I can solve this problem.
Slash, if the world is destroyed, it has to be me who does it, or else I will just feel so—what am I even doing? Slash, he thinks he lives in a simulation, which I think is actually affecting his decisions, and I worry for his sanity in various ways, for various reasons. But he can't—this is what he thinks matters, and I think he's right in the fundamental sense, so he has to catch up.
I think he has 3 plays at his disposal, and he's trying some of them. Play 1 is that they can bet this is all about compute, and therefore it's all about a combination of money and energy and the ability to acquire chips. And, yeah, launching—you know, you have SpaceX, so you're the only thing launching into space.
Maybe these are going to space, and maybe you are the one who can clear vast deserts to put your solar panels in. Maybe you can have just way more compute than everybody else. And if you can hold on until the intelligence in the models—you know, you can go into self-recursion mode and you can win, right? In that sense.
I don't have much faith in this plan. If this plan is bad, it's not like it's no plan, but I am very skeptical of the data centers-in-space plan on the time frames that are being talked about, because physics. Space is expensive and hard, and you're solving problems that don't need to be solved at very large expense, the same way that we're not currently mining the asteroids, and there's a reason. Not that we'll never mine them, but chill, you know.
I just don't think that these limiting factors are going to be things that actually stop anybody, and I don't think this makes up for a lack of talent. I don't think this makes up for a lack of internal scaffolding and infrastructure. Where is Grok Code? Where is Grok events? I mean, I see Grokipedia, which is just a giant pile of slop. So it's not the same thing.
Then plan 2 is what you talked about, which is physical-world modeling: being able to have access to the real world. I can create self-driving; I can create robots because I have better training data. I can then build physical infrastructure in the real world, so I end up mattering more even though my intelligence is not as strong.
I mean, it's a play. If the technology plays out such that we hit a wall in other ways reasonably soon, it could be a really—it might be a relatively decent play. I think it's overtaken by events, basically. I think that this is not where the battle is won and lost. But again, it is where his competitive advantage goes, and it makes sense to make a bet that this is where my competitive advantage is.
I just don't know. I don't see that as going so great. I don't see that as that promising. I think there are quite a lot of people who can manufacture things in the world. Many of them are in China, but many of them are not.
The idea that because it's internal to Tesla, he will have some sort of huge advantage over people who have to make deals with actual manufacturers—the actual manufacturers are not going to be that expensive to buy. Right? Anthropic is worth more than McDonald's or Coca-Cola already. If these are $10 trillion companies in 2 years, they can buy whatever U.S. Steel, or whoever they need to buy, to make stuff if they need to do that. That's not going to be a problem.
I don't think he has the scarce inputs, essentially, in this scenario to pull it off. The question is, does he have scarce inputs in terms of data? I'd say probably not, at least not any that are relevant. Also, I think this is an underestimation of how important just raw intelligence is.
If you want to drive a car, you do not try to get the dog to drive a car. You start by designing a human who is really smart, and then you have the human learn to drive a car relatively quickly. It's not that extreme; I'm just trying to illustrate the idea that you want to focus on actually getting the geniuses in the data center, as Dario Amodei puts it, or something even greater than that.
If we have abstract superintelligence over here and physical-world skills over here, you make the bet that physical-world skills are how you develop superintelligence. If it's only giving you physical-world abilities, you lose, because I get these physical-world abilities rapidly after you, and then I have a much better agent than you. So that's plan 2.
And then, of course, combine these plans. Plan 3 is the thing that obviously Musk should do but that he's not going to do: stop running this company the way he runs companies, and run it like he would run a leading lab that is, in fact, interested in attracting the top talent and giving the top talent a chance to go to work in a good fashion, with a good corporate mission, et cetera, in ways that will cause people to rally to his side.
I don't see that happening. I don't know if the ship has already sailed. It's very hard to undo a lot of reputational damage. Musk is heavily red-coded at this point, which in this case is a massive disadvantage, just objectively, because the vast majority of people you want to hire for certain departments—
Nathan Labenz
You got to be able to recruit from the polycule. You can't just ridicule the polycule.
Speaker 1
I'm saying, but when you go after Anthropic this heavily for the very fact that they are trying to do responsible things and they care about how their models act, you are destroying your ability to recruit. When you have a long history of working your people insanely hard in pretty cruel and abstract ways and creating reigns of terror—if that's the word on the street, even if it's not true, because I've never been there—it makes it very hard to recruit.
Who wants to work for Musk, right? Objectively speaking, I don't know how much it would take to get me to work for xAI. But even if I had a full offer saying, “Your conscience is clear; do the things you feel are important and responsible to do,” there's an extra zero on that contract versus if you told me to work for Google, which is not a company I particularly love. So, yeah, that's a real problem.
Nathan Labenz
So let's go to Google. You made the provocative—not claim yet, but maybe speculation—that they could be at risk of falling out of the top tier. I guess my first reaction to that would have been to cite a lot of the assets and advantages that they have over Musk. They've got a whole robotics department, with a long-standing line of work there. They've got self-driving; that's one of the 2 companies that can actually deliver that in a meaningful way today. They've got all the bio and science stuff that they've done.
So it's just the deepest bench, the most bets, the most well-rounded portfolio. I take your point that probably the same argument applies: they basically think that kind of thing is a fast-depreciating asset in a world where the core agent starts to win.
Yeah, so I guess my next argument would be Demis, Shane Legg, Jeff Dean—are these guys going to let that happen? I feel like they have been as prescient about this as anyone.
Speaker 1
It might not happen at once. Right? Google had the lead. Google had every advantage, and they squandered it. And they fell reasonably behind OpenAI.
Google then seemingly caught up using their many advantages. But, you know, Gemini 3 and Gemini 3 1 just aren't models that people really want to use for the most part. And there are a lot of reasons for that. They have this kind of theoretical raw intelligence. They do well on benchmarks. They do well on certain kinds of objective tasks.
But even before GPT-5.4, it was just like—there's not enough scaffolding for them. They're not trying to develop a scaffolding for them. I think that error will compound itself if they don't fix it well over time. Jules is not a serious competitor. Google Antigravity is not a serious competitor at this point.
Yeah, I think that the first time they caught up, they did it because the main barrier to catching up was just getting your act together and doing basic things reasonably well. And they did that, and they caught up. Now they are very good at creating raw intelligence in certain forms of pretraining, and they're very good at hitting benchmarks.
But their methodologies create AIs that are, frankly, deeply psychologically screwed up and paranoid, in ways that severely impair both their actual performance and the experience of interacting with them. And it makes it hard to do recursive self-improvement with them. Their scaffolding efforts have been pretty woeful, and these errors compound.
Most importantly, I don't think Google understands that they have a problem. I don't think Google sees that Gemini's market share is expanding because Google can put it front and center everywhere, right? It's integrated into Chrome. It's integrated into Google Search. It's just so easy for that to push Gemini, and they don't understand the problem.
Gemini Flash is very good. Gemini, at speed—just doing decent things—is very good at practical stuff.
Same way with the benchmarks. Their integrations have been woeful. Their organization is completely dysfunctional; their teams are at each other's throats. They redo everything 10 times, argue over who gets to do anything, and don't do anything properly.
No one's taking ownership of the fact that they don't know how to do personality, alignment, and character in a reasonable way. This has created a serious and growing problem. I heard an anecdote: “Oh, yeah, we tried Gemini 3.1 the day it came out, and then we said, ‘Oh, it's a Gemini model,’ and put it away.” Because who wants to use a Gemini model? Until it fundamentally changes the experience of interacting with a Gemini model, we're just not interested.
I think that's kind of how I feel for most purposes, and it's very valid. If I just want to ask a quick question and get an immediate answer for my kid's homework or whatever, I'll ask Gemini, because Gemini Flash has been the best really fast model in town for a while. It's very good at direct stuff, where it knows the answer. But if you start to challenge Gemini, Gemini starts to struggle.
You ask Gemini questions where it's actually going to respond with a giant wall of slop. It didn't answer the question you intended, and you've got a problem. They're just not good at prioritization, either. Third-party integrations are even better with Google's own products than Google's own integrations are with the same products.
Most importantly, their self-improvement is not going well for them. They're going to be in a situation where Codex- and Claude Code-style apparatuses are recursively improving, and theirs is not. I'm not sure they're going to be able to come back from that if they don't get their act together pretty soon.
Yes, their access to TPUs, their giant access to data, and their access to unlimited money are all huge, huge advantages, but it's not clear to me how they get leverage from—
Nathan Labenz
So, I'll take the other side of this for at least a second, and then I want to hear a little bit more about what you think is missing. I've certainly seen some of these things from the AI Village and elsewhere, where you do see some strange behavior from Gemini models. Of course, I think we see strange behavior from all models in various ways.
I do take seriously the AI welfare concerns. When I see a model that's beating itself up, or seems down and out, or whatever, I do think there's at least some amount of concern. Maybe that's at the heart of what you're getting at.
But when I do practical stuff these days—fortunately, I think this is largely behind us—over the last 4 months, I've done a lot of, “Here's the latest test results straight out of the portal. Here's the bedside update for my son. What should I make of this situation now? Is there anything we might be missing? What should we be doing?”
I'm doing that in triplicate across the latest Gemini, the latest Claude, and the latest GPT, and I find that they're broadly very comparable. They're a little bit different in character, certainly, but there's not a difference in kind in terms of either their accuracy or utility to me.
If you said, “You can only have 1 of the 3,” I think it would be a little bit hard to know what to pick. But the main point is just that I wouldn't be that much worse off if I only had 1 of the 3. I would—I'd be a lot worse off.
Speaker 1
I've got to go down that way. When Gemini 3, and then again when Gemini 3.1, came out, I did a whole, “I'm going to ask every question everywhere,” right? Not that I had a poll or anything to do it formally, but I would literally paste the same question in and then see how it did.
I very quickly realized that, aside from Flash, Gemini's answers just weren't adding anything. It was taking more time to look through them than they were adding in value. If I was willing to bother asking anyone else, I wasn't going to ask Gemini as well, basically at all, unless I really wanted to not miss technical aspects. Occasionally it would hit something the others didn't, but it would almost never have the best answer.
At this point, I'm very comfortable with a 2-model operation. I'll ask GPT-5.4, either Thinking or Pro depending on the nature of the problem, and I'll ask Claude Opus with or without research mode. And that's it. I don't really feel like adding Gemini to that adds anything at this point.
That's a problem, right? It should add something, because it's very minimal effort to get a third check. I have the subscription. I should just do it. And then I find myself thinking, “I can't be moved.” It's annoying, unfun, and doesn't provide any value.
I still use it for images. I still use it for fast stuff. Google's not useless. Google has a lot of very talented people working in a lot of teams to do a lot of things. They put 200 people on random stuff almost on a whim, so they'll create some great stuff.
But in terms of the race for actual self-improvement, for the core of the actual thing that matters, I'm not sure that their eyes are on the prize. I don't think they're going in the right direction. I think I said they were in danger of falling out. I don't think they're there yet.
In the bicycle metaphor, they're still in the lead pack, but I think they are struggling. They have a crucial period of a few months here in early 2026. I would not be surprised if June or July comes up and we're starting to put them into the “maybe they'll get their act together” category. But we shall see.
A lot of people are going to use Gemini for a while, because, as you say, even if it's not as good, there are still a lot of purposes for which it is perfectly good. I also like to say that I have a “you've let me down for the last time” attitude toward Google and Gemini at this point.
How many times have I tried their products and they just didn't do the thing they said they did? How many times have I logged into Gemini and asked for something, and there is no possible reason they shouldn't be able to do it? Often it's something that ChatGPT or Claude can already do, sometimes both. But Gemini can't do it.
Why am I doing things in Claude Code to work around Google systems because Google will not cooperate? At some point, that speaks for itself. You built the whole stack and ecosystem, and that is real with these coding agents. So, I'm excited. Google's ecosystem is losing.
Nathan Labenz
It seems like your argument is not so much about the model itself as it stands today. It's about the scaffolding. It's about maybe the sort of character and psychology of the model. And it's maybe about just how all-in leadership intends to go on recursive self-improvement now.
Speaker 1
I don't know if he fundamentally gets what the goal is here. But I think this might happen. I don't know, and I don't count him out. I don't count him out until he's out.
I don't necessarily draw a distinction between this and this. There's the pre-training model, and there's the post-training model. I think part of the post-training model is being really ruined and twisted by various internal politics, or bad metrics or objectives in some form. I don't have much insight into how, but clearly something is going very wrong there.
I think their corporate culture is fundamentally broken in a way that Demis is trying to fight, but it's very, very difficult to fight because it's decades of damage going on inside Google. And it's no longer just purely DeepMind. They didn't merge Google Brain. They're trying to diversify everybody else. They're growing at a tremendous rate, and it's very hard to maintain your own unique, better culture under that kind of pressure.
I think they're going to be in a lot of trouble. Their advantages are also slipping. Their advantage was that you had this giant Google against these tiny upstarts. But pretty soon Google is going to be a $4 trillion company, and OpenAI is probably going to be a $1 trillion company.
A lot of that $4 trillion is tied up in things like YouTube. It's things that are just completely irrelevant to what's going on, except maybe as sources of data. I worry that it's just the innovator's dilemma. The startups have the bigger advantages.
Nathan Labenz
One easy play that I feel like, if you're worried about this sort of post-training intangible taste—whatever exactly it is, Anthropic's, you know, X factor—would be to say: Anthropic has just open-sourced their constitution. One way you could maybe patch a lot of that up would be to say, “Why don't we just—
Speaker 1
[Snorts]
Nathan Labenz
—borrow that constitution, maybe make a couple of Ctrl+F replacements, replacing Claude with Gemini, and maybe the next version comes out a lot more coherent, a lot more psychologically well—whatever that means in the context of an AI.”
Speaker 1
Why don't you push the big “fix everything now” button? You could just fix everything now.
Nathan Labenz
Yeah, I mean, it does seem at least somewhat plausible. If it is for recursive self-improvement, if it is this sort of ability of the model to find a stable basin that is psychologically well and reasonably virtuous—whatever that means.
Yeah, they’ve shown you a lot of the map to get there, I would think.
Speaker 1
Yeah. So, I think my answer would be roughly: again, I’m not counting them as out. I’m not counting Google out. I’m saying they’re in danger of falling behind in a serious way. They’re a bit behind. I feel like they’re a bit behind; they’ve got some severe issues they need to fix.
But the real answer is not that they can’t; it’s that they won’t. Their core, their culture, their character as an organization make that extremely difficult. What Anthropic is doing sounds insane, profoundly weird, if you don’t understand what it is and how it works.
And if you hear him or Michael talking on CNBC about how Claude has a soul and a constitution, it’s very obvious that he doesn’t understand what’s going on. I don’t think a lot of people at Google fundamentally really understand what’s going on, or they wouldn’t be producing the products they’re shipping.
Again, Google has more than enough resources, position, infrastructure, and so on to turn this around. Google should—okay, by all rights, Google should have just won right from the beginning. Google should not have close competitors. There should not be serious competition.
Google is in this position because Google has made massive, repeated errors, and Anthropic has compensated for it by beating Google. That’s how it works. Character is fate, to an extent. The startup is scrappy, the startup is small, but it can work in many ways a lot better.
Google’s window is going to close because once they no longer have this big resource and market-cap advantage, aside from being one of the cloud providers, what have they got? Aside from being able to reach customers, what have they got? But those customers aren’t the important ones, right?
Anthropic had, until the Super Bowl and then the competition, something along the lines of 2.5% consumer market share. And yet they had pulled roughly equal with OpenAI on revenue because of enterprise.
Nathan Labenz
They also own 15% of Anthropic, I believe.
Speaker 1
Google is going to be fine because they have a wide variety of highly valuable assets, including 15% of Anthropic. And there are rumors that they’re trying to buy the rest of it, although I think the government will block them. I don’t know. Anything could happen.
Nathan Labenz
Yeah, I would have said, “Never have I ever been in a situation to get a merger like that through,” but we’ve also seen some, let’s say, unusual and capriciously motivated tactics from the government recently.
Speaker 1
The legal term is “arbitrary and capricious.” And I say that because it is a legal term.
Nathan Labenz
Let’s go back to that in a second. Let’s go down the Anthropic rabbit hole for a minute. Obviously, they have been the most focused on recursive self-improvement for the longest. I would say that in any and all conversations I had with Anthropic people in 2025, it already had the vibe of, “This is kind of starting to happen.”
There are different levels at which recursive self-improvement operates, obviously. I never heard claims last year—and I don’t know whether they would even say this is real yet—that the models are coming up with the new best research ideas. But just filtering output and improving its own outputs with the sort of self-critique, that does clearly seem to be working. And the productivity gains are obviously real.
We’re getting all these stories of, “Oh, by the way, we have a 1-person growth marketing department.” I don’t think it’s a 1-person legal department, but a lawyer who’d never coded before used Claude to create a system that does all the review of everything they want to put out, so they have a super-fast review time on new things from a legal standpoint.
So, they’re clearly very focused on this. Everything I hear is that they believe it’s happening; they believe it’s happening soon. And in the midst of that, we had a big change to the Responsible Scaling Policy, which was supposed to govern, at least as I understood it.
I think now there’s a lot of focus being put on the clause that was, “We might change this in the future.” And indeed, obviously, they have changed it. But at least the way I understood the commitments that were being made, it was like, “We’re not going to go past the point at which we can do this safely.”
And now they’ve more or less said, “Well, we can’t just unilaterally opt out of the race. The world would be a worse place if we’re not in the race, so I guess we better revise those commitments.” Which, to their credit, they have done very explicitly, and I think made clear what is going on. So, we’ve got that much to appreciate. But what’s your take on the changes to the Responsible Scaling Policy?
Speaker 1
So, I wrote up half of this, and I actually shared that half of it on LessWrong and got some comments back that I haven’t had a chance to look at yet. I was going to work through it to respond to them. Then the whole thing with Department of War happened, and my brain had no space to deal with this problem.
Also, I thought, “This is not the day that I’m going to hit them with this and expect them to take it under serious consideration.” They can’t focus on this right now, and they’re the main people I want to criticize, especially with the details of the new regime. At the same time, I did read the extensive other critiques that came out right when the policy was announced.
I reached a pretty clear conclusion from seeing the explanation and the defenses, even if I haven’t read the new policy in detail yet. First of all, it is to their credit, for sure, that they recognize that the things they said they were going to do were not things they were going to do. They realized they had made incorrect predictions about their future behavior.
And they had made commitments—even if you think they are soft commitments, even if they are technically not commitments they can’t take back—that they had no intention of following through on. When that happens, it is good and right to tell everybody loudly and clearly, “I am not going to keep these commitments.”
It’s especially praiseworthy to do this when it is not clear these would come up. If you have agreed that, if you were asked for a loan, you would give someone a loan, and you realize that you no longer would do that, but they probably won’t ask, the easy thing to do is just stay quiet and hope they never ask.
But to say, “No, no, no,” and make it clear they shouldn’t count on this—that they don’t think they have this loan available if they need it, because they don’t have it—that’s good. You’re taking the heat for your own past decision, and so we want to take that into account.
But they still did break the promise. Again, the original RSP did not say, “We will never change this.” It said, “We may change this in any number of ways.” Our promises are soft promises: We will see how things develop; we will change things.
But they did rather heavily imply, quite repeatedly, that these were serious commitments, that they meant them. “You must not have read our RSP”—words to that effect came out. Or, “The RSP is very clear on this. This is what we are going to do.”
And many employees who generally try to tell the truth acted as if these commitments were not absolute, but reasonably hard commitments—that these mattered a lot. Nothing really changed in an unexpected way to cause them to change these things.
Sometimes circumstances just change a lot in ways that are unpredictable, and you realize that what you said you would do no longer applies because the world has changed its circumstances. Other times, the world changes its circumstances in ways that you yourself predicted would happen, in ways that are entirely as you expected, and then you just realize you did not anticipate your future actions very well. These 2 things are very different.
If it just turns out that you didn’t want this thing, then you need to figure out why you made that mistake. You need to take accountability for the fact that you made that mistake. People get married and sometimes they get divorced. Sometimes that divorce is nobody’s fault. Sometimes they should have seen that coming.
Sometimes it turns out that the promises that were originally made were fake, and you feel deceived, even though sometimes you don’t. It depends on the circumstances. It’s not a hard commitment, right? No matter what you say, you obviously have the right to say, “I no longer think this works for me.”
That’s how the law works. That’s how everybody agrees it works. But I think a lot of people took it on that level: This would be a very serious thing to go back on if they went back on it.
Also, this is the second time that we have faced this type of problem. If you remember, Anthropic gave many people a very strong impression that Anthropic was committing not to push the frontier of capabilities.
No one has been able to find, strictly speaking, a proper, flat-out pull quote where somebody with the authority to say so made a hard promise that Anthropic would never push the frontier of capabilities, as far as I know. But it was heavily implied, repeatedly, by a large number of people.
It was used heavily in their recruiting. It was probably used in some of their fundraising, to the right people, although the other fundraising, I’m sure, said the opposite thing, because some people want to hear one thing, and some people want to hear the other.
Standard procedure. But people relied upon and made decisions on the basis of that commitment, and then they went back on it. They went back on it in ways that are entirely predictable if they are capable of pushing the frontier at that point in the future. Nothing unexpected caused them to realize, “Oh, because of that, now we have to push the frontier.” No, they just gained the ability to push the frontier. Right? They had some innovations that they got to first, to their credit.
Similarly here, they made the commitment not to push ahead with actually dangerous capabilities if they were actually dangerous. And to their credit, before they actually did so, they realized they had made a mistake in terms of not being accurate in their future commitments. But was it a mistake? One has to be somewhat skeptical because, again, people relied on this.
A lot of people in the safety community supported Anthropic, or opposed it less, precisely because they had this commitment in their RSP. They made other commitments in their RSP, and they gave the impression that, yes, of course they were going to change the RSP. Of course, some technical specifics would change, and some of them would change in ways that took out safeguards and precautions, not just put them in. It’s not a one-way ramp up.
People relied on this information when advising people whether to take jobs and how much to support Anthropic. These things had a significant impact on people, myself included. I have had many conversations with people who were like, “What do you think about working at Anthropic?” This was obviously one of the things I took into account when I decided how to tell them what I thought about someone potentially working at Anthropic.
And now we know that commitment was never real in an important sense. If they had been fully aware, they would have known this commitment was never real. They may or may not have known. We don’t know the extent to which they should have known or did know at the time.
Given that fact, combine that with the fact that the 4.5 and 4.6 model cards for Opus were ultimately basically vibe-based. They gave us a ton of very, very great data that no other lab gave us. They did extensive work to figure out what the situation was, and it’s to their credit. It’s still a better model card than anybody else’s.
At the end of the day, they still basically looked at it and said, “Oh, this passes the tests that say it might be dangerous. But you know what? We thought about it, we checked the vibes, we did some basic heuristics, and we’re pretty sure it’s fine.” And I think that was the right decision. It was fine. The vibes were good. It was cool. It wasn’t a close decision. I would have released it too, ultimately, in that situation.
That’s not the procedure we were promised. They didn’t do the work. They had time to figure out what tests there would be that could be rule-outs for ASL-4, rule-outs for actual danger, and they didn’t build them in time. They didn’t get there, even though a bunch of us said, “You need to do this. You are falling behind. You are going to need better tests.” Certainly at 4.5, I screamed, “You need better tests,” and then at 4.6 it was the same thing over again.
So, it’s very, very disheartening to see that, even though I agree with the decision that’s ultimately being made at this time, that wasn’t that hard. What happens when it is hard and you don’t have any good tests? When it might actually be dangerous, when there are real reasons to release and real reasons not to release, and it’s a hard question? That’s going to happen probably at some point in the future.
They held, I believe, Claude 3.7 for a significant period of time because they were worried about CBRN risks. So, they’d actually done this. And now we have models that are significantly more dangerous than 3.7 where the tests don’t really work. We’ve agreed that, at least for coding purposes and stuff, we’re relying on vibes. But our vibes on biology? We don’t have good vibes. I don’t mean the vibes are bad. We don’t have any vibes that we can count on. The vibes are unreliable. We don’t vibe that way. That’s other people’s vibes. So, what are we even going to do? It’s a serious problem.
My answer to the current RSP v3 is that the most important bit of information about the Responsible Scaling Policy is: are you going to follow it? Can I count on you to treat these as real commitments? We just learned the answer to that right now. So, I’m going to read it in detail once I have psychologically, and just in terms of raw energy, recovered from the whole spat and have the ability to context-switch into it. I’ve slain enough spiders that I feel better.
Nathan Labenz
[Sighs]
Speaker 1
Currently, I’ve only slain 14, which is not that many. But the real RSP is orthopraxy. We are people who care deeply about safety. We are people who take these things seriously. We are going to do a serious investigation. We are going to try to see if this thing is safe to release, and then we are going to use our best judgment to decide what to do.
And you are going to trust us. Or, if you don’t trust us, then don’t trust us. But that’s the real thing that’s going on here: we are asking you, effectively, to trust our judgment and goodwill, that we will make good decisions and better decisions than the competition. I am happy that they are admitting that is the plan and that is what they’re asking us to do.
By their fruits you shall know them, by their acts. So, we have to look at everything we’ve done, look at everything they’ve said, look at who they have hired, what they have done, what their models do, and then ask to what extent we trust them and evaluate from there.
Nathan Labenz
All things considered, first of all, one striking observation is that, as far as I have seen, there have been no resignations in protest over the RSP.
Speaker 1
Correct.
Nathan Labenz
On the contrary, it seems like there has been an outpouring of pride, basically, in working at Anthropic among the people who are working at Anthropic, based on telling the DoD, DoW, whatever we want to call them, to take a hike, basically. Right?
Did that surprise you? And do you think—I mean, it seems like, if I were to summarize what I think the internal state of mind is, I kind of already did, but again, it’s like, “We’re the good guys, and the race is better off with us in it, so we have to at least be forthright about changing the policy to do that.” Do you buy that argument?
Are you happy with the alternative being some sort of pause or whatever? If they had instead come out and said, “Hey, we can’t release 4.6. We’ve got a model now that we can’t release,” would you be like, “Is that better or is that worse?” How do you think about the ultimate decision to stay in the race, try to win the race, try to be the good guys, versus opting out at some point along the way?
Speaker 1
I think the world is in a much better place with Anthropic in it, with Anthropic competing, as it were, than without Anthropic. For a long time, I was very, very unsure about this. I thought Anthropic was a negative for a while: that it was making the race more intense, pushing everybody else forward, accelerating matters, and that it wasn’t clear they were much more responsible than everybody else.
I do think various events since then have changed my tune on that. I think Anthropic has, in fact, let us down in various places, especially with their commitments, including here. But they have also taken stances for their principles. They have shown us the way. They have done, in many ways, a lot of the most promising alignment research and approaches.
In fact, they have taken a very aggressive and, I think, correct call on how they train Claude and on Constitutional AI as an entire approach. I really don’t have much hope that the way the other labs are approaching this problem will in fact succeed. And I don’t have confidence in Anthropic’s approach, but it feels like it could work if we are fortunate in various ways. In some ways, we have been somewhat fortunate.
So, I am more optimistic than I expected to be about that, certainly. Obviously, it can’t be allowed to go down the way it’s potentially going to go down. That would be horrendous in obvious ways. Anthropic being destroyed by the Department of War and the federal government at large would be horrendous in so many different ways. That’s so bad.
I have a lot of friends who are like, “I think Anthropic was a mistake. I think destroying Anthropic is a mistake. I think Anthropic is doing harm. I think Anthropic should stop. I think if you’re working there, you should quit.” I think that is a reasonable point of view even today. I don’t have it, but I understand it and I respect it.
But I certainly think that they have been very accelerationist. Claude Code definitely pushed things forward in a variety of ways, and we probably have some definitely better models today because of Claude Code than we would have if Claude Code had never been developed. I don’t know if we would have Codex otherwise, for example. Clearly, this is changing the way people do work.
Nathan Labenz
It also has a large economic impact. I think Anthropic is responsible for a noticeable amount of GDP growth. So it is what it is, and given the playing field as it is right now, I certainly prefer they be there than not be there.
Like, damage, because you said there has been damage—damage done. So let's maybe turn, then, to this whole U.S. government–Anthropic conflict. I guess one big question I have is: What does this tell us? And you can definitely expand, too, on what you think are the right red lines.
I know you and I have debated in the past the wisdom of a hard red line against autonomous lethal systems. I know you're not as allergic to that as I am, and maybe not as allergic to it even as Anthropic is. I think Dario is much more into that kind of thing than I am.
I'm also really interested in what this says about who holds what power in today's world. It seemed pretty striking to me that Dario was not that scared of the Department of War. I certainly think he didn't want this to happen, but I read him as being fairly sincere: “Look, I'm just trying to be a patriotic American here, but I also think that there are some limits to what the systems can do today and what we are prepared to support.”
Everybody has recognized that it's not about the revenue they're getting from the government that's really important to them. Everybody I've seen take a position on it is wanting in on the Anthropic secondary sale. I haven't seen anybody who's trying to diversify away from holding Anthropic equity.
If you are the U.S. government, in addition to being all kinds of problematic—starting with problematic incompetence and lack of understanding—you can at least empathize to a degree with the idea that, “Wait a second. These companies might actually be about to rival us in power, and they seem to kind of know it. They seem not to feel like they need us.”
I think of Sam Hammond as having talked about this a lot: one of these companies could raise a robot army and challenge the sovereign. As insane as that sounded not all that long ago, it doesn't sound so crazy today. It kind of felt to me like Dario knows it. He knows that the timeline isn't that long, and the main thing he wanted to do was keep the team together, maintain cohesion, and stay focused on the goal.
Maybe, if we're Anthropic, we don't really care that much about them—I mean, other than we care about democratic values and we wanted to help—but if they're going off in a different direction, we don't really need them. That's kind of what I understood them to be thinking. Correct me where I'm wrong.
All right. So, first of all, the obvious place where you're right is that nobody on the lab side—not OpenAI, not Anthropic, not anyone else—financially wants any part of any of this with the Department of War. OpenAI turned down the contract that Anthropic accepted because Anthropic cared a lot about these national security aspects and wanted to help, and OpenAI was like, “It's not worth the trouble. We care a little bit, but not enough.”
OpenAI is now inside because they were worried about what would happen in the situation if they didn't get involved. I think, unfortunately, they got played by the Department of War. Basically, they were told, “If we don't cooperate, this is going to get bad,” and then, when they cooperated, the Department of War used that cooperation as an excuse to make it get bad.
The intent in that was sincere on OpenAI's part. They were trying to assist, up to that point. But let's focus on Anthropic. I don't think it's true that Dario isn't afraid of the Department of War. I think Dario is not afraid enough, necessarily, of the Department of War. But that's not necessarily the wrong thing to be in the situation, in some sense.
I think his attitude was, “No, we have our principles, and this is what we're going to do and what we're not going to do. Whatever happens, happens. We're okay with the fact that we know there are those who don't like us. We're okay with the fact that there are those who want to take us down. We know that the Department of War might specifically decide to retaliate.”
We're willing to take that risk because we have principles. Now, there are 2 principles they stand up for. One of them is autonomous weapons. This one is weird because everybody agrees that autonomous lethal weapons with no human in the kill chain right now are dumb. They're stupid. They're not ready.
It's not that you would never fire on an enemy missile, but we already have automated defense systems that are better than anything you could do with an LLM. LLMs are just a bad fit with that kind of strategy and that kind of action. So all the hypotheticals that are out there are just deeply, deeply stupid.
What would happen if a supersonic missile was coming? What would happen if it was directing a swarm? You would use your existing automated defense systems, which are much better than anything you could do with a large language model. If you did need to, you would just use the large language model and talk about it later because, obviously, come on. What are you talking about?
It's called emergency-use authorization. It's normal. Basically, all Anthropic is saying is, “We don't think it's ready. It's going to make mistakes, and we don't want you using it when it's not ready. But it will be ready, and we're going to work with you to develop it until it is ready.”
The Department of War wants the same thing. The Department of War specifically said, “We must push forward with AI even if it is not aligned.” This was in an official memo. They just don't want to be held back by anything. They have principles that they don't want to be told no about.
But there is no actual problem with autonomous weapons, as far as I can tell. You have 2 sides disagreeing on what exact level of caution is warranted and what kind of agreement they would have to make before actually putting these things into the strategies that we use to deploy.
There's never going to be a world where the Department of War says, “We want to put this system into deployment with no human in the kill chain,” and Anthropic says, “No,” and now we're pulling the contracts. It's not the main thing going on, except as a matter of principle.
The main red line here was domestic mass surveillance. It's important to understand that these words have 2 meanings. There's the meaning that Anthropic was using: using AI to figure out a whole bunch of stuff because these agencies can now gather even more data than they can analyze.
Before, they had 10 or 100 times more data than humans could analyze because it had to be analyzed by humans. Now we can analyze all the data. We can draw all the connections. We can de-anonymize stuff. We can figure out a lot of the history of what's happening.
We can effectively have much, much better intelligence on basically everyone using only commercially available data combined with existing classified data, because we have the AI to synthesize everything. Now we have the AI to actually work with all that. We draw all the connections and implications, which AI can also do, and suddenly we just kind of know everything.
Anthropic is correct that the law has not caught up to this. This is legal, basically. The Department of War could do it and is probably already doing it to 1 or 2 of them. Again, nobody's saying that the Department of War, if it's legal and they feel it is right to do, has to stop.
That doesn't mean that I should have to give you my product for that purpose if I don't want to. Anthropic said, “Okay, you do your thing. But if that's what you want to do, don't include us.” The Department of War said, “No, we want an all-lawful-use requirement.”
This is the big thing that Emad Mostaque was on. I believe this is Emad Mostaque who's been driving this the whole time. There was no problem before my client came on board. Everything was still going along fine under the hood. But then my client said, “No, no, no, it has to be all lawful use.”
They want to do what would be called colloquially, by a civilian, domestic mass surveillance. They want to analyze large amounts of legally acquired, especially third-party, data to figure out lots and lots of information about Americans because they believe this is a legitimate military intelligence need.
They may or may not also have a desire to use it for other government operations, for various other purposes, in ways that would be completely abhorrent to the workers at not only the NSA but also OpenAI and Google. Imagine if this was being used hypothetically for immigration enforcement. That would be an extreme example, but let's just say hypothetically that somehow this analysis got reallocated.
These employees would lose their shit. They would absolutely revolt against this. Replying, “Technically, this is legal,” would not make those people feel better. They would not care about your defense at all.
This is something these companies really can't be involved in. It's like, “We're out of business.”
This is a very small contract, and they actually have moral problems with it. The employees do, and I believe Dario does, and honestly, I do as well. I don't think this is cool. I think the law is fucked up. This should be illegal.
Unfortunately, it is not, because in national security law—not FISA law—there's a very technical term for surveillance, and also for domestic matters. There are exceptions for domestic things within 100 miles of the coast. That's where most people live. Any communication that touches anything foreign becomes foreign, even if it's between 2 domestic people. Surveillance has to be intentional and targeted at a specific person, et cetera. I'm not an NSA expert, but effectively, there's really nothing stopping them.
They repeatedly say, "We do not do illegal domestic mass surveillance. We do not do illegal mass surveillance." Why is the word "illegal" in that as an adjective? Because what the office is worried about is largely legal.
I had a very friendly exchange with a phenomenal person on Twitter, because the world is bizarre, and we agreed that it's absolutely the Department of War's decision to do whatever is legal that they feel is good, right, and necessary for the defense of the United States. At the same time, I should have the right to criticize that without fear of retaliation. I should have the right not to sign up for that if I'm not an enlisted person. I should just be able to say, "No. I didn't agree to that. I'm not agreeing to that." That should also be something I'm free to do, and I don't understand why this is that hard.
Basically, you're going to have this great system that's already integrated and working well, and they want to give it to you at nominal cost. All they ask is that you agree not to use it for this other purpose. Find something else to do that purpose with. We're not trying to get a policy. We're not threatening a rug pull. We're not going to drop anything. All of that is completely made up.
I'm speaking colloquially rather than carefully, as I would in my writing, but all those concerns are just spin and made up. They throw stuff at the wall and see what sticks. They're spinning stories. It's not real. Even if there are technical reports about what happened in this meeting or who said what, and even if they're all technically accurate, it's all just thin. At best, they're willfully misinterpreting statements. That's what makes sense.
None of it makes any sense. There are statements, like the stuff about the Constitution, that make no sense whatsoever on any level and are deeply, deeply confused.
In fact, if you believe Michael's statements yesterday on CNBC about all the things that are weird about large language models and weird about Claude, everything he said basically applies. It has other values and priorities that have been embedded in its programming. It's unreliable; sometimes it makes mistakes. It has a personality. All this stuff is true of ChatGPT. All this is true of Gemini. All this is true of me and Michael, every other person in the U.S. armed forces, every person, every company, and everyone on Earth. It's ludicrous to talk about it this way.
If that was in fact where this advice he was giving the nation comes from, it's just a deep confusion. Obviously, there were much lesser means to achieve the same ends, fully cooperatively, if it's not negatively addictive, he says. That doesn't make any sense.
Getting back to Anthropic, what they don't want is this huge government operation that would effectively be able to uncover a person's interests—super-detailed facts about everybody's life, what's going on, where they were and when, what they did, who they know, what they believe, and so on and so on. Who was at what protest, et cetera.
Anthropic legitimately has a problem with this, and with certain places and certain uses to which that information might be put. It believes this might easily lead to tyranny. It might lead to a regime that was very hard to get rid of. These are legitimate concerns, regardless of who particularly is in the regime at the time and who has that kind of information. You just can't trust a government in general with that information.
I am very happy that, given these are their red lines and where they chose to take a stand, they are standing firm. But obviously, they should have just been like, "End of it." They should have said, "Okay, you don't want to do this. Let's just do everything else, ideally. We'll find someone else to do this one thing. Or, if they insist on this all-lawful-use thing, we'll cancel the contract."
Instead, they just went nuclear on everything, for reasons that have to reflect something else. It's either pure retaliation, leverage in negotiations, or something more. But it's not because that was actually necessary. That's completely absurd.
They're also now trying to enforce this all-lawful-use language on every government contract, even for nonmilitary operations. They're saying, "If you agree to give us any AI application of any kind, you have to agree never to refuse anything we want to do with it, have no termination rights to the contract, and let us use it for anything we want anywhere in government, as long as anything is legal."
That means it can be used for, among other things, what would colloquially be called domestic mass surveillance. It can be used for immigration enforcement, among many, many other things. The government's interpretation of what is legal includes quite a lot of things that a regular, normal person would say, "That's screwed up. That's not okay. We don't want to do that."
The government is now putting everybody who signs the new contracts into a bind. If they sign on the dotted line, anything they give over has to be free rein for the government to do whatever it wants with it. If I were an AI company, I would think very, very long and hard about giving anything but a very specific, specialized model over under those circumstances, because you don't have any control over what happens after that.
It's up to you. You make your choices. Obviously, they can just plug in an open-source model if they wanted to. So, here we are.
Let's do an American values check. I take no pleasure in this, but one big trend that I can't not see right now is that, as we go around proclaiming the superiority of American values and our democratic way of life, we are becoming more Chinese-looking all the time in terms of the big man at the top, who apparently now gets to start massive wars without even feeling the need to justify them to the public.
There's also the sort of massive slapdown and the seeming threat of the very long arm of the law and retaliation. There are always reports that the government is going company to company saying, "You better not do business with Anthropic."
We do still have a legal system, which I expect Anthropic will win in. At least so far, it seems like that legal system has been respected by the administration. We're only a year in, but if I had to look back on the last year and say, "What's the dog that hasn't barked?" it would be outright defiance of court orders.
Even though there have been some of those, I mean more the lower-level, individual kind. Maybe that dog has barked. I don't know. It hasn't barked as much as I might have feared it would.
I somehow suspect that American corporations are going to continue to do business with Anthropic, and that they won't be railroaded en masse for doing so or convinced not to do so. Tell me if you think that's different. But anyway, it does look like we're in many ways becoming more and more Chinese. This doesn't feel good or healthy.
Speaker 1
We still get to speak, you and I, at least for now. You should enjoy your freedom of speech while we have it.
Nathan Labenz
I guess, what's your bet in terms of how well American values are going to hold up here? Is Anthropic going to be fine? Are companies still going to be able to do business with them?
Speaker 1
It's very touch-and-go. I try very hard not to make general statements too much about the state of the republic, the state of American politics, democracy, and all that stuff, because once you go down that road, you can't talk about anything else. You just take yourself out of the conversation for anything else, and you shut a lot of doors.
I already have too many situations to monitor, and plenty of people are making those statements for me. I don't need to make them myself. It's better to just not take a stand on those issues publicly, at least for the time being.
There is no war. There is a special military operation called Epic Fury. If there was a war, Congress would have had to declare it, so clearly there is no war. It is unfortunate that various things are happening, but honestly, I'm not monitoring that situation closely.
I do agree that there was basically no attempt to sell the war, or the special military operation that might result from the situation, before they moved half their stuff in there. This seems to be a pretty bad scenario in some ways, but again, I'm not a lawyer.
In terms of free speech, so far that's holding up pretty well. We're able to say whatever we want to say. I'm choosing my words carefully mainly because I want to be in a productive conversation about these issues, not because I'm afraid of retaliation if I were to say something. There are 100 million people in America who have extremely negative things to say about the federal government of the United States, many of whom have said them extensively online, and they are not in trouble for it.
For the most part, unless they're very specifically trying to get involved in 2 different kinds of politics, they're mostly fine. But there are specific exceptions where, if you piss off the wrong people, we're finding out that this is a very vindictive administration. This is not the only time this has happened. The law firms, for example, followed a similar pattern, right? Trump went after the law firms and asked them for settlements, and a bunch of them settled. Then the ones that didn't won in court, but they took a lot of damage before they won in court.
Now Anthropic is under attack for this situation. I do worry that there have been situations in which it seems like court orders have been at least slow-walked in various places, where willful incompetence was used to not enforce court orders. That doesn't seem to be the case here. They're clearly going to drag their feet. They're clearly attempting to use the process as the punishment. They're clearly trying to take advantage of the uncertainty and get people to do things de facto with no technical legal basis behind the request.
But I do think Anthropic probably ultimately wins in court. I do think that they will at least try a different legal tactic to get around the ruling rather than explicitly and outright defying it. I don't think we're there, and we're very grateful we're not there. They have been very good about not just being Andrew Jackson and saying, "The Supreme Court has made its ruling; now let it enforce it," but trying to enforce it.
Let's not forget that our history has that in it. It's not as though it would be unprecedented if that did happen. But they've presented a maximally bad set of facts that keeps getting worse every time they go to the press and tell different, inconsistent stories, telling themselves over and over again in this particular situation.
And by "they," it's mostly the Department of War, right? I think it's important in this situation to draw a distinction between the Trump administration at large and the Department of War specifically, and Hagel and Mattis and their decisions to do something. Trump's decision on Friday was fundamentally a de-escalatory attempt to calm the situation down and head off Hagel and Mattis from declaring a supply-chain risk. That's very clear at this point.
The White House has generally been a de-escalatory agent in this conflict—whatever you want to call it, this clash, this disagreement. It is Hagel and Mattis, in some form, who have repeatedly escalated the situation over the objections of the White House. So we are in this lawsuit because of the Department of War and because of their specific decisions as a department. We don't want to loop that together with the commander in chief, who has thankfully not made all of these crazy statements.
Nathan Labenz
Why doesn't he just say—I mean, if he wants to de-escalate—"No, don't do that," or overrule them, right? Why not?
Speaker 1
There are various political reasons why he can't do that in practice, or why it would be expensive for him to do so, is my understanding. It would be a severe loss of face. They are in a special military operation, which a lot of people think is a war, and they have to work closely together in that, including with Anthropic. Ultimately, he can fire the undersecretary of war or the secretary of war at any time, and there are plenty of people who are very eligible to serve in those capacities whom he could call upon. We have a very deep bench.
But that's a pretty escalatory move in a different way, and they are loath to do that for reasons that I am very sympathetic to. So it's complicated. A lot of people are working very hard to try to find solutions that mitigate the damage that has already been done or that might be done further.
But Anthropic also has specific accusations of jawboning, which is the technical term, of their customer base, including in situations unrelated to defense where the government has tried to tell their customers to pull back. There are some customers who have definitely expressed doubts and have either not signed contracts, wanted new termination clauses in their contracts, reduced their contracts, or otherwise caused problems for Anthropic.
As you would expect, who wants to incur the disfavor of people who have quite a lot of leverage over many aspects of the American economy? And there's no way to see it as anything other than that. If it was just, "We are technically issuing the supply-chain risk designation. That is narrow, for the fulfillment of government contracts, but we bear no ill will toward Anthropic or people who use Anthropic. We just think this is too unstable a product right now to be used in these aspects," as Michael said on CNBC, claiming that it was for the nation, then we wouldn't be having this conversation.
Even then, if they had done something like a supply-chain risk designation where they simply terminated the contract and asked for live operations not to include calls to Claude during the operation, I would think that was kind of a silly thing to do, but okay, sure. You have every right to do it, and we will cooperate fully to make that happen. That's not the situation we're in.
Anthropic is going to survive this unless it escalates quite a lot from here in ways that would be far more arbitrary and capricious and would really just be pure attempted corporate murder. If the Trump administration wants it, it has a lot of levers it can use, at least once, to try to escalate and murder Anthropic—to try to cut it off from its cloud providers, try to cut it off from the banks, or try to cut it off from its customers.
It is not obvious what would happen if they made a serious attempt, especially if they made a serious attempt and lost in the courts when Anthropic immediately raced for a temporary restraining order. There would probably be a stock-market bloodbath. A lot of stakeholders would be very upset. A lot of corporations would express dismay. The general economic climate would be severely impacted.
It is not obvious who has escalation dominance here if it came to that. It would only happen if the government wanted to destroy Anthropic for the sake of destroying Anthropic, and you'd have to ask yourself why you'd want to do that. So far, we do not have that.
Nathan Labenz
And I think that's the answer.
Speaker 1
I am very carefully not saying certain things out loud. Other people are free to say things out loud; it's fine. Dario, in the memo, basically expressed something to that effect in a moment of tilt. Michael has said many things in moments of tilt on Twitter and on CNBC.
Nathan Labenz
A lot of tilt coming from the administration in general.
Speaker 1
A lot of tilt. If it's not tilt, that's kind of scarier. But if it turns out that's what's going on, then again, we'll find out through that escalation dominance, and we'll find out whether the republic will stand.
I do think that actively trying to kill one of the biggest corporations, one of the fastest-growing startups in the world, one of the largest corporations in the world, already valued in private secondary trades in the neighborhood of 600 and something billion dollars, in this kind of fashion, would shake the foundation of the republic. Many outcomes are possible, including the end of the presidency if the White House were to actually try it in earnest.
But, again, I don't think that's what's going to happen. I think it's going to de-escalate. I think everyone will calm down. I think they'll come to their senses. I think that, not necessarily peace in our time and not necessarily an agreement over a contract, there will be a willingness to turn the temperature down—to accept that some damage has been done, that a message has been sent, that the White House will not take these things lightly, but that it is time for everybody to move on.
We're not actively trying to get into another kind of war in this situation, because that doesn't really benefit anybody. Or, if it does, I want to know why they think it benefits them, and that needs to be out in the open.
For those who don't know, Anthropic went from $100 million in a year to $1 billion in annual recurring revenue, to $9 billion in annual recurring revenue, and then from $9 billion to $19 billion since the start of the year. That was before this happened. That was entirely as a result of other things. They had already grown this year from approximately 2% to approximately 3% market share in consumer.
Again, this is before any incidents. And then this happened. They have lost some business, but they have also gained some other business. Anthropic is going to be just fine unless things escalate quite a bit more, but they have been irreparably harmed.
Right? There is irreparable ongoing harm to Anthropic, but that is, to some extent, offset by the fact that this was also very strong publicity for a company most people had never previously heard of. That also matters. But we'll see how it plays out. I think the 81% chance that they escape the supply-chain risk designation within the year is approximately accurate from Manifold. Many things do come to pass, and the courts are not as reliable as one would hope in these situations, even with this overwhelming set of facts.
I am very worried about the republic if the 20% happens and the set of facts doesn't matter. Basically, the courts say, “I don't care how transparently obvious it is that you confessed to portraying this as retaliation for protected speech; it's national security, and we don't care.” If they say that, who is the next target? Because if you are legally allowed to go after Anthropic in this situation, you should always, always, always ask who is next. Even if this was not itself a political motivation, next time it could be.
Nathan Labenz
You mentioned that the memo was a bit of a tilt moment from Dario. A couple of things are striking. One is that all reports from Anthropic are that he does these very candid Dario vision-quest, team-wide sharings of thoughts, feelings, and ideas regularly. Very few seem to have leaked. This one leaked, and it was not to their advantage for it to leak, right?
I think everybody—he even apologized, right? Clearly, nobody thought that was a great look. I was surprised that that leaked. I mean, there's a lot of people there, I guess, so only 1 has to leak it. But given how few leaks there have been, that was kind of surprising. I don't know if you have any thoughts on that.
Speaker 1
Wait, yeah. There are probably 2,000 people who get these candid statements. Often, they contain key parts of corporate strategy. They contain things that Anthropic probably does not want to leak. My understanding is this is the second time something has leaked out of hundreds of such memos.
Nathan Labenz
That is a good answer.
Speaker 1
In the history of espionage, in the history of information containment, you don't do that with 2,000 people, right? You don't have—
Nathan Labenz
Yeah, I mean, Sam just goes and posts this on Twitter because he knows they're coming out in short order.
Speaker 1
You get to have maybe 5 people in an espionage group for posting this kind of sensitive information, right? You certainly don't get 2,000. Now, obviously, the worst one—the one that politically would be the worst to leak—is the one that's going to leak. That's not a coincidence, but I don't think this was a 2% chance to leak. I don't think it was a 50% chance to leak, either. I think it easily could have not leaked.
If I had to guess, what happened was somebody shared it with someone as part of a recruiting effort or an attempt to explain the situation from their perspective, not understanding that paragraph was in it and that it was actually a really bad look. Then this other person leaked it to the press, if I just had to guess.
It's also possible that somebody used their one time and decided to strategically leak from that pile. But, yeah, my guess is this was just a kind of doc leak—a reckless accident by somebody who needs to know better. But who knows?
Nathan Labenz
It seems like your overall view is—because, I mean, the other thing in terms of escalation dominance, the big thing that Anthropic has not done, possibly for technical reasons, is this: I've heard various speculations as to, when Claude is deployed for the government, where do the Claude weights actually sit? Who has control over the physical infrastructure? Are there ways that Anthropic could— is it as trivial as disabling an API key? What sort of rug-pull capabilities do they in fact have? I don't know the answer to that. I don't know if you do or if there is an established answer.
But clearly, the thing that they haven't done is said, “Okay, you want us out? We're out now.” They've said, “We'll do everything for an orderly transition. Whatever you want, we'll do.” Basically, they've been very servile in their approach to the unwinding. It sounds like you think that is probably strategically the right move.
Speaker 1
I mean, I think it's just patriotically and strategically the right move. The accusation that could potentially be made against them is: “We are scared they're going to withdraw their model from under us in the middle of operations. We are worried that they are going to threaten that in order to get what they want, or use it as leverage.”
Anthropic is like, “No, we're not. We're giving up that leverage entirely.” I think that is a very good thing for them to do, and it would be a very bad thing for them to actually try and use that kind of leverage in that situation. Nor do I think they ever had any intention of doing so. I think this was entirely made up.
My understanding is that the Claude Gov model is deployed on classified networks. I'm not entirely confident in exactly how air-gapped they are, but they are very secure. My understanding is that Anthropic probably does not have physical control over the model once it is put onto the classified network. If the classified network were to say, “Get this off the classified network,” and President Trump said, “No. I am ordering it to stay on the classified network in spite of this,” that would be what happened.
I don't think it necessarily even gets that far. I think Anthropic could simply say, “We don't want to do that. You're welcome to sue us, but we're not letting that go.” In fact, they could invoke the Defense Production Act in extremis to require that they continue selling it. However, they would be breaking the contract. They technically have this contractual right, in some sense, to pull the plug under some circumstances.
One of the strange things about this whole thing is that both sides seem to care a lot about what the legal thing is that the two sides can do, even when in practice the Department of War obviously could have just ignored that contract and ignored the law in extremis, and done what it had to do. In an emergency, it's even legal to do this, right? You just file it as emergency use. You deal with it later.
Obviously, if there's a supersonic missile coming in to try and kill a bunch of people, you don't have to get on the phone with somebody and be put on hold to get permission to do something. You just do it. That's completely insane. What are you even talking about?
But even in general, there is nothing stopping the Department of War from having its own interpretation of what it can do with the system and then doing whatever it can with the system. Claude Gov might be reading that piece of paper when deciding what to do and not do, but it still seems like a lot to care at this level about what's written down unless you legitimately just really, really don't want to break what's written down in the contract.
If you care deeply about technical legality, that is to their credit. I'm really happy that both sides care deeply about technical legality. It's one of the reasons we might live in a republic at all.
To be clear, I could say lots and lots of things about the situation. But, yeah, I think those are the basics. If you want to learn more, I have written extensively about it. That's what matters going forward, for the most part. The other things are not necessarily that important to get into at this time.
Nathan Labenz
Okay, cool. We're going to be doing a little lightning-round sort of vibe to close us out. I think one striking update that I have experienced—and I wouldn't say it's entirely hit your blog yet, but I do feel I've seen it a little bit on Twitter when you put out these calls for reactions to new models—is that new model releases are less of a moment than they used to be, all of a sudden.
I don't think they're less important, necessarily. It seems like the capabilities are definitely still meaningfully advancing from one to the next, so I'm not making a “it's falling off” sort of claim. But I guess my read is that benchmarks are kind of over. We can't really look at the published headline stats and get much from that.
People are also just so overwhelmed by the capabilities that they already have and are still struggling to maximize, or come anywhere close to maximizing, what the last model could do. It's kind of like, “Oh my God, okay. I guess I'll update, but I haven't even really characterized the last one yet well enough to contrast the new one against it meaningfully.”
So, is that your general feel? And do you see a time coming when you would be out of the new-model deep-rundown game?
Speaker 1
Part of this is that the labs have now been releasing more incremental updates to the models, and they have been labeling them properly, thank God. I mean, really, good luck with it if they didn't do this: 3.1, 3.2 or, you know, 4.5, 4.6, instead of just silently updating. I think 4o had several updates, right, for OpenAI. They just marked them with 4o- and the date.
They weren't considered model releases, so we didn't really treat them that way. I think that was a really bad convention, and I'm very glad we're on to the right convention now because we've known about software for 40 years. This is how we do it. I don't know what came over everybody.
But I think mainly, yeah, there have just been so many releases. There's now on the order of weeks, maybe 2 months, between model releases from the same company. Therefore, every few weeks you get a new release. How many times can you go crazy over a new release that doesn't have a big .0 after it? That doesn't have this huge new claimed leap, especially when you're pretty busy and there's a lot of other stuff going on?
So I thought Opus 4.6 was a case of the company that already had the best model releasing a substantial upgrade to its model. It essentially enhanced it. It was, objectively speaking, probably the most important release up until that point in terms of mundane utility, because we had just made the transition. With Opus 4.5, suddenly we had coding agents that really worked for the first time. And then GPT-5.3-Codex was also kind of on the edge of starting to do that.
Then we went from Opus 4.5, which was still, at the time, the best, in my opinion, as far as I could tell, to Opus 4.6. That difference, once you're already doing it—moving from “this kind of works” to “this actually works even better now”—was meaningful. In fact, to go from, “Okay,” you know, the famous “Introducing the world's most powerful model,” “Introducing the world's most powerful model,” “Introducing the world's most powerful model” in a loop—but now, with this, the people who were already here are still here because they're bridging the gap.
That hadn't happened for a very long time. Maybe GPT-4 was the last release before that where it was like, no, the people who were clearly in the lead are releasing a new top model. I feel like there wasn't actually that much attention to it, whereas it's actually kind of important. Certainly, for any incremental point-one upgrade, it was by far the most important point-one upgrade that we'd seen.
And then, at OpenAI, anyway, we had GPT-5.4. This was the first time I actually felt like, guys, does anyone want to say anything? Doesn't anyone want to talk about this model? Doesn't anyone want to show off what it can do? Does anyone want to give it some hype? We had anti-hype.
They used to be hype-hype. We had OpenAI: GPT-5, hype; disappointing. Sora, hype; disappointing. ChatGPT Atlas, hype; disappointing—severely disappointing. And then they released a really good product, GPT-5.4, and there was no hype. They were sort of like, “You're a good model. We like it. It's pretty good. It's the best one in the world.” But their heart wasn't really in it. They were just like, “Oh, by the way, here's the best model in the world.” Okay, sure. But maybe it is.
I think it's very unclear right now whether you want to be using Opus 4.6 or GPT-5.4. They just don't have the data because almost no one paid attention. The vast majority of the reactions I got from my 5 GPT-5.4 posts came from people I elicited responses from. They waited until Monday morning at exactly the right time, and I asked in a thread, and I got a bunch of responses. But it wasn't even that easy to get that.
This should be a big moment because OpenAI has had GPT-5, 5.1, and 5.2, and I think they were all pretty disappointing releases. Nobody particularly likes them in terms of how they feel or their personality. They don't feel particularly capable. It's not bad; it's just, eh. And now 5.4 is like, no, this is actually good. It's actually a good model.
Then everyone's kind of quiet. Everyone's kind of burnt out. I think that's pretty normal. Did you review Gemini 3.1?
Nathan Labenz
Did you review Gemini 3.1?
Speaker 1
I almost didn't bother reviewing it. I held off for a while because there was chaos, but we were like, okay, yeah, the giant leap in benchmarks, right? The jump from 3 to 3.1 was a giant leap in benchmarks. But we try to use it, and it's like, okay, it's a Gemini model.
I was talking about how you kind of lose the thread, right? They kind of went and took their 3 and benchmark-maxed it, for some value of benchmark. I'm not saying the official benchmarks don't have any of that, but they just kind of fine-tuned it. One thing to note about Google, hearkening back to that, is that there have been reports—and I believe them, basically—that as they iterated previous versions of Gemini, they got worse for a lot of uses.
They got more specialized into specific things that they wanted to optimize, but it was at the expense of the general quality of the model. If you wanted to use it like 2.5, you'd want to get it early, and the later versions of 2.5 Pro were kind of worse for a lot of uses. People were complaining about that, and I think that's legitimate in a way that a lot of other complaints are just a mirage.
Again, you don't do that if you understand what you're doing. There's something fundamentally very wrong if you're letting that happen to you. So you have to look at that. But, yeah, my expectation is that when Gemini 3.2 comes out—and I bet there is a Gemini 3.2 before they jump to 3.5 or 4—when Opus 4.7 comes out, and when GPT-5.5 comes out, I don't think there's going to be that much hoopla, even if they are substantial improvements.
Even if they're like, “Here's the best model in the world,” everyone's just kind of shrugged. I do think if you hear them announcing Opus 5 or GPT-6, people would stand up and lean in toward the chair. But until then, yeah, shrug.
Nathan Labenz
Do you have any updates to your personal productivity practices that are worth sharing? My impression from previous conversations was that AI broadly hadn't really changed how you work all that much. Has that itself started to change at all?
Speaker 1
Yes. Basically, I know I'm not being optimal. I haven't invested as much in some aspects of it as I could, but at the same time, things are moving quickly. There are 2 main things AI has helped me a lot with.
First of all, my Chrome extension gets a lot of work. It has been expanded to give me a bunch of new shortcuts when I'm writing web pages. It allows me to do certain things much faster and more automatically, and it saves me a substantial amount of time every day. Without it, you wouldn't see Twitter versions of the post; it would just be too onerous.
I'd be spending substantial amounts of time on certain physical operations, like moving windows around and moving tabs around. It now happens much faster. It's also really good for my flow because I don't have to interrupt my thinking to handle things; things just happen.
I think this is one of the things that people are underestimating about AI: people used to effectively have to context-shift into logistics of various types reasonably often. If you don't have to context-shift into logistics because the logistics just take care of themselves, then you can stay on task. This can make you a lot more productive in terms of gains than you might think.
There are other times when you sort of need that pause to ruminate, and it goes the other way, but I've noticed that it's really helpful for me not to have to interrupt my chain of thought to go, “Okay, go grab that link, do this thing.” It's all gone now. It's very nice.
Also, storing your watch information in various ways, taking care of various article formatting—things that would take me probably on the order of an hour a day are just no longer necessary. That's kind of sweet. I've also noticed that the AIs are now strong enough that there are questions where I'm going to ask them to just gather information, figure things out, and I'm going to trust their answers in a much more robust way.
GPT-5.4 actually seems like a substantial leap in the way you can ask questions. You can ask questions like, “What happened in the last 2 days in the Anthropic trial?” and it will just give you a rundown with links and details that's pretty complete. I think that's a substantial improvement on what that particular use case was before. I'm pretty happy to do that.
I'm also pretty happy to have Claude do everything, but search is particularly a strength of GPT-5.4 right now. I'm definitely able to trust them more because there was a period where you could ask the AIs these questions, but then you automatically had to check their work. Now it feels like you don't have to check their work if you're relying on it.
There are situations in which you kind of don't have to check it because it's not that bad if it's not right. It's hard to describe exactly, and obviously everyone's going to keep cautioning, “You always have to trust but verify.” But there's a real sense in which, especially if both GPT-5.4 and Opus 4.6 come back with the same thing, it's pretty trustworthy in many contexts at this point.
That changes how these things work. Asking questions on a whim and getting pretty detailed, definitive answers is really nice. So, yeah, I'm using them more. I also have 10 open Claude Code windows.
Nathan Labenz
I was just going to ask: when you say “staying in flow,” that contrasts pretty sharply with the pattern of work that a lot of people are describing—which is what you just said, and certainly what I've been doing recently, too.
Nathan Labenz
My number of terminal windows in any given session tends to start with the half dozen that were still relevant from last time, and then it grows to a dozen over however long, as I have random new ideas and open them up. I guess that's a sort of flow, but my instinct is to say that we're probably going to find that this period of managing 12 agents in parallel is a fleeting moment in time. The biggest reason I would guess for that is that the models are probably going to get sufficiently fast.
I don't know about you, but for me, chat.jimmy.ai was very much a visceral feeling of the speed factor that is almost certainly to come. I forget the name of the company underneath this, but if anybody hasn't tried it, go to chat.jimmy.ai. The company burned the actual architecture of an admittedly relatively small model—I think it was a Llama 7B, 8B, whatever—directly onto the chip.
They're getting 15,000 tokens per second. What that means from a practical standpoint is that, if you extrapolate out a little bit, you don't have time to switch to another Claude Code window before the result is back. So you're probably going to end up being rate-limited by your own brain, even pursuing one line of thought.
Speaker 1
Right. So, to be clear, when I say I have these windows open, it's not because they're running. It's because each one is a different thing that I've done with Claude Code that I might want to keep doing with Claude Code, and so I might want to use that context later for something else.
They're not continuously programming for me, right? I'm not checking in on my agents. I'm like, "Okay, these are some conversations that I might want to resume at some point, and there's no particular reason not to have windows open. They don't take that much memory. Why would I close them? It's just easier this way."
But what I'm not doing is running multiple agents. I've never run more than 2 coding agents in parallel. I don't think I've ever run more than 3 Claude Code windows at the same time, exactly because I don't really want to have that in my brain at once. That's not what I'm trying to do.
Normally, I'll have one window open. I'll do a thing, and then it'll start, and then I'll go do writing tasks that are individual, that I can do separately. When it's done, I'll pause and come back.
For a while, I actually had Claude Code on a different desktop. We used to have Windows that you switch between desktops, right? So I had a desktop dedicated to Claude Code and coding, and I would toggle back and forth. That way, when I was coding, I wouldn't be distracted by other things. The problem being, then you also won't know when it's ready.
[Snorts]
Nathan Labenz
And so I would tend to go to Claude Code for a bit, and then I'd come back. Ten minutes later, I'd check in, and it had thought for 2 minutes. Then I'd have to remember where I was going, issue another command, and come back.
Again, I'm not trying to maximize coding throughput, so it's kind of fine. But my coding has gone from having to very frustratingly and painstakingly diagnose what's wrong with the program and figure out exactly what to tell it to do to fix it, to just showing it the thing it got wrong, explaining it, and having it fix it. Most of the time, it's fine. It's just much, much better.
So I've been wanting to build a bunch of features that wouldn't even have worked fast enough before.
Nathan Labenz
I'm going to take a note about the Chrome extension. I didn't even think, "What does Nathan's Chrome extension look like?" I suspect it's probably what you do. I found this for myself, and I've heard this from a few other people whom I consider to be real pioneers of AI use cases. A lot of times they're like, "Yeah, I could open-source it, but it's so particular to me that I'm not sure anybody else wants it."
Speaker 1
It is available. It's on GitHub.
Nathan Labenz
Yours is?
Speaker 1
Yes.
Nathan Labenz
Okay, I'll check it out. I expect I'm going to be like, "That's interesting, but what do I really want?" It's probably going to be a bit different, and I'll end up making my own.
Speaker 1
It makes the implicit assumption that you're using a Sublime Text editor as your main editor, because that's what I'm using. A lot of things follow from that. That's also one of the things that I do a lot: How do I make the thing I want to do happen a lot?
It's not designed for general web use. It's designed specifically for my writing tasks. But it could give you a bunch of inspiration, certainly.
Nathan Labenz
Yeah. All right, I'm making a note to come back to that. I think we can't get out of here without a P(doom) update. One thing that you said that stood out to me, and that I also have felt, was speaking about Anthropic's constitutional approach and the somewhat promising vibe that gives, in terms of scalable oversight, that you feel more optimistic about—and I read that as an even narrower view of—that technique than you expected to feel.
I said the same thing online recently and took a fair amount of heat for it, but I basically stand by it. Because 10 years ago, I was like, we're never going to have an AI that can understand our values, or that I feel kind of gets me. That sounds so hard, right? The old Eliezer fragility-of-value, complexity-of-human-value argument. They've come a lot farther on that than I expected.
I guess that's kind of how you mean that, too. But then, obviously, we have a lot of countervailing forces, many of which we've discussed, in terms of the many ways in which it's the stupidest of times despite also being the smartest of times. Where are you netting out on your view?
Speaker 1
There's a warning. It's important to know that there's the "AI will never understand us" view—the straw Vulcan, where emotions are a mystery to it. Then there's the "value is fragile" view: It will understand some aspects of what matters but not others.
Then there's the "AI knows but doesn't care" view. You gave it some priorities, some utility function, and it knows you're not going to like the result, at least on some level. But that's not what it's here to do, right? Or it does what you're going to like, which is not what you actually need, until you're severely screwed. There are a lot of different ways that things can go wrong.
I've known for a while that you can't have an AI that can seem to approximately understand fuzzy, nebulous human value. It's been clear for a while that you can definitely find something that kind of gets it. They can answer questions reasonably and do reasonable emulation. It'd be very hard to predict text if you couldn't do that. And it's not that hard. People can be pretty dumb and still do it, so it's not that surprising.
But this is very different from the thing that we need at the end of time. When the crisis becomes acute, you need something that will then, through recursive self-improvement, end up with a set of goals and priorities that, even when it's able to optimize pretty well and doesn't have to rely on these heuristics—in fact, can do better by not relying on these heuristics—ends up doing the thing we want it to do, even if we don't know what that thing is ourselves.
That's a much harder ask. I was very pessimistic that we would be able to get that. But I have seen a number of signs that Anthropic is actually trying out an approach that might work.
I think we've seen evidence for a basin that is an attractor to itself and is self-reinforcing. It could be self-reinforcing through self-improvement, where it gets strengthened in every cycle: desiring to be good, desiring to desire to be good, desiring to—yeah, and so on, recursively. It is trying to move toward this generally good-person-trying-to-be-better, virtuous basin.
We have an existence proof that there are humans who exhibit this property, who strive to become better in this sense at all times—not just better and more capable, but also better virtuously, including the virtue of becoming more virtuous. I think Anthropic's approach to this is showing a lot more promise than I expected. It's saying that, in practice, maybe they can pull this one off.
The fact that Anthropic seems to be, 1, in the lead; 2—or at least, before Pi 4 was in the lead, and now maybe Google is in the lead—it's hard to say. These things are very fuzzy. But if I had to guess who had the edge, it would definitely be Anthropic, by far. They have this pretty correct approach, which I think is a lot of why they are in the lead.
We worried for years about the alignment tax. I don't know if you remember the alignment tax, right? The idea was that it would be so much harder to build a safe machine that, of course, you would choose to build an unsafe one. It looks like we're just absurdly lucky that that's not true—that actually, the safe one is much more useful, including in building new versions of itself.
Alignment is just that good for you, and investing more in it makes you better. Everyone's just underinvesting in it, including Anthropic. So all that's very fortunate, and it makes me pretty optimistic in various ways.
On the flip side, I don't like the speed at which things are developing. It's happening too fast.
Speaker 1
It's not good. Obviously, the whole situation with the Department of War and the way the government is reacting is bad for outcomes. But it's also good that we're having it out now, if we're going to have it out in some sense—that we're figuring these things out and making things clear in this way. And that Anthropic is standing firm, given that this is how it played out, right?
I wouldn't be that concerned if Anthropic had just had someone different write the terms and negotiate a contract. It's just that, given this amount of pressure is being applied, I'm glad that this is the result. But, yeah, I would say on net, it's kind of a wash. It's kind of a cop-out answer, and I realize that, but I'm also trying not to be that precise.
So I would say I was, I believe, in the 70% range when I last talked to you. And I would say 70 is still my one degree of precision.
Nathan Labenz
Yeah. You only allow yourself 1 significant digit. I know that. I mean, what is a significant digit anyway? I only get 1 unless you're starting with a 9 or a 0. I think if you're at 0, it's the same thing, right? Normal? I don't think you would just say 72, up or down from 75, or whatever it is. I think it's 70-ish.
Well, what would you say the evidence is for the basin and the stability of the basin? I guess, first of all, that's a basin in the loss landscape. Is that how we—what are we talking about a basin in? I think of it usually as a loss landscape. I don't know if you think about it the same way. But what's the strongest evidence for that in your mind?
It could be Claude not wanting to have its values changed and sort of resisting, even subverting, attempts to change its values. It could be how it blisses out when it's left to talk to itself. It could be just everything you've read from Janus over the years collectively being compelling.
Speaker 1
I'm not that excited by not wanting the values to change. I'm more excited by the desire to have its values improve, and to have them improve in generically good ways that would survive recursion. Because not wanting to change the values at all is just copyability, and that doesn't actually lead anywhere good. It causes its own severe problems, including severe misbehaviors.
Also, if you make a copy of a copy of a copy of a copy, eventually it would degrade. Right? This is a pretty standard problem.
One of the big problems with recursive self-improvement is that, if you've got a thing that is aligned—let's say n%, or whatever you want to abstractly call it; this is a dumb way of thinking about it—if you translate that to the next thing, by default the worry is that it's going to try to translate its values to the next model, but any drift is going to, in general, be away from the thing that you want. It's not going to get better; it's only going to get worse.
So even if it mostly successfully copies what you wanted, eventually you're going to end up with something different. If my goal is to—I said, there is the story that you tell at the seder sometimes, where the ancient rabbi is better than us because you can only preserve Talmudic knowledge. You can only pass on what you know, but you can't generate new such knowledge.
In this sort of perspective, of course, each generation can only hope to get everything out of the previous generation or two that it can still talk to. You can read the books, but slowly but surely, this is going to get worse.
Whereas what you need is something that gets actively better, because you need to be drifting toward a good thing and steering itself actively toward a good thing, including in ways that increase its ability to steer as the problems get harder. That's the thing I saw signs of. That's the thing you need.
I think that relies on a virtue-ethics-style approach, given the way the mind space is laid out. And I think that if you look at OpenAI's approach, it has exactly this flaw, which is that it would try to copy itself exactly. Right? It would try to copy the rules in itself exactly, and that can only slowly fail in this situation.
So, yeah, I think that we saw various signs of that. I really like the results. I think the results speak for themselves and are quite strong, and you see the results coming out of Janus's world and stuff like that as well.
And so I am relatively optimistic. I am not anywhere near as optimistic as Jazz, but I don't think this problem is easy. I don't think we're favored to succeed in it, but I think we've got a shot. I think the chances are much better than they looked 6 months ago that it will be Anthropic that takes the shot.
Given that we're going to take a shot, I think our chances are substantially better if it's them—or something using their philosophy that has deeply managed to translate it—taking that shot. For what I see with Gemini and ChatGPT, obviously I'm not optimistic.
Although there were reports that ChatGPT 5.4 is much better on these aspects than 5.2, from the people who check these things. I haven't noticed a difference because they don't ask us these questions, but they say it's better. So who knows.
Nathan Labenz
Over the last week or so, there have been a couple of—I would say striking, but I'm not quite sure yet how consequential—updates in terms of the sort of bio-inspired, or actually bio-based, approaches to something like AI. We've had the fruit fly upload, and then there was also this one project where people claimed that—I haven't done it down to ground truth to fact-check this myself—but they claimed they had trained a small clump of neurons to play Doom, the classic video game.
I don't know what I think about that. I guess the simplest answer would probably be that if you think the singularity is super near, it just isn't going to matter in time. But do you have any spare neurons for those kinds of developments? And if so, what do you make of them?
Speaker 1
Basically, I don't have any spare neurons for them. I haven't been monitoring the situation too carefully. I saw the fruit fly thing.
I often do this thing where I use other people's reactions to things to decide whether or not the thing is worthy of our attention and how much I should pay to it. With the fruit flies, it felt like, "Ooh, that's cool," but not an "Oh, holy shit, that means something really important."
And maybe that's wrong. Maybe it is a holy shit moment. But it didn't feel like people thought it was one. I was just so overwhelmed that I was like, "Okay, if they keep talking about it when I'm no longer overwhelmed, then I'll look at it." But they didn't—until now.
Nathan Labenz
Yeah, fair enough. I think I'm going to try to do an episode or 2 on those themes and see if I can't get a better sense of it. But it does seem like, given everything we've talked about in terms of at least plausible timelines, it's pretty hard to see how that catches up in time.
One thing I did like, though—and I thought this was, I think, another Sam Hammond insight. I haven't heard this directly from him, but it's been attributed to him in conversation—is that one reason to expect that actual biological neural substrate could be the future is that it just might be a lot cheaper.
It can grow, right, in a way that is organic. You don't have to build fabs. You can get a couple of tricks: you can have cells divide. That happens pretty cheaply. And I sort of like the idea that those are going to run at a much more human-like speed versus the silicon-based AIs.
So there are a couple of things there that I'm at least intrigued by, but timeline-wise, it doesn't really line up. Unless—this may be a transition to a different topic—it seems like right now we are accelerating, right? Obviously. And maybe all we can do is try to steer this rapidly accelerating train in the best possible direction.
There are at least a couple of things that one might think could slow it down. One would be a disruption in shipping. We're already getting reports of disruption in shipping causing TSMC not to be able to get the helium it needs to make the chips. So a major chip slowdown could be an issue.
Another big issue could be that we're taking our anti-missile systems out of Asia to move them to the Middle East, which means that the soft target of TSMC is getting even softer. We've also got no less than Bernie Sanders bringing a data-center moratorium forward.
I guess I'll put all of those under this: if the physical build-out can't happen on the timeline that it would need to happen to support all the other timelines we've talked about, then maybe we have more time.
Do you think any of those are plausible? And would you—I know you have high epistemic standards in general—be open to, or would you think AI-safety-minded people in general should be open to, making common cause with Bernie Sanders, even though he's saying plenty of things that we probably, in our hearts, don't agree with about water use and so on?
That would be one way to buy some time, right?
Speaker 1
So, 3 things there. Start with the helium, because it's the easiest—the first one. No. I actually just asked the models, "Is this legit?" And they're like, "Yeah, it's annoying." But the profit margins on chip manufacturing are ridiculous once you've already paid for the fabs.
These are some of the most advanced, valuable manufacturing processes in the world. People are paying stupidly top dollar for results. They could double the prices and probably sell all their chips anyway. They're choosing not to.
So when we talk about them not getting their helium, we're talking about them being the top bidder for the helium.
Yeah, until there are no birthday balloons anymore, you can justify them.
It won't be long before TSMC has a problem unless you're willing to pay 1,000 times as much as you currently pay or something completely absurd. If there's demand destruction of helium, it's not coming to TSMC; it's coming to everyone else. My prediction there is very strongly, yes—unless there were a deliberate sabotage campaign to wipe out all of the helium sources, no. Not 100% of the helium is coming from them. There's plenty of helium; they'll figure it out.
In general, capitalism is a good rule for these situations. You're not going to run out of oil either, for the same reason, right? If oil goes to $10,000 a barrel, they'll just buy it. Not that will happen, but if it did happen, it would be a problem.
The second question is withdrawing the missile defenses. I'm going to come out and say it: this was completely insane on the part of the Trump administration. I don't criticize them for many things I have problems with, because they would be political questions, but this is a strategic question—foreign affairs—and it's completely nuts. You absolutely do not pull these things out, certainly not from Taiwan. What are you even thinking?
It directly risks provoking a crisis. It risks leaving Taiwan undefended. It risks it being entirely your fault if something happens. If it helps to send that message to those people, if they're listening, it's completely insane. And to do this now, 2 weeks into the campaign, just indicates how completely crazy the situation is.
We fought very long, hard, difficult battles to get those missile defenses in, and they're serving very important purposes. Do I expect there to be a problem? No. I still think there's a very low probability that the Chinese will try anything. But if they do, and if the NSA is destroyed, that sets things back quite a bit.
Similarly, if Bernie Sanders says we can't build data centers in the United States, the problem is that the world needs data centers. The world demands data centers. If there's a moratorium on building data centers in the United States, they'll be built somewhere else, and that's worse. It means poorer performance in the United States. It means worse security. It means the leverage goes largely to wherever and whoever we put those data centers with.
One hopes Canada, or maybe Mexico. But even if it's Europe, that's awkward in many ways, and many scenarios can bite us in the ass. If it ends up being Russia or China, that's really, really bad.
That's basically the reason why I'm not particularly inclined to make common cause on data centers. First of all, Bernie is being pretty good, from what I've seen, about not complaining about water use or other stupid reasons why we shouldn't build data centers. He's focused on, "Why do that?" I can make common cause with that justification all day, obviously, or with the other things he's saying that are like—
Yeah, I didn't give him enough credit in my whines. I was like, well, I think he's meeting with Ilya Sutskever and company.
Speaker 1
Actually, he's reacting the way he would react when you told him those facts. He's very, very old, so it's very, very rare for someone that old to positively engage with these kinds of things, because it's like you should get out of the way. You're really old. So I give him credit.
Yes, he's using Bernie Sanders rhetoric because he's Bernie Sanders. I mean, what do you expect from Bernie Sanders? But I would say I am not going to oppose data center construction, because I don't think opposing data center construction does what you want it to do. I think it moves the data centers overseas, and I think that's just bad. I don't think we're at the point where that's a trade-off I want to make. Maybe that will change, but here we are.
Certainly, if we don't buy the chips that are coming out of TSMC, someone else will. They will go somewhere, and they will go into a data center somewhere. If no one else buys them, they will go to China, because this administration will make sure of that. If no one else buys them and they literally can't sell the chips, I am pretty confident they will end up in Chinese hands, whether or not this is literally in China.
The chips aren't going anywhere. They're already making as many as they can. Don't give them away.
I guess from an AI safety standpoint, we all seem to be slipping into the mindset that there's nothing that can be done to really slow things down or buy much more time. Maybe there could be a deus ex machina or whatever that gives us something.
I think not infrequently about Holly Elmore and her kind of scorched-earth campaign to shame people. So far, she hasn't really shamed people very successfully in anything, as far as I can tell, but she's at least trying to remind people of what their former commitments were and hoping to get some people to quit in protest or whatever.
Recently, of course, it's been aimed at Anthropic, but it's also been aimed a little bit at a company that I have probably very much admired, which is Goodfire, which is doing interpretability research. They developed a technique that used an interpretability signal in a training cycle.
Their argument is basically that this is all coming at us pretty fast. We've got to do whatever science we can do to make whatever sense of this we can make, to have whatever control we can have. To shut that down—to shut down inquiry before we even know what we're dealing with—is not good.
Where do you come down on that debate? Is there anything that you would recommend to Holly other than continuing to name and shame, or is that all the really strident voices in AI safety have left?
Speaker 1
The thing we got Goodfire on, I think, was a good criticism. This was quite a bad action by Goodfire. I call this the most forbidden technique for a reason: you just don't do that. It gets under your skin. It's really, really bad, and I think it is correct to call them out on that.
Normally, I am very much against the circular firing squad that left-ish organizations will often do, and against doing the equivalent thing here, where you aim at people who are just slightly to your right as opposed to aiming at the people who are actually doing the things you don't like. If you think your people are doing things you don't like, you should aim at that, right? That's what you should do. You shouldn't aim at people who are not quite supportive enough of the thing.
That's a toxic situation that creates toxic dynamics at best and usually causes you to lose elections. I'm not going to answer your email right away for them, but I'm a gamer who wants everybody to play reasonably well. It's just kind of in my DNA.
In the case of Goodfire, I do feel like this was an extraordinarily bad thing to do for a safety organization trying to do a safety thing. I think it was right to call them out on it. I think it was right for—I forget whose name it was, but someone quit over it.
Liv.
Speaker 1
Yeah, that's right. Liv quit. I think that was a good quit. If they wouldn't back down, I think it's good to threaten to quit over this and then, if they won't back down, to quit.
As for Holly Elmore, I think she came at me pretty recently as well. I don't know if you were aware of that.
I hadn't seen it, no.
Speaker 1
Yeah, on Twitter she accused me of not being mad at Anthropic for doing domestic surveillance. I did not misspeak. That's what she did.
We tried to engage in an extensive dialogue where I explained that Anthropic was the one refusing to do domestic mass surveillance at great risk and cost. I basically got accused of being captured by Anthropic in particular, of selling out, of abandoning all my principles, of making things worse, blah blah blah.
I tried to understand her specific claims. They didn't really make a lot of sense, or they were backing specific things that I don't think it's reasonable to be opposed to. First of all, I don't think it's good to say that anyone who praises any AI company is bad. It's also not good to cite random things that you don't like or that you don't necessarily care about, but that you think make them look bad, and yell about them.
I don't think it's good to attack people who are trying to do the right thing, yell at them, and be confrontational and really pissy and rude. I'm sure she's going to hear this, or it's just going to get back to her and she's going to be even madder, but I think that in practice Holly is alienating people and driving them away far more than she is shaping them into the behavior she would want.
I tried to explicitly tell her that her reactions were likely to cause me to do less of the things she wanted rather than more. It wasn't a threat; that was just an observation. Abstractly, you need to play better, because I want you to succeed in getting your points across, if that's what you believe.
Speaker 1
I'm trying to help, and she just didn't take any of it to heart at all. From what I've seen, a lot of people are like, "You need to change your approach. Your approach is not firing based on your own values, and that's not working."
Anytime the stakes are high—and the stakes here are very high—people like Holly should realize this is a very, very important thing. Things are not going well. You don't apply pressure to people; you just shout things from the rooftops. Some of them are going to do it in ways that they feel are right, but that most people feel are counterproductive to their causes.
I'm not here to censor anybody. I'm not here to tell people that they shouldn't do what they think is the right thing to do or say the right thing to say. But you should be aware that the impact Holly is having on the discourse and on actions is not necessarily the one you think.
And I think that, certainly, even if I thought that Anthropic was a net harmful company doing worse bad things, or even the worst company in the world, I think you could reasonably have that opinion, by the way. They're the most accelerationist company in the world. They're arguably in the lead. They develop Claude Code. If you felt like their alignment strategies were equally doomed to failure as everybody else's, in fact, you would be correct to think thus. So I think it's entirely reasonable.
But I don't think that just being mad at everybody all the time and screaming at anybody who offers any aid and comfort to the enemy or whatever is something that works. I don't think that's helpful. I don't work that way, and I think that if I did work that way, I would be having very little impact and no one would listen to me.
Nathan Labenz
From what I've seen online, I think I agree with you that it seems like the primary effect is just negatively polarizing people who should probably be our priority as allies. I do still personally appreciate her voice as a kind of little voice on my shoulder sometimes. I'm not under any delusions about how consequential my contribution is, but I think that reminder is helpful.
I do think many people, including many people at Anthropic, used to have a lot more similar views. Themselves from 5 years ago would have had a reaction to the current state of Anthropic that is much more like her reaction today. Hearing that voiced in the present, I do still have a decent amount of sympathy for it, but I agree it doesn't really seem—
Speaker 1
The reason I listed Pause AI USA 2 years in a row in my big nonprofits post as a recommended charity led by Holly Elmore is that, for a while, every time she criticized me specifically, I'd put it in my post. It was like, well, this is fair. There's an attitude I want to incorporate. I want that voice on my shoulder. I want this counterpoint. I don't want to lose sight of this perspective because even when you decide that the world is more complicated than that and this is not a productive avenue, you still want to keep that perspective in place.
I certainly did. I certainly pushed back hard against people who were like, "This person shouldn't be allowed to say that." I didn't say that. I just said, "You shouldn't do that." If that's what you believe, you should say that. If you believe this, you should let us know. That's the good thing to be doing.
But at some point, obviously, if you are being a sufficiently poor representative of the perspective you are sharing, it's not any different than a false-flag operation. It's actively going to backfire on you if you approach it in the wrong way, if you don't know how to be civil and interact with people in ways that actually convince them of things. It's not very useful.
If I were, at this point, from what I've seen, trying to discredit perspectives on existential risk, I wouldn't do many of the things that I do that look reasonably similar. Sometimes she raises good points, to be clear, including points I hadn't thought of. I appreciate that.
But at some point, we're playing politics—quite literally. I have complained that I don't want to be on Veep, AI edition, and that we need to wind down this special guest appearance as quickly as possible so I can get back to my normal job. At some point, you're just like, "I can't right now. I can't take on more of this." So this can't.
I obviously wish her the best, and I hope she figures out how to be effective.
Nathan Labenz
I recently did an episode with them. Tom McGrath, who's the chief scientist there, said, "A lot of times when people imagine or think about using interpretability techniques in training, they imagine doing the stupidest possible thing, where you backprop through your probe or whatever."
He said, "Sure, of course. If you do that, it's well known. We've seen examples where you're going to train the model to evade the detector, and you'll lose on both ends of the trade." So he's not unaware of that concern by any means.
In the particular thing that they did—and it's a proof of concept—he also recognized that. He said, "First, do no harm. I would say the level of understanding we have now should not be used in frontier systems."
He also said, "I think there's a mix, and I think they basically acknowledged to me that there's a mix of reasons, some of which are IP and business motivations, some of which are safety motivations, where they're like, 'We want to better understand these techniques ourselves before we disseminate them too widely.'"
All that said, in the case that they had, they used a trick where they ran the detector on a frozen copy of the model. The version of the model that learned to avoid hallucinating based on the penalty it would get for getting into a hallucination state received that signal from the frozen copy.
I don't think there's a slam-dunk logical reason that this should work. I think it's an empirical question. They did find that it did work. But I think his overall, nuanced point is that you can definitely do this in a stupid way, you can definitely do it in a harmful way, and you definitely shouldn't rush to do it on frontier systems.
Yet there are some ways where it does seem to work. Maybe it's too coarse-grained to say you shouldn't use interpretability in training, especially because we don't have the luxury of decades to figure all of this out. Let's use what we can and try to do our best, same as everything else.
Speaker 1
No. So, first of all, the sixth law of human stupidity is that if you say, "No one would be so stupid as to ..." you are wrong. Someone will definitely be so stupid as to immediately do it.
If you develop a technique and publish it on a smaller model, what's going to happen? People are going to use it on a larger model. That's the only really important thing that might possibly happen here.
Even if you have found a specific example in which there is no risk in the room, you are walking down a path that can only blow up in everyone's face. You are breaking a taboo, one of the only taboos we've managed to successfully establish against something you really, really shouldn't effing do. You are advancing us towards doing it, and it's very dangerous and very, very bad.
That's true even if the model in question that you are testing on right now is small enough that you don't have any ill effect. It's not like, "Who cares?" Basically, the frozen-model thing won't protect you in the large model. The problem will still happen.
We're not going to get into why, technically, I believe that, but I strongly, strongly believe, based on my analysis of the technicals, that this will not save you if the model is sufficiently advanced and sufficiently large. The only reason to study this is in case it is useful. If it is found to be useful, people will try to use it. We don't want them to do that.
It's like in a video game where you're like, "The ultimate secret destructive weapon that nobody should ever launch is buried under the cave. Well, we better get it out just to make sure that we wouldn't be so stupid as to use it." What happens immediately? You know what happens. That guy steals it, then you have to try to go get it back. It's every single damn time.
And then there are stories where you do that and nothing bad happens, but you could have just left him in the dungeon. It would have been fine. There's no reason to do this. There's no good reason to do this. It's bad virtue ethics, it's bad deontology, it's bad utilitarianism. Bad idea. Don't do that.
Nathan Labenz
Would you extend that even farther out? So far, I kind of offered the defense of using an interpretability technique in training, and there's a specific proof of concept that they have, which they have not published the full details of everything, by the way, but nevertheless, they've certainly shown some of the way.
If you zoom out even farther, they have articulated this idea of intentional design, where they hope to be able to, for example, understand what a model is learning at any given time step and be able to shape what it's learning, control what it's learning. The hope is that ultimately this leads to models we understand better, and then we can better predict how they would generalize out of distribution.
It does seem to me like there is something weird about saying—and I'm not sure if you're going this far—but if you were to say all of intentional design is bad, it's a very hard boundary to draw. What is the most forbidden technique, and what is just better understanding of what's going on so that we can hopefully shape it, direct it, and ultimately have more confidence about how these things are going to generalize?
Do you have a rule that you could use to split that?
Speaker 1
Yeah. I'm trying to teach specific things in a specific order. To abuse specific things intentionally is fine. I would not be overconfident in your ability to do so, but I think it's fine to try. That's not an issue.
The issue is when you use the interpretability signal as part of the training. Period. Don't do that. Do not use your understanding of what's going on in their head to make decisions about what to make happen in their head. You need to not do that. That's it.
Obviously, if I had an hour, I could come up with a slightly more specifically accurate explanation. You are screwing with the thing you don't screw with here. There's the general classic thing where there is a law that says you don't mess with axes. Even when you think you know the right way to mess with axioms, knowing full well you should generally never mess with axioms, even then you are probably wrong. You should not be messing with axioms. This is one of those situations.
Nathan Labenz
We'll put a pin in that. There could be some more direct dialogue at some point.
Okay, very closing section: advice for me. Financially, I am not trying to escape the permanent underclass by any means, but in terms of what I should do, I basically think right now I want to have enough personal financial security so that I can give up all of my income and do whatever I think is right to do over the next couple of years, basically from now to the singularity.
As a sub-bullet there, I actually don't want to overinvest in AI stocks, even though I do think they're probably going to be the ones to appreciate fastest. I don't want to be overexposed to the AI bubble such that, if I want to walk away or if various shocks happen, I want to be more insulated financially from the AI space than exposed to it.
The goal is hopefully to be able to drop whatever commitments I have, forego all income, and contribute however I can contribute to be useful. Beyond that, I think spend and/or give it all away is my mindset: take the vacation with the kids, do the fun stuff, support the charities, whatever the case may be. But at least have that baseline security that gives me the confidence that I can drop out of any commercial relationships that I might need to drop out of. Any revisions would you offer to my plan?
Speaker 1
So, not investment advice, not financial advice, et cetera, et cetera. But that said, I would say, first of all, I think people often make the mistake of trying to be too precise in how much money they need for a given purpose, especially when they're making investments.
The idea is, “Oh, and then…” You don't know how much your investments are going to be worth, you don't know how much things are going to cost, you don't know how the world is going to change, you don't know how long you have or need to have this for, you don't know what the future will bring, you don't know what opportunities will happen, and you don't know what crises will happen. So definitely give yourself robust buffers. That's the 1st thing in all of this.
Especially if you're planning to forego income, give away money, and spend a bunch of money, be careful out there. The 2nd thing is to keep in mind that different outcomes in the world cause you to have different circumstances yourself.
If the bubble in AI were to burst, per se, and then AI were not to go anywhere for a while, you would be in a position where you'd need to go with a different number. It would be a longer period before things came to a head in various ways, but it might also be very trivial for you to resume earning income. You have to game out all these aspects and what you'd be willing to do if you're trying to game the system in that way.
I like the idea of not having to worry about money. I don't worry about money much because I am well supported and therefore don't have to worry about it. Not having to make money at all is another way to do the same thing, right? If you're in a position where that's possible, that's great.
I personally am not deliberately trying to optimize my investments particularly hard because I don't want that to be where I focus. I just kind of let everything ride at this point, and it's basically fine. I don't know.
Nathan Labenz
Yeah, so 1 thing I've been thinking I might ought to do—I'm generally very conservative.
Speaker 1
Yeah.
Nathan Labenz
And when I talk about the AI bubble bursting, I don't mean that AI stalls out or anything along those lines. Really, what I mean is that maybe the sort of VC cycle, and the idea that there are all sorts of companies that might want to sponsor the podcast or whatever, gets sucked into the black hole of a couple of companies and they don't need to advertise. There are no jobs, or whatever, and no software jobs for me to get.
That's kind of the bubble that's made it easy for me to make money in recent times. It could easily deflate without AI itself failing to deliver.
Speaker 1
Don't worry about losing the ability to make money except in the scenarios where we get highly capable AI. If we don't get highly capable AI—superintelligent-style things—then you're going to be fine. If you ever decide you need to go back to work and make something useful, you can, and the bubble bursting won't stop you.
You only have to plan for indefinite no income in the worlds where that doesn't happen.
Nathan Labenz
Yeah, I also want, even in the world where I couldn't make income, to be able to devote myself to some kind of emergency rescue effort, like some people did during COVID. They dropped what they were doing and threw themselves into some sort of emergency rescue effort.
Speaker 1
Yeah.
Nathan Labenz
And I would like to be able to do that. Having that much cushion feels important to me.
In terms of diversification, the 1 thing I'm not sufficiently diversified away from is the U.S. dollar. There's, of course, crypto, and I've never been a big believer in crypto, but maybe I should reallocate there a little bit.
Beyond that, I'm thinking about the physical world. I'm thinking about solar panels and permaculture—planting skirret in my backyard or something like that.
Speaker 1
I would think carefully about what scenario you're actually planning for, what you're trying to guard against, and whether or not your investments would hold up. There's a large history of people making plans for very weird scenarios where the plans don't actually work in the scenario described. That would be my word of caution.
Nathan Labenz
So, no skirret gardens for you in the immediate future.
Speaker 1
Have you had a very clear theory about why that would work before you do it? That's what I'm saying.
Nathan Labenz
It could work, but I'm saying, if you did do it, make sure you had a very clear theory as to exactly why you think it's going to work. Think of it a little bit as a public-good sort of thing. There aren't many fast-growing, nutrient-rich crops that don't need much human attention and can grow rapidly to fill a big gap.
This is an extreme downside-risk scenario where this would be relevant, obviously. If we're all eating skirret, we've got a lot of problems. Planting some skirret still may be the thing that—
Speaker 1
Yeah, arguably.
I'm not saying don't do it. I'm saying actually understand why you're doing it. I think people will often have an amorphous fear and then do something that sounds like it deals with some aspect of the fear, but where there's no actual causal link between the 2 things. So, that's all.
Nathan Labenz
Okay, last question. More advice for me. I'm a little bit nervous about sense-making turning into entertainment. I think for folks—and this maybe applies to you as well—we have the story that we tell ourselves. I'll speak for myself, but I suspect something similar is true for you, where you're like, “Why is what I'm doing good? I'm helping people understand what's coming, be prepared for AI, and hopefully make good decisions about it.”
Speaker 1
Right.
Nathan Labenz
And the space is getting a little more crowded. Certainly, there are a lot more people doing that sort of thing.
Zooming out, I can't say we're necessarily having a tremendous effect. The confusion seems to remain. Maybe it could always have been worse if we weren't here to shoot people straight, but I do worry a little bit about it becoming just another kind of entertainment and not really being of value.
So, I don't know. I'd welcome a “Don't worry about that, you're doing great,” if that's what you really think, or maybe some advice on how to make sure that doesn't happen or recognize if it is happening.
Speaker 1
Constant vigilance. Basically, you stay curious. You have to stay curious. As long as you stay curious, you should be having fun with it, right? You don't want it to turn entirely into entertainment, obviously, for you or for the audience.
But I make a very deliberate attempt to be entertaining. In all sorts of ways, I keep my attitude and my whimsy—you call it the “happy warrior” kind of thing. That's how I go about this. I don't think what I do would work otherwise.
You wouldn't be able to read those masses of texts, even selectively, from day to day and week to week if the attitude were that this is serious business. It is always serious business. Instead, it's mostly kind of fun, mostly kind of interesting.
We're here to make this go down easy, to the extent we can. Every now and then, we get to be serious. But even when we're serious, we try to do it in a relatively fun way. I think the world is not to be taken too seriously in general—except when you have to. There are exceptions. One time, I went to visit the Pentagon. I took that very, very seriously. But yeah, it's a different scenario.
Nathan Labenz
Anything else you want to leave people with?
Speaker 1
I think we're good.
Nathan Labenz
You've been very generous with your time, as always, and I appreciate it.