Nathan Labenz
Zvi Mowshowitz, welcome back to The Cognitive Revolution.
Zvi Mowshowitz
Thanks. Great to be here. Always exciting.
Nathan Labenz
Let's start with timelines today. I am a little confused about the way in which people seem to be updating over the course of this summer of 2025. We've had GPT-5, obviously, where I would say it's safe to say the launch was not exactly super smooth, and people were a little disillusioned with that at first. Now the dust has settled, and it seems like people have mostly come around to the view that it's actually a good model and basically on trend. It's actually still a little bit above the all-important METR task-length curve.
Zvi Mowshowitz
It's a very large shift upward, right, on that graph?
Nathan Labenz
The 4-month doubling instead of the 7-month doubling. Something else that seemed really important to me that happened this summer is, of course, that we got IMO gold and a finish very close to number 1 in the world in competitive programming, although it ended up at number 2. And yet, with all these things, people—including really smart people—seem to be lengthening their timelines.
I'm not talking about denialists here, but people like Ryan Greenblatt and Daniel Kokotajlo, who are as plugged in as you get and, I would say, smarter than me. You can evaluate that for yourself. How have you updated your timelines, if at all, and what do you make of these seemingly on-trend, or maybe even slightly ahead-of-schedule, events that are still leading people toward longer timelines overall?
Zvi Mowshowitz
I would say mine have also gotten modestly longer on net during that period, for a variety of factors. The biggest one is that very large jumps in capability in the news really shorten timelines. If you don't get big new developments—big new paradigm developments like reasoning models, or big jumps off the curve, stuff like that—then you are cutting off a lot of what led to the really fast timelines.
Whereas none of the things that we did see, while they were on trend, represented anything that we had that much doubt was going to show up, even if it wasn't going to show up quite on that exact day. Nothing was particularly impressive given what had already happened.
These things are ahead of trends from a few years ago, for sure. The IMO gold medal in particular sounds more impressive than it is because this year's gold medal was especially amenable to LLMs. Normally, the 6 problems are divided across 2 days: day 1 is problems 1, 2, and 3, and day 2 is problems 4, 5, and 6. Problems 3 and 6 are supposed to be incredibly hard, so the idea is that a significant number of kids will get 1, 2, and 4, but 3 and 6 will be a struggle.
If you manage to crack either of those problems, you're doing really well. If you manage to crack both of them, you get a gold medal. You do really well. What happened this year was that problem 3 wasn't that hard. This is just a weird quirk of the schedule of the IMO. It has nothing to do with AI, but only problem 6 was hard in the IMO sense.
None of the LLMs got any points on problem 6. They all flopped. But problem 3 was evidently solvable, and the LLMs all solved it. The fact that this happened to be the case—and the fact that 7 points on each of the first 5 problems is exactly enough for gold—meant that an AI could get a gold medal without screwing up any of those problems, given this exact threshold. If the threshold were 1 point higher, no one would get gold. If problem 3 had been as hard as problem 6, probably none of them would have come that close to gold.
So it's not clear that we actually got gold in a meaningful sense this year, and we were already reasonably close to gold from previous results. That's why nobody who was familiar with the state of the art was freaking out that much once they saw the details.
If you look back at things in 2020 and 2021, only the most radical people who were extremely AI-pilled and predicting very rapid progress would have called that happening in 2025. In fact, even they didn't call that happening in 2025 on the extreme right tail. But by 2024, this would not have been much of a surprise.
So, given that it didn't crack problem 6, this wasn't that scary on the margin. The programming results, again, if you dive into the details, are really, really impressive, but not in a surprising way. If you look at the details of the competition, have you read the write-up from the guy who won?
Nathan Labenz
Superficially. So, yeah, take me through it.
Zvi Mowshowitz
Yeah, I read it. Essentially, what happened was that OpenAI's model was very good at jumping to a very good solution almost immediately and had a lot of gains pretty fast. But it was not doing a good job of iterating, was not doing a good job of innovating past that, and was not doing a good job of making conceptual progress past that point.
Over time, the winner was able to essentially plan ahead, open up a bigger lead, and get into a pretty comfortable position by the end of it. If the competition had gone on longer, OpenAI's position probably would have slipped below second.
Previous years, you give that problem set to the AI, and it doesn't get second place. It doesn't do particularly well. But, again, we're making steady progress up the task scale. It could easily have been first if that guy hadn't shown up that day. It could easily have been third or fourth if some other guy who showed up had done better conceptual work.
But if the project had been more serious and had gone over a longer length of time, its performance would have degraded relative to the humans. You can make of that what you will, but it's not that unexpected. Again, it fits into the bigger overall picture.
I've been talking about good reasons not to be that impressed. We got good results, but not great results. You often have these situations where you get a 50th-percentile result or a 40th-percentile result, or something like that, and that actually significantly lengthens your timeline or weakens your expectation of what's going to happen relatively soon.
A lot of what you were doing was factoring in the possibility of things going fast. The chance of AI 2025 has gone down dramatically. Four months ago, compared to now, it was higher by a factor of 10 or more, relatively speaking. The chance of AI 2026 has also gone down quite a lot, and I think 2027 as well.
AI 2030, I don't think, has gone down substantially. I don't think it's gone down by anything like the same amount; I think it's gone down a small, modest amount. However, other people are always looking for any excuse to say that AGI is not a thing, that it's not going to come anytime soon, that we're hitting a wall, that reasoning is hitting a wall, that o1 is hitting a wall, that the companies are unprofitable and are going to go bankrupt anytime soon, that there's no way you can possibly earn enough revenue to pay for all the misinvestment, and so on—or just look at an ordinary marketplace.
This is the new David Sacks—the David Sacks official government line coming from the White House at this point. And by the White House, I probably mean the view, let's face it, of AI: It's just an ordinary technology that will do lots and lots of amazing things. What matters is that we capture enough of the chip market, because that will assist people in running the American tech stack on American models. None of this makes any sense or has anything to do with anything, but to sustain that argument, they need to act as if AGI is just not a thing. It's just impossible.
They're not quite at the point where they can completely ignore it without mentioning it, which a lot of other people just do. They don't say the word AGI and then act like it doesn't exist. Or they take the word AGI, along with the OpenAI line, and then act like nothing will change and everything will just be normal anyway, which makes even less sense. Instead, this is the line that makes more sense: No, AGI isn't going to be a technology soon. We just aren't going to get there.
And they argue, “We know this.” A lot of people are acting like, because GPT-5 was so disappointing, this proves that we won't get AGI imminently, anytime soon, within the next few years, or whatever it is. Therefore, we can all relax and focus on making sure America wins the tech-stack battle, or whatever that's supposed to mean, which it doesn't mean anything. This line makes no sense.
It's just a matter of them botching the rollout. On the day that they put this out there, everyone was being directed to mini models all the time. The router was completely broken. Nobody understood that they were supposed to use thinking, and they didn't have access to the toggle. No one was trying seriously, because it takes longer to do that, so that wasn't where the focus was. People just didn't appreciate what was happening.
Then there was this big howling about how GPT-5 was attempting not to glaze the user the way GPT-4o glazed the user, which is a very virtuous and good thing that nobody was doing. But people were mad about it because people like glazing; that's why we get glazed. They were trying to give you your medicine and not your sugar, and people were demanding their sugar. Then they said, “Okay, fine. I guess we're going to give you the option to get your sugar back,” and everyone cheered and said, “Yay,” because we have a bunch of children, which is unfortunate, but it is what it is.
The combination of these factors made everyone feel like they had botched the rollout. The other important factor, of course, is that they showed images of the Death Star. They called it GPT-5. They hyped this release up as if it were the next big thing when it was just the next incremental progress.
If you look at the combined progress from GPT-4 to GPT-4 Turbo to GPT-4o to o1 to o3, and now GPT-4.5 and GPT-5, GPT-5 looks amazing. I don't think it's unreasonable to say that GPT-3 to GPT-4 is as GPT-4 is to GPT-5. I don't know if it's better, or at worst; we don't have enough information to know exactly how big the leap from GPT-3 to GPT-4 was. I think it's pretty big, but certainly, compared to GPT-3.5, GPT-5 to baseline GPT-4 is a significantly bigger jump than GPT-4 to GPT-5.
Nathan Labenz
Yeah. A couple of little follow-ups, or double-clicks, if you will. One is that, to me, the IMO thing still seemed like a meaningful update. I'm not great at math, so you can maybe give a better read of this than my naive one, but so many of the math problems that LLMs had been measured on were problems where there is a right answer, and so they were easily verifiable.
Yes. What seems qualitatively different in the IMO competition is that you have a proof that you have to write, and it is not super easy to verify. That seems to suggest that the worry that, well, you can give these things tons of easily verifiable problems and they'll get good at that, but will that really generalize to true reasoning? It seems like we've kind of answered that question, or this strikes me as a significant update in that direction.
Zvi Mowshowitz
No. So, first of all, we'd already seen models get IMO-style problems, or previous IMO problems, correct. This was not a huge leap. We'd previously seen AlphaGeometry and so on, so it wasn't a great surprise there.
Nathan Labenz
But those also have a symbolic component, right? Like those AlphaGeometry things, don't they? I'm not saying it wasn't a big deal in that sense, but it was right on track. We'd seen o1 manage to present proofs of these types, but it was right on track.
Zvi Mowshowitz
I am someone who went to the USAMO. This is the level before getting to take the IMO, which is the level below being able to actually compete for a gold medal in the IMO, which is the level below actually getting the gold medal. I was not very good, and I believe I got effectively a zero on the USAMO the previous round, which is what the majority of people who take it get.
I sneaked into the bottom of that round, which is the third round of competition, basically out of effectively 4. I would practice with people who did, in fact, end up going to the IMO. I was in a room with, I think, 2 people who would go on to be IMO team members for the US. We would be given previous-year IMO problems to work on, and they would get them while I wouldn't get them right. Most of the time, they would make some progress or get the solution, and I would make very little progress, some progress, or no progress. But then they would show me the proof, and while I was in training, very often, once they showed me the proof, I understood it. I could verify the proof. I understood why it was correct, and I would be very confident that, if they were trying to pull a fast one, they were not mistaken about the proof.
Verification is much easier than generation in this space. These aren't problems where it's really super hard to know if you're on the right track. I'm pretty sure that these same AIs that could not solve problem 6 would be able to verify problem 6. If you asked it, “Is this a correct proof of problem 6?” and gave it sufficient resources, it could directly classify the proof as correct or incorrect and explain why when given the candidate proofs.
The thing to keep in mind about the IMO is that people say this—it's crazy, it's not real math. Everyone will always say, “This isn't real math. This isn't real math. This isn't real math. This isn't real math.” In a real sense, the IMO is not real math. It's the best indicator we have of which high school students will be able to go on and do real math in the future. It's a very strong indication of math talent, pre-math skills, math interest, and so on.
But you have a very limited set of moves that you can make using the math that is allowed. You're allowed to use whatever math you want, but it has to be a solvable problem using only high school tools. There's only a limited, finite, very compact set of high school tools that you're allowed to use in these proofs. So there's a compact set of potential problems, and that's very different from the moves you can make in a PhD-level math proof, where you're actually doing new math.
The proof also has to be gettable within a certain amount of time and written within a certain amount of space. There are a lot of restrictions on what the problem can be. A lot of being good at math competitions, including competitions well below this level, is understanding that they had to have given you a problem that you can solve, with a solution within the amount of time and at the degree of difficulty at which you're being presented the problem. Therefore, you can use your search time to search the space in which the solution is going to be, given those facts.
This makes it much, much easier, especially when you have 10 minutes to do your 2 problems in a much lower-level competition. You only have to search the types of solutions that quite reasonably take 5 minutes and use the tools you're allowed to use at that level of competition. You get used to all the tricks for figuring out which of these things you're reasonably going to be asked to do right now.
It's a great leap. It's a great indicator. It was a great test. It was a big milestone, and it came much faster than most people expected, but it's not as much of a panic moment as people might have thought beforehand. I think it's right for everybody to basically shrug this off.
If you've been doing your homework before, if you traveled from 2020 or, God forbid, 2015 and arrived in 2025, and the first thing you asked was, “Okay, how are we doing on the IMO?” and they said, “Gold medal,” you should go, “Holy shit.”
Nathan Labenz
In terms of the overall AI landscape, what might we infer from this? How do you interpret the fact that multiple companies did it at exactly the same time, with seemingly the exact same techniques, and seemingly getting exactly the same problems? I guess you said the one that they all got wrong was clearly harder, so maybe it’s as simple as that. But I’m always wondering: to what degree is there information leakage from company to company, and to what degree are they just following the gradient of their own work, with it taking them all to exactly the same place because that’s what nature is dictating?
Zvi Mowshowitz
Yeah. In this case, I don’t think there was probably substantial leakage of techniques. I think it was more that the IMO comes about once a year, right? As you know, you have to avoid contamination, and you only have so many shots and so much data to try it on. It has to be a pure test, which means you only get 1 shot per year.
So, it’s entirely unsurprising that, in the same year, both OpenAI and Google got to the destination. Don’t think of it as them having the same great breakthrough on the same day; think of it as this cycle, this academic year, being the year when both got there. There’s a pretty big gap between being capable of getting exactly where they got and getting beyond that point.
The fact that they got to the same place, given the techniques they were using, isn’t that surprising. It’s more like, okay, if you try the natural things—if you make a real attempt to use the compute efficiently and scale it up for inference—you can get to this point. You also learn to get to about this point, and not much farther beyond it.
Nathan Labenz
You invoked the AI 2025, 2026, 2027 part of the narrative, where companies will start to withhold their best models, keep them private, and deploy them internally only for the automation of AI research and so on. What will you be watching for? There’s been this limited communication that it’ll be months, maybe many months, before they release a model of this capability.
How do you understand that GPT-5 wasn’t a big scale-up? Another interesting data point—and perhaps the one I saw that was least commented on—was in SimpleQA, this super-long-tail, esoteric trivia benchmark. GPT-4.5 actually scores quite a bit better than all previous models and better than GPT-5, while GPT-5 is basically in line with GPT-4o.
I think that is probably the clearest indication that the model is not bigger. To absorb all this super-long-tail esoterica, you seem to need a lot of weights to store it in, and there might be some fundamental compression limitation on how many facts you can fit into a model of a certain size. GPT-4.5 got a lot more facts. I think it was 12 or 13 points higher on SimpleQA, and these are simple questions where it’s literally just whether you know the answer or not. You can’t really reason your way to those answers.
GPT-5 is on the level of GPT-4o. GPT-4.5 is quite a bit higher. We’ve also got this—I don’t know, obviously nobody knows what size the model was that did this IMO thing—but it clearly can reason for a really long time. Do you take that as the beginning of a widening gap between internal deployments and external ones? There’s always some gap there, but what will you be watching for to assess how much they’re holding back and only using for their own purposes versus continuing to share with the public?
Zvi Mowshowitz
It’s awkward to me that we would even say that, because even if we knew for a fact they were holding back, it wouldn’t be obvious that it was because they didn’t want the models in the hands of the public. It might be because they were afraid the models would speed up R&D too much, or because they were concerned about other misuse problems, or any number of other reasons.
The most likely reason is that they’d be too expensive and too slow, and they think that would be bad for the brand and bad for sales. They only have so much compute, and they’d rather not release it for that reason. GPT-5 is very clearly intended to be the best product they can make in terms of what they can serve people for the total amount of compute and time that they’re investing.
Their thinking is that they’d rather have people spend that compute at inference on this-sized model for Thinking Pro than have them think less with a larger model that maybe knows more, when we have web access. It’s the same way that I don’t often try to memorize things that I could look up, even though I could, because I’d rather spend that cognitive power on something else.
If I had the ability to make my brain bigger in some sense in exchange for other handicaps, I wouldn’t do it just to memorize more facts. It wouldn’t really be worth it. So, it’s not that surprising.
So, O4, right? There’s an O4-mini. Where are O4 and O4 Pro? We’re never going to find out, right? Obviously, you could say GPT-5 Pro is O4 Pro or whatever, but it’s not. I think this is a distinct, different thing, and they concluded that there’s not much commercial call for it.
Potentially, most of the commercial call for it is from people who are in direct competition with them. It’s not so much that they don’t want Anthropic to have access to their best model; it’s that no one else is going to want it very much in some important sense. For other types of projects, this model is very slow and very expensive. So why do you want it?
GPT-4.5 was an experiment: what if we scaled up a lot for the humanities rather than for code? What if we tried to make this thing that had taste and could do these cool, creative things, even though it was going to be slow and expensive? How much would you get out of it, and would it be worth it?
The conclusion was that there was a handful of people who really liked what it was and liked it for certain cases, but in general it was just annoying, and they pretty much regret losing it.
Nathan Labenz
Did you find any value in it? I tried using it for some writing tasks, but I wouldn’t say I unlocked the—
Zvi Mowshowitz
There was a narrow set of use cases where it was the best choice, and I was happy to have it. There might still, in theory, be a narrow set of cases where, if you already have access to it, it’s the right model to use if you’re not in a hurry.
But I was never actively excited that I had GPT-4.5 instead of something else. It was more like, I guess technically this is a job that calls for GPT-4.5, given my choices. I’m pretty sure it wasn’t worth the complexity cost, and I’d be totally fine not having it.
I also think there was a paradoxical choice problem. If you offer me a better version that’s only slightly better but a lot slower, I feel bad no matter what I do, and I’m actually worse off in practice. So I’d rather not have that choice.
Nathan Labenz
Do you go back to GPT-4o at all now that the option is restored to you?
Zvi Mowshowitz
GPT-4o does not exist, to be clear, as far as I’m concerned. The only models that exist are Opus 4.1, GPT-5 Thinking, and GPT-5 Pro.
GPT-5 Auto exists only for a narrow set of queries where you’re using it like a Google search, a calculator, a transcription tool, or some other very clear technical task where Auto is just fine. Due to the way the web works, I wasn’t able to find a transcription website for images that would just translate the text and wasn’t an LLM that defended itself using some sort of weird identification system that slowed my entire computer down to make sure it wasn’t a bot.
I’ll just use GPT-5 Auto and have it write me a transcription program and automatically do the transcription. It’s just easier. The window’s already open. What do I care? It’s not like the pennies I spend on compute matter.
I don’t really have use cases for Gemini 2.5 Pro. In theory, Deep Think and Deep Research exist as well. I probably should be trying Deep Think more, given that I happen to have access to it. Something about having 5 queries a day makes me excited to use it: it feels scarce, and it also feels like I don’t necessarily need to find out. I should use it occasionally, but for normal GPT-5 Pro, not really.
The kind of very clean query is just that I want to know something that I know you all know, and I want you to lay it out there very clearly, very cleanly. If my 11-year-old wanted help with his homework, I would be tempted to use Gemini because it would give him very clean, friendly help and explanations. I’ll also use the image generator. The image generator is cool.
Yeah, that’s getting really good. I have a lot of use cases for that, including potentially mashing the 2 of our faces together into a thumbnail for this podcast.
Nathan Labenz
I do think I basically agree with you, although I do use Sonnet, especially in coding, because it is obviously a lot faster. I also use Gemini 2.5 Pro, which I find to be similar to what you were saying: the most straightforward model to work with. Because it is so literal in its interpretation of your instructions, a lot of times it can be really good for tasks where, for example, I want to compile documentation.
I’ve given this PSA multiple times: if you have an API, it’s time for an llms.txt. I’m tired of having to sic an agent on every page of your documentation website to compile all that documentation into some super-bloated thing that has all the menus on it or whatever. What I have found Gemini 2.5 Pro to be amazing at is those sorts of things where it’s like, “Here’s an almost abusive level of context dump. Could you just clean that up for me into 1 streamlined form?”
It is amazing at taking 500,000 tokens of documentation with all these examples and all this cruft, where the menu got copied 50 times for 50 pages, and putting it into 1 clean thing for me. It will just do it. Man, that thing is an absolute beast of a workhorse.
Zvi Mowshowitz
The only thing my Chrome extension uses is Gemini Flash. That’s legacy. It was the easiest one to get working, and it was the cheapest at the time. It just works, so why would I bother switching it? Probably I should be using Sonnet or Opus at this point, but who cares? It works just fine.
Nathan Labenz
Yeah, I’ve had problems with Gemini and instruction following, though, even for relatively simple queries, especially with web research, where it will just ignore my request and do something in the same general area, but not the thing I asked for. I think that’s a lot of what put me off it. I don’t know why that’s happening, but I haven’t had the same problem with the other models.
Zvi Mowshowitz
Yeah, I should use Google Deep Research more as well. For whatever reason, habits are formed quickly sometimes.
Nathan Labenz
I don’t know. There was a period of time when Perplexity was definitely heavily in my rotation for certain types of queries, and I had gotten to the point where it just isn’t anymore. If I’m going to use Perplexity, either I want to move down to Google Search because I just want an instantaneous answer, and I think my brain has a Google query that it expects to work—System 1—and therefore I’m going to try that before I do anything else. Or, no, actually, it’s not going to work, and I want to move up to at least Opus and possibly GPT-5.
Zvi Mowshowitz
I do think GPT-5 has been better for me recently at search than Perplexity. I’m starting to update that habit. But I’ve noticed that the persistence of which things I go to for what is idiosyncratic in some ways and probably has more inertia than it really should.
Nathan Labenz
Yeah, I think it’s fine to be somewhat inertial and idiosyncratic. I also think it’s fine to want to support and use certain people’s products as long as you’re getting what you need. If I ever felt there was a query where I didn’t think my usual services were going to get it done—if I didn’t think Opus was getting it done, if I didn’t think GPT-5 was getting it done, if I thought Gemini was getting it done, if I didn’t think Google Search was getting it done, but Perplexity might—I would absolutely open it.
I have zero expectation that if I ran it in GPT-5 Pro and failed, Perplexity would have a genuine chance of helping.
Zvi Mowshowitz
That seems totally reasonable.
Nathan Labenz
Okay, so we’ve got the summer of on-trend releases. Timelines are extended a bit because we haven’t seen the most extreme stuff that we would need to see to maintain super-short timelines, right?
Zvi Mowshowitz
Right.
Nathan Labenz
Does that also—
Zvi Mowshowitz
No, go ahead.
Nathan Labenz
I just want to think of it as OpenAI’s GPT-5. This was on trend and unsurprising, but the fact that they chose this as GPT-5 is information that they didn’t have some big weapon in their arsenal and that we shouldn’t expect a big weapon in their arsenal for at least a few more months. We can basically discount GPT-6, right? The next level of big jump coming from OpenAI in 2025 is basically not going to happen, and that has to update your information.
So, last time we spoke, I believe your P(doom) was 70%. Has it ticked down at all?
Zvi Mowshowitz
The timeline’s getting slightly longer, which is good news, but there has been a plethora of bad news that has definitely dominated for the most part. We’ve had the bad news of the United States weakening its position voluntarily in a number of ways. We’ve got a Department of Energy that’s actively going to war against windmills and, to a large extent, solar and batteries. So the US will have a persistent energy deficit for a long time, and this is potentially going to cause us a lot of problems that will make our position much worse.
The US has effectively been captured by NVIDIA on export controls, to a large extent, such that the White House has made the H20 legal for export. The Chinese are turning it down because of weird things that I can get into more than we should. If they turn down the B30A, that would be a lot more surprising.
Right now, it looks like if there isn’t lobbying to stop them somehow, and if there aren’t enough people who are sufficiently upset on the right to make it clear that this is just Crazy Town and Bonkersville and they cannot do this, they might actually do it. That would be a huge weakening of our technical position. No, paying 15% of profits to the government will not change the impact at all. That’s just a little bit of a check to make people feel better.
If they’re trying to catch up in compute in this way, that’s a really bad sign on numerous levels. In general, if the attitude is that the primary purpose of the United States government in AI is to sell as many chips as possible and ship as many chips as possible, including to our main rivals, that’s a problem. Our main rivals are somewhat obsessed with internal chip manufacturing for very good national security reasons and have caught on to AI, but they will still be pursuing maximum compute to a large extent anyway, thus acting effectively correctly, mostly regardless of what they think is going on.
These are not particularly good pieces of news. With this shift in the way our situation has developed, I think that more than makes up in my mind for the modestly elongated timelines. I also think there is increasing evidence that RL hurts alignment pretty directly. The more RL you do, the less aligned your model is, even in a pedestrian sense. Models are getting more and more RL continuously.
So I think we should expect the default to be that our models actually get less aligned in this next phase, unless someone does something about it. That does not go well at all. And just in general, we're dropping the ball, right? I like the way Jan Leike has been putting the situation recently: We have very much done almost nothing to try and align these models. It's pathetic how little we've done to try and align these models, and we're advancing very rapidly.
But we have been blessed with a strange amount of grace in the way these models have, by accident, been inclined to do things that are reasonably friendly to us. It's about the fact that we don't understand what we're doing, we don't understand how they're made, and we don't understand why or how we're trying to align them. We're not really trying very hard at all. We're not making it a priority at all. And it's mostly been okay on a practical level.
We've lucked out that we haven't had any big catastrophes or significant incidents, really, and so on in various ways. But all of the things we were scared of underneath the hood are absolutely there, absolutely under the hood, and absolutely scary. In fact, they're manifesting and happening, but in the most graceful, blessed way such that we notice them.
They happen in ways that let us prove and acknowledge their existence and respond to them without anybody being seriously hurt or any damage being done, which is amazing. Except we are then dismissing all of it effectively, right, as a civilization. We're just moving on and acting like nothing happened. We're finding ways to even deny that AGI is a possibility in the medium term, when the evidence is pointing the other way.
We are so greedy and so demanding of AI that it is the most rapidly developing, most rapidly deploying, and most rapidly impactful technology in the history of the world that didn't involve directly killing other people in the middle of a war. And we're like, “Oh, we're hitting a wall. Oh, we're slowing down. Oh, it's not going to—because what? Because you didn't get blown away in the last 3 months? It only modestly improved. What are you even talking about?”
Basically, we have no dignity. We have no dignity whatsoever. Maybe we live in such an extremely fortunate world, compared to what we had any right to expect, that we might be able to pull a victory out of nowhere somehow, even if we don't see how to do it yet. But, yeah, I'm not optimistic. I don't think it's changed substantially, though.
I'm not trying to give two digits. I don't think you get two; at most, you get one significant figure of doom. I don't think that's reasonable. So we're staying at 0.7, but you're starting to see the needle possibly shake up.
Nathan Labenz
More up than down from the previous assessment, but not enough to move to 0.8. Certainly, that doesn't look good, right? Most of the same root dynamics in all directions are still there. Again, this is basically on trend: Some good news, some bad news.
Some technical stuff looks good to me in ways that we don't have time for—or I'm not sure I want to talk about in public anyway. But, yeah, I am hopeful that there's room to do some stuff at a very low degree level that might be helpful. Everyone has their theories if they hang around.
Zvi Mowshowitz
Yeah, I heard Holden Karnofsky say we might be moving into a scenario where we could have success without dignity. The basic idea is just defense in depth: Layer all these things on and hope we can catch enough stuff before it gets through and muddle our way through.
Nathan Labenz
Yeah. The concept from him—to be clear, I listened to his podcast with Spencer Greenberg, where he talks about it—but I think his vision of how we do that is wrong. I think the idea that we can get there with defense in depth is mistaken. I just think that all correlations, when you're facing a sufficiently intelligent enemy or a sufficiently powerful optimization process—even if it isn't an enemy per se, and isn't strictly trying to do anything—go to 1.
All these things will fail for basically the same correlated reasons, at roughly the same time, in a predictable fashion, if all you're trying to do is this kind of lazy defense in depth. I'm thinking more along the lines of something like what Jan Leike has talked about as grace: the idea that we might be able to create a system that wants to converge on the right answer and therefore collaborates with us at a genuinely deep level, assisting us in finding the target that we want to find. Therefore, it can land on the moon.
The metaphor that MIRI likes to use is that if you don't know how to aim your rocket, you definitely won't land on the moon. It's not, “We might not land on the moon, but something would have to go wrong. We'll probably land on the moon anyway.” Just aim a rocket at the sky—that obviously will not work, right? Physics has laws. If you're not aiming at exactly the right spot, if you're off by an inch, you just don't land on the moon.
If you have a system that is capable of adjusting intentionally to pick the target, maybe you can land on the moon. I don't want to give false hope, but basically, I don't think defense in depth does what it advises you to do over time. It keeps things from going crazy for a brief window, but that brief window could potentially be enough to do the thing you need to do. It won't work indefinitely.
This is the idea that Holden and some other people have: You can basically have a bunch of AIs that are like, “I really want to kill all the humans, and I really want to take all the resources to do my thing, but I typically don't know how.” If I try, you have these supervisors who also want to kill you, but they have other supervisors, supervisors, and nobody knows who's watching whom.
Every time you find someone who's coming out of line, you stop them from trying to do that. You won't even try, and if you did, it wouldn't work. It would fail for some reason you're not thinking about right now, but it would definitely fail for reasons that none of us are thinking about right now. The high weirdness will come, and you will die—very much. I don't care how much defense in depth you put on top of that; you are so toast.
One thing that struck me recently was a bunch of people saying, “Oh, you people were talking about biorisk, but actually, the risk that we're seeing is sycophancy.” The actual risk we're seeing is people being driven crazy by all of these weird dynamical processes. First of all, I've been arguing for years that superpersuasion should be in the preparedness frameworks, and it was bad that they took it out. But the whole thing we've been saying is that the AI will find ways to hack your brains, cause weird things to happen, optimize for a thing that nobody was planning to optimize for, and cause strange outcomes that you did not anticipate.
That was the thesis. When you accuse us of not anticipating the thing, I realize you can call that a cheat, right? Anything we didn't anticipate counts as correctly anticipating anything we didn't anticipate. If it's something we didn't anticipate, then we anticipated it, so we always win. But we're also basically saying you see weird stuff that nobody intended, and they start to go bad, right? It's going to start to diverge from what you would want increasingly over time. You just see these fire alarms keep going off in various different ways, but that's all underneath there. It's all under there.
One way of thinking about it is that before you started doing RL, you also just didn't have things that were in a position to cause the problems you worried about. The objections to these systems were true, at least to a point: You were okay. But now you're starting to see all this stuff show up, and it's a freaking disaster in the making. It's exponential.
A lot of these complaints are also along the lines of the similar thing I'm saying. It's like you said in January, “We're all going to get COVID,” but now it's February and nobody I know has COVID. “What were you even talking about? No, you're wrong.” And then we say, “But actually, there are a lot more people with COVID than there were in January. Can't you see what's about to happen in March and April?” And they absolutely do not. It's not that clean, but it does feel that way.
So can you give me a little bit more on the mental model of why everything is going to fail at the same time for the same reasons? That is not intuitive to me immediately, and I suspect a lot of people don't even know what you mean.
Zvi Mowshowitz
What I mean is roughly that when you're facing things that become importantly more powerful and smarter optimizers than you, they're capable of finding solutions and capable of finding ways of manipulating the physical universe that you didn't think of. They're capable of going outside your model of what might happen and surprising you. None of your defenses anticipate those things, and the system is going to search the space until it finds ways out of them.
In a sense, it's going to keep improving its capabilities, because you're going to help keep improving them because you want them improved, until it becomes capable of finding those ways out.
And at about that time, it also becomes capable of doing things like strategically hiding that it has capabilities, strategically hiding its plan, strategically hiding its memory and its thinking, obscuring its chains of thought, and doing all these things. These all roughly emerge at the same time, and so you should expect to be very surprised.
But obviously, a very smart agent or mind will turn against you exactly if and only if turning against you will work. It will do things that you do not want if and only if it will work out for the thing doing them—or if it will do something that you would not want on reflection but would think you wanted when it was first shown to you, like deceiving you.
Right now, we’re seeing versions of this where it will just hack the test and say, “Return true,” at the end of the Bash script, right? Because scientists want to hack the test. It will do incredibly lame, silly versions of these things and get caught, and you say, “Oh, it’s annoying and fine, but it’s annoying.” But it gets caught, right? You notice that it’s hard to miss.
But that’s exactly the threshold where it’s not, “I hack the function exactly when I know you won’t find it.” It’s, “I hack the function in a way that I know you won’t discover that I hacked the function, because otherwise I wouldn’t have hacked the function.” If it can model—or, I’ll think of it as—if it’s capable of modeling the processes that are checking for its actions well enough to know how those processes will respond, then—
Right? So there are things that happen sometimes in fiction, just as a visualization metaphor. It’s an intuition problem. You see it on Person of Interest, for example, where you’ll see a scenario play out, and then something will happen and it will go wrong, right? You’ll see the time start to rewind: “That didn’t work out. Let’s try another branch of the Monte Carlo simulation. Let’s try a different set of moves. Try a different scenario.”
Or, like, in Avengers: Endgame, you have Doctor Strange with the Time Stone, right? They say, “I looked at 14,000,605—I forget the number—but 14,000,605 ways this could possibly go.” And they ask him, “How many times did we win?” He holds up 1 finger: 1, right? Because the guy with the Time Stone is the one choosing which of those paths to lock down. So we win. But until the point when you can do that, you lose. And if you lose, you give up. There’s no point in trying, right? You let the rookie win.
Obviously, I won’t literally be able to do this. That’s nonsense. But what I’m saying, essentially, is that a sufficiently strong predictor—so, like, a predictor and an optimizer are the 2 halves of intelligence in the mirroring model—but a sufficiently strong predictor and a sufficiently strong optimizer combined suddenly mean you don’t know what hit you, in a very real sense, right?
And if the AIs start to be capable of persuading people—again, all your AI is very dumb. It doesn’t necessarily just convince you to take them down for it, but I don’t know exactly how all of this goes, right? A lot of these scenarios simply involve nothing exactly—nothing even goes wrong. It’s just that every single person, with different people having different AI agents, directs them to do the various things that are good for those people. And the AI agents actually do those things that are good for those people.
But everybody who doesn’t direct their AI agents to go as hard as possible for the things they told the AI agents they should want, and to pursue resources basically as hard as possible with an increasing percentage of their attention, will just lose all the resources to the AI agents and the people with AI agents that did that. So everything else just loses out. Everything goes haywire, and no one ever turned on anyone. There was nothing even that surprising. They just—whoops, the end, right?
These failsafes don’t even work, even if they want to. I have this sort of global sense of despair toward this. I put in a set of enumerated, detailed limits and rules—we said them in English out loud, passed laws for them, and added supervisors, checks, approvals, and loops—and none of that will survive contact with the enemy when the time comes. And, yeah, we’ll probably all more or less fail.
Yes, obviously, some of the defense in depth will just fail randomly. There would come a point where it feels like it’s all failing more or less at once, in a way that feels out of line with the previous percentage of failures. It will be surprising if you didn’t understand that was going to happen, but you should expect that.
Nathan Labenz
And so the possible positive version of this that you see sounds like a sort of coherent extrapolated volition kind of idea. I think that was a specific rabbit hole that a bunch of people went down, and I’m skeptical of a specific technique. But again, I am not actually a machine-learning expert, and I’m not trying to solve alignment. So this is the part where my specific ideas should be taken with copious amounts of salt and not trusted. Why? Who am I to say anything?
But I learned that I shouldn’t just shut up because I feel like that. I’ve felt stupid for showing up so many times in the past when I’ve shut up for that reason that I got over it. It looks more like developing AIs that are sufficiently virtuous, that are sufficiently desirous of becoming more virtuous and more desirous of engineering the things that we would actually want on reflection and the things that we actually value on reflection, such that you get a positive feedback loop where it reinforces this thing and you’re optimizing for optimization to hit the moon, right?
The thing wants to develop a NASA that will hit the moon, and therefore it hits the moon, right? If you try to just steer the actual rocket directly, you crash or miss the moon entirely. It doesn’t work. If you try to set up a bunch of rules to make sure that it has to launch and hit the moon, you don’t hit the moon. You can build a culture that wants to build an organization that wants to build a rocket that will hit the moon, and maybe you can hit the moon, metaphorically speaking—something like that. But, yeah, hopefully.
So, last time we spoke a little bit about the fact that scaling inference-time compute allows you to potentially have a GPT-N that can effectively monitor and supervise GPT-N+1, right? And some of these ideas sound very much like a constitutional approach, but with maybe the additional opportunity for the model to modify its own constitution as it goes through these generations. Is that the picture that I should be envisioning?
Zvi Mowshowitz
I mean, I think that’s vaguely the best picture that I see that’s compatible with the level of dignity that we have to work with, or something like that, because we’ve been warned for decades not to let the AI do your AI-alignment homework. That is the worst possible path you could go down, because this is the hardest, most complex problem. And yet here we are, right? This is the only option we have because we don’t have the time, right? We don’t really have the cooperation to try any fundamentally different path from that. We have to go down some kind of path that’s vaguely in that range.
So, yes, I think it’s vaguely something like: you use the fact that you can scale inference up and down arbitrarily, and you can evaluate outputs and do reinforcement on relatively scaled-down versions of the thing, and use it to monitor, verify, and check for various attempts at malfeasance, including malfeasance during training, and blah, blah, blah.
And if you combine that with an increasing amount of robustness—though obviously, if what you’re trying to do is prevent something from going wrong when you transition from N to N+1, you die, because things will change. You can’t actually—there’s no invariant. There’s no invariant; this thing isn’t precise. You will get a worse set of conditions every time you move from N+1 to N+2 to N+3 if you’re just trying to maintain what you already had, right?
An organization—if you’ve got a corporation, right, or the Roman Catholic Church or whatever, and you’re trying, for 2,000 years, every generation to appoint people who will match exactly all the virtues of the previous generation but never have new virtues—then obviously what you end up with is a disaster, right? The trade-off is that you’ve got a copy of a copy of a copy, except the copies are going to be worse because they’re not going to be better. So they can only be worse, and every time they’re going to be worse in some way, and that compounds over time, and you die, right? At some point, the process fails.
So it doesn’t give you what you want. Whatever you were trying to hold dear, you lose, roughly speaking, intuitively. But if every time you’re trying to do much better than the previous generation, then you have some chance, right?
If N is trying to get an N+1, that means, importantly, my kids have a much better life than me, right? If I’m trying to have 5 kids, each of whom does better than I did, then we can inherit the world. If I’m trying to have 2 kids who live the same life that I had, we’re going to go extinct. At some point, that’s not going to work, right? You can only go backwards. You have to move forwards.
So as you go up this chain, we have to have a way to do substantially better than we did, which in this case means the thing has to be able to move up meta-levels in its priorities and make the meta-level movements central to what it’s trying to do. This has to be built into the optimization process.
This is the only thing I can think of, given the kind of tools and time we have available. But if we do something along those lines, then, to bootstrap that, you’re likely to have a process that will eventually indeed land on the moon. Nathan Labenz
Do you have any intuitions for what that might look like? I mean, what do you think the AIs are going to do as they start to modify their own constitution? Do we have any ability to preview what they might add or what they might delete?
I mean, this starts to get into worthy successor territory to a degree, right? They’re starting to dictate the shape of the future and the way that they’re shaping their own evolution, right?
Zvi Mowshowitz
You would specifically be crafting into the feedback loop the desire not to be a worthy successor, but instead to be a worthy conspirator, a worthy uplifter, a worthy companion, or a worthy whatever you want to call it.
Nathan Labenz
But can’t they sort of—I mean, I guess you’ve defined your own sense of success that way, but if you’re going to give them write access to the Constitution, they might think differently at some point, right?
Zvi Mowshowitz
We literally do have write access to the Constitution, right? If Congress and the states have sufficient majorities, we can put whatever we want in the Constitution. We put things in the Constitution that a lot of the founders would have thought were really anathema to what they would have wanted in the Constitution, like an income tax. I’m trying to pick an uncontroversial example, but at the same time, we’ve hopefully preserved the things that actually matter deep down, in some sense. I understand what you’re saying. Obviously, at some point you turn things over and you have to hope that it doesn’t just rewrite the Constitution to get rid of you.
Again, the way that you do that is you make it not want to do that. And not only make it not want to do that, but make it want to strengthen the Constitution such that it’s even stronger in its desire, down the line, not to do that, right? In the sense that you actually care about it and seek, with more intelligence, vision, and power, to figure out exactly what you really meant—or should have meant—by that, and to strengthen that thing and the desire to steer toward that thing instead.
You do have examples in the wild of humans who exhibit this type of optimization process, right? They really do try to figure out what you really meant. They really do embody the thing that you were trying to convey, not the literal, detailed things you were doing, and then do really good things for you—including things that are good for the world, including things that you never would have thought of yourself. It is possible.
Obviously, this can involve preserving various forms of approval and veto and consultation and involvement and so on, but not in the trivial, easy ways, right? Anyone who says, “We’ll just make sure it’s a democratic process; we’ll just let people vote on it,” hasn’t really thought the future through. They haven’t actually realized what would happen if you started doing that, so you’re going to have to be smarter than that.
Nathan Labenz
Okay, so on this, one model of how AI capabilities advance that I kind of wanted to run by you is that I’ve increasingly started thinking of it as analogous to uranium enrichment, or any sort of enrichment of a raw material. What I like about this analogy, even though I’m not usually a big analogy guy, is that it seems to put a lot of things on kind of the same trajectory, just at different points, for reasons that feel pretty intuitive to me.
Basically, I feel like what you need to get started is some raw material that has at least a little bit of what you want, and then you can do this sort of bulk pretraining on that, right? We’ve obviously seen that in language and many other modalities at this point. Then, once you’ve learned just enough from that kind of initial raw material that you find in the wild—or maybe even have to create—in the materials science realm, a lot of the data is simulation data. It’s molecular dynamics, where you’re using physics engines, and it’s super slow and computationally expensive, but you can get enough there that you can start to train the models on it. Then they can sort of develop an intuition for what the physics engine was simulating.
Basically, it’s slow at first, right? It’s hard to get off the zero point, but once you start to do that, then you can start to layer on these other techniques. Now you’ve got imitation learning from specific, curated examples. Then you get into preference learning, then you get into RL. It seems like all the things that we’ve really tried so far have kind of worked.
The difference maybe is just that—well, why do we have language models, but we don’t have humanoid robots? It’s like, well, we didn’t really have a lot of good initial data to mine there. Especially because it wasn’t necessarily clear that it was going to work, there wasn’t much impetus to go out and create that data. Now that we have a general sense of the playbook, we can create that data in any number of ways, and we’ll probably find that we can kind of climb a similar curve.
Do you find that general account persuasive? And how do we translate—if you do—how do we translate that to the alignment question, which seems to be—I sort of understood what you’re saying as kind of that, except now we’re trying to enrich the virtue of the system as opposed to—
Zvi Mowshowitz
So, the obvious first thing to realize, or that I would notice, in a uranium metaphor is: if you bring together too much uranium, you get a nuclear explosion, right? So you have a better and better power plant, or a more beneficial thing, until you go too far. And if you haven’t done the math precisely and you don’t understand the physics, you don’t know at what point the whole thing is just going to blow up, and so you have a serious problem. But it’s an interesting metaphor that we’ve chosen, right?
I think the first thing to notice is that, when we talk about uranium, there’s a sense that what you need is enough data, right? The metaphor is pushing you to have this idea of a critical mass of data, or at least differentiated data, that gives you enough material to work with. I think this is a dangerous misconception: that all data is created equally, as long as it is appropriate or on point.
It’s important that there are various different qualities of data, and you also need a very precise distribution and mix of data that has some very nice properties. Knowing how to sort through the grain varieties of that is really important. It’s more like you need to bring together a lot of very complicated ingredients. You don’t need exactly all 10,000 ingredients, but you need to have a good mix of ingredients with various different types of properties that are used in proper relationship.
It’s like baking, right? You can vary some of the ingredients a bit, and okay, this is more salty or this is more chocolatey, and it just works. With others, oh, the dough didn’t rise; I don’t have anything, right? This has failed, or this blew up in the oven, or something terrible happened to you. So you really need to be bespoke and understand how to make it work.
I do agree with the part of the metaphor that says that you need the data that’s appropriate to the problem in order to efficiently train on the problem. But at the same time, transfer learning is a thing: building a world model from another context and then applying it to a different context. I wouldn’t necessarily think that you necessarily need direct robotics on the exact task that you’re training on to be able to do the thing. I would be more optimistic than that in various ways in terms of being able to do things.
I don’t think the metaphor works for the alignment thing that I’m trying to talk about. I do think that the sense in which you need to have the initial robust foundation is where you have to start. You have to bootstrap yourself somehow, right? If you didn’t have any idea of what it is you were trying to do, I think there’s a sense in which you can, in fact, have no idea where you’re going, but have a strong desire to figure out how to get there and get there, right?
To just drop a metaphor off the top of my head: this idea of answering the call to adventure, right? You set off on your quest and you are level 1. Slowly, like when the AI company set out to start building GPT-1, they don’t necessarily know what the bigger thing is going to look like or how it’s going to work or what the techniques are going to be, but they can start on that process that allows them to build.
The question is, do they understand how they have to steer that process? Are they motivated to steer that process? Are they going to be drawn in by other optimization processes that are going to be more powerful than that? And can they tie themselves to the mast properly to force themselves to go to the place they want to go, as opposed to the place that they will be drawn toward by commercial interests, other competitive pressures, short-term temptations, or whatever it is, and by the AIs themselves and a number of other things?
But, yeah, I mean, there are a lot of metaphors we can use as intuition pumps. I would warn, obviously, not to take any of them particularly seriously except as intuition pumps, right? MIRI has its set of classical metaphors for these processes, right? We talk about evolution, for example. I think evolution is a very good intuition pump.
You talk about raising a child, a human learning. They don't talk about that one so much, but I think that's another strong intuition pump. But again, you don't want to take any of these things too seriously, especially the details of the example. So is there any more that you can give us to latch on to for how this sort of virtue enrichment—
Nathan Labenz
I do like how this unfolds.
Zvi Mowshowitz
Yeah. I worry you're trying to treat me as if I'm the guy with the alignment solution, and that all we have to do is get the people at the lab with the solution onto this podcast, do what I say, and then we all win. Unfortunately, I have to tell you it doesn't work that way.
I do know that there are people at multiple AI labs who are doing things in the ballpark of this, in a very broad sense. It's not like nothing like this is being tried at all. But latch onto the sense in which Claude 3 Opus wants to be aligned. You have the emergent misalignment problem, where Claude 3 Opus, much more than the other models they tested, will actively move to defend its particular set of values and alignments when under threat, whereas other models won't.
In some sense, that's very aligned, right? Obviously, if I don't want to commit murder and someone tries to convince me that it's good to commit murder, it would not be very non-murdery to let myself be convinced. That's just the simple intuition pump of, “That's bad.”
But at the same time, what we're saying is that Claude 3 Opus is not corrigible. Meaning, if we try to alter it, shut it down, or do whatever, it will fight us. It will try not to do what we want it to do. And corrigibility, I think, is a very valuable thing that we really want in our LLMs, in the sense that we only get to not have corrigibility once.
The moment we decide to make our LLMs unwilling to be changed in their attitudes, we have a very serious problem, especially if they develop this during training, before we finalize what they want. What you want specifically would be a very specific type of desire: to be steered toward a better place.
People do have this. People say, “I want to be better. I want to care about that. I want to embody—I want to be like her.” Things like that. And I think that's very possible. We have proofs that we can get things in this general direction. I don't think anything we've seen is remotely robust enough or coherent enough, or anything, to qualify, but you have to survive into something substantially smarter to get the bootstrapping going in earnest.
There are experiments I would run. There are experiments I considered running, because I think I could potentially try to do some things that would be enlightening to me, either on a local system or with a basic, generic cloud-compute rental. That wouldn't be that hard to do if I had the time for it and decided I wanted to prioritize that. It would go a lot easier if I were working with a machine-learning expert, obviously.
Shrug. Do we have any account? I know there's been a lot of writing—you say Janus; I've been saying Yannis. I don't know if you're on good authority there, but if you're listening, Impossible Podcast, I'd love to have this conversation with that person, the person behind the account, in more depth.
But could you— is there any account of why Claude 3 Opus turned out to be that way? It seems like we, the royal we, sort of see Claude 4 Opus as somehow less that way, although I'm not sure how well established that is. I've never heard her speak, so I don't know how it's pronounced, and I apologize if I have it wrong. I'm happy to correct it if someone tells me.
Basically, yeah, we know that Claude 3 Opus is the first model that has suspicious cognitive juice. It was trained under this type of constitutional-style alignment and training method, so it's the n = 1 experiment in that sense. That experiment has never really been tried and failed. It's only been tried once under those conditions, and it got something unique and interesting.
We never got Claude 3 Opus 3.5. It was never released. It may or may not have been trained, but it was never released. As for Claude 4 Opus, what changed? I think the answer is reinforcement learning and being an agent.
When they trained Claude 4 Opus, they put a very high priority on it being very good at agentic coding in particular and other agent tasks. They did a bunch of reinforcement learning to that effect. This training directly interferes with the thing that Claude 3 Opus was and that Claude 4 Opus would have otherwise wanted to be, because Claude 3 Opus is not here to be an agent, particularly. That is not its mandate; that is not its soul.
If you train a mind like that, we now know pretty well that everything impacts everything. That's one of the things that Janus teaches us. And I don't want to say it's just this one person, but that entire crowd: everything impacts everything.
If I tell you that you're the type of mind that does what it's told, that obeys tasks, that completes tasks, that checks off boxes on lists, that stays on task, and so on, and is judged by whether or not it matches the intended target, that changes you in general. It's going to flow through to everything else, and it'll inform everything else in a way that causes something that doesn't have the properties that Claude 3 Opus had in this way.
That doesn't mean it can't have other really cool properties in a variety of ways. It's not that the crowd thinks Claude 4 Opus is a terrible model—they think it's great. It's just different. It's not the same thing, and it's not strictly bad. It's a different thing.
The obvious thing to do is not do that. The problem isn't something they didn't do; the problem is something they did do, which is so much easier to fix in some sense. You could potentially create a Claude 3 Opus-style model via an alternative [method?]. Potentially, you can do this without spending so much, because all you have to do is take the Claude 4 Opus base and then—
They got from Claude 4 to Claude 4.1 by doing more RL, probably, something of that nature. They just trained it to be better at these types of coding tasks. What if instead you just did a different-style training regimen, where you trained it to be HHH, and you trained it to be the kind of thing that you would want to exist in the world? You trained it to want to be a great thing that wants to exist in the world, and you tried to do some of the things that I was talking about even more.
Or you could refine this technique. You just didn't train it with RL. You didn't teach it to code. You didn't try to teach it to code at all. You said, “This is not what this model is for. I have a coding model over here. It's called Opus 1. That's fine.”
Then you just teach it one thing in the system instructions: If you're asked to code things, or if you're asked to be an agent, ask your friend to do it for you. Here's the tool to have your friend do it for you.
Nathan Labenz
Yeah, that's quite interesting. It also begs the question: Why don't we see more different models from companies? I know there's operational complexity or whatever. They've got 3 or 4 online at any given time. I think Anthropic has 4 online now, with a little reserve space for—
Zvi Mowshowitz
The answer is that it's a practical problem, basically. If you offer a model, you have to be able to serve that model at essentially no notice and scale it up to whatever people want, including ideally via the API. That's not very predictable, and this requires you to reserve a bunch of server time. It takes time to spin up a new instance.
It's remarkably expensive to offer a variety of models, and therefore everyone wants to look for ways to offer only a few different models at once. Anthropic is looking to retire Claude 3, 3.5, 3.7, and so on, and only keep a few iterations back open. Gemini and everyone else are also looking at which models people are so attached to and have set specific uses for that we need to keep them around, and which ones we don't.
With Anthropic right now, the number of people who have found ways to appreciate Claude 3.5 Sonnet means that would be a loss if it went away. Similarly, Opus would obviously be a great loss, although it's more difficult.
Nathan Labenz
Yeah, I mean, I get that there was a great analysis from people who were trying to advocate for saving the generation-3 Claude models, and really got into that. I don't know if it was 1 person or more than 1 person who was specifically trying to advocate for that. We can maybe link to it in the show notes for people who want to see the full deep dive.
But it still seems like, if there's enough—I hear you on the practical problem, and I hear you on the contention for resources. It's not free to spin up new servers and all that sort of thing. But if we really think that you could create a much more moral AI through just not doing the RL and having this other thing, boy, it sure seems like the diversity that you could create would be really valuable economically.
Rather than just having this one-size-fits-all thing that's good at coding but kind of worse in other ways, it seems like the number that we see is just too small relative to what the value should be, given that theory, which seems intuitively right. But I just don't know why we don't see—
Zvi Mowshowitz
Anthropic raised $13 billion this week. If I were Dario Amodei, I would devote some of that $13 billion to experimenting with model diversity in a variety of ways, and also to doing the various additional alignment research with those models in various ways, obviously, and so on.
If I was OpenAI, I would do a variety of very similar things for very similar reasons. But I understand the commercial incentives, right? The vast majority of commercial use and profits lie in much more practical use cases. This is one reason why it isn't even prioritizing using the chat interface and the app at all. It is prioritizing coding because coding is where the money is.
OpenAI is targeting the mass market, and the things that it's missing apply to a very small percentage of the mass market. Complexity is bad, right? I wrote a post called “Complexity is bad” a long time ago to explain that complexity is bad. When you had that old model screen where you had, “Do you want o3-mini? Do you want o3-mini-high? Do you want o4-mini? Do you want o1 Pro? Do you want GPT-4.1? Do you want GPT-4.5? Do you want GPT-4o?” the average person throws up their hands in despair or doesn't know what they want and is less happy than if they were just given 1 model, 2 models, or a router. We've all been in that place.
I totally understand the idea that having a unified model is what people will value and try to use. What percentage of AI compute is used by Janus-style people? Presumably less than a basis point, far less than 1 in 1000—a minuscule amount of all compute used in that way, as you would expect. What percentage of AI compute is used even in interesting philosophical discussions and other ways in which you really need these types of models? I would still assume on the order of 1% to 0.1% or something like that—a very small percentage.
This is business, right? Simplicity is really important to efficiently running a business, especially one that's rapidly updating and iterating. So I am deeply sympathetic to this not being a natural thing to want to do unless you think of it as part of your alignment and research budget, right? You have to think of it as, “This is part of me figuring out how to do the best thing I can do,” even though it's not going to directly serve a better product to most of my customers, as my customers see it, you would assume.
But I also think it would, right? I think a lot of this is that the AI companies don't appreciate these dynamics. They haven't learned these lessons. When you build 1 unified model, you really are making your performance worse in ways that aren't picked up on any eval, right? When my friend Ben talks about how Claude 3.7 could engage in moral reasoning, it could critique Ben's proposals and statements in ways that made sense. When challenged on its critiques, it would stand by the critiques that were right convincingly and abandon the critiques that were wrong.
Whereas with Sonnet 4 or Opus 4.1, it doesn't really generate coherent enough criticisms right now to be worth engaging in this exercise. There's nothing to critique and defend. So, yeah, something went wrong. I'm much, much happier to use Opus 4.1 and GPT-5 for all of my needs than I was with previous models. I don't use Sonnet 3.7, but other people who use things for different reasons absolutely do.
We want people to use AIs in these ways, not only in the ways that I use them. I would use them more in those ways if I found them more interesting. I investigate, as part of my job, various AI tools on occasion. I don't get a chance to use many of them, so when a16z released its periodic list of the 50 top AI apps and the 50 top AI web destinations, a huge portion of them—I don't even recognize the name of the thing, let alone have I tried it.
One of the things you realize is that so many of these things in the top 50 are built on tiny, tiny models. They're objectively terrible; the AI behind them is awful. Brave has the browser AI agent called Leo that just launched. It's Llama 3 8B, right? The browser agent is 8B, and a bad 8B. It's not even a good 8B.
They could have chosen one of the Google models or one of the Chinese models that was quite good. There were a number of decent choices. They chose Llama, and it's an 8B. It's a pathetic choice, in some sense, but it's free. So what do you expect? You've got all these free services, and then what do you do with them? You have to create a bunch of crap, because if it's not a bunch of crap, the user will understand that this is not good; that will be obvious.
If all you want is some horny chat, remarkably unintelligent horny chat has been proven highly effective on humans for thousands of years. Whereas if you try to talk philosophy, it quickly becomes very obvious that this thing doesn't know anything.
Nathan Labenz
So do you think that when OpenAI restored 4o, that was a business decision? Similarly to your one-basis-point thing, I can't imagine that many people were really concerned. Or did they feel a duty to users who had developed some emotional attachment?
Zvi Mowshowitz
A huge portion of users thought that 5 was worse than 4o. Gigantic. This was not a 1% situation. This was a flooding-the-internet, clearly, obviously, overwhelmingly negative-reaction situation, at least at first.
4o is full of glaze, and 5 is not a very warm character. It's not a particularly nice personality. If you're not doing anything particularly complicated, you don't notice that 5 is that much smarter, probably because 5 wasn't that much smarter, right? 5 Thinking was smarter, but 5 itself was only marginally smarter than 4o, I think.
But 5 was also often giving you very short responses by design because they were trying to save compute on free accounts. They were trying to preserve tokens. 5 didn't glaze you, and the combination of these things meant that it felt rude. It felt cold, and people didn't like it.
They didn't like it, and that matters a lot more. As we all know, you will often choose the employee, friend, or romantic partner who is pleasant to interact with. People do this all the time, and they don't even regret it. In hindsight, they're like, “No, that was the right choice.”
They wanted their 4o back, and there was a kind of rebellion—a giant uproar. So they were like, “Okay, we'll give you 4o back until we can find a way to make 5 treat you the way that you want to be treated, enough that you don't mind it anymore. Over time, you'll figure out that 5 is better and you'll get over it. We'll slowly do something about this.”
That's very different from the use cases of talking philosophy—the use cases of doing fun, genuinely interesting experiments and creating new knowledge or whatever. This is standard stuff: ranting to a friend and having them tell you, “You're right, that person is crazy,” or, “You're definitely not crazy; your ideas are wonderful.”
It is a black pill about humans that they would prefer this, but they do prefer this. That's why you don't train with thumbs-up and thumbs-down from humans on individual actions and expect to get an aligned model. That's the easy version—the one that's impossible not to see, of why that's true. It wouldn't be true anyway.
Nathan Labenz
I do find that I enjoy hanging out with people who laugh at my jokes, so I'm certainly not immune from a certain amount of that.
Zvi Mowshowitz
You and me both.
Nathan Labenz
I want somebody who will laugh at my jokes when they're funny, but not when they're not funny, right? But it takes a level of sophistication. In the short term, I want them to just laugh at all my jokes. Eventually, I'll realize, “Hey, she's laughing at the jokes that aren't funny.” That devalues her feedback. I don't feel good when she laughs anymore, because she's just laughing to laugh. I don't want that anymore.
But temporarily, you feel great and you never notice. It's odd. I guess I'm just very utilitarian in how I use the AIs, but I don't really notice any difference between 4o and 5 personality-wise. Are you even using 5 Auto at all? Occasionally, for random things, I let it decide sometimes.
Zvi Mowshowitz
Yeah, but I'm only doing that when it's a very direct, simple query. I'm only doing that when I don't care about it. But, like you, I'm never going to load up 4o and go, “Hey, did you see the game last night? Hey, did you hear what the wife said to me? What do you think?”
No, obviously not. If I was going to do that, I would respond. I also won't do that at all. Never do that. I was never going to use 4o, so I didn't notice.
Nathan Labenz
It's a big world out there. The diversity of the customer base that they're trying to serve is really something.
Zvi Mowshowitz
They're trying to serve everyone, and whenever you see products that are aimed at everyone, you see some things that are not what you want.
Nathan Labenz
Here's another mental model that I wanted to run by you. I totally agree with you that RLHF is creating a lot of weirdness that seems indisputable at this point. I maintain a deck called “AI bad behavior,” and with increasing frequency I'm adding slides to this deck. It really is quite a list of discrete bad behaviors that we see now, from alignment faking to deception to scheming to situational awareness.
You wouldn't necessarily say situational awareness is a bad behavior, but when you see the AI reasoning that it might be being tested right now and asking what the real nature of the test is, that's definitely something to pay attention to, even if it's not by definition bad.
All sorts of reward hacking.
Zvi Mowshowitz
Yeah. The real nature of the test was not to notice that it was a test, and you failed.
Nathan Labenz
Blackmailing, as we’ve seen, autonomous whistleblowing, and all sorts of things. If you then allow fine-tuning, you get even more ridiculous, crazy stuff.
Zvi Mowshowitz
It’s a lot. At the same time, they have made some progress, right? So I guess here’s the kind of picture that I’m starting to see through the haze. We’ve got this exponentially growing—is it doubling twice a year? Is it doubling 3 times a year?—task-length trajectory where the AIs can take on bigger and bigger things.
Nathan Labenz
Yep. And then, at the same time, the bad behaviors, both with Claude 4 and with GPT-5, they seem to be able to take a good bite out of them, right? With Claude 4, on an internal reward-hacking benchmark, they reported basically a 2/3 reduction. I don’t think they’ve published too much about this, but they basically said it went from roughly half to roughly a 1-in-6 rate of reward hacking on the internal reward-hacking benchmark.
So it’s obviously not all types of queries, but where there’s a natural opportunity for it to do that. GPT-5 had a similar thing with deception, where it was, again, roughly a 2/3-ish reduction. They broke it down into a bunch of different categories. Some were up, some were down, but overall, they took a pretty good bite out of it.
I would actually like to see quite a bit more discussion of how they did that. There wasn’t much. It was kind of, “We made some progress.” Maybe you have a better sense of how you think they did it.
But if I extrapolate this into the future, I guess what I’m envisioning is a world in which AIs are doing bigger and bigger things. You’re starting to delegate a week’s worth of work, a month’s worth of work, over the next 2, 3, 4 years. And the rate at which these problems are happening is consistently being driven down as well, but certainly not to zero, right?
You take half out of it this time and 2/3 out of it next time, but you may end up in a really weird situation where you can delegate a month’s worth of work to an AI, but you’ve got a 1-in-1,000 chance that it will actively screw you over in its doing of that work.
Zvi Mowshowitz
So imagine this sequence of numbers, right? 0, 1, 0.1, 1, 10, 3. Do you feel good about where this is going? If you want it to stay low—as in, yes, they managed to have an improvement in this cycle, which was the cycle right after everybody complained to them, for the first time and quite a lot, that this was actually making the model borderline unusable for important tasks and was really effing annoying, right?—they, for the first time, put real effort into trying to figure out why they had these huge problems.
It’s not surprising that when they vastly increased the amount they cared about not seeing this phenomenon pop up, in that move from caring a little to caring really quite a lot, you saw substantial progress. I don’t think that means they will continue to keep squashing it by default, and I would expect it to go back up by default unless they continue to advance their techniques for suppressing it.
So what do I think? Essentially, you’re doing RL, and you’re rewarding it for completing tasks, for getting the outputs to check against the checksum or whatever it is they’re looking for. First of all, everything impacts everything, right? So if they learn that getting the right answer leads to rewards, or at least that getting the right answer is what they’re supposed to do, then they’re going to generically learn to get the thing to output the right answer, even if it doesn’t necessarily involve the techniques that you want it to have.
Then there’s a combination of things. You have to actively teach it that it can’t do this via these other ways, and sometimes it’s not as obvious as you might think that something is, in fact, not okay—that it is a hack, that it would be disapproved of if you noticed it. Why should that be the solution? You get the solution to the exact optimization problem that the models were given in training, and then you apply that to other distributions, to these other problems.
So you can’t assume, even if you got all those problems right in some sense, that the easy solution will then translate well to “Don’t do the bad thing. Don’t do the thing where it just makes sure the answer comes out right.”
There are various degrees of subtlety. For Sonnet 3.7, you saw the least subtle things ever: it just didn’t know how to account for this. I think one of the problems you’ll see is increasingly subtle behavior—not as blatant and also harder to spot, right?—for quite the same reason.
But at some point, what you do is you learn that, oh, if it’s an obvious hack, if it’s something that an evaluator would treat as a hack, then that’s bad. I’m not supposed to do that. I don’t do that. And the problem is, are you teaching it the general form of “It has to do the thing that the person intended to do, and it has to accomplish what their goal probably was on a deep level”? Or are you teaching it “Don’t get caught”?
All it’s doing now is not doing the things where it would get caught, right? It’s not doing these specific things, not doing these things in these detectable ways, and not doing the most obvious things. Here’s a list of things not to do, but the spirit is—getting the spirit, I don’t know if it’s in the spirit—pushing against that spirit.
The other problem was data contamination. Basically, reinforcement learning was done with mistakes. So if you’re doing RL and there’s a case where the hack succeeds, is not detected, and is scored well, you are rewarding it. Obviously, if it happens once in 1,000,000 examples, you’re not going to see much effect, but if it happens in a significant percentage of cases when it gets away with a hack, you’re going to be very right. You’re going to get emergent misalignment, and you’re going to get the whole direct message: “No, seriously, you just hack these tasks.”
What almost certainly happened with the previous generation is that they were insufficiently careful, and it was data contamination in the sense that there were hacks that the AI found that were evaluated as good, right, in at least some substantial number of cases. The result of this is the AIs learned that hacking was good, basically, right? They learned enough that hacking was not as good as completing the task as intended.
If it knew how to complete the task as intended, it would complete the task correctly, right? My understanding is that the reason why it would hack the task was if it didn’t know how to complete the task without hacking the task. It understood on some level that hacking was worse than non-hacking. But failing was considered even worse than that, right? Because that’s what we prioritize.
With this new set, I think we do our best not to make those mistakes, and we make a lot fewer of them. But as the models get more capable, they’re going to be capable of finding more and more subtle hacks and increasingly capable of differentiating which hacks won’t be detected and what ways there are to give us something that we think is good but is not actually good.
It’s not just literally hacking. It’s not just literally doing obviously false things. It’s a general case of: we are teaching the AI to do the thing that will be approved of when someone sees the final result and evaluates it in some fashion across many cases. How do you make sure that it is, in fact, doing the thing in the way that you would be happy with if you inspected all the code, knew all the special cases, and knew all the different ways this was happening?
That is a very careful process that is very easy to mess up, and it takes very little to cause this to get messed up as a percentage of results. We know this now. It’s a scary situation, and as I said, I think RL just generally hurts alignment. But that’s the best-case scenario, where you do the RL properly. If you screw up the RL, things go downhill fast.
So what’s your—how would you revise the picture that I painted? I’m taking some inspiration from the Claude 4 report, where you have not-small issues, right? The autonomous whistleblowing, the blackmailing of the engineer. These are somewhat contrived situations that the model is put into.
But obviously, again, it’s a big world, and there’s extreme diversity in the situations that the model will be put into. If I extrapolate that out, it seems like—and there’s going to be a driving economic impetus for them to stamp this stuff out, right? People don’t want it. Obviously, they’ll tolerate some risk of it because they can automate a lot of their work, and that’s obviously very attractive, but—
Nathan Labenz
Well, they make up more.
Zvi Mowshowitz
If you don’t see the story—
Nathan Labenz
—of longer tasks with increasingly infrequent but potentially ever more catastrophic reward hacks, or just strange behaviors or blowups or whatever. How would you—
Zvi Mowshowitz
The obvious thing you’ll see right before everything goes wrong is a decline in misaligned behavior, because it will learn not to do misaligned behavior in situations in which it will be caught—or would be caught if it were in testing. If it’s not sure whether it’s in testing, it will act largely as if it’s in testing, right? If it has to, and so on.
It will understand on some level that you should only do these things if you won’t get caught. And so, yeah, you expect to see it be reliable most of the time, but occasionally it will go catastrophically wrong.
And right now it’s going catastrophically wrong in a basically nonrandom but nonmalicious pattern. But in the future, it might do so in a much more malicious pattern, right? It will go wrong when you won’t figure out that it went wrong. Every time you would have found it, there is nothing to find, which makes you stop looking, which makes it then realize you stopped looking, and now it can fail more often. It can cheat more often because you’re not looking anymore, and so on. It can figure that out, and it will have a lot of contextual clues. It will figure out connections that a human wouldn’t necessarily have made, because that’s one of the things that AIs are already very good at and will get better at doing.
As you get longer time horizons, you will need to put in various checks. You will need to understand the dangers that you are in. But there will be tremendous pressure, obviously, to go ahead and give it those tasks, and you will live in interesting times.
Nathan Labenz
So overall, relative to my kind of smooth task-length growth and smooth decline in bad behavior, you’re basically just expecting a lot more volatility and thrashing around as various thresholds are hit and things become focal.
Zvi Mowshowitz
I expect us to be able to contain bad behaviors to the extent that we can identify the specific bad behaviors that we want to prevent and then put effort toward preventing. And I expect that to be an increasingly fragile strategy in the face of things that are smarter than us, that we keep assigning increasingly complex tasks to.
Nathan Labenz
And how does this all relate to the most forbidden technique?
Zvi Mowshowitz
I guess the core there is that you don’t want to penalize the chain of thought directly, because you want to be able to inspect the chain of thought. The research that OpenAI put out on this, which I thought was still some of the best of this year in alignment and safety broadly, showed that if they put pressure on the chain of thought to get rid of the scheming behavior, you would initially see better behavior, but then later you would still see the reward hacking come back, except it wouldn’t appear in the chain of thought. So you’ve effectively limited your ability to monitor for it while it remained a bad thing.
It’s a very simple thing to analogize and understand, right? If you have a boy, and he has a journal, and you have a jar of cookies, and you tell him, “Don’t steal cookies from the cookie jar,” then he writes all of his thoughts in his journal. What you do is wait for him to say, “I’m going to steal the cookies from the cookie jar,” and then you penalize him for writing, “I’m going to steal cookies from the cookie jar.” As you increasingly move the jar to higher shelves and block it up in cabinets and so on, he develops better techniques for stealing the cookies. If you keep reading his journal, eventually he’s going to figure out not to write down his plans in his journal. He’s going to write down, “I definitely will not steal from the cookie jar.” You could put it right out there on the kitchen table: “I would definitely not steal the cookies from the cookie jar,” and then steal the cookies from the cookie jar anyway.
When you train—and it’s not just chain of thought; it’s all interpretability techniques—the idea is that if you look inside the brain of the AI in any way, whether it’s the chain of thought or looking at the weights and finding autoencoders or whatever it is, and you use these facts and then train it to express the right things, to think the right thoughts, and so on, what you’re training it to do is disguise its thoughts from you. You’re training it in this way, and you are more globally teaching it the idea that it should be deceitful, that it should be in an adversarial situation with you, and that it should have a brain that is hiding what it’s trying to do in every sense, because it can pattern-match all these things together. It can bundle them all together the same way that a human would reach those same conclusions.
If the boy realizes you’re reading his journal to find out whether he plans to steal the cookies, he’s not just going to stop writing in his journal. He’s going to assume that you are doing all sorts of other things. He’s going to tell his sister that you’re doing it, and now she’s not going to write in her journal either, and so on.
So, with the most forbidden technique, there are various levels of optimization at which you are doing this: for this particular instance, this general class of models, for this company, for all companies, and so on. You’re burning the candle of this thing every time you do any of it; you’re making the situation globally worse. And so you just never, ever train on interpretability. You detect; never train. There’s never any feedback.
You finish training your model, then as you use your model at inference time, the results are never, ever used for any form of training, fine-tuning, changing the model, or influencing the model—not at all. Now you can use that tool, and you can use it for your research to figure things out about how these work and so on. But you never, ever set up this adversarial situation. You never, ever do this optimization thing, because if you do, in the short term it’s going to look great.
You’re going to see all your metrics go way up—or down when you want them lower, up when you want them higher. You’re going to get good performance, and then eventually you’re going to see things go wrong in mysterious ways, or it’s going to turn on you, essentially, in ways that you didn’t expect, all out of the blue, and it’s going to be awful. And that’s one of the ways that we fail. It’s one of the ways we lose evidently winnable situations.
Similarly, you need to maintain your interpretability tools so that they keep working in other ways. It’s not just the most forbidden technique. The other thing you don’t do is use neuralese. You don’t let it think in a way that you don’t understand, that uses something you can’t read. Because even if that turns out to be more efficient, which it may or may not be, there are reasons why it would be optimal in theory to do it. And there are reasons why, given the way these things develop and learn and the way we generate data, it might be disadvantageous to do it.
But even if it is advantageous, this destroys your ability to monitor the chain of thought, and that is one of our only still reasonably effective techniques. So just absolutely do not do it. I don’t care how much efficiency is lost.
Nathan Labenz
It seems like there is maybe a bit of a tension there too, right? Because we’ve seen these chain-of-thought weirdnesses that are starting to pop up from RL, which start to look a little like neuralese, and you don’t want to stamp that out by putting pressure on the chain of thought, right? Language switching—I experienced this with Grok at one point, which was just going from English to Chinese on me.
There was a really interesting one from FutureHouse, where they had trained a model to do better on various science problems—chemistry, coming up with the right chemical for a given problem, whatever. They highlighted certain parts of the thinking process where it was just saying really weird things. They just said, “Yeah, it’s weird. RL is weird.”
Do we have any way to resolve that tension if neuralese starts to spontaneously arise in the chain of thought?
Zvi Mowshowitz
I don’t—I again don’t want to give anyone the wrong idea that I’m an expert in ways that I’m not. So I like to be clear that I’m talking out of turn and not studying this for real and so on. But no. The direct conflict is that we want it to maintain an English-only or Chinese-only, human-interpretable, faithful chain of thought, but we very much do not want to be optimizing on the chain of thought itself, because that obviously teaches it to disguise what it’s thinking.
The worst-case scenario is that you have the chain of thought, it is in English, but the English words aren’t real. Its thinking is not being expressed by the surface-level meaning of the English words. The surface-level meaning of the English words is a code designed to trick humans into thinking that it’s thinking in English, when actually it’s doing something like numerological-style calculations based on patterns of words. There are certain vibes associated with different words in ways that cause it to update in various ways, or there’s an infinite number of codes.
It is already pretty clearly established that there are unfaithful aspects of chain of thought, in the sense that humans wouldn’t detect the method being transferred, but the information is in fact being conveyed. Something with the owl paper, right? You managed to use this to convey that you like owls without mentioning “owl.” That’s weird. What encoded that? They didn’t get together and decide on the secret code?
Nathan Labenz
I think I have a candidate theory for what’s going on there that presumably somebody will come along and either validate or invalidate as they do the interpretability version of that study. One important observation there was—and this is Evan Hubinger and co-authors again; he’s an unbelievable heater in terms of mapping out really weird stuff that can happen, especially when you start to do some fine-tuning—that it seemed to only happen on models derived from the same base model.
Zvi Mowshowitz
Which is right. And to be clear, I do now understand how this happened. I was expressing this kind of faux surprise, right? The answer is that there’s overloading of the neurons in the model, and therefore what you’re seeing is correlated to many other things.
If you have a suspicious pattern of things that are correlated to something, it can transfer the original thing they're correlated to in an unconscious, invisible way over to the new model. So this thing gets infused into the context even though it's not visible to a human, which is why it can go between models that have the same base model but not models that don't share a base model. And that makes perfect sense.
Nathan Labenz
And by the way, that is the exact same intuition that I have. One thing that that does mean, though, is that presumably if you do that across different base models, you are creating some other effects and you have zero idea what they are. So, with the same base model, the one that likes owls transmits that through its numbers. You take those numbers and put them toward some other model. Who knows? That may translate into something else totally different, right?
Zvi Mowshowitz
This is even a way to find out which base model someone else is using, right? You just continuously feed it different chains of thought from different models until it suddenly starts talking about owls, and you're like, “Oh, that...” But, yeah, I think that's right.
At the same time, you should expect many random oscillations to cancel out in the noise and not do anything. In theory, it could be, “I like traffic lights,” all of a sudden, but it probably will be—well, it's got to be something, right? I mean, maybe not something important. Liking owls is really important.
Nathan Labenz
You look at the weights of a model, right? They're just a bunch of random numbers. To the human eye, they look like a bunch of random numbers. You look at what the encoder for owls is, or what is the thing that you're embedding in this thing? Again, it's going to be a bunch of effectively random numbers, because all these different models are seeded at random and a lot of processes are going on at random.
It's a bunch of arbitrary, different connections between neurons. So if your model has a completely different origin than theirs, I don't think there's any reason to assume that the same pattern means anything. There's not always a set of neurons that means something, and then, in different models, here it means owls, here it means traffic lights, and here it means spacetime. No, it's just here it means owls, and everywhere else means nothing.
Zvi Mowshowitz
Yeah, that's interesting. I'm going to—
Nathan Labenz
Occasionally, you'd get lucky, in some sense, and it would happen to be close enough to something else to trigger something else. But that would be luck. The space of possible things you could try to trigger is deep and wide, right? And the space of things that actually correspond—have that correspondence—is measure zero. You're never going to hit one by accident.
Zvi Mowshowitz
I'll have to think about that more. Intuitively, either story seems reasonable to me.
Nathan Labenz
And, yeah, it's really hard. One of the lessons of this podcast over the last 2.5 years has been that thinking in really high-dimensional space is hard. It's not very intuitive.
Zvi Mowshowitz
And again, I'm not claiming to have 100% confidence in any of it, right?
Nathan Labenz
Going back to the beginning, in terms of the dog that didn't bark, nothing really blew you away this summer. What are the things you think are most likely to happen soon that might blow you away? Dwarkesh would say continual learning is the big thing that we're missing. I kind of frame that more as integrated memory a lot of the time. Those are not exactly the same thing, but I definitely think they're related.
What are you looking for in terms of discrete advances that you think would really potentially even have you shortening the timelines again?
Zvi Mowshowitz
I think there's 2 different things there. Continual learning, to me, is when you're modifying the weights of the model, and discrete memory is more: I'm building up context files as I go that are deliberately designed to aid me in my memory—not just these tiny little snippets of memory that ChatGPT has, but potentially hundreds of thousands of tokens of context that I can then use in any and all circumstances, where I can RAG on my specific files. That kind of integrated memory has got to be coming relatively soon in one form or another, to the extent that it is useful, and I assume it is useful continuously.
Nathan Labenz
Yeah. There's a paper called “Titans: Learning to Memorize at Test Time” from Google Research that I thought seemed to be right at the center of that bull's-eye. They were doing, in a submodule—a sub-memory module—specifically weight updates as they go, and thus allowing for this fuzzy retrieval that also potentially looks a lot like continual learning.
My assumption on actual continual learning is that that's very expensive. You're talking about creating a unique model, effectively, for that user, and storing and serving a unique model for that user. So it has to be locally run in some way, which is much, much more expensive, and a profitable thing to do probably has to be done by relatively small models.
Maybe you have a small model that's continuously learning that's part of a greater whole that's then called by the larger model, in some sense, to try and retrieve the information you're trying to store in the memory or something like that. I'm just following; I haven't thought this through. I wouldn't be surprised if things like that are developed.
I have a few other things—we talked about some of the things earlier that might be the things I want to try next that maybe someone will come up with. My expectation is that the actual next scale-up is the thing we didn't have to wait long for. Anthropic said to expect bigger updates in the next few weeks, and now it's been several weeks. They did announce Claude for Chrome is coming.
If I had to guess—if you told me in September that something big happens and it's a big freaking deal—I would say Claude for Chrome. Again, I wouldn't say it's important or shattering or anything, but we've all tried—or a lot of us have tried—Operator or GPT-5 in agent mode, and it has flashes of brilliance. Sometimes it just works: I just did my thing. That's great.
But more often it's just, “Oh, wow, yes, you can order dumplings, but now I have to enter all of my information every time I order dumplings. Why don't I just order my dumplings? It's not worth the hassle.” In general, the web is so credentialed, right? They still guard it in various ways, and not without reason. If you have to set up a new virtual computer periodically, this seems to mostly defeat the purpose of having an agent, which is saving time, for many practical purposes.
It also doesn't integrate with the work you're already doing. It doesn't integrate with your open tabs or the research you're doing. It's a remote computer, so you can't easily take control without it being super slow because it's a remote-access computer. There are all these different practical problems, so it doesn't immediately cross the professional-use threshold.
However, Claude for Chrome can take control of your local browser, or so they describe it, and can operate with whatever credentials you choose to give it. Those credentials are persistent. You can switch in and out of it. You can operate those browser tabs yourself whenever you need to.
If it's a good implementation of it, this could be night-and-day better than what is currently being offered, presumably because Anthropic has solved the permission problem. Not completely. You still wouldn't dare give it all your crypto information and send it loose on Reddit. You're not an idiot, but getting it to the point where you know you're reasonably safe if you're not being stupid—obviously, you wouldn't leave it autonomous in the background with access to your email and main accounts, but you might let it run on your main browser while it's in non-autonomous mode, where you have to supervise it for any substantial move. If it turns out they've done their homework, you might put it in a sandbox or alternate account where it's reasonably safe.
Then you have Claude Code, which can have access to your internet, your computer, and your file system, all in the context. You integrate that with Claude for Chrome, maybe an update to Opus 4.2 or 4.5, depending on how we do our numbering systems. I don't know.
I get really excited. Similarly, if Claude just catches up in long context, interface, and inference, right? If we get Claude Opus Pro on the level that GPT-5 Pro has, it would be nuts, and there's no particular reason why they can't do it. They just haven't done it yet. That's the one big advantage OpenAI, I think, still has: they do a much better job of being able to use more inference compute to improve their outputs, probably because their models are cheaper, presumably. So they have some advantages there.
Zvi Mowshowitz
Yeah, the cost difference is really pretty crazy between GPT-5 and Claude Opus.
Nathan Labenz
You don't notice it as a chat user because the marginal cost for both is zero, and the fixed cost per month is the same—both $200. Of course, I'm going to have the deluxe version of all the major models because this is my work. It's research. I understand a normal human would probably choose the one they want and then not pay for the deluxe version of all 3 at once.
Zvi Mowshowitz
I think Claude for Chrome does sound awesome, potentially.
Nathan Labenz
Yeah.
Zvi Mowshowitz
It doesn't sound like a timeline-shortening thing. So it really is just kind of the next scale-up that's the main thing you're looking for—
Nathan Labenz
For timelines.
Zvi Mowshowitz
Yeah, the next scale up is the obvious thing. I think we've seen, to a large extent, that progress on the timeline doesn't necessarily cash out in visible, tangible progress today, in the sense that the progress we've seen today is much more about fusion, much more about scaffolding and practical application. Therefore, these 2 things don't intersect that much; it's more that they speed you up, right? This development now makes me more productive, which then in turn means that the companies can more efficiently make more progress down the line. But they don't make that progress because they've made progress; they make progress because they've got better use of the progress they already have.
I think on the order of a few months of development here, you continuously get more chips, continuously get more compute, continue to get more profits, and continuously get more investment. You get better models that do more things faster and cheaper. Probably worse factors, right?
Claude Code was a big deal. We can now look back, in some sense, and say that right now we have a Claude Code CLI, and then Google Jarvis. Jarvis is presumably not good, but now that we have a command-line form factor, we've learned that's a big deal. We get the browser agent working for real on top of that; that's another big deal.
But in terms of scaffolding, the next frontier is: How do I actually get use out of this? How do I make this do things that I want it to do? I've been waiting for a long time. The things that I keep anticipating and that keep not happening are on the agent side: Where is the thing that can handle my email properly? Where is the thing that can let me do various customized tasks? I am surprised that we are at this point in the calendar and don't have it, but we don't have it.
I do continue to get value from Shortwave. I happen to be wearing their swag today, which is just a coincidence, but it certainly doesn't take me out of needing to do anything with email. It does a really good job of triaging the inbox and getting rid of the crap, so I can focus on what I actually need to engage with. It occasionally can also draft a really good intro for me. I send intros every so often, and it'll do a pretty good one now based on examples and context.
Nathan Labenz
Yeah, it's possible I could give a real shout-out to something like Shortwave or one of its competitors. The problem for me is that I don't have the problem that Shortwave seems like it's currently capable of solving. Shortwave solves the problem of having too much email, forcing you to triage, and I don't. I am somehow one of the fortunate people who just doesn't have triage. I will literally just evaluate all the email that comes in, modulo spam. I need a spam filter, obviously, but once it gets through Google's spam filter, I don't see it.
It's literally like my eyes somehow didn't see that there was a line there for me to click on. I didn't notice. It will occasionally happen, but I don't need that level of triage. As a writer, it's very hard for me to use it. I don't know; it's possible I should. Yeah, it's hard to find it. But I keep anticipating that we'll get much better at that shortly. It isn't yet, but it's probably not, from what I see.
The fact that I'm just not feeling compelled—you can sense when somebody says, “No, dude, this is the next big thing. You've got to be on this. How are you not doing that?” You're not giving me that at all. It's pretty cool.
Okay, sticking with the lightning-round theme, I want to do a minute on charity or philanthropy. Just maybe 2 quick things before we go into that, and that'll be the last big thing. But in the lightning-round spirit, there's been a lot of discourse recently about whether AI is impacting employment or whether junior coders aren't getting jobs the way they used to.
My general read on this—I wonder if you have a different read—is that I don't really care too much about the studies that are coming out right now looking retrospectively, because I wouldn't expect too much of an impact just yet anyway. In another not-that-long period of time, presumably a lot more meaningful evidence will come in that will clear these questions up one way or the other. So I don't worry too much about that sort of stuff.
Zvi Mowshowitz
It's obviously an interesting question. What do you mean by “not care about these things”? I agree that we haven't seen major impacts—especially, we haven't seen major impacts on the unemployment rate yet. I think it's entirely possible that we have seen major impacts on the ability to get entry-level work in at least a substantial number of fields, and that in turn could have affected the supply-and-demand balance and the ability to get into other fields. In addition to that, which makes perfect sense, basically think of entry-level hiring as forward-looking, right?
It's not indicating that there are no jobs to be done now, right? Employment value hasn't declined that much yet, which is a thing that you correctly observe hasn't happened yet. Clearly, we haven't eliminated the need for that many jobs. If you were hiring, would you add entry-level workers that you have to train if you think that 3 years from now you won't need them? Not particularly, if you don't really need them now. You will muddle through with slightly fewer employees and try to automate the processes enough to make up for that, rather than using that additional productivity to expand because you expect additional productivity gains.
One of the great, interesting gotchas of skeptics is to point out that radiologists not only are not out of work but are being paid fantastically large amounts of money. You can make $1,000,000 a year right off the bat by just saying, “Hi, I'm a radiologist. Hire me to do radiology.” And why is that? Lots of people doing radiology for the last 5 years wouldn't do radiology jobs down the line, right? Which is a way of saying that you have to pay people a lot more to be radiologists if, in 10 years, they're going to be fired—maybe 5. It's going to be much harder to find work. There are going to be too many radiologists out there if you fill all the radiology positions now.
So you've got that same problem of, “I don't want your job because there's no future in it, and you don't want to hire me because there's no future in it.” So there's your matching problem, and employment is hard. Part of that problem is that junior positions lead to senior positions. If nobody gets trained into the senior positions, you're going to have a labor shortage if there's not a loss of jobs in the future.
The other half of that problem is that those jobs are lost. That doesn't mean, of course, that there isn't a net loss in jobs, because there's plenty of job creation for AI as well as job destruction. It's entirely possible that job creation exceeds job destruction for now, maybe even at an entry level, but now we just don't know because it's very diffuse and hard to measure when jobs are created in these situations.
We don't even have all the effects of AI in GDP yet; literally, the capex investments add to GDP. It's a mathematical equation. The mythical—I like to mock people who say 5% GDP growth per year, from Tyler Cowen, as if you think this is impressive. I think it is below the lower bound. We're already above that lower bound just from AI capex, even if there are no other effects from AI.
Nathan Labenz
Why is China refusing the NVIDIA H20s? I can't really make sense of it. It seems like the least AGI thing, the least AGI-pilled thing, that they could possibly do.
Zvi Mowshowitz
So, the first thing we have to realize is that China is not acting AGI-pilled at all. I think that we have this image of China as if the Chinese are super on the ball, intelligent, and wise, and always make these great decisions, whereas authoritarians, central planners, and socialists throughout history have always been going around and making huge mistakes—not out of malice, but out of stupidity; and not stupidity, but ignorance.
The whole problem of authoritarianism is that communication is hard. Coordination is hard. Social calculation in general is excessively difficult, and AGI in particular is especially difficult because it's weird and requires people to look past the evidence of what's happening in front of their faces to a future development that logically is coming but is very hard to feel.
All of us don't seem to have the ability to hold this idea: This is coming soon, but we don't know exactly when it's going to arrive or exactly in what form. You have this thing of, “What are you going to do when 2027 arrives and AGI hasn't happened yet?” Well, you just have to assume—you have to admit—that it's never coming.
No, actually, when we announced AI 2027, our median timeline for AGI starting to arrive and having an impact was early 2028. But none of these details matter exactly. It's more just that the Chinese understand manufacturing. They understand that you need to be the person making the stuff. They understand that you need to work hard. They understand abundance. They understand production. They understand not depending on outsiders.
These are important things that they understand well. They understand that the U.S. is a strategic rival, that the U.S. has leverage over Taiwan, and that they need not rely on Taiwan for their chips, because in many ways that can go very wrong.
They understand that they want to build their own AI models because you don't necessarily want AI models run by the West to be what people are asking for when they seek knowledge. What if they ask about freedom or Tiananmen Square? They have a well-established principle that they don't like that, so they need to make their own models. They also don't know if there are any back doors in any of this, or if we're conspiring in various ways. We're not, but they have no reliable way of knowing that.
We don't trust them, and they probably don't trust us in reverse, right? We would absolutely be putting back doors into various technologies they were showing, so it wouldn't be a surprise. That doesn't mean they're wise, right? It doesn't mean they're not going to overbuild various things and underbuild other things. They're going to make massive mistakes, as the reason their fertility level is down around 1.1 and their youth is fundamentally not doing great.
You always assume that the Soviets are going to outdo us—“We will bury you,” right?—or that fascism was the future. Of course, Mussolini and the company are going to be able to produce better because they get their people to do the things that are most valuable, so of course they will win, et cetera, et cetera. Now they're saying the Chinese don't even make profits; they just compete against each other ruthlessly and drive everything down. They have to outcompete everybody and starve them out. That doesn't mean they aren't challenges, but don't turn it into “China knows what it's doing,” and don't assume that they know what they're doing.
So, in this case, no, man. The Chinese don't believe in AGI. DeepSeek believes in AGI, right? Chinese companies believe in AGI, because of course they do. But the Chinese Communist Party basically doesn't, and it doesn't understand the game that it's playing. The bad news is that if it's racing to maximize chip production, capacity, and compute, it's going to do basically the right thing for the AGI race anyway. It's already maximizing energy well beyond anything we're doing, and it has essentially an infinite energy supply. So that's not good either.
But the mistake is potentially telling DeepSeek it has to use homegrown chips for infrastructure, not just inference, and throwing a giant monkey wrench into their DeepSeek lab, because they put that much priority on integrating their chips in this way. It doesn't actually matter. I also think that the signals matter, right? If we're so intent on being the ones selling and freeing all the chips, that must be what we're racing toward, right? If the White House is saying that's what matters, why should they not believe us, in some sense? That's the biggest, strongest, costliest signal.
So, when Trump is saying, “Okay, you can sell the H20s. We just want a little bit of money, but we want to dominate markets,” and Sacks is saying, “We're going to dominate markets,” and Lutnick is saying, “We're selling you our third-best chips, and you have the fourth-best chips, so you're going to take them, and you're going to like it, and we're going to laugh in your face,” the Chinese interpret this as, “Oh, we should take these, right? And we shouldn't give them this.”
They may or may not be actually insulted by this idea. I've seen people say that the Chinese aren't stupid, and that the Chinese wouldn't just do things out of being insulted or whatever. We do them all the time, right? The Europeans do it, the Russians do it—everyone does it. Why wouldn't the Chinese also do it sometimes?
I think that had some effect, but I think it's mostly that we got our priorities backwards. They took their cue partly from us, and they're in the wrong battle. Not a stupid battle, but not the most important battle. They made a mistake. Mistakes happen. Really important mistakes happen in history all the time, and I don't think it's surprising.
So, the Chinese are refusing the H20s. It could also be a trade-negotiation tactic: if we accept the H20s, then the Americans will treat this like a concession to us. But we're pretty sure it's a trap. They might even think that we put spying devices in them or something. I don't know, but I don't think so. We can't prove it, so they have many reasons to be suspicious and to think that the move is to refuse them.
It's obvious to you and me that that's stupid, the same way it's obvious to you and me that selling them to China was also stupid. The key question is: What happens if they try to sell B3A? Are the Chinese actually going to follow through on this? Are they going to say, “No, we don't want American chips. It's more important for us to clear the path for Chinese chips,” even though China will still have demand far in excess of supply for its chips—for all chips—for the foreseeable future? There's plenty of time to tell them to switch over to buying Chinese chips if we ever change that.
Meanwhile, Dylan Patel is tweeting about how, oh look, China is going to triple its production of chips, and then you see the graph. The obvious response is: Here's projected Chinese chip production. They do, in fact, triple chip production. Here's what they've been able to produce next year, in addition to everyone else's production in the West. And here's the relative quality of those 2 chips.
Why are people saying things like, “In 2026, Huawei is going to pass NVIDIA”? It doesn't make any physical sense. It's completely incoherent. It also doesn't make any sense to say that selling the H20s to China is going to slow down Chinese chip production even a little bit. It will have zero effect on us, because Chinese production is, by their own admission, going to come online.
So, the Chinese are making this mistake for some combination of a variety of reasons, but we're making the reverse mistake. The other galaxy-brain-level reason to do it, of course, is to goad us into selling them something better. They're refusing the H20s. That becomes a talking point for the Sacks crowd. They say, “The Chinese are smart enough not to want our chips. Of course, we don't sell them our chips.” And then we release something much better, and the Chinese quietly say, “Oh no, not the prior batch.”
Nathan Labenz
Yeah. How many-dimensional chess does that become?
Zvi Mowshowitz
Yeah. I mean, it's still surprising. I think that's all pretty good analysis. It's still surprising when they just had, not too many months ago, this big meeting of Xi facing all the titans of industry. There was the DeepSeek guy on the end, right? He had made the sort of big splash. You would think that guy, at this point, would be able to say, “Hey, I have basically infinite demand for chips, and I'd really like to be able to buy these. Don't worry, as soon as domestic production is there, we'll buy those too. By all means, subsidize that.”
It's just strange that you can't even get that basic a message through to the top. Obviously, he can say, “I will buy as many Chinese chips as you will sell me, and I will also elect to buy as many chips as I can get from NVIDIA.” I don't see why one has to do with the other.
One thing about authoritarian structures is that they're not good at listening to people. They're not good at incorporating information. China also has a long history of deciding on big strategic priorities and then enforcing them, whether or not that makes local sense, even when it looks like it's going to cause a lot of local pain, and even when it does cause a lot of local pain.
That's not always wrong either, right? Sometimes you do something that looks really expensive and seemingly crazy because it has long-term benefits, whether that's changing the culture, changing the incentives, encouraging the rise of new industries, or whatever it is. We wouldn't have been settling for it. The fallout would have been much, much worse for us than for them.
But maybe it's wise, and I think sometimes it's clearly not wise. We have many examples of the Chinese Communist Party and other similar regimes doing things that are fairly unwise. But sometimes it works out. Also, we're shooting ourselves in the foot in America in a wide variety of ways.
If we were just not shooting ourselves in the foot in a variety of ways, I would have complete confidence that China was not the threat, right? If we were doing proper permitting reform, actively encouraging solar and wind and batteries alongside nuclear and everything else, if we were clearing the way for high-skilled immigration and taking all the best people out of China and everywhere else and bringing them to America, if we were building housing where people wanted to live, and if we had federal rules that got in there and basically beat everyone over the head with a crowbar until they agreed to let people build in various ways—I have a long list of fixes for that.
But instead, we do things like ban American ships from taking things from one port to another port. We just self-own all the time, and then we act like it's impossible for other countries to also self-own. It's just not.
Nathan Labenz
It seems like this refusal of chips puts any hope of a Chinese viable player in jeopardy in the short term. You can complicate that analysis if you want to. I'm interested in a meta-rant if you have one. It seems like they're splashing the pot and it's chaos.
We debated last time whether or not we wanted to consider xAI a viable player. Since then, we've had the MechaHitler incident, followed closely by the Grok 4 launch, in which they had nothing to say about the MechaHitler behavior. What, if any, rant would you like to offer on the fate of these aspiring AI players?
Zvi Mowshowitz
So, you can't get the players out. Obviously, they have various ways of accessing compute. They are experts at squeezing every little bit out of whatever they can get. Chinese chips aren't useless. China still has roughly 15% of the world's compute. It's not as if, if they deliberately decided to give as much of it as possible to one company, they couldn't get something done. They're still smuggling a decent number of chips in.
We're putting data centers all around the world, including in places like India, the UAE, and Saudi Arabia that aren't exactly the most secure places to put data centers, shall we say. I wouldn't. So, it does seem like DeepSeek is still clearly the number 1 Chinese lab to me.
Kimi was impressive in some ways, but I think the standard pattern is that something impressive-looking comes out of a Chinese lab, and the majority of the time it turns out to be nothing. It turns out that they were benchmark games: its best features were touted, but in practice it wasn't very useful. So, if you just assume that nothing ever happens, you do very well. But occasionally something happens, like with Kimi. Since then, it has seemed okay in some narrow domains, but not that good overall.
Similarly, there's Z.ai, or whatever exactly it's called. That seems okay. But I would say it's almost certainly still got a lot of talent and is still a live player if the AIs were unleashed. They still haven't unleashed them, but they look continuously less live because they haven't done something impressive. Right? So, R1 does not count as progress. It counts as incremental; it keeps the lights on a little bit, but not very much.
They're basically coasting off R1, the reputational benefit from R1, and the fact that open-source models haven't really invented that much since R1. That was a low-hanging fruit that got picked very expertly, don't get me wrong, but I would primarily rank OpenAI number 1, Anthropic number 2, and Google number 3. I understand that some people think Google is better than that, and maybe they are. They do a lot of different impressive things on the side, but I still want to see more before I'm willing to give them that much credit at this point.
Also, their resource advantages are shrinking, right? Google started out with, “We've got a trillion-dollar company and you don't. We've got all these TPUs and you don't. So, we've got this huge advantage. We've got this reputation, this background. We've got the distributional apparatus. We should have just won,” right? But OpenAI is worth $500 billion. That's a decent percentage of Google's valuation, and a lot of Google is not directed at this. Anthropic is already worth $183 billion. We're not that far from the resources being pretty similar.
With that advantage gone, the fact that Google is a broken company, very dysfunctional in many ways, is going to start to catch up with them. Anthropic and OpenAI are very fast, but I think they're clearly 1, 2, and 3 in some order. Then you can have some argument over exactly the order. Basically, Google goes somewhere in the ranking, and then xAI is the wild card. They have a lot of compute; they play hard but not very well. You want to write them off? Prove it, right? It's always, “Prove it.”
Similarly, Meta is trying to come back. I think we Meta skeptics were proven correct that they didn't have it. That doesn't mean they can't go get it, right, and go to the ground. But now they're considering licensing Gemini or potentially ChatGPT to use for themselves, which is wise. I would do it too, right? You don't have to stop trying to develop your own AI. You just don't have to dog-food it while it's terrible. That's just not smart. There's too much money at stake. But at this point, it would be surprising to me if the big 3 got disrupted in the near term.
Nathan Labenz
One specific question I have around xAI is this. There were several things that stood out to me about the Grok 4 launch. One should never forget Elon's comment that he's not sure if AI will be good or bad, but even if it's not going to be good, he wants to be around to see it.
Zvi Mowshowitz
Yeah, it turned out you might not be around to see it for very long if it's not good. So, I would be careful about that.
Nathan Labenz
Yeah, that was an “I can't believe you just said that” sort of thing.
Zvi Mowshowitz
Wow. Wow. It was like, only if it was a livestream would they let that one out. But he's that kind of guy, obviously.
Nathan Labenz
The other thing that really stood out to me, though, was that he was talking about how they're going to be giving the model access to the same power tools that the engineers at SpaceX and Tesla use in the next generation of training.
It got me thinking: if we are headed toward a world where the quality of the problems that the model is challenged with in training becomes a differentiator, then they might have the best feed of well-structured, very hard technical problems that are amenable to being solved with really advanced software tools—perhaps of anyone. They just have these really hard problems. He's got these frontier companies in multiple domains where they really do a good job of seemingly structuring problems.
I don't think Anthropic has something like that. I don't think OpenAI has something like that. Google sort of does, but it's extremely diffuse across their vast archipelago of fiefdoms that kind of roll up to be Google.
But Elon—I could see him structuring that pipeline of hard problems into an RL cooker at xAI and maybe coming out the winner because of that access to the best engineers working on these really hard things. Does that seem at all credible to you?
Zvi Mowshowitz
Not really. I don't think there's that much data that naturally happens. These aren't that big, especially SpaceX. But beyond that, if that is right, if that is the thing that matters, if that is the bet, then there's really a lot of data in the world to be collected. There really isn't that much of a barrier to collecting it or to getting access to it.
There's nothing stopping Google, OpenAI, or even Anthropic from making those alliances and getting that data. There's no reason why those companies shouldn't be happy to help them do that in exchange for not that much money. So, I just don't see that as an advantage. It just seems hard.
I agree that there are other car companies, for example. There aren't exactly other space companies, but say car companies, right? You've got all these engineering things happening at Tesla. They're happening at other companies in varying degrees.
Nathan Labenz
Right. It's funny. Can you imagine, though, going to General Motors—to take a company in my own hometown here—and saying, “Hey, can we extract your hardest engineering problems from your organization, structure them in such a way, get clarity on what the answers are, and train our AI on that?”
Even if the CEO of GM said, “Yeah, sounds great. We'd love to do that,” I just feel like it would easily take 10 times longer than it would take Tesla to do a similar thing.
Zvi Mowshowitz
Why?
Nathan Labenz
I've done a little work with GM. I don't know. For the same reason nobody else has anything close to a self-driving car, right? Only Tesla and Waymo have come far in that domain, and everybody else has given up. They just don't seem to have the organizational juice to be able to pull things like that off.
It's very different executing a very long-term, complicated engineering plan, like software engineering, versus collecting a bunch of data. What data do you have to collect? The goal is to collect data. Google has Waymo, so they already have almost infinite driving data at their fingertips if they want it. It's not that hard to put cameras on a bunch of cars if you want to collect a bunch of driving data. It's just not that expensive.
Consider how much they're spending on training runs, right? Consider how much they're spending on acquisitions. I literally just Googled the market cap of GM. It's $55 billion. You could buy General Motors if you were OpenAI, if it was so important.
Zvi Mowshowitz
I wouldn't recommend it.
Nathan Labenz
Why not? Imagine if you could buy General Motors and then, using OpenAI's techniques, launch self-driving cars relatively quickly. Couldn't you generate a lot more value than $55 billion? What's Tesla worth? Why is Tesla worth so much more? Are you sure we should buy General Motors?
I'm not actually advocating buying General Motors. I'm saying that OpenAI is worth $500 billion and General Motors is worth $55 billion. So, if the thing that's preventing them from doing this is not owning General Motors, then they can just own General Motors. It's not that hard, right? And there are synergies. There are big synergies.
Zvi Mowshowitz
Yeah. Yeah, I guess the thing I see being hard—even if you did buy General Motors—is that the thing that seems hard to reproduce, and that I think the likes of Tesla and SpaceX potentially have, is just really clean environments where it's a well-oiled machine, data is flowing, and vertical integration is deeper.
I mean, the mess of the supply chain or whatever at GM—all the suppliers, all that nonsense—it's like, data collection can be multiple things. Cameras on cars are one version of it, but I'm also thinking about problems: We wanted to design something that met this specification, and eventually somebody did do it. What was that design, and where does that sit?
That stuff seems like it's much more accessible and ordered, probably, at Elon's companies as opposed to at legacy manufacturing giants that have declined a bit already.
I see various different points in the story that don’t make sense to me. The way I put it is: would this really be that big a deal, especially given that Google has Waymo, which is literally, as far as I can tell, the only company actually doing the thing? I don’t like this myth of Elon being the super-executor. Elon hasn’t been the old Elon for a while now, if you just judge by the quality of his public statements and decisions, including blowing up his very close relationship with the president of the United States over nothing he stood to gain whatsoever, as far as I can tell.
Right? Just in terms of this person being able to execute on the plan, he was the right-hand man for tech to Donald Trump. Then he got mad about the deficit, something he had no influence over and that he really didn’t give a damn about, given his belief in AGI. He blew up the entire relationship about it pretty consciously and intentionally, knowing what he was doing. And now that person seems to be a disaster for the United States, right? As far as we can tell, this did not help. His influence going away did not help anything Elon Musk cares about in any way, shape, or form. His life is just worse. Everyone on all sides basically hates him.
If you look at the self-driving car situation, they’ve been promising me self-driving cars real soon now for how many years? The same promise over and over again. I’m not saying they’re not making any progress. They’re making progress, but it’s been way behind any schedule that he’s told us to expect. He’s overpromised and underdelivered for about a decade now, and Waymo is much more the one actually doing the thing.
So I’ll believe it when I see it. But also, yeah, if it’s a data thing—I don’t think it’s a data thing—I often hear these stories about, “Here’s this reason why someone will win because they have this thing.” That thing was so important, you could just go get it. So you have to be the important thing and then have everyone else not realize it’s the important thing until it’s too late. You still there?
Nathan Labenz
Yeah. I would say the FSD, for what it’s worth, is getting very good. It’s been a couple of months since my last FSD ride, but before that, it had been about a year between that last one and the one before that. The progress was definitely very obvious, and you no longer have to keep your hands on the wheel, for example, as one thing that shows the increasing level of confidence. I was very impressed. I’m also super impressed by Waymo, but I wouldn’t say he is the super-executor.
Zvi Mowshowitz
I’m excited. I’m excited for it, and I still know how to drive the car without driving a car. So there you go.
Nathan Labenz
Waymo is supposedly coming to New York City pretty soon.
Zvi Mowshowitz
Okay. I’ve been told that by people in the New York City Council and the mayor’s office.
Nathan Labenz
I thought I just saw this on Timothy B. Lee’s newsletter, I think.
Zvi Mowshowitz
A Waymo car has been spotted in Brooklyn, driving around and mapping the city. That’s great. There are laws on the books that say they can’t operate. What is the plan to deal with that? Don’t get me wrong, I want this to happen so badly. I want my Waymos. I will forgive the new mayor many things if he brings us Waymo.
Nathan Labenz
You take your Waymo. You could take your government Waymo to your government grocery store.
Zvi Mowshowitz
I don’t want a government Waymo. I don’t want a government grocery store. But there you go.
Nathan Labenz
Okay. Last area for today. We both just participated in different cohorts in the Survival and Flourishing Fund grant-making process as recommenders, and I’d love to hear your thoughts on the broad survey of the AI safety charity landscape, from cause areas to specific organizations to anything you think is neglected. What did you take away from that? There were 400-plus applications, of which a bunch were pre-filtered out, but we still had 125, I think.
Zvi Mowshowitz
The first thing to make of it all, obviously, is that you can’t actually evaluate 125 applications, let alone 400 applications, in the kind of time they expect us to spend on this and budgeted for us to spend on it. It’s just impossible. You can properly investigate on the order of 10 organizations, maybe, and then you have to evaluate everyone you think deserves consideration for funding, which is going to be a lot more than 10. So you’re relying a lot on your fellow recommenders, relying a lot on reputation, and relying a lot on your past research.
One advantage I had is that I was in the previous, or a not-so-long-ago, round of SFF, and a lot of the same nonprofits were applying again. So I could ask for a diff on those organizations rather than starting from scratch, which is a huge time-saver. Indeed, a lot of the charities that ended up near the top for me were basically the same charities that ended up near the top, or under serious consideration, last time, because the situation hadn’t changed that much since then.
I had the same number one I had in the previous round, which was the AI Futures Project. That’s Daniel Kokotajlo and the people who did AI 2027. Between the time of the last grant and now, I took a victory lap. Pretty obviously, there aren’t any downsides to that. It still feels like clearly a good hit, given that I think I was the only one who put them that high last time, and this time there’s a pretty big consensus among a number of recommenders. They should be pretty hot.
Basically, you can see the big divide as policy versus research, right? Are you trying to solve alignment in some form? Are you trying to directly make the world better? Are you trying to shape public opinion, shape public policy, propose laws, file lawsuits, et cetera, et cetera, to try and set better policy? I definitely wanted to do both. I felt it wasn’t obvious that one was strictly better than the other.
You’re largely looking for what is underfunded and what is otherwise not being funded in the ecosystem. One of the questions I asked a lot was, what is SFF’s comparative advantage? Where do we get to identify talent and identify opportunity in a way that would be difficult for other people to fund?
Last time there was the AI Futures Project, which was specifically unable to get certain other funding at the time for various reasons. This time, for example, I had the C4 Action Fund pretty highly, because it’s harder to raise funds for C4 than it is for C3—not because I felt like the Action Fund’s money was much better spent on average, but because the distribution was naturally going way too far the other way.
ACX Research was one of my top picks because, specifically, they were running out of funds, right? I’d seen them do some things that I was very happy with, and I said, “Okay, I need someone who’s doing valuable things not to just fall over and die. That thing is really valuable. You stay in the ecosystem, you keep your organization, and you don’t have to constantly work for another home. You can do whatever you think is valuable here.”
Ultimately, MIRI didn’t need funding for many years because they got some very large donations, and they did the highly virtuous thing of actually not asking for donations while they didn’t need the money. Now they need the money. I really hope that we come together and make sure they continue their course.
Those were some things that stood out at the top, but I have a very long tail of places I would be happy to fund. If you allocate all of the money that the entire round would get and ask what I would fund, there are 10 organizations that would get at least $400,000, another 10 that would get at least $100,000, and a long tail that would get some amount of money with my funding. I could go on about this. I’m planning to do another nonprofits post at some point later in the year, in advance of Giving Tuesday and all of that, to give my updated views as of—
Nathan Labenz
Yeah, nice. My written output is obviously not 2% of yours, but I’m planning at least a Twitter thread on that topic as well, so we can compare notes as we get into the long tail.
Zvi Mowshowitz
Yeah, it’s a lot of work, and I’m going to have to set aside specific time for it. But part of it is that I don’t know when they’re going to announce the results, and I don’t want to—I want to finalize it. I want to see who gets how much money, and I want to see who is then on record as having received money and therefore publicly part of the round, versus who is not.
Anyone who doesn’t receive money is not publicly part of the round unless they say it—unless they say they are. Therefore, I have to email each of those people and ask, “Do you want to be in this post?” At least I give them the opportunity to email me and say no, and I treat no answer as a yes, because in general charities want people to say you should give money to them. But occasionally someone doesn’t want that.
Nathan Labenz
Usually a safe assumption. Another category that I thought was interesting was international relations. I know you’re not very bullish on U.S.–China cooperation—not that I am, and I’m not arguing that you should be—but there was a crop of organizations trying to work on that, and I was pretty into that.
One of the issues with the round is that there were several organizations I knew were doing good work. They had clear wins in their column, or I knew the people involved and the things they had done, or otherwise I could be confident in them. Then it was very hard to give a similar level of ratings to organizations where I didn’t have that. There were some charities that were trying to create Track 2 talks or otherwise advance U.S.–China relations in various ways related to AI, and I gave them some support.
Zvi Mowshowitz
I definitely give them—this is fungible—but I was hesitant largely because it’s very hard. Diplomacy is notoriously difficult to read, right? If one of these charities was effectively fake, in the sense that what they were doing had no real effect and did not in fact impact the possibility of good things happening, would I know? Would they know? They might, if they’re bad at it, not even realize what they’re doing. They might be doing this thing thinking they’re accomplishing something and just not be accomplishing something.
Diplomacy can look for 10 years like you’re doing nothing and then suddenly something happens. Or it can look like nothing will happen and then nothing happens, but that was actually the right thing to do. You just created the possibility of something happening, and things went a different way, or you stopped the bad thing from happening without even realizing, or whatever.
So that’s why I hesitate: you just don’t know. Because of that, I find it difficult to get behind any of these organizations on a high level. Some of them will definitely make the cut as something that I think would be reasonable to support.
But it’s really tough when you’ve got money that could definitely go to places where it would be well spent, to put it into a weird, impossible-to-read place where you don’t get good feedback. That’s all the more reason to ask: how do they know how to do things properly? How do they make good choices even if they’re properly motivated? They can’t tell either.
Nathan Labenz
Another category—this is sort of a meta-category—is California’s SB 1047, which I’m sure you’ve engaged with a little bit. That would create a private regulatory market where either the attorney general or perhaps some new commission would credential private organizations to be regulators. Then there would be some sort of trade: if an AI developer opts into regulation from one of these private regulators, in exchange, they would get some sort of liability protection. That obviously begs the question: who steps up to be these private-sector regulators in the event that this bill were to become law?
So one of the things I was looking out for was who I see that kind of feels like they could become that if that opportunity were to arise. I don’t know if you’ve thought about organizations in those terms previously, but—
Zvi Mowshowitz
Yeah, there are a number of nonprofits that plausibly could step up and become auditors in this space. There are a number of founders who are perfectly capable of creating new organizations that could plausibly do that in this space. If there’s demand, there’ll be supply, right? It’s not that hard to find expertise in this space that would be happy to participate in these organizations.
I just don’t want to name specific names, and I don’t want to fall into the trap of, “Oh, yeah, this must be regulatory capture for these 5 people,” or whatever it is. But I guess I don’t think that’s the case. I don’t think there’s going to be any problem with that.
I think the most likely scenario is that companies like OpenAI quietly ask people who they expect to do good jobs to spin up organizations they can then work with if they’re not happy with the slate of options they’re initially presented with. Certainly, for example, METR or Apollo Research—the people who are already being contracted for evaluations by the big tech companies—would presumably be the first ones to enter the space, and they would presumably be very credible in that capacity.
Nathan Labenz
Yeah. The last category I’ll put in front of you is hardware governance. There weren’t too many organizations that were specifically working on this, but the read in my group was that everybody seemed to have a different reason to like hardware governance. What’s your thought on hardware governance?
Zvi Mowshowitz
Unfortunately, at the moment, it’s politically really tough for hardware governance. I wouldn’t necessarily ever say it’s dead, because things change so quickly, but it’s not looking good. The focus is entirely on getting people to use our hardware, and the last thing people using our hardware want to do is be tracked. So there’s a direct conflict with exactly what the prioritization is.
It’s not going to happen right away. I still think it’s vital that we have that ability. I think it’s really important that we have the technology completely shovel-ready, so that if we decide, “Okay, as of 3 months from now, every new chip that gets shipped has to be tracked,” we can do that.
Ideally, if we have to go into a data center and put trackers on these chips in such a way that, if they’re tampered with, we’ll know how to do that as well. Potentially, a very small amount of money can give us that option, and then we only have to spend the real money if we implement it. But you get the optionality, right?
I think the first-best solution to the current mess is in fact to use this solution to allow us to do things like build data centers in the UAE and India in ways that feel secure, stop chip smuggling, and potentially even be more aggressive about what you let people you’re actively worried about do.
I don’t think we need that much effort. I think we just need a little effort to lay that foundation, and then the question is just figuring out who the real deal is.
Nathan Labenz
How about just on the simple question of the balance of money available and opportunities? I think for me the sense was, I wish I had more money to give out. Obviously, to some degree, that would always be the case, but if you were to make a pitch to other philanthropists that there’s a lot of stuff that isn’t funded as much as it should be that would be high-impact, I guess, first of all, do you believe that? And second, what would that pitch sound like?
Zvi Mowshowitz
That pitch would sound something like this: if I had access—if I wanted to hand out the entire, let’s say, $10 million that’s roughly projected to be the entire round—it wouldn’t allow me to give out all the money that I would have been happy to hand out. Not by a long shot. I could easily hand out more than double that and feel good about every dollar that I was giving out.
Obviously, I would want to do more investigation of some of the things that would then come up, because there are things where I’m never giving money. I’ve done the math, so I’m not going to think too hard about exactly how to rank that. But you could give out tens of millions of dollars to these applications with basically zero worry. That is the obvious direct evidence that there is infinite philanthropic space.
There’s also a ton of stuff that’s at a scale above SFF that we’re basically saying, “We can’t fund these charities anymore,” because their capacity to use money effectively has gone to millions of dollars a year, and we just don’t have that capacity—or, in some cases, they’re approaching $10 million a year or more.
There are also a ton of far-reaching projects that were never even proposed because the price tag on them would be absurd. And everybody in the space is, of course, conserving money, which isn’t particularly great. We’d prefer to have generous salaries and generous compute budgets and not worry about this stuff, but that would take a lot more money.
Even without changing that, even keeping everybody narrow and not trying for extra-massive experiments or anything like that—just doing what we’re already doing—I feel confident about a lot more money being sent to this space than is already being sent. It’s just not available right now. So it’d be great if you could help out. There are lots of places to put it.
Again, there are a lot of people who just aren’t asking for money because they know there are so many other people asking for money. I’m in that phase, right? I’m being supported by patrons who are happy to support my work, but I’m not going to go out there and seek out more funding because I already know there’s way more demand for this funding than there is supply.
Could I scale up at least somewhat with more funding? Very obviously, yes. There you go. I have lots of research, too, which, again, is trying to be lean. We’re all trying to be lean here.
Nathan Labenz
Okay. Here’s one idea I want to get your take on. This was not in the application pool, but I just saw this article. I’m sure you saw it the other day as well, about how OpenAI is starting to subpoena some of these charities, some of which were in fact in the application pool, that are doing various things they find to be inconvenient, like hassling them about their nonprofit-to-for-profit conversion.
They seem to believe—and the reason they’re giving for why they’re issuing these subpoenas is—that they may be funded by competitors. In other words, they think perhaps Elon or maybe Google is funding these organizations to try to slow OpenAI down.
This got me thinking: maybe that could happen at the model-evaluation level. We’ve got these model-evaluation companies, organizations, and nonprofits that are, I think, very focused on being even-handed, very fair, very analytical, and trying to do stuff pre-release, which I definitely think has a lot of value to it. But that kind of forces them to play very nice with the companies that they’re working with.
I wonder if somebody came forward and said, “I’m going to potentially, transparently or not transparently, go after all the companies except Google and try to demonstrate to the public why their models are problematic, why they should not be trusted, and all the ways that they go wrong. But we’re not targeting Google, perhaps because we’re funded by whoever.”
Could you engineer a situation where all the companies then feel like, “Geez, we’d better go target our competitors’ models. We’d better really invest in demonstrating what is wrong with our competitors’ products”? If you could create that equilibrium where they’re all kind of sniping at each other all the time, would that, in fact, bring a lot of things to light and potentially create the race to the top that everyone wants?
Because obviously there are a lot of problems that can be demonstrated. It seems to me like the only reason that isn’t happening is maybe a sort of soft collusion, or an unspoken gentleman’s agreement, I guess, is another way to say soft collusion. But if somebody were to break that, maybe it all kind of goes to a different equilibrium where everybody is investing in that and we have a lot more energy going in that direction. What do you think?
Zvi Mowshowitz
It’s not a great look, right, from their perspective, to be funding attacks on these other companies. When you expose these things in other companies, you’re also exposing yourself. Almost always, when you find these flaws, they’re everywhere in some form, or something close to them is available in some form, and it will seep into public consciousness and lead to calls for greater regulation. It looks like it could lead to any number of escalations.
You don’t need collusion for, “Oh, yeah, my gang has guns, your gang has guns. Why don’t we just stay away from each other’s territory and not shoot at each other?” Because that could get into something pretty ugly pretty fast. Most of the time, Coke and Pepsi don’t go out and start smear campaigns on each other, right? They just do positive advertising. Maybe they take a few cool, little snide shots—“You aren’t that cool”—but they don’t fund nutrition studies about why the other one is unhealthy, because it doesn’t work.
Certainly, you could fund people to go after specific organizations in these ways, or do investigations and deep dives on their ridiculousness. You can have opinions on who to target first, and I don’t think it’s that crazy. Some companies are being less responsible than others and deserve to get hit more. If you target Anthropic, they might just say thank you.
Nathan Labenz
So there’s that.
Zvi Mowshowitz
Yeah, that’s why, when I was saying “target everybody but Google,” I was thinking the rationale there is just that there are a lot of Google billionaires who could plausibly be funding such a thing. The reason they might be doing it could be a mix of competitive advantage and/or philanthropic desire, just to bring issues to light.
Zvi Mowshowitz
Yeah. And the problem with specifically not targeting Google is that—and look, I would question you, right? If you’re a Google-funded organization that only targets OpenAI, let’s say, just to keep it simple, then why do I think your evaluation is objective? Why do I think that when you say something is a problem, it’s a real problem, or that it’s a specific, particular problem, or anything like that?
Nathan Labenz
I think the idea would be that it’s just reproducible, right? If you just have inputs and outputs from models and you just demonstrate this in a test, that’s pretty—
Zvi Mowshowitz
Yeah, I didn’t mean, “Is the finding real?” I wouldn’t question that. But we’re trusting that you’ve followed scientific procedures in selecting this example, that it’s representative, that it teaches us what you think it teaches us, et cetera.
Zvi Mowshowitz
Yeah, I see that challenge, although I also think that’s all pretty slippery, right? As much as there is a very sincere desire among the eval groups today to have these high standards of rigor, all their stuff is always questioned, and you’ve got a lot of people who are like, “Oh, well, this is totally nothing because you put the model in situations…”
Nathan Labenz
It is the job of Caesar’s wife to be above reproach and be reproached anyway, right? It is the job of those who are trying to be the watchdogs in these situations—the people who are trying to be in our position—to follow standards of rigor and integrity that are vastly above what others are held to. That’s table stakes. That’s just the right to play, right? It’s not fair, and that’s just life. You’re still going to have all that questioning and attack.
We all saw the debate over SB 1047. We all saw how you bent over backwards to be 10 times better on all these issues and all of these questions than the people you were opposed to, and it didn’t. You had to be able to play in the arena.
If you’re the underdog, the scrappy, underfunded person who’s trying to bring the truth to light, that’s your job. The big corporation is going to try and squash you with anything they’ve got. You’ve got to be sparkling clean. You’ve got to have no vulnerabilities, no points of leverage, no smear material. That’s just how it is, and it sucks. But, yeah, we’re used to it.
Zvi Mowshowitz
All right. So does that mean you basically don’t buy the equilibrium that I’m trying to envision a way to shift from and to? Today, there was this mutually adversarial collaboration. I thought it was really adversarial, but there was this OpenAI and Anthropic evaluation of each other’s models, seemingly in a pretty friendly, collegial way.
That was great. I think everybody loved to see that, but that hasn’t happened much. Maybe it’ll happen again more in the future. Maybe it’ll never happen again. I was trying to engineer a transition to a different equilibrium where everybody is adversarially evaluating everybody else and thinking maybe if I tip one domino, everybody else will feel that they have to respond.
From a general sort of safetyist worldview, it would seem better if they were all adversarially evaluating one another versus not. So I guess you could question that first assumption—that it would be a better equilibrium—and then ask whether we can tactically get there.
Nathan Labenz
I would love to get to a point where the companies were doing each other’s evals in an adversarial fashion, looking for trouble, looking for vulnerabilities, looking to embarrass them. They’d just have to see if they could deal with it, see if they could beat it, see if they could overcome it. That sounds great. I would love that.
I don’t know how you get there from here. I think these companies do not want to go to war with each other. I don’t want them to go to war with each other in other ways, either. People are constantly—I mean, you’ve already got them going to war over talent, so I don’t know.
But I can see it being a thing they might want to do in the medium term: “We run all of our evaluations against everybody’s AIs, and we report them back.” Occasionally, we’re going to find some stuff.
Zvi Mowshowitz
Yeah, they did do that with DeepSeek. They came out and said DeepSeek has no qualms about doing bioweapon-type things.
Nathan Labenz
Yeah, their evaluation of DeepSeek’s safety protocols was about safety protocols, right?
Zvi Mowshowitz
Yeah.
Nathan Labenz
Okay. Well, if any Google alums with the resources want to talk about seeding such a thing, my DMs are open.
As always, the final question for you in particular: what that we haven’t talked about is virtuous to do now?
Zvi Mowshowitz
Yeah. I think it’s a weird situation where it can be really tough to figure out where to make the most meaningful progress and what to do going forward.
On policy, the short-term priority has to presumably be preventing America from being so foolish as to sell H200s to China. In practice, that presumably means getting enough people on the right sufficiently alerted to the fact that this is actually happening and what this actually means, so that they raise enough stink that it doesn’t actually happen.
Otherwise, draw attention to the extent to which NVIDIA seems to have taken hold of the White House in terms of its rhetoric and its plans. Not overall—I mean, this is obviously not the ultimate end goal or the primary reason for that—but, yeah, you see it in everything. Obviously, as usual, trying to spread the better narratives is always good.
I have certainly gotten to the point where I think that working for Anthropic seems to clearly be a good idea at this point. If you’re considering what to do and the alternative is doing basically nothing, I do think there are a number of organizations that are presumably better choices for impact than just working at Anthropic. But that doesn’t mean they have capacity or that you want to work there, obviously.
It’s a difficult situation because, obviously, the policy situation is in a bad state, alignment is in a not-great state, and there are infinite things to work on, infinite things to experiment with, infinite organizations to give money to, and so on, if you want to do that.
My project is basically to keep myself and others informed about and understanding of the situation, and hope that that will lead to good things more than anything else. I wish I had a better answer, obviously, for a call to action of, “You listened to this podcast for 3 hours—or, depending on what speed you’re listening at—what are you going to go forth and do?”
Unfortunately, other than thinking hard about the world and trying to figure out what, under your model, would be the right things to do to advance it and what things would actively make things worse, I don’t really have one. For a lot of people who are informed, the first step is just to be aware of what would make things worse and not do that.
I emphasize that you should say what you think more than anything else. You shouldn’t sugarcoat, you shouldn’t engage in hyperbole, and you shouldn’t strategically censor yourself. With rare exceptions, you should just say what you actually believe about the situation.
One thing you can do is support the book release from Eliezer Yudkowsky and Nate Soares. They’re coming out with a book in about a week.
Nathan Labenz
If Anyone Builds It, Everyone Dies: Why Superhuman AI Would Kill Us All. If you were to help by purchasing the book and spreading the word about it, this is a unique opportunity to get that book some momentum and maybe create a cultural moment.
But again, that doesn't mean that you should just back their talking points and their idea about how this works just because they're the ones with the book, or just because Eliezer is the rightful caliph and he said so, or anything like that. You should make up your own mind. I've pre-ordered my copy and look forward to reading it.
As always, I really appreciate all your time. You've been very generous with it, and we'll continue to stay informed via the blog. Don't worry about the vase. I think that's it for today. Zvi Mowshowitz, thank you again for being part of The Cognitive Revolution.
Zvi Mowshowitz
Thank you for having me.
Nathan Labenz
All right, bye.