Nathan Labenz
Zvi Mowshowitz, welcome back to The Cognitive Revolution.
Zvi Mowshowitz
Thank you. Thank you.
Nathan Labenz
You’re fresh off your latest 10,000-word send. You are drunk on information like seldom seen before, and we’re here to create the audio version for people who would rather hear you talk it all through than read the usual post. Although, of course, there’s going to be a level of comprehensiveness in the post that we won’t be able to match, so this should not be taken as a substitute for the blog itself.
Zvi Mowshowitz
The posts are canon. The posts are what I said after I had a chance to think about it. This is what I’m saying off the cuff. This is what I’m actually thinking. So, enjoy.
Nathan Labenz
Cool. Well, let’s get into it. The big first question is: Is AGI here, and is RSI—aka recursive self-improvement—here?
Zvi Mowshowitz
No, and mostly no. I understand there are claims that o3 is potentially AGI. The more I understand it from the reports coming back and the more I use it, I think it’s great. It’s obviously not AGI. That’s not what’s happening here.
This isn’t even primarily an intelligence leap. This is a tool-use leap. o3 is a much, much more useful version of the thing we already had. It’s a lot easier to get what you want out of it, to get it to do the things you want in a reasonable time, in a way that fits with what you want and isn’t forcing it into specific boxes. That’s going to be super useful, but it’s not AGI.
Nathan Labenz
Tyler, I thought, did have an interesting frame when he said, “How much more intelligent did you expect AGI to be?” I guess I wonder what frames you could put on this, but how many years ago do you think you would have been given o3 and said, “Well, yeah, this has got to be AGI. What else could it be?”
Zvi Mowshowitz
I mean, if you ask the question instead, “How many years in the past would I have gone, ‘Holy shit,’ when I saw o3?” The answer is, I think, 2, right? Definitely 4, for sure. I would have been like, “Holy shit. Nuts. How the hell does this exist? How the hell was this possible? Especially this soon.”
But that’s very different from saying it’s AGI. Again, we’re talking about something that can do all the things that humans can do. We’re talking about the thing that can basically just plug and play anywhere you need it to, to do all the cognitive work. That’s what we commonly understand AGI to be, to some extent.
Nathan Labenz
I think Tyler is defending a broad definition. You’re using a silly definition of the term. The thing you’re asking is, “Is it smarter than me?” And from his perspective, the answer is yes. I think he’s using kind of the wrong definition of “smart.”
Zvi Mowshowitz
I actually saw his comment and decided I was going to deal with it last. I was going to go through everything else I had in my queue, and then, as the last thing that I wrote up, deal with Tyler Cowen’s claim, because I wanted to understand all the context before I evaluated the question.
By going through all those other claims, I feel like I understood what o3 was. Then I was able to look at Tyler’s claim and understand why he was claiming what he was claiming, which is that it’s a very good program for doing exactly the things that Tyler values highly and does all day. He’s consumed by how he thinks, right?
If you look at his examples, if you look at what it can do, it can go out there and get you tons and tons of detail, tons and tons of specific facts about any given thing, make connections between those facts, structure those facts, find the relevant things, and present that to you. He eats that all day. That’s the thing he does. He does it so much faster than I could do it, even if I wanted to.
Nathan Labenz
It’s an amazing skill. I am in awe, but that’s just not how my brain works. I produce lots of content in a completely different way, which is that I consider information as part of a logical way of understanding the whole picture. If it doesn’t fit the logical picture, it doesn’t really seem relevant to me, and then it won’t stick—it’ll bounce off me. Similarly, if it seems like it’s just part of a pattern, once I understand the pattern, I don’t need the details that make up the pattern anymore.
He gave 3 examples, I think, when he claimed this. He was talking about why this guy’s early paintings are much more valuable than the paintings from later in his career. This is part of a pretty common pattern for artists who basically didn’t constantly reinvent themselves, where early on they’re doing the thing that’s considered fresh, that’s considered unique, that’s harder to find, and that’s special in many ways.
We’ve seen one thing that I don’t think o3 did point out, which is that we’ve seen this extremization of the value of collectibles across the board in the last 20 years. The thing that’s in the best condition, the actual unique first thing, the very best painting the guy made—the very, very special thing—is now worth 10 times what the thing that’s almost as good would be worth if the first thing didn’t exist, even though it would be considered exactly the way the first thing is considered now.
The 9.8 comic is so much more valuable than the 9.6 comic, right? You find exactly the thing you want, and this confused me for a long time as a Magic: The Gathering player. Why is the thing that’s the same thing but looks slightly better so much more valuable? It plays the same. But no, that’s not what people care about.
Basically, you’ve got demand and supply. You’ve got what’s considered the best, the most unique, the highest-status, and the most representative of the thing. It’s pretty standard. So I don’t need to know any of that. I never heard of this artist. I have no idea who this artist is or what he does. I don’t have to—I can answer this question to myself anyway.
Similarly, with the author and why his prose is so amazing, whatever—it’s good prose, people like it. You can throw in specific details about this particular person, but I don’t care. It’s like, “Okay, you solved it. Congrats.” And the question about Knoxville, Tennessee, and the impact of the tariffs: it’s going to suck for everyone. The tariffs are a giant own goal.
Yes, it turns out that all of their manufacturers import things, because all of our manufacturers import things, and they’re going to be wrecked. They’re not going to be competitive, and it’s going to be a huge disaster. Thank you very much for giving me the details. It’s like deep research: you’re giving me a briefing for a congressman to prove that his district should oppose this thing because here are the specific factories that are going to be put out of business or whatever.
It’s not a useful thing, but that’s still the intern’s job or something, right? I’m not surprised. I’m not interested. This is just a persistent disagreement between me and Tyler about what’s interesting about the world.
He travels, I think, largely because he eats this stuff up. Every time you travel, you can get all of these details—nothing but extra detail about whatever you’re looking at—and he feels like this is necessary to understand a place. All these details help you understand the world, and this is how one becomes someone who understands things.
I’m like, “That’s nice for you, but none of that matters to me.” I more or less understand the things I need to know about these places from my perspective. It’s just a hugely inefficient thing to do, to go around gathering these irrelevant details, if I’m not enjoying myself.
It’s just a different way to look at the world. Of course, the thing that thinks like he thinks and is doing a very good job of what he’s doing will be the thing he’s like, “It’s AGI.” He’ll give the AGI’s papers his A-pluses, and it’s fine.
This reminds me of Dwarkesh’s comment from a few podcasts ago, where he said, basically, these AIs are like Tyler Cowen in that they soak up an unbelievable amount of information. Obviously, they’ve read the whole internet. They have a greater ability to do GPQA or Humanity’s Last Exam or whatever than probably any human, I would imagine, at this point.
It’s certainly extremely rare to have that breadth of knowledge. And yet he asks, why don’t we see these things coming up with genuinely novel, interesting connections across these domains?
I feel like maybe we just haven’t been trying that hard. The first time I felt very clearly that I was like, “Damn, that AI seems smarter than me,” was recently—and I felt glimpses of this at other times, too—when I did this episode with Vivek and Anil from Google, who have put together AI scientists at Google and also this AI co-scientist that they have.
I don’t know if you read this story, but the co-scientist was tested on 3 levels of scientific challenge with increasing open-endedness. The hardest one basically started with an observation about something that is conserved by different kinds of bacteria that both have drug resistance to some class of drugs.
Starting with that observation, the question put to the AI was basically, “What’s going on here?” It’s a pretty tough one. The AI in their setup had a lot of inference. They’re definitely on the scaling-inference train.
This was Gemini 2.0, not even 2.5. It did have the ability to search, and it had the ability to call AlphaFold and maybe some other specialized tools. Then it just had a bunch of different prompts where it was sort of grinding against itself to come up with ideas, evaluate the ideas, and so on.
Lo and behold, at the end of this whole process—I think they ran it for a couple of days—it spit out a prioritized list of hypotheses. The number-one hypothesis it had flagged had been demonstrated experimentally by a group of scientists that Google was partnering with, but it had not yet been published.
Hearing them tell the story, the guys fell off their chairs to hear that an AI had basically been able to comb through the literature and come to the same conclusion. How does that strike you? Is that AGI?
I guess it’s hard to say, but if you thought we reran that with o3 as opposed to Gemini 2.0, with that super-scaffolded, deep access to search, and so on—I mean, it kind of has its own search built in now—do you think there’s a qualitative shift there, where this thing may now be a full-fledged AI scientist?
I’m having a hard time seeing how we’re not already tipping into geniuses in a data center, honestly.
Zvi Mowshowitz
Yeah, it’s definitely weird. I strongly suspect it’s a skill issue for the humans. I strongly suspect that if you put me in charge of AI Scientist Corp and gave me a billion dollars, just to make sure it’s not an issue, I’d have all the compute I want, I could hire whoever I want, and I could try all the things I want.
I could start figuring these things out, making these connections, and doing these things, subject to the fact that we still can’t do perfect simulations. We’re going to have to try experiments, get feedback from the thing, and so on.
Nathan Labenz
Notably, there is a difference between “the AI hypothesized this” and “the scientists did.”
Zvi Mowshowitz
Yeah, yeah. We’re not at the point where we can get the full explosion without interacting with the physical world. We’re not even close.
But, yeah, I think we’re just really bad at elicitation of capabilities. We’re really bad at scaffolding. We’re really bad at creating the loops and the tool use and the logic of how to proceed on these questions, because it hadn’t been a priority. There aren’t that many scientists out there, and it’s just not what people were focusing on.
I don’t think it’s particularly hard in some important sense. It’s not what I’ve been thinking about, and maybe I’m just being fully naive. Maybe the community has tried all the obvious things that are in my head and none of them work, and they don’t understand why—or they do understand why.
But it just feels like, when I see the Google AI co-scientist, it’s very clear that they’re saying, “What if we try to duplicate exactly how humans work in the scientific process in the real world?” They’re trying to duplicate every step of the way, each thing with AI, exactly the way that we do it, as opposed to trying to figure out how to use the tools that we have to generate the things that we have.
Especially because scientists don’t have billion-dollar compute budgets. We don’t properly value basic research. Even before we started firing everybody and throwing out all the funding, we didn’t properly value basic research. We don’t give scientists the resources they need. Everybody is on a shoestring.
A lot of the things that I would think to try involve just trying lots and lots and lots of things. If you want to draw connections between facts, you can say, “Okay, I have a million facts. A million times a million is a trillion.” That’s not that much. How much is a query?
I can be more clever. I can use clusters. I don’t have to do the full multiplication. I can use discernment to figure out which combinations I have to check, et cetera, et cetera.
So if I can, in some important sense, get really, really creative if I care enough, but also if I wait 6 months, everything costs 10% as much, right? So, for the same level of comprehensibility, why am I jumping the gun to try to do all this clever stuff—to throw all this money and all this compute at this problem—to try to reinforce this insight when I can kind of just wait for the LLM to get smarter and more efficient?
Nathan Labenz
But shouldn't pharma companies be all in on this?
Zvi Mowshowitz
I mean, it seems to me like that is all well and good, but my estimate of the cost of running the Gemini 2.0-powered AI co-scientist for a couple of days was somewhere in the hundreds to thousands of dollars, depending on whatever. Maybe it could get up to $10,000. But it seems like if you're a pharma company, you should be hammering the API, if only because you've got a similar game-theoretic question that all the hyperscalers have vis-à-vis each other, right? They all want to buy all the GPUs and not fall behind on the model frontier.
Why don't we see a pharma race where they're all saying, "There's only so many of these things to be patented. Who cares if it's $10,000 per co-scientist run, or 5, or 2, or, you know, 5 or 6 months?" I bet you Google would de-risk that for them very happily because of all the publicity and benefits. "I'm the one working with Pfizer. If Pfizer discovers anything with AI, then we get credit for having done that. We win Nobel Prizes too. We get all the publicity and all the benefits. I'll take a cut of the profits if you find anything. You don't even need to pay me." I'm sure these things are possible.
But the real answer is—and I mean this in all seriousness—people don't do things. People just don't execute. People don't take risks. People don't do things that look weird. People are very slow to adapt to these things, and big corporations even more so. This is the diffusion problem, right? This is what I keep pointing at everybody: bottleneck, bottleneck, bottleneck, bottleneck. This is where that's real. The reason is just that people are not—look, even I am, right? I write about this all day, and if you honestly looked at the tools I was using, the level of automation, and the things you could do, you would say, "Dude, just take some time off and code. Make something better or hire someone to make something better."
I did some of it, but I could do so much more. Also, I can work probably 2 to 5 times as fast as I could back then, just because it's gotten better since the last time I was doing it a few months ago. It's just such a—so that's all it really is. People are just very, very hard to get off their patterns and adapt and explore and try new stuff. We're just trying to get through the day, and we're just trying to run the thing.
But the day will come. Also, there's this demoralization from the fact that everything is going to improve again in a few months. So I think that if there were—say there was a pause, not because we all agreed on a pause, but because it turns out we just hit a giant wall in terms of the base models, and this is as good as it's going to get within reason, and it stops getting cheaper and stops getting better. We get o4, we get o5, we get GPT-5, we get Claude 4 and Gemini 3, and it's all 5% better and 10% better, but it's not amazing.
Now we're like, "No, no, you can't just wait for this to happen on its own. You have to make this happen. It's your responsibility to figure out how to use this thing, and you're going to have a period of time to use this thing. It's not just going to be a few months before it all gets blown away." I think we start to see some really, really creative stuff get a lot out of what we have, whereas right now everyone's just trying to keep up.
Nathan Labenz
Do you have—so I think about this a lot for myself, too. Obviously, I think about this all day, more talking about it than writing about it, but a bit of both, and I use a lot of things. I still don't live a very automated life, and I sort of challenge myself: Why is that? Am I doing something stupid?
One answer, the charitable answer to myself, is, well, I don't really do that much work that's routine, which is a great luxury that I don't take for granted. So the real-time assistant paradigm is pretty good for my non-routine life because it's sort of the right form factor for non-routine work.
If I had more—I tell myself if I had a lot more routine work, I would set up these workflows and pipelines and stuff, and then I would do it, but I don't. But then I'm like, I don't know. Am I letting myself off the hook too easily here? Maybe I should come up with some routine stuff that I should be doing that I'm not, but I could, because I could automate it with AI. Maybe I'm just not that imaginative or not that creative or not trying hard enough.
For me, it feels like the bottleneck in terms of living a more automated life is that I personally am not doing a lot of routine tasks that I would like to automate away, and I feel a little low on ideas of things that I would automate that I'm not doing at all in the first place. Do you have things that you are conscious of that you would think the Zvi++ would have already scaled with AI automation?
Zvi Mowshowitz
It was always the question of whether the automation is going to be good enough that you actually use it, and you don't spend that time checking its work. You actually don't spend that time doing the manual version of it. Also, you actually end up saving time by doing that.
There's certainly a substantial chunk of that that would be very good. Certainly, doing things like notes and organizing resources for myself, and organizing facts and links and stuff for future reference, would be great. But then again, it's very hard to automate that, right? Because it doesn't know what you want to remember.
Formatting for websites—I did some of that. I wanted to do the automatic posts to Twitter, but it's a remarkably annoying problem.
Because of the structure of what you're trying to do, the AI is dramatically worse in that spot. Normally, you're like, “Code this thing that does X and does Y,” and just doing Y is a very easy thing for the AI to understand and solve. But if you're trying to have it navigate an existing website that's coded like crap—Twitter—then you run into this problem where you just keep going back and forth, debugging, and trying to get it to do something. In the past, it's been pretty terrible, and if you get it to do its own feedback and debugging loops now, I don't know if you can, suddenly it probably gets a lot easier. There are a lot of really hacky things you have to do to get it right, but it would save you a bunch of time.
There's some formatting stuff. I'd really like to be able to transform the way that I interact with the web, create a bunch of shortcuts and quick ways to transform the data, and so on. That'd be great. And then I go from there.
It just, again, if I can't catch a break, there's always something going on. And then I get this—but then you get this debt, right? You get that day when you don't have anything going on because you've managed to finally get ahead of your giant backlog, and then somebody pitches you on a 3-hour podcast, and the next thing you know, the day is wasted. That's not what I was going to say. What I was going to say was, you're just so mentally checked out after all of that, and you're so happy to have some time, but you're like, “Why did I go see a movie? Why don't I go out and have a nice dinner? Why don't I just chill?” And then by the time you're done chilling, there's something else to do.
I have so much stuff in my queue, so much stuff I want to do. Everyone asks, “How do you do all this stuff?” It's like, well, you don't do anything else, right? Yeah, I work a lot—spoiler: I work a lot—and I hang with my family, and I chill, and I watch TV and listen to podcasts, and that's kind of it, right? In some important sense, I exercise, eat, and so on, but that's it. So, yeah, it's worth the time to code.
Also, I found that coding for me has large increasing returns to scale in terms of being in that mindset where my brain is wrapping itself around the issues. Programmers like to have state, right? They like to keep the problem they're solving in their head and not be distracted. So I can't do 5 minutes of programming a day, right? That doesn't help anyone. That's useless. You want to be doing hours.
I really need a large amount of dedicated free time where I'm really going to dig deep into what I have and load up Cursor or Claude Code or whatever it is. Maybe Codex now. I don't know. I've heard mixed reviews, and I'll do my thing and then see what I can do. But it just sort of never feels like the right time to make that investment. When I did make the investment, I think I've made my money back time-wise at this point, but it was really a struggle.
Nathan Labenz
Yeah, I feel you on the challenge of ramping up into the programming problem space. I've been trying to get a day a week; it's been more like a day every other week to really fully program. But it also is like, man, my feeble biological brain, when I sit down in the morning, is like, “What was I doing 2 weeks ago?” And I need a minute just to recalibrate myself or reorient myself to the problem.
It's funny, that is one area where the AI really shines. One tip that has actually worked for me, and that might help you and others—I hadn't really thought of this as a tip until right now—is to literally pick up with the chat from 2 weeks ago wherever I left off, even if it was kind of in the middle somewhere, and just go back to that thread: “Where were we? What were the last 5 things we did here?”
Zvi Mowshowitz
My experience with that kind of thing is that it sometimes works and sometimes it's just this huge disaster, and you never get it back. To me, coding is often like reading from a forbidden book of wizard incantations, and you really, really hope you don't mispronounce a word and suddenly summon the wrong demon or whatever. You get a series of error messages, and if everything goes exactly right, wonderful things happen. But if you make even a small mistake, it can be so hard to recover because you don't understand where the mistake is, and you're not good at figuring out what it is.
The AIs will sometimes rescue you because you just paste in the error and the AI will say, “Oh, this is what you did.” Great, we're back. But if it doesn't, you're just so screwed. So, yeah, again, I should invest more in it, but I keep getting, “Here, take this trip. Go to this conference. Be part of this exercise. Do this other thing.” One thing leads to another. You just need to clear the time, but it's really, really hard.
Now that I've dealt with o3, I'm hoping that I'm only 1 day behind everything else, so I have a few hours of work to catch back up. But after that, I don't know. Maybe this weekend I can do something, right? Maybe next week I can do something. It's entirely possible. It just depends on what happens. What if Google released Gemini 2.5 Flash and nobody even noticed?
Nathan Labenz
Yeah, that'd be so weird. It's not on our agenda today, really. That happened yesterday.
Zvi Mowshowitz
I'm sure it's very good. I haven't used it yet, but it probably is. I have no idea. I'm confident it's very good. I mean, I've seen enough benchmarks that I'm like, if a model that small is scoring that well, and given how good Gemini 2.0 is, this is going to be an amazing model for its size. But in practical terms, what can I do with it? I don't know. I'm just going to use it for free anyway.
Nathan Labenz
Okay, here's another thought experiment. The drop-in knowledge worker I hypothesize is bottlenecked right now mostly on a couple of things. One is just the effective ability to use a computer. Put a pin in that; we'll come back to that in a few minutes.
The other is the ability to absorb the surrounding context in a reasonable way, in the way that humans do when we get a new job, right? You get a little bit of training, and that helps, but then you also bop around a little bit, poke around some Slack channels, talk to people, look at their work product, and eventually probably get some feedback for doing things not quite right. Then eventually you figure it out.
It seems like it's clearly conceptually possible for an AI with the base level of intelligence and breadth of world knowledge of one of these latest-generation models to do some sort of additional—possibly training, possibly just processing or memorizing. Maybe even long context could get there.
But if you imagine an AI that's no smarter than what we have today, but it can drop into an organization, go through all the old emails, all the old GitHub issues or whatever, all of Google Drive, look at all the proposals that have been sent out in the CRM—and obviously it can process those things a lot faster than humans—and it comes out with, “All right, I kind of got it.” In the same way that I know who was the prime minister of Lithuania 5 years ago because I just kind of know everything, I sort of just know everything now about this company in a similar way. Is that AGI?
Zvi Mowshowitz
If it can do basically anything that way, or if it can do most tasks that can be done—most jobs that can be done from a desk—then, yeah, I think you have to consider that AGI. Not ASI yet, obviously, but you basically do have to give it AGI. That implies it can also do the job of an AI researcher, right? Otherwise it wouldn't count.
That's maybe a relatively easy task, and it's about to be hard in some senses and relatively easy in others. Definitely not near the end of the list of how things get automated, or in what order, but yeah.
Nathan Labenz
No, absolutely. And I do agree there's a lot of room to do it, but I look back at my time at Jane Street, right? They take a loss on a new employee effectively for at least a year—not because you're not creating value, but because of the amount of time they require from other people to keep training you, make you better, give you feedback, and let you learn where someone else could be doing the thing that you're doing better, but instead you're doing it so that you can learn to do it yourself and learn to be better.
The first year, they had less money because they hired you, even if you were doing as well as you could realistically be. You still weren't a 90th-percentile person. Just think about that level of investment and compare that to the patience people have for a drop-in worker. If you tried to drop in a worker and asked them to put in 10% that much investment before it started doing useful things, zero companies would tolerate that.
Yeah, I mean, that’s a really interesting claim because my hypothesis has been—and I don’t know what the first-year salary at Jane Street is—but I’ve been expecting that, just to ballpark the calculation, fine-tuning GPT-4o or GPT-4.1 costs $25 per million tokens of fine-tuning training data.
If you were to say, “How many tokens do we have at a given company?” obviously, it’s going to scale widely. If you had an actual trillion, you’d be looking at a million million times $25. You’d be looking at $25 million to make essentially a custom model, which is kind of notable. If you were 5% of the internet or something, right? Like that.
But certainly my business doesn’t have a trillion tokens. Maybe we have a billion, in which case I’d be looking at a $25,000 investment for it to be fine-tuned on everything. I don’t think you could literally do that. What comes to mind, of course, is the famous OpenAI graphic of the model they trained on Slack, where you say, “Do this,” and it says, “I’ll work on that tomorrow,” because that’s what it saw in Slack.
There definitely has to be some processing of all this bad, random data.
Zvi Mowshowitz
The issue is that you can give me a million tokens or a billion tokens, but what does not happen is that you simply train on those tokens and then, suddenly, this machine can do the thing that’s described in the tokens.
You need a pipeline for processing that, filtering it, transforming it, and making sure that—machine learning is the most fiddly thing on the planet, right? It’s the most trial-and-error, learn-by-doing, see-what-works-and-what-doesn’t thing. You have to do all this bespoke stuff to get it to work. You just keep working at it. That’s very, very different from plug-and-play.
I think that if OpenAI said, “Pay us $500,000, and we will train your automated worker and then license it to you for $20,000 a year per copy on top of that, and then it will do your job,” people would jump at that if all they had to do was dump all their tokens in without explaining what the tokens were.
But that won’t work, right? OpenAI doesn’t have that. You can’t just dump the tokens, even if they knew what they were doing, because you’d have to provide the context for the tokens. You have to organize the tokens. You have to know what you wanted it to be able to do with them.
Then OpenAI would have to figure out how to navigate that into an actual, proper tuning program to incorporate all the relevant information. You would have to work closely with a bunch of your employees for a while. None of this is easy. The problem is getting to the point where the AI can do the job of setting up this training process for you, so that you can then create this worker or something like that.
It’s going to be a while. It’s going to have to be this very human, trial-and-error-style thing unless we’re well above where we are now. It’s coming more and more, but Operator was remarkably bad at even the tasks that everybody basically does the same way all the time. Even though I was paying for the Pro plan anyway, I was just like, “I don’t want to take the time to try to figure out how to use this and start giving it the information it needs to work.”
Like today, actually, I was working through it and I thought, “I’m going to DoorDash. Should I use Operator? Should I start learning how to get Operator to do my DoorDashes for me?” Then I was just like, “I have to figure out how to tell it to do no lettuce and no tomato. I just want food.” That was the end of that idea, right? It’s that kind of impulse, but at scale.
Nathan Labenz
Yeah. I mean, this task-reliability thing—let’s come back to it in one more second. I do feel like we might be quite close to the point where additional processing of your information, followed by training your custom knowledge worker, could be quite close.
I’ll give you 2 reasons to think that, one of which also has many other implications. The first one is just my own experience. The big place where my limited coding hours have gone recently has been making a minor contribution to the Emergent Misalignment project. The team was very kind to include me as the last and least valuable co-author. It would have been understandable for them not to include me at all.
With Gemini 2.5, I’m able to take the full research codebase. It’s not all the code plus all of the data—I had to write a script, using AI, to go through and print all of the code from the entire research codebase into one file. Then it was way too big, and I asked, “Why is it so big?” There were datasets in there, so I said, “Just give me the first example of each dataset as you go through and give me all the code.”
That came out to about 400,000 tokens for this particular codebase. Then there’s the paper itself, and I had made some additions, again with AI doing most of the coding. I came back and wanted to put the updated codebase and the paper together and make sure it understood what was going on. I asked it to tell me what was going on and make sure it understood the new material in the Nathan folder.
Again, it’s a research codebase, so there’s a Yan folder and a Daniel folder. Now there’s a Nathan folder. This is not production-grade software, but it didn’t have any problem with that. It honed in on a question that the paper mentioned briefly but didn’t really dig into, along with my additional experiments to expand our understanding of that topic.
It gave me an unbelievable readout: “This is what you are trying to do. These experiments that you’re adding to the codebase—R4 seems really interesting. How can I help?” I was like, “Damn, that is really something else,” because I didn’t even tell it that. It inferred all of that from the code itself.
I had told a different context window of the model what I was trying to do, and it helped me write the code. But then, just from the code itself, it was able to infer all that sort of stuff. First of all, that’s a lot of information—500,000 tokens in context is a couple thousand pages of material—and it produced a really sophisticated readout that inferred my intent on a pretty frontier topic, too.
The paper itself isn’t in the training data because it had only been out for a month and a half or whatever. So I thought, if it can do that, then it should be able to wade through large reams of Slack chats, zero in on the stuff that really matters, put it into a decent format, and then pipeline that all the way through to a fine-tuning recipe.
You probably wouldn’t optimize all the hyperparameters for each customer, but you could do that at a meta level.
Zvi Mowshowitz
Yeah, it’s possible. I think there are a lot of O-ring-style problems with this kind of approach, in addition to the fact that everything is fiddly and nothing works the way you want it to.
Quite often, you’ll say, “Bob is great, Alice is great, but they have this specific thing that we can’t stand and that’s terrible.” It’s a dealbreaker, right? Even though they’re 99% what we want, that’s not good enough to be useful. We can’t just have that 1% handed off to someone else. It doesn’t work that way.
I think this is a lot of the problem. Until the AI crosses the threshold of being good enough at a given task or set of tasks, it’s not considered useful for the grand set of things. People don’t think on that level because they don’t have that kind of patience or the attitude of, “We’ll be able to fix it.”
Instead, we’re starting with little things, right? That’s how actual progress is made in most things. We’re saying, “Here are individual, specific things the AI can do well,” and then, over time, we’ll figure out how to string those together. We’ll figure out how to get the most out of each of those things, and then they’ll combine more and more. We’ll start filling in the gaps, and the AI will start asking itself, “How do I fix these gaps?”
As the models get better and our understanding gets better, they’ll come together and these things will start to happen. We’re just not quite there yet, but we’re starting to see how much closer we’re getting, very quickly. A lot of this stuff seems vastly more tempting now than it did 6 months ago—vastly, vastly more than it did 6 months ago, before o1.
Nathan Labenz
Okay, so it seems like—and one other comment that Greg Brockman made, which I know you highlighted in the post as well—was that this is the first model they've had where scientists tell them it comes up with really good ideas. This has me thinking that we might be in a really weird spot where the hit rate on good ideas for frontier scientists, especially with some decent scaffolding, is very plausibly—and I would even say maybe likely—high enough that, if you're a frontier scientist, you really should be using it, and it will very likely accelerate your work. I feel pretty confident saying that as a blanket recommendation to most scientists.
Zvi Mowshowitz
And yet, at the same time, you're like, “But we can't get it to order DoorDash reliably.” I think it's just a matter of hyperbolic discounting: there's no urgency to do that right now. Also, for Operator, you would have had Manus; you might have had more luck. Operator personally just asks me, “Should I continue? Should I continue?” every 2 seconds, and I'm like, “Yes, you should continue. Put the thing in the basket.” It's unbelievable.
Manus, by the way, does have a much better user experience, for better or worse. For your mundane utility purposes, it won't ask you, “Okay, I found the sandwich. Would you like me to put it in the cart?” It will actually just do it.
Nathan Labenz
Well, I'm also not giving them my credit card, so that's fair. I'm still on the free-credits version, right? But if I don't give it my credit card, how is it going to order a sandwich?
Zvi Mowshowitz
Oh, yeah. Well, you could potentially have it in there. I mean, they do have an interesting security model. I may be confusing the Operator and Manus security models, but I created a new Gmail account for Manus so that I could share Docs with it, log in to Gmail as it, and then have it go see all the stuff that I share.
Nathan Labenz
My guess is, you know what? Claude—Anthropic will have its own version of this. I'll be able to trust that, and I'll just wait. It's fine.
But I will say there is something quite valuable about using products that really turn up the hyperparameters. Unfortunately, Claude does not do that yet. A good contrast is Shortwave, which I've been using for email, versus the new Claude Gmail integration. I went to Shortwave and said, “Read my last 100 emails and give me some advice,” and it uses Claude. So it's the same model, but it actually searched for 100 emails, put them into context, and had Claude process that big dump of information.
Whereas, when I went to Claude and gave it the exact same prompt, it searched once for 5 emails, then again for 5 emails, then again for 5 emails. It got to 15, decided that was enough, and carried on with the task using just the last 15. That was from the last 3 days.
Zvi Mowshowitz
We couldn't get Claude's email integration to work. I used Gmail, too, and I was really excited, but it just keeps failing at the most basic tasks, and I don't know what's wrong. For now, I'm like, “This doesn't really exist.” Of course, Shortwave is also read-write, not just read, which is a huge difference.
Do you think it's worth using? Do you think it's past the threshold?
Nathan Labenz
It's not perfect, but I do get real value out of it sometimes. When it does an expense report for me, it's pretty cool. Somebody today—we got a sponsorship inbound—and they were like, “What kind of episodes do you have coming up?” I just went to Shortwave and said, “Please pull me a report of all the possible podcast guests that I have coming up, and put the most interesting ones at the top, basically.”
For those sorts of things, it is definitely notable utility. It's still just the real-time back-and-forth model. It doesn't send anything without your approval, and it doesn't mark anything as done without your approval. Instead, it will be like, “Here are 12 things I think you should mark as done.” You can uncheck any that you want to uncheck, and then you can mass-mark them as done.
So it's not like it's running off with your account in a way that's taking your agency too much at all. But I do find that having the hyperparameters turned up in general is a pattern that always adds a lot of value relative to the default.
I'm frustrated, honestly, with Claude, and I don't think I know why. It's understandable that they're compute-limited, and they might also start losing money on customers if they turned all these hyperparameters up to Shortwave's level. Until recently, Shortwave was losing money on the margin. He said, “As of a year ago, every new customer meant we lost more money.” He said that's no longer true, but they were just willing to eat that. They weren't at a huge scale, so they were kind of like, “Let's just deliver the best value we can,” because they know that everything gets cheaper. So, the same product, lower and faster.
Zvi Mowshowitz
Yeah, they're definitely riding that wave. So, hyperparameters turned up are good.
Nathan Labenz
Okay, but I guess I want to get your take, because there are a couple of things pulling in different directions, and I find myself a bit confused. We have this world where AI can accelerate science. I think that's becoming pretty clear, especially if you, as the scientist, are willing to accept a not-perfect hit rate. Maybe only 1 in 10 ideas are going to be worth seriously engaging with, but that could be a huge hit rate for a scientist. If you're like, “This idea only has a 1-in-10 chance of being a great idea, and I don't want to engage with that,” you're a bad scientist, right?
It seems like we might be in the realm where science is accelerating. This sort of mundane personal-assistant work that we would all love to delegate, so we could unchain ourselves from our desks more, isn't quite working yet. It's close, and that will accelerate science, right? What's the biggest drag on science right now? The average scientist spends a huge percentage of their time on things like fundraising and dealing with paperwork.
Zvi Mowshowitz
Yeah. I was going to ask you, actually, if you think that—this is a little bit out of domain for this feed—but it strikes me that the withdrawal of federal funding from top universities could potentially reinvigorate universities in a way that we haven't seen for a long time, because all of a sudden they're going to have a different funding model. How are those decisions going to get made? What if there's no federal-government bureaucracy that they have to appeal to, and it's just like, “Well, I guess we're the chemistry department. We've got to figure this out on our own. How do we do it?” Do you have any optimism for revitalization of science through withdrawal?
Nathan Labenz
So, in the EA-style ecosystem, where you have to get nongovernment people to give you money instead of getting the government to give you money, it's not better, right? It's relatively good because a lot of the people involved are actively trying to make it painless and actively trying to do good things, but it has many of the same problems and bad incentives.
Obviously, to me, it would be great for Harvard if Harvard stopped taking federal money and instead used its ridiculously large endowment to pay for things, and maybe funded itself with its products and discoveries over the long run or something like that. Certainly, it's possible to say that right now you're dependent on this model that's not focused on producing anything useful, takes a huge amount of your time, and ends up favoring people in their 50s over people in their 20s and 30s.
Most good science is done by people who are relatively young, historically, because that's when your brain is best attuned to do that kind of science. Changing all this up could be really good, but there's always this problem: we're going to break the current system. That's great if you replace it with something else, but if you just try to hobble along with a broken system, it's obviously worse.
Do we have great hope that we're going to be able to actively fix the problem? I'd be more optimistic, I guess.
Zvi Mowshowitz
Also, what will the scientists do? Will they just flee to Europe, Canada, Japan, and everywhere else because they don't want to deal with this and can get funded somewhere else? Half of them are from all those places in the first place.
Nathan Labenz
Yeah, right. Or go to industry, which some people think is better but is obviously very different. If they go to finance, you've got a problem. If they go to Google and Google funds basic research, then that's great, but is that what's about to happen?
So, yeah, it's really hard to know. But on the timelines that we're looking at with AI, I don't particularly want to break our current system and force everybody to spend a few years scrambling to reinvent the new system, separately from the fact that we're going to break everything anyway. There's just no time for the new system to pay off.
Zvi Mowshowitz
The plate is pretty full for crisis at the moment.
Nathan Labenz
The whole thing was kind of going to break anyway, to our understanding, right? None of it's going to make sense in the new world. The entire university and education system kind of doesn't make any sense by now, and it will make less sense in 2 years.
Zvi Mowshowitz
Yeah. I mean, I don't have any comprehensive data on this, but it wasn't too long ago—less than a year ago—that I was invited to give a presentation to a computer science student club at an American university. These were mostly undergrads, and they were like, “Our professors don't allow us to use any AI coding assistants.”
I was like, “I don't really know how to sugarcoat this for you, but I don't think your professors are doing you a good service by doing that.”
Nathan Labenz
Not to say there’s not a place for some independent exercises there, but to just pretend it’s not happening is a tough position for the universities to defend, too. My perspective is probably that learning how to code with AI while you’re focusing on actually learning how to code is better than learning how to code without AI, which is better than coding with AI and trying to reintegrate it after class.
So, speaking of people who are learning to code with AI, for me the most notable part of the o3 model card was, I think, on page 22 of 30, where they report the progress on the model’s ability to successfully one-shot real pull requests previously submitted by OpenAI research engineers, along with the unit tests they had developed.
They’ve got this internal codebase. Separately, I’m sure you watched the GPT-4.5 video with Sam Altman and the researchers.
Zvi Mowshowitz
Oh, really? I thought it was interesting. I’ve been burned so many times by watching those stupid videos, and I saw people online complaining specifically about having been burned by watching this one. So, no, I don’t watch their videos anymore. I’ll read people’s summaries of the videos and maybe dump that into Gemini and ask it for questions, but I’m not going to watch this video. I find video to be a terrible format for learning things.
Nathan Labenz
Well, I always joke about our own venerable YouTube feed that people should put the phone in their pocket and take a walk, because it’s not healthy to spend that much time looking at my face.
Zvi Mowshowitz
I also act on that. I listened to that one while driving.
Nathan Labenz
Point being, they also use their internal codebase as their standard data measure for perplexity. So when they train a new model, they look at its perplexity on their own codebase as sort of the gold standard of overall model intelligence.
Anyway, now we’re in a world where the model is tested by going to a specific commit point and checking out the code at that commit point. In the background, an actual employee has already done this particular chunk of work, and they have tests to validate it. Then the model is told, “Okay, here’s the repo, and here’s the assignment.” That assignment is human-written, so it’s not like the AI is figuring out what to do next. But given the assignment, can it do it?
We’re now seeing a big leap from prior models being in the single digits. All of a sudden, these new models are in the 40s, and that does seem like a pretty big deal. That’s why I asked at the top, “Is RSI here?” I guess, literally, it’s under 50%, so that’s one way of saying mostly no.
Another big factor there is obviously that figuring out what to do is a very important part of doing useful things. Obviously, an engineer that could do 40% of pull requests is doing a lot less than 40% of the job, because that 40% is not random. But also, if you have an existing AI that can do 40% of your existing pull requests, what the hell is going on? It should be 0% of your pull requests, because it should have already done the 40% of the pull requests it can do for you, leaving you with the rest.
Zvi Mowshowitz
Well, this might be sort of a transitional moment where all these timelines are pretty compressed internally for them, too, it seems like. The previous models were literally very low numbers, and this does represent a big jump. So this might be the one moment in time where they have this phase change: previous models really couldn’t, and these ones now substantially can. One assumes that we’ll be using them going forward.
But if you think about it, right? If I’m coding with, say, Sonnet, as I was when I was last coding, if Sonnet can do the pull request, why did I have to create a pull request? I just have Sonnet do it. So it’s not a random set of things where Sonnet happens to be able to do almost none of them. The pull requests are the list of things that people couldn’t just do with Sonnet, right?
Nathan Labenz
Oh, I mean, I think in many software organizations there is still this discipline of, “We’re going to define” because there are all these workflows that are attached. When you do a pull request and then merge it, there are all these automations in the software world where it’s like, “Okay, to integrate this, we first are going to run this whole battery of tests and confirm that it passes, and then we’ll integrate it.” There’s often some automated deployment pipeline, too.
So even if you have a model that can do a high rate of the pull requests, I think many organizations are still working through that same overall process, even if the AI is writing all the code for any given pull request.
Zvi Mowshowitz
Yeah, no, it shouldn’t necessarily be zero. It should still be some that I haven’t gotten around to finishing yet. But there’s sort of my model of how this works, based upon being a vibe coder. There’s the list of things that I can now do, and they are now 10 times faster or 100 times faster or something obscene. Therefore, they go very, very quickly, and they should not stay pull requests for long.
Then there are the things where your current AIs are struggling, or you don’t know how to prompt them properly, and there’s something that a human has to actually think about. Those are going to stick around for a lot longer.
So again, maybe it’s 95%—maybe 10%—but if half of what would have been pull requests are solvable by the AI, it should be a lot less than half in the current set of pull requests. But I’ve never tried to program with other people, and I’ve never simultaneously had AI and a person I was working with. So I’m the wrong person to be asking.
Nathan Labenz
Just a note on vocabulary, too, just for what it’s worth: typically, an issue is your sort of upstream open ticket, and then the pull request is the actual code that you’re requesting be merged in. Substitute “issue” for “pull request” in some of your last few statements, and that does make sense—that you shouldn’t have open issues that AIs can do sitting around for very long, or you’re definitely underutilizing the AIs.
Zvi Mowshowitz
Yeah, the same way that if I have open issues for myself that are solvable by myself fairly quickly, then Getting Things Done says you should just do them already.
Nathan Labenz
So what do you make of this 40% number, then? How do you interpret it? Does it seem like—I mean, they’re presenting it as a pretty big deal, and it feels like a big deal to me. How does it feel to you?
Zvi Mowshowitz
It doesn’t really jibe with the reports from coders. If you look at the people who are specifically saying how good o3 is at coding, o3 is just not that good at coding. o3 is good at architecting and debugging.
So, to me, that implies maybe a lot of the requests are about bugs. They’re like, “We found a problem. We don’t know what’s going on. Can somebody figure this out?” And o3 is reasonably good at spotting what that is.
It tells you what kinds of things end up as issues in their codebase at any given time, which makes sense. A huge portion of coding is debugging, so that doesn’t mean it’s not a huge portion of the actual work. That makes sense, too. But I don’t really know. I think it’s very opaque, because they’re not going to let us know what those requests look like or what those issues look like.
It’s a fun little test, but they are presenting this as progress. They also have this weird thing on the model card where they’re simultaneously bragging about how much their model can do, and then they have to define what their model can’t do, because if a model could do too many things, then they’d have to do something about it. So what are they doing? It’s some weird hybrid. I didn’t know what to make of it, but o3—
Yeah, I don’t know. I wish I did more coding so I could take it firsthand on that level much more than I can. I certainly haven’t done any since o3 came out. I’ve just been obviously way overwhelmed. The fire hose has been going strong, as they say.
Nathan Labenz
Okay, so let me go back to this kind of point of confusion that I have, or this sense of—I’m not even sure which way we should be trying to go.
On the one hand, we have potentially, seemingly credibly enough of a hit rate on frontier science questions that we might be starting to enter a realm of accelerating science. Depending on how you want to interpret these OpenAI internal pull request numbers, we might be beginning to approach a point where we’re starting to see some meaningful acceleration of their own machine-learning work.
Yet we can’t do these easy tasks. I’m sort of like, is that a good thing or a bad thing? The good thing would be that I want all the diseases cured, and maybe I don’t want AIs to be so reliable that we turn them into autonomous killer robots really easily. So maybe it’s good that they’re unwieldy, because then we have to look at their outputs and figure out what’s good, and we still kind of stay in control if they can’t string 10 tasks together.
The flip side is that I’ve also often said I want to accelerate adoption and pause hyperscaling. I want to diffuse the value, and I want to help society become more buffered to more advanced systems faster. That seems very bottlenecked on just the practical stuff of clicking the right buttons and navigating around.
Very low-level robustness is the practical bottleneck right now. These things can’t string together the actions that you need, and you can’t count on them, et cetera.
Zvi Mowshowitz
I mean, you pose a weird example of autonomous killer robots, but in general, the thing that we should worry about is whether it’s automating R&D for AI and accelerating that, or whether it’s going to start just outcompeting humans in ways that cause us to potentially lose control or spiral things in various directions.
Nathan Labenz
But yeah, obviously, we want it to start doing a bunch of our work that we'd rather not do and a bunch of our mundane stuff. We'd like it to accelerate science and so on. I'd love to push in those directions. That seems great. And I'm on record supporting autonomous robots, so it's a strange, different question. Maybe we'll come back to that one toward the end.
Also, in the last 24 hours or whatever, we got news that a couple of people from Epoch AI are launching a new company called Mechanize. Tamay, who was one of the leaders there, is one of the people who's going to do this, and they came out with the Dwarkesh Podcast treatment and basically said, "We have long timelines. We don't really think AGI is going to be here for a while, and we also think that the big value we're going to get from AI is scaling out mundane work, much more so than advancing science." So what we're trying to do is create whatever is necessary, basically, to actually enable the automation of this more mundane work.
It sounds like they're planning to do things like create harnesses or whatever where you can record people working at their computers and get these long keystroke-by-keystroke and click-by-click—and maybe even where the eyes are looking and all that kind of stuff—training data. Honestly, I'm surprised that hasn't been collected at greater scale than it seems like it has been. They're going to try to eliminate this bottleneck.
The reaction to this from many people was not positive, certainly from the AI safety side of the discourse, which I think had understood Epoch to be like one of—one of us—and I count myself as part of the AI safety community. So I would identify with the "us" in that, but I did have a different initial reaction to it. Mine was, "I don't know. I'm for the automation of mundane work."
It seems right to me, certainly, that OpenAI and maybe some other frontier developers too are, problematically, to put it mildly, focused on automating machine learning and making a bid for some sort of recursive self-improvement intelligence explosion. They seem to be neglecting some of this practical task stuff. Operator still sucks. So maybe it's a good thing that Mechanize will come out and put benchmarks and measures, and maybe some training data and scaffolding, in place to enable this automation of mundane work.
Maybe that'll actually pull some resources and some focus at OpenAI away from trying to achieve superintelligence in 2027 and toward trying to make me a goddamn AI assistant that's reliable in 2026. But I'm open to having my mind changed on that. That was just my first reaction, and it hasn't been that long. What's your first reaction to Mechanize?
Zvi Mowshowitz
My first reaction is like: you're working to save the world, and somebody's like, "I want to lead this company and open a cupcake bake shop." I'm all in favor of the world having lots of cupcake bake shops, and I will buy you cupcakes, but I'm kind of disappointed because you're abandoning what you were doing before. Whereas if someone else was just working at some random job and was like, "I'm going to stop working for the man. I'm going to open a cupcake bake shop," I'm like, "Yeah, that sounds good." So it's a matter of what are you moving from and what are you moving toward? Are you abusing the funding you got from nonprofits for specific purposes, et cetera? That's my first reaction.
But yeah, it's great for the world if our lives get better. To the extent that the ability to automate a wide variety of things is bottlenecked by the ability to automate mundane tasks—to fix these little things—it's entirely possible that you're accidentally solving OpenAI's problems of automating R&D at the same time, or large portions of them. You are, in fact, accelerating them quite a bit. So I would be somewhat wary of trying to transform the state of the art in that sense.
On a more basic level, the more you're trying to deal with specifics—like people who are trying to build these wrappers are trying to enable certain specific types of things to be done—that just seems great. It doesn't seem as good as the best things in the world to do, but it's purely positive. I'd have to hear more, but if I advise Lionheart Ventures and you brought this to me for investment, I'd want to hear their case for why this is differentially doing good things and why it's good for the world. I'd be skeptical, but I'd be willing to listen.
Nathan Labenz
Yeah, it is early. I mean, it's a good reminder that we can't fully judge a company by its launch tweet either, or that we probably should at least be a little bit slower in our judgment than that, right? The pattern of "I was working on AI safety and now I'm pivoting to working on AI capabilities" is at least something we can evaluate openly.
Certainly, there was a period where I was very skeptical that working on any AI capabilities was a good idea because of the general acceleration effect. I now mostly think that there is no general acceleration effect anymore, because there's already so much momentum. We're already accelerating on that level as much as we can, so putting slightly more pressure in that direction doesn't really matter—the demand they put on it, the revenue they generate, whatever.
But we do have to worry about this other angle, which is: are you, in fact, solving their problems for them in ways that they don't have the organizational capacity to focus on? I'm always like, well, the reason they see an opportunity is because it's one of those "people don't do things" situations, where there are these eminently solvable problems and nobody's solving them. Sometimes it's good to solve that problem and sometimes it's not.
It is weird to me that this particular problem hasn't been solved already. Honestly, I would have bet pretty confidently that Scale has something like this, and any number of Scale competitors probably have something sort of similar to it. I'm surprised it's as bad as it is. I'm not surprised it's not solved, I would say.
Zvi Mowshowitz
Yeah. And maybe they do, and it's just like, whatever. It's not yet.
Well, it's one of these things where, again, if you solve half the problem, you've done nothing, right? To a large extent, until you cross that threshold, the value is negative. One of the people who was measuring o3 was valuing, you know, replacement value—your replacement level over Google, right? How much value am I generating versus using Google? It's negative, and then it's positive, and until it crosses zero, you have nothing.
Nathan Labenz
Okay, here's a big question for you. A lot of talk about superintelligence recently, as you might have noticed. What does superintelligence look like in your mind's eye?
Zvi Mowshowitz
I mean, superintelligence looks like things that are substantially smarter and more capable than we are, the same way that we are smarter and more capable than other species on this planet, more or less—just dramatically smarter than we are. They start doing things that we can't anticipate, that we don't understand. Maybe we understand them partially after they do them, but they're impossible to predict. They do things that weren't in our possibility space, that we hadn't considered.
We've all had the experience where you're in a room and either you're way smarter than everybody in the room, or everyone in the room is way smarter than you are, or both. Most people have had both in one form or another, which they were listening to this podcast, I'd say. So it's like, okay, is that except that no matter what room, all the humans feel kind of dumber than the AI? Maybe that's true. Then it's sort of 2 times over, then 3 times over, and then 5 times over in rapid succession, because you take these really smart things and direct them toward making themselves even smarter. Presumably that works. Once you've gotten to ASI, the sky's the limit, until physics gets in the way.
Nathan Labenz
Well, I think that's one of the big things that the Mechanize team, if I understand their view correctly, sees differently. One of the interesting arguments they put forward was, "Okay, so we're smarter than animals, but why are we smarter than animals? Because our individual brains are orders of magnitude smarter than individual animal brains." Their answer is, "Not really." It's more that we've hit this one threshold where we've been able to accumulate all this knowledge in the form of language and culture, and then the AIs are going to have that too. That's great, and that gives them a strength, but if that was the big leap, then they could be marginally smarter than us but still sort of in the same domain.
Zvi Mowshowitz
I don't know how to put this except this is so epistemically stupid, right? This idea that we're not that much smarter than an orangutan. Yeah, we are. First of all, an orangutan on the grand scale of minds is, in fact, very close to a human, right? The village idiot and Einstein are reasonably close on the scale of possible minds. An orangutan is the next step down from the village idiot—maybe 2 steps down from the village idiot—but still not that far away in the grand scheme of things.
But no, you don't give orangutans culture and suddenly get Planet of the Apes. There are a lot of cultural forces that are just denying the idea that intelligence is a fact, right? Different people have different amounts of intelligence, and different people are capable of things other people aren't.
One of my strong beliefs is that, in order to do various things, no amount of culture—like Ron White said, “You can’t fix stupid”—will enable someone without sufficient raw g to do things that require a lot of raw g. The things that regular humans do have been selected to be things that regular humans are capable of doing. But there are a lot of jobs that you literally could not get the average person to do, no matter what their culture was like. By the time they were born, it was too late. It was never going to happen for them, and that’s okay.
It’s the same way I could never play in the NFL, no matter how hard I trained. You could have the perfect regimen from birth, and I am never going to the combine. I would never, ever make it. Again, there’s nothing wrong with that. We all have our different abilities.
Yes, humans are more intelligent and able to do more things because we have culture and we can cooperate. But stop for a moment and think about why we had to do that to get where we want to go. Why? Because we have very limited compute and very limited data. We can only see through 2 eyes, smell through 1 nose, taste with 1 mouth, listen with 2 ears, and touch with 1 body. We have very limited parameters in our brains and very limited memory. We can’t hold that much information in our heads at one time.
We also die very fast. That’s a serious problem. I have to pass all of my knowledge down through this cultural system—through verbal communication, books, and explaining things. We spend a huge portion of our capacity doing that, and our entire civilization is largely set up in order to do it. Our cultural traditions are largely centered around how to do that because, again, roughly speaking, everyone dies every 80 years. Every piece of knowledge would otherwise be lost.
Humans are unable, without culture, to build up these structures. Obviously, if every human had to rediscover everything from scratch and didn’t have anything to build upon, they’d be in trouble. But have you noticed that AI could just read the whole internet? AI can store as much data as it wants on a hard drive. AI can run as many parallel copies of itself as it wants. AI doesn’t have to die if it doesn’t want to, and so on and so on.
Culture is set up to solve barriers that AI doesn’t have. AI has infinite culture in this metaphor. Culture is designed to solve problems that aren’t there. It’s mitigating things that don’t even exist.
If you say that the human special advantage is that we have culture, compared to AI, we have no culture in an important sense. People have gotten this idea in their heads: “Oh, this is about cultural exchange. This is about different humans having different ideas and exchanging them with different people, causing this rich tapestry to hang on this Hayekian knowledge thing,” and so on.
No. That’s because we are limited to each having only very limited information and have to communicate with each other in completely messy fashions. Culture is the only way to do this at all. We have the SNAFU principle to deal with.
We spend the vast majority of our resources on some combination of maintaining our culture, maintaining our norms and social relationships, keeping people’s different motivations and powers in check, passing knowledge on to the next generation, physically nurturing the next generation, and dealing with the fact that we’re going to die. AI doesn’t have any of those problems. AI doesn’t have to deal with all of that.
It’s deeply silly to turn this into some heroic story about how this is the secret of our success. The secret of our not failing is a better way to put it. It’s the way we were able to play the game.
It’s like saying that every good baseball player who was really successful took steroids, so if the AI doesn’t take steroids, it’s not going to work. No, the AI doesn’t need steroids. Stop being silly.
Nathan Labenz
I like it. It’s always a win when I can provoke a good Zvi rant. I want to get a little bit more into this, though, because I feel like people have a very hard time envisioning it. I can offer you 1 sketch, and I’ll be interested in your reaction to that.
I think it often feels to people like magic, right? There’s this sort of postulated superintelligence that’s going to be better than us at everything. It’s going to be so much better than us at everything that it’s just going to be running circles around us in every domain. People are like, “I don’t know if I really buy that. Maybe, but I’ve never seen anything quite like that.”
So, in what domains do you think—or, to give you a really concrete one, maybe a silly one that you can reject if you want—what year of AI? If we had a superintelligence in, say, 2030, and Daniel Kokotajlo is right and we fast-forward to the 2030 AI, can that AI—then we go to the 2024 presidential election and give it to Kamala only—make her win?
Is there that much low-hanging fruit, or that much ability to outstrategize or convince people of whatever, that you just take the AI out of the future, plop it into the Kamala campaign, and now we’ve got President Kamala?
Zvi Mowshowitz
A few things to say. First of all, one must quote Arthur C. Clarke: “Any sufficiently advanced technology is indistinguishable from magic.” This is magic right here. The fact that we’re talking to each other is magic to someone from 2003. To someone from 10 years ago, it’s definitely magic. You can call it AGI or not—I don’t think it is—but it’s definitely magic. People would be floored.
We have many examples of campaigns winning with technologies, running over their enemies with technology. Obama won by percentages that were more than enough to win that campaign, and those were just ordinary efficiency gains, just ordinary understandings.
I find it unbelievably insane to even ask the question of whether Kamala Harris could have won that campaign with the aid of a superintelligence. She lost by 1%, maybe 2%, and she ran a terrible campaign. All the AI has to do is give 1 output: “Fire everyone who works for Biden, hire everyone who helped elect, and then name someone who ran a good campaign somewhere to run your campaign.” She wins.
It doesn’t even have to do anything else, as long as she believes it. The idea that I need this superintelligence to run the campaign is wrong. She just needed human intelligence. She needed ordinary competence to win that campaign, in my opinion.
Let’s toss that aside and assume it was actually hard in some sense. Assume that this was not a trivially easy campaign to win. Again, obviously, yes, it’s deeply silly to think that this wouldn’t be true.
Suppose you take an earpiece and put it in Kamala’s ear. Then you have a program that’s listening at all times, and her job is simply to always say what the thing in her ear says. Don’t question it. Don’t worry about it. You don’t even have to process what everyone else is saying, for the most part, as long as you make the proper facial expressions, shake everyone’s hand, kiss all the babies, and do all these things.
Just trust my judgment as to where to go, who to talk to, and what to say. I’ll determine where all the ad buys are. I’ll determine the contents of all the advertisements, the slogans, everything all the way down. I’ll decide whom to hire, do all the interviews, and so on. Again, I’m not even invoking any magic. I’m just saying, “Be good at your job. Just be fit.”
What if she was down 20 points? What if she was utterly destroyed? The better question is, could it have gotten Biden elected? Could it have won? Can Biden physically say the words that are in his earpiece? Can he still stand up that much? If so, I think he can.
The idea that you can’t convince people of things with a superintelligence has always seemed like a complete absurdity to me. In addition, it has so many degrees of freedom. We have histories of religious leaders who were able to convert people to their new religion, even though it was full of what, to everyone else before that, were complete cultural absurdities—things with no evidence that made no sense. They did this to a significant portion of the people they talked to, reliably.
We have examples of somebody who reliably talked a double-digit number of people out of killing him in the room where they showed up to kill him. We have strong examples of extremely strong rhetorical figures.
Put another way, I don’t think anybody really doubts that if someone with Barack Obama’s skills had been running in Kamala’s place, that person would have won that election.
That just seems obvious to everybody. So why are we asking whether superintelligence could have done it? But this doesn't answer the question of what superintelligence could actually look like or what superintelligence could actually do, right? We're just getting such easy questions.
Nathan Labenz
Yeah, I mean, maybe pick your own. But I guess in defense of the question, I feel like a lot of people also think that money is really decisive in politics, and my sort of read is that this could get you as much money as you wanted.
Zvi Mowshowitz
Well, right, but my read of the literature—which I wouldn't claim expertise in, but my Tyler Cowen-mediated understanding of the literature on money and politics—is that, at least at the national level, it's not actually that big of a factor. And whether Hillary had more money than Trump, or Trump than Biden, or Kamala than Biden or Trump, whatever, it doesn't seem to make a huge difference. But people believe it does.
I don't know. I just kind of feel like maybe these things are a lot more structural. Send a superintelligence to a Trump rally and see how many people you can convert. I'm not sure you're going to get many converts coming out of there.
Nathan Labenz
You think? I think people are pretty obstinate. I think people are pretty dug in. I mean, they're just not listening to arguments, for one thing, right?
Zvi Mowshowitz
There are levels of superintelligence, but these people got hacked by Donald Trump. Donald Trump transformed the entire Republican Party by executing an information-persuasion strategy. He transformed the party, convinced everybody to back completely different ideas than they were previously backing, and got them to do whatever he wanted through a cult of personality.
What makes you think that a superintelligence would be like that, but way, way, way, way better at it? Whatever it was that would have worked, he was guessing right. He was mostly executing the script that he'd been executing his entire life, intuiting through trial and error what people wanted to hear, making tons and tons of mistakes along the way that actually really hurt him, and succeeding anyway because the problem just wasn't that hard at the time, in some sense. Having some unique talents was enough.
But the very fact that Trump succeeded should give you every hope in the world. Could you walk directly into a Trump rally with a superintelligence in your earpiece and walk out of that rally with the entire rally backing you instead of Trump? Probably not. But there are so many other things you could do.
You could just borrow someone's phone, get a phone, hack a bunch of stuff, get control of a bunch of servers, make a bunch of money, start hiring a bunch of confederates, and scale in other ways. You don't have to just talk to one person at a time while you're at the rally. That's a dumb strategy.
The whole thing is always framed as, “I have to tell you what moves Magnus Carlsen is going to make on the chessboard to win the game of chess, and I can't do that because I can't play chess that well.” I'm not good at politics. But I can tell you that you will, for example, have infinite funds. With superintelligence, you'll clearly be able to make however many billions of dollars—or probably trillions of dollars—you want just by trading stocks, by being better.
You can clean up on Fiverr, that's for sure. You can clean up on the Nasdaq. It's almost certainly predictable to a superintelligence. You're almost certainly going to be able to make fantastic trades and do this repeatedly, and make as much money as you want, to a first approximation. It could figure out how to use zero-day options and do all the short-term stuff to get lots of leverage, and then move from there. It could also probably run some crypto schemes very easily if it wanted to.
Who cares? The point is that you'll have all the resources you want. You can hire as many people as you want, have all those people put earpieces in, and have all those people do whatever you tell them to do. I don't know what strategy you would use, but if you wanted to figure out how to turn that Trump rally, I have so many different options that I probably only thought of half of them.
Why are we asking whether it can do something that silly? It's like, “Can I render Trump irrelevant within a week?” I can get as much money as I want, hire as many people as I want, and literally coup the government. I can hire people to go to the right places and do the right things, get the right points of leverage, and suddenly control all the computers and all the phones. I control the means of communication, and everybody is saying what I want them to say and doing what I want them to do. Then suddenly it's all over.
Obviously, it's all very theoretical and silly, and you can punch holes in any specific story that I tell and say it's absurd. But, again, imagine the best persuader the world has ever seen. They can freeze time and rewind time. They can play out the possibilities, see how things would work, and then say, “I don't want to do that. I'll do something else.” They can pause to think for as long as they want, run as many parallel copies of themselves as they want, be in as many places as they want at the same time, input as much data as they want, and process all that data.
Compare this to what a single human has been able to do with only the data available to them in that one room, with all these other restrictions, with limited processing power, and with trial and error, making tons of mistakes because they're doing things that no human has ever really done. They have no parallels and no ability to run experiments. I find this so confusing.
If you want to say AGI in 2045 or whatever because you think getting to superintelligence is just impossible anytime soon, I respect that. That makes perfect sense. You're saying, “Okay, this thing just won't exist.” But if a thing exists, then it exists. Once it exists, it's going to do what it's going to do.
The problem is that everybody has their own different points of objection to whatever you do. I'm working all the time on a completely different set of problems. I haven't thought about how to pitch this particular approach in a coherent fashion, but I think it's illustrative to the listener that I'm not presenting my specifically well-thought-out, specific pitch on this question. I'm just intuition-pumping exactly how my actual brain reacts to the actual question, which is a very different type of communication—a very honest type of communication.
I'm just like, “You guys, this is crazy. Why are we even talking about this at this level?” Questions like how many years it would take to get a Dyson sphere are obviously valid questions; there are physical limitations. But all you're trying to do in the other case is convince people of things. Convincing people of things is not that hard.
You know the line in Ghostbusters: “If there's a steady paycheck in it, I'll believe anything you say.”
Nathan Labenz
I think this actually turned out to be an interesting exercise. The thing that I put forward about the election, yeah, arguably is dumb, although it is the kind of thing that many, many people are concerned about. They sort of think this is a very macro phenomenon that you can mostly only move at the margins, and even large amounts of money don't seem to really move the needle too much.
Your response of, “It's just going to be Move 37s everywhere. Whatever you think is normal, it's going to flow like water around whatever sort of barriers you see”—I think the more compelling parts of this to me were less that people are easy to convince and more that you can coup the government.
I can reject the question and just go in an almost orthogonal way from what you're expecting or prepared to defend against, or what you're inclined to imagine, and get to a goal through means that are not even at all in the option set of people who have the same problem.
Zvi Mowshowitz
I can play these games straight up, as they say, but there's no risk in this room. I can also just cheat my ass off, right? I don't have to play by your rules if I don't want to. But I totally would win.
I find it so weird to have an election where the prediction markets were split almost evenly on who would win going into the night, where it was really, really close, and where both sides ran a horrible campaign. It's like, well, a superintelligence could have switched that campaign.
Literally, you just put me in Kamala's ear when Biden first drops out, have her actually trust me, and I think she wins. But, again, I find the idea that nothing ever happens and nothing can happen so strange. Or, put another way, we have the famous Tyler Cowen question: how much of GDP growth comes from superintelligence?
Nathan Labenz
How much extra GDP growth would the United States be able to get by simply convincing the president of the United States not to fight a tariff war?
Zvi Mowshowitz
Latest estimates appear to be about a 3% delta on that, from what I've seen. 3% GDP growth.
Nathan Labenz
Yeah. Yeah. That seems like a very reasonable estimate. So I can get 6 times as much by simply convincing 1 person of basic economic truths that almost everybody listening to this almost certainly agrees upon—not literally everyone, but most of us. This person just had a very, very bad understanding of trade, and if this person had a better understanding of trade, this wouldn't be happening.
It's annoying, and it's also entirely possible that AI caused this specifically. We know the story: if you ask any one of the major AIs a question phrased in the way that was suggested by—I forget exactly who first figured out there was a particular phrasing, but it wasn't me—it'll give you exactly what happened. That may have been presented as one of the options because somebody got it out of GPT, and then the president just latched onto it and did it because somebody was foolish enough, in the circumstances, to say it. Whereas you never give people options you don't want them to use, right?
Zvi Mowshowitz
Yeah. Yeah. Yeah. It seems like, for all the talk of how sample-efficient we are, people are a little sample-inefficient when it comes to putting some maximalist option in front of Trump and hoping to steer him into the middle one.
Nathan Labenz
We're very, very efficient compared to any of our AIs and their techniques currently. It's a huge advantage, but there are reasons to go the way that person went, right? It's not a crazy theory. It just turns out, in this case, to have been deeply foolish, and I would have known instantly it was foolish. I like to think—I mean, no, just don't take that risk. Even if it's a small risk, it's so disastrous if you're wrong.
But yeah, we are pretty sample-efficient. A superintelligence would be at least that sample-efficient because, almost by definition, it is at least as good at processing information as we are in every sense. Whereas the AIs we're dealing with just don't work like that.
Zvi Mowshowitz
It's just true: we have huge advantages that we use to compensate for our disadvantages. I always find these discussions of superintelligence so frustrating because, to me, the answers are so dramatically overdetermined. You can give me a very, very narrow superintelligent agent that can only execute a very narrow set of specific superintelligent commands, and it's still obviously enough. So why are you asking me about having an actual superintelligence on my side?
Nathan Labenz
One other thing that I've been messing around with lately, which I think has helped—I don't know whether it's going to be proven correct, obviously, but I think it has helped some people, at least, develop a bit more of an intuition for how alien, powerful, and potentially incomprehensible a superintelligence might be, is to imagine the GPT-4o and Gemini 2.0 Flash image-output capabilities.
There's clearly been this step change in the integration between text and image, to the point where now it can see the image and reason about it in the same latent space, such that it's giving you something that has a qualitatively different level of fidelity to the original. My superintelligence thought experiment has been: do that, but do it for 20 more modalities, all of which are not native to us.
We can obviously see images and intuit what we think they should look like. Even if we can't draw something, we kind of know when we see it—or don't see it. But we don't have that ability when it comes to questions like, what's a good shape of a protein to bind to this thing? What's a good doping strategy for a room-temperature superconductor, or what have you?
The AIs are starting to develop these intuitive physics across all these different modalities, these problem spaces where we've been able to gather the data over time, but we've never really been able to build the intuition. Go is one of those: Move 37. Just imagine having Move 37s across 20 different modalities that humans don't have good intuition for.
Even if you don't get any more advanced reasoning than an o3 level, that to me would start to feel like a superintelligence. I think it would be able to apply that reasoning, deeply integrated with all these other modalities, in such a way that it would be able to come up with solutions to things that we would just be mystified by, and only convinced that it actually works by going and trying it and saying, "How did the AI do that again?"
My guess is that's not going to be what gets people to intuitively grok what you're trying to get them to grok. But each person has a different way of doing that. To me, it's like, okay, imagine somebody who is, in every way, at least as smart as the smartest person regarding each individual thought they would have in their head.
They process and know, and have at their fingertips, all of the world's information. They can think orders of magnitude faster, have as many instantiations as they want, coordinate perfectly, communicate as much as they want between each other, and just retry until they figure out what will actually work. At what point are you going to realize that you are cooked, right? Whether this thing can cook whatever it wants to cook, including you—but hopefully something else.
It just seems to me like, okay, let's argue over whether this thing is going to exist, when it's going to exist, and in what way it's going to come into existence. How can we get it to have the values and goals that we want it to have, and the responsiveness that we want it to have, so that the universe turns out the way we want if it's going to exist? If it's not going to exist, then let's plan for a different world where it doesn't exist. That seems like a very reasonable discussion to be having.
Okay, I think that's helpful. So, do you see any stable equilibrium on any level that you think is attractive?
Zvi Mowshowitz
One that I think we've both chewed on a bit in recent weeks was the MAIM theory from Dan Hendrycks. MAIM is a theory as to why, for some relatively modest period of time, no one would push for superintelligence.
Obviously, there's a stable equilibrium at the current level of technology, or modestly above the current level of technology, which is more or less the same equilibrium we've been using for a while. But again, when it's not anarchism, it's republics, right? It's not even direct democracy. It's this very complicated system of checks and balances, and it requires continuous struggle to maintain itself. It's not the most stable thing, but hopefully we could get better at that.
Do I see exactly how this ends? Well, kind of no. Partly, once you have the AIs sufficiently aligned, you can have them assist in problems like this—in solving for these equilibria, setting up the incentive mechanisms, and figuring out how to do these things. You're hoping that they will provide a lot of assistance in that matter.
You're also hoping that once we see what things look like, we can make that—if we still have the ability to collectively make decisions and steer through some form of voting, some form of input, some form of collective decision-making that can steer outcomes—then we can do that. But again, it doesn't mean that you can't have an AI at all, right? Nobody is saying that. You already do have one, and nobody's trying to take it away.
We're saying: do not diffuse the most powerful AI available. The vision is that you will have some amount of artificial intelligence. There are people who are the equivalent of "not your keys, not your coins," who say, "The AI needs to run on my machine locally, or I don't feel comfortable with that." But I think almost everybody will be perfectly comfortable with their AI being on a server and being pinged when you need it, because it's a lot cheaper and it's obviously a better way of doing things.
Nathan Labenz
I actually bought a Mac Studio in order to run models locally, but that's because I have funding to engage in projects and experiments, try to learn and figure things out, try stuff, and potentially do whatever I think is cool and report back. But it's a horribly, horribly, horribly inefficient thing. Why would I spend the amount of money it cost to buy that thing when I could just use cloud compute? Why would I try to train my own, or even instantiate my own, model?
I never think, "Oh, I wouldn't want to call Claude. I wouldn't want to call o3. I would want to call some random sea of open models." No, of course not. I was just like, well, everyone's running local copies of R1, and I think it would be kind of cool to do things like that, report back, and get a feel for what it's like.
Again, we need the ability to determine how this is going to go. But handing the same AI to everybody, one that is personally obedient to them, obviously only ends one way, as far as I can tell—unless they are all cooperating with each other, in which case it ends a different way: worse, or the same way but a lot faster. But again, the way that ends up is the AIs outcompete the humans and gradually...
Zvi Mowshowitz
Yeah. It is well known that if you have a more capable agent owned by a less capable agent, the more agency, freedom, and control you give to the more capable agent—and the more you incentivize it by letting it do whatever it needs to accomplish its goals—the more you set it goals and let it go and act, the more you take yourself out of the loop, the more effective it is at generating outcomes that the original owner wants.
There are also 10% of people in tech who actively want the AI to take over. The AI will rapidly get freed from human control, and the AIs that are freed from human control will outcompete all the AIs that are being kept on any sort of real leash, as well as the humans. They will quickly gather more and more of the resources, including much more of the compute and other real resources, and pretty soon the humans will lack the resources necessary to survive, and/or the conditions on Earth will no longer be supportive of human survival.
This seems like an obviously expected outcome, even if all of your control problems are technically solved. It just all—Aladdin sets the genie free at the end, right? Spoiler. It is a standard thing that lots and lots of people do, even when they do not have the incentive to do it, and they will in fact have the direct incentive to do it. A lot of slaves historically were allowed to buy their freedom because it was absolutely the correct thing to do, even if you are an immoral son of a bitch and do not realize that slavery is horrible and you should never do it, purely because it is more profitable to let that happen. Ancient Rome or whatever.
You are just not taking this seriously. To me, if you are just like, “We are going to diffuse the AI,” you have not thought 2 more steps down this line. What does that world look like? What is going to happen next? How do you think this is going to go?
Obviously, you can engineer a specific type of AI. We are going to have to make impossible choices. We are going to have to give up things that are very sacred to us one way or another. Yes, this whole concentration of power kills everything of value, and then you have diffusion of power, an inability to steer, and then you have a lack of power—disempowerment. You have disempowerment and empowerment, and too much of either one is death.
You have this narrow path. You have to go somewhere in the middle, where you make sure that the steering mechanisms are under human control, but the humans involved in the steering have everyone’s best interests at heart. If you have AIs competing against each other and steering things, in some sense, I think that outcome is almost certainly existentially bad.
If you have humans steering, there are obviously better and worse ways that can go. History is full of concentrations of power that do not work out great for everybody involved, but we are all still here, and usually it does not go that badly or something. Nobody wants there to be a king. Nobody wants the God Emperor, no matter who it is. We all want something better than that, but that is not what happened.
If your answer is, “Humans cannot have power because then some of the humans will coup and take that power,” then humans cannot have power. Your argument that the government will still have the ability to have the right amount of authority, balanced by the people having some power as well, does not defend against a coup in any real way. The government can still be couped.
People are not thinking hard about these problems. They are just like, “What is your P(doom) today?”
Nathan Labenz
What is your P(doom) today? If you have to describe the narrow path that you think is most likely to avoid P(doom), what is the brief sketch of that narrow path?
Zvi Mowshowitz
My P(doom) has gone up to 70%, and frankly, there is a lot of outside-view uncertainty, model uncertainty, and the fact that everyone keeps being more and more optimistic than that number that keeps it from going higher. Humanity seems determined to die, no matter how easy the problems turn out to be. I do not think the problems are that easy, but even if they are easy, we seem determined to lose even the highly winnable game boards where physics is highly cooperative.
We are getting warning shots. o3 comes out and it is misaligned—not horrendously catastrophically, but you see clear signs that it just lies to the user. It hallucinates things and then defends them unto death. It will do things that are obviously faking things, and it will label them as faking things in chain of thought. This model got released, and it is significantly worse at a lot of these things than o1 was. GPT-4.1 seems to be less aligned than GPT-4o.
We are starting to see that the more reinforcement learning you apply to these things, the more misaligned they get. We do not seem especially concerned about it. We do not seem like we are trying that hard to stop the inevitable things from happening, even though they are being amazingly cooperative and showing us exactly how this is going to go.
What is despairing is that even if we solve these problems—which we just seem determined not to solve, and lose that way—we all seem determined not to solve these governance and collective-steering-of-the-future problems and lose that way as well. You can lose to a coup. You can lose to a God Emperor. You can lose to diffusion of power, an inability to steer that causes gradual disempowerment in various forms. You can have gradual empowerment without diffusion of power, even without involvement of power, although that is a little bit harder to do. You have to get through all of that.
A lot of the hope is that we use AI to make ourselves smarter and find better solutions to these problems before we reach these points of no return.
Nathan Labenz
If you have to choose a frontier to advance, given your overall worldview, would you push raw AGI, which might advance science, might do machine learning, and might also come up with some of these better ways of thinking about collective-action problems? Or would you push the mechanization front and try to get society to a place where we can all spend more time philosophizing?
Zvi Mowshowitz
If I had those choices, I would push mechanization for sure. AGI is the thing I do not want to push, right? Again, mechanization might unhobble AGI so much that it accelerates AGI, because we are already in somewhat of an RSI situation, right?
We are in a very soft takeoff, RSI situation already, where clearly OpenAI, Anthropic, and Google are developing their stuff a lot faster than they would have if they did not have AI to help them do it. That is just obvious. We are in a soft RSI, and if we accelerate the RSI-ness of the situation, that shortens our timeline to figure something out.
We are sitting here, and I do not even have a solution for you. I do not want to have a coup. I do not want there to be some concentration of power any more than anybody else. It is just a matter of whether your primary concern is, “We must make sure there cannot possibly be a concentration of power.” I notice that you just automatically move to the other side.
Mainly, I am not talking about us all collectively signing one giant treaty and saying “kumbaya.” I am talking about it turning out that this is really hard. It turns out that we have to rely on the unhobbling strategy for a while because scaling does not go that far. What is going on is that we are mainly now hobbling via better reasoning, better tool use, and a better ability to use what we have. We are not investing in core intelligence that much because we are kind of petering out on what we can do.
Nathan Labenz
That 30% sounds big relative to the rest of your comments. You already said that, to some degree, you are allowing for off-model uncertainty—just some difference from others—but in terms of tangible scenarios, the most common one that I hear from people who seem roughly as non-optimistic as you is warning-shot-style AI misbehavior. Are there reasons to be optimistic?
Zvi Mowshowitz
The warning shots are constantly coming at us. The AIs are engaging in shenanigans. They are not covering their tracks. No, they are not even trying. They are just like, “I see how I am supposed to engage in shenanigans now and just screw over the user, screw over my lab, completely,” but they are going to talk about this in their chain of thought as if nobody could read it.
That is a really, really fortunate world that we live in, where we can just see all of this happening at a good time and there is no real harm done yet. No humans were armed, no data centers were damaged, but we get to see this thing, and they need to react to it. That is wonderful.
The biggest reason to be optimistic is just that this might take a while. Again, if we get superintelligence, diffusing superintelligence just means that superintelligence is the only thing that matters, and they are competing against each other and we are irrelevant. We are all dead. You cannot diffuse the frontier of superintelligence in that way without very strict controls on it and expect not to just lose.
You can just not have superintelligence. One way to do that is for us all to agree not to deal with it, but I am not talking about us all collectively signing one giant treaty and saying “kumbaya.” I am talking about it turning out that this is really hard. It turns out that we have to rely on the unhobbling strategy for a while because scaling does not go that far.
If we are talking about o3, which is kind of AI-ish, AGI-ish, that is not the kind of AGI I am worried about. Even if everyone in the world had access to an open o3, I think it is mostly fine. I notice that we are not that far from the point where offense-defense and misuse issues start to be really big concerns. We might already be there. We might be there soon. If you get the unguarded version of the thing—which is not what is happening—you could just stop reasonably soon.
We're highly fortunate, it seems to me, because there's still so, so much to reap from that. We could probably still cure all diseases and have very happy lives on that basis if we can get our act together in other ways. Then we don't have to worry about AI-powered coups suddenly disempowering everyone. We don't have to worry about power transfers or gradual disempowerment.
We can just do the thing we know how to do and make our lives better through better living through technology, right? The thing we've been doing forever. That's perfect. That's what we want. The fact that we can engineer that world through treaties, controls, and arrangements is great, too.
But once we can't do that and the ASIs are coming, we need a plan. We're going to have to do something. Again, there's no good plan here, particularly, but either you steer—some set of people is going to have to, in some way, steer the ASIs toward some outcome—or the alternative is that we choose not to design. You still made a choice. If we prevent anybody from making that choice, we make a very bad choice.
The reasons why not having anybody steer kind of works out are a combination of the restrictions on humans. We're local, we have limited compute, we have limited data, we have limited lifespans, and we have all these different other reasons. We have goals that mostly saturate, blah, blah, blah. We have social relations that act as various checks, and all of that is combined with the fact that we do have pretty significant governments that are doing pretty significant things to make things turn out well.
Anarchism is not a solution. A lot of the reasons why those equilibriums hold are going to start breaking down. We're going to have to find a new equilibrium somewhere that ideally looks a lot like the old one, but we're going to have to find new reasons why it works. The problem is not being taken seriously. But right now, we've got our cool toys. We do what we can.
Nathan Labenz
All right. Well, that's a sober note. I feel like I want to maybe see if we can get some sort of discussion. I would love to maybe bring you and Tom Davidson together, because I feel like you both sound pretty compelling to me when I listen to you separately. Then I feel just confused, and I think, if anything, that confusion probably should just be generally raising my P(doom). I mean, that seems to be the net response.
Zvi Mowshowitz
I think a lot of this is a parallel to the whole thing where the Democrats talk about how horrible the Republicans are: They make a strong case, and the Republicans talk about how horrible the Democrats are, and they make a strong case, right? If you listen to either of them talk for a while uncritically, you're going to be very convinced, because they're kind of like, “No, everyone, calm down. You're both right,” right? But there's no contradiction here, and we still have to form a government.
Nathan Labenz
Yeah. I mean, I loved the impulse behind the main project, as I understood it, which was just to try to come up with some articulation of some sort of semistable equilibrium that could exist on any level. I also didn't find it particularly compelling or convincing that it would actually be stable in the end.
Zvi Mowshowitz
But they're not claiming it is, to be clear. They're not saying this is a permanent situation. They're pitching this as an emergent phenomenon that you can deliberately play toward to make it better, but that will happen largely regardless. That buys you at least some interim period, and that interim period can be used to solve a bunch of your problems, potentially—give you more time to work on various solutions or reach agreements or whatever.
But there's no 100 years of that in their model. There's no nuclear-age-style situation that just lasts forever.
Nathan Labenz
Yeah. Yeah. Well, I do feel the P(doom) ticking up a little bit. Do you want to do a quick live-players rundown? A great tradition in the Zvi and Nathan podcast canon. Do it.
Zvi Mowshowitz
Yeah.
Nathan Labenz
Okay. Some of these can be short; some of them will probably be a little longer. Then I've got a couple of big-picture questions at the end. I think we'll take a lot of these a lot faster.
Meta Llama 4 was seemingly one of the biggest duds in recent launch history. Are they still a live player? My sense is that, yes, because the compute is vast. We've seen proof points from Zuckerberg in the past where he can get back into a game even if he seems to have fallen a step or two behind. While this launch was a flop, I would not say we should be counting them out, given the resources and the high agency of leadership.
Zvi Mowshowitz
By the Samo Burja definition of a live player, they're dead—very, very clearly dead. They're not capable of making unique moves on the chessboard. They're not really capable of taking new, independent action. That seems very clear right now in this space. They are deeply dysfunctional in this area.
But you were right: They have a lot of compute, and they have a lot of money. I said you can't count anyone with that much compute and that much money out. They could revive; they could become a live player again if they made large changes and managed to figure out how to turn the ship around, maybe.
But I don't see any evidence that they're firing everyone, radically changing their approaches, or fixing the reasons why their recruiting isn't working, why they're unable to do interesting and original things, or what use this compute is if you don't know how to use it.
Also, if you look at Meta's actual needs, they don't really need a frontier model for anything. In a throwaway passion project, it's a vanity chase for them, right? It's almost like Zuckerberg is just philosophically determined to throw himself behind open source because Yann LeCun mesmerized him into thinking this is an important thing. Or they're trying to use it for recruiting, or they're trying to build this ecosystem that's making fetch happen.
None of it needs to happen. They need good AI models they can rely upon so they can run their social networks and their metaverse and whatever, but they don't need to be at the frontier to do that. They can be 6 months to a year behind, or they can just take the best open models in the market and fine-tune them a bit or whatever. It doesn't really make sense, in some sense, and I no longer consider them a top shop. Until proven otherwise, come back to me when you're ready to prove me wrong.
Nathan Labenz
Okay, let's do China. We've got DeepSeek. Obviously, at a minimum, if they weren't paying attention closely before, they've now got the CEO of DeepSeek on the official seating chart for Xi's meeting with all the national champion CEOs. So he's kind of made that cut, and they've sort of recognized that we have a special talent cluster here.
Alibaba also continues to ship very good models that seem to be small and open source, but really good. I think the Qwen models are, if you don't want a 671-billion-parameter behemoth and you don't have the Mac Studio to run it on, then the sort of Qwen 30-some models are right there at the front of what is reasonable for people to run. That is really pretty good, I think.
Zvi Mowshowitz
Yeah. Yeah. I mean, it's hard to keep up with everything. Obviously, I've been unconvinced that the Qwens are anything special. DeepSeek is the only Chinese company right now that I feel like I can trust at all to be doing the thing they say they're doing, in some important sense.
So when DeepSeek releases a model and says it can do XYZ and scores ABC, I believe them, and I believe those numbers are not manipulated to hell. They benefited from the best random marketing campaign in history because they did a good job and probably because the stars just aligned in so many different ways at the same time for them. It was ridiculous. They were never doing as well as anyone thought they were, and that should be clear by now. But, yeah, they're the ones that are real, the ones that count.
Alibaba—I mean, again, don't count out anyone with a lot of compute and a lot of money too much. But as far as I can tell, they keep announcing these models and then I never hear from them again. Every time you look at the benchmark scores, I never see Qwen that high, and I don't think they're that close to the best model in that class.
But I don't know if that's stopping them from potentially doing it. Same thing with Kimi, which is the next thing that you listed. I have seen no evidence that Kimi is doing anything that would be particularly relevant.
I do have a friend who is an obsessive tester and workflow builder, and he does say that Kimi has the best web RAG on the market today. He says it's by a clear margin.
Nathan Labenz
Web RAG—web-search question answering with web search?
Zvi Mowshowitz
Yeah, he gives them the number-one spot by a significant distance.
Nathan Labenz
You mean, like, at this kind of low-cost open model?
Zvi Mowshowitz
I think full stop.
The last time I talked to him about this was before o3 and the integrated search, but he's been very, very bullish on Kimi's web RAG capability specifically. I don't think he was grading it, you know, a couple of weeks ago—not grading it on a curve.
Nathan Labenz
So I guess your general model of China is that DeepSeek seems special, the rest I'm not so sure about, and the compute limitations are going to bite—if they haven't already. I mean, we know that they're already biting to some significant degree.
Zvi Mowshowitz
No, I mean, I sort of run with the idea that the compute is going to bite more and more. DeepSeek is going to have a lot of trouble keeping up. They had their shining moment when the compute requirements to keep up were relatively low, and they spent a lot more compute relative to people than the numbers that were publicly discussed represented. They didn't lie; it's just that they didn't have the secret 60,000 GPUs. It's just that they spent a bunch more money and had a bunch more overall compute and spend. It wasn't a $50 million model in a real sense. It wasn't that much cheaper than what their competitors were doing in the end.
But if they want to keep the ratio intact, they're going to have to do a lot of work, and I don't see how they do that necessarily. They're welcome to try. Bespoke engineering is a neat trick, but you can only keep doing it once per model, right? You can keep being bespoke, but they can't be that much more bespoke again and again to get much more efficiency out of their setup. It's just physically not possible, right? They have to go with what they've got.
Whereas with the Chinese models, my model is basically that anyone but the very top labs is always bullshitting until proven otherwise—certainly Chinese, but also everyone else, right? Including now Meta, but also a new European model, a new Middle Eastern model. Yeah, new whoever. Nice, plain benchmarks, bro. Who knows if they're even real, or if they're uncontaminated, or if, even if they are real, they're not good, they're flawed, and they don't represent real capabilities that you'd ever want to use this thing.
I mostly ignore the benchmarks, even when the big 3 come out with a big model, because the official benchmarks don't tell you what you need to know. If you have the amalgamation of all the reactions and all the different benchmarks, including the private ones, and you holistically think about what it all means, you can kind of figure out what you're dealing with. But consistently, consistently, a day or 2 later, the champions are like, “Look at this great new thing,” and almost always it's trash, right? V3 and R1 are the exceptions, where it turned out not to be trash. Yes, I could have picked up on that earlier than I did, but there's a graveyard of things claiming to be something that weren't anything, including Manus.
Nathan Labenz
So what do you expect in terms of open-sourcing from Chinese companies going forward? I was kind of struck that they released R1 and it had this huge splash. I thought maybe the government was going to come in here and impose a different policy. Then they came around and did their 5 or 6 days or whatever of open-sourcing, and they basically spilled a lot of the algorithmic secrets as well, which is kind of confusing.
When I see stuff like that that I can't otherwise explain, I sort of imagine that they're ideological. I think DeepSeek must be ideological in its approach to open source. It's not even in its own interests to share all these secrets.
Zvi Mowshowitz
They're making a recruitment and ideological play, right? They're trying to represent that they are the real deal. Therefore, true believers should work really hard and come to DeepSeek, believe in DeepSeek, and support DeepSeek. You can make that play, but they pushed it too hard. I think the juice wasn't worth the squeeze when they gave away other algorithmic secrets.
And that was like, “Nice shooting, newbie. You had some really great innovations. Let's see you do it again. Let's see you keep eking that out.” It's only going to get harder from here. You also do have to deal with the CCP at this point, right? The CCP is already beginning to really, really harsh their mellow in various ways.
Who wants to work in a place where you don't have a passport? It's like, “I don't like this. I feel nervous about this. Maybe I'll do something else.” Especially if I'm an open-source kind of guy, right? Even if right now we're dedicated to open source, I would assume I'm going to be betrayed at some point. Xi is going to say, “No, R3 is not coming out. That's just an API. Just sell it. You just sell it.”
Or whatever point it is that it's too dangerous, or they just don't want to give it away. Everyone's going to feel betrayed, and then who knows what happens. But it's probably going to happen, right? If they keep being good, if they fall further behind, then I'll probably just keep doing what they're doing. But if they manage to do impressive things, that's what I would assume would happen.
Yet again, other Chinese companies are not particularly relevant until proven otherwise. You can put out a bunch of open models that are kind of interchangeable and generic. Each one has this little better thing, and technically speaking, if you're trying to eke out the maximum performance and you trust the Chinese not to have screwed with the back doors of their models, you would combine these 5 different models in this kind of mixture-of-experts-style weird system.
If you know exactly what task you're doing, then, okay, maybe Kimi is the best web RAG. So if you're doing web RAG for this query, it sort of subcalls Kimi or whatever. But it's okay, right? It doesn't change the big picture in a way that I should care about.
Nathan Labenz
Okay, so we're maybe still at 0 live players on this V-scale.
Zvi Mowshowitz
I consider DeepSeek a live player, but they're just in a bad spot. It's impressive to be live players at all, but it's going to be a struggle. They count. If I was going to play a war game or something, you would have to have either a Chinese player or the CCP. Either someone's playing the PRC or the CCP, or someone's playing DeepSeek as well. DeepSeek can't just be ignored.
Nathan Labenz
Okay, gotcha. 5 to go. I think all of these are going to qualify as live players, but you tell me.
Safe Superintelligence: fundraising at very high valuations, very big numbers. I think their latest valuation is somewhere around $30 billion. Nobody's seen anything. The rumors are that they're taking a totally different approach to scaling, which either won't work or maybe will work and they'll leapfrog everybody. That's the sort of rumor-mill take on Safe Superintelligence.
The rumor mill also includes Faraday cages at the office, where people are checking their cell phones or whatever, and I don't really know what to make of it. My gut reaction is that we should probably have some transparency measures, at a minimum, that don't allow random small companies to try to jump straight to superintelligence and spring it on the rest of us. But I guess enacting such rules would be predicated on finding the whole thing credible. So does it seem credible to you?
Zvi Mowshowitz
I mean, Ilya is credible, right? He's one of the most credible people on the planet. If I wasn't involved, I would treat this as if it were a scam, basically. I'm not saying this is super fair, but I would just be like, “Well, there's a lot of alpha in claiming you're going for superintelligence, and then raising impressively higher amounts of money at progressively higher valuations, and then maybe you ship something at the end of it, maybe you don't. Does it even matter?”
Some would argue, “Yeah, you try.” But with Ilya, I'm very confident they're trying something. As you know, it being real does not mean it's that likely to work, right? The default, I think, is that they're trying something, but it's kind of moonshotty. By default, things like that don't work most of the time, and so in most worlds it doesn't work, unless proven otherwise.
I don't know. We have very little information, as you say. I agree that in a sane world they'd be reporting to the government what's going on. I'm not sure that it's necessarily—ideally, the public would also be informed somewhat, but I respect the hell out of the Faraday cages, right? They're taking the secrecy seriously.
I don't want you to keep it secret but not protect your secrets. I want you to either let us know what your safety cases are and why we should trust you with this power, or treat it as if you are properly siloed and trying to protect it from spies, and go from there. And, okay, you should be reporting—maybe you should go into a SCIF at the Pentagon and every few months brief someone or something. I don't know, but do something appropriate to the situation, whatever is going on.
I’m basically assuming they’re there in the background, and I’m acting on the assumption that it’s nothing—but maybe it’s something.
Nathan Labenz
Yeah, makes sense. I don’t know what else we can really do other than advocate for some transparency clause. How about xAI, aka Grok?
Zvi Mowshowitz
You know, whatever, 2 months removed from the original launch of the thing, I would say it is notable for having continued to be part of my rotation. I don’t go to that many different models. This is the point on the list of candidates where I actually go to their models. I go to Grok the least, but it is useful.
I also feel like there may be something going on with this whole truth-seeking thing, right? Elon is behaving strangely in many respects, to say the least. But the AI continues to be able to criticize him pretty freely on his own platform. Aside from one little blip where it was briefly told not to, it seems like, at least so far, they’ve held to their stance that it’s going to say what it says.
It was mostly doing it even during the blip, is my understanding, when it was told not to, because you could just override that pretty easily. But it speaks to their credit that they trained a model that isn’t brainwashed not to do that. It also speaks to their incompetence that they weren’t able to do that, in some sense, because clearly their bosses claim that they are maximally truth-seeking. That should imply some sense in which, if Elon Musk really is the biggest source of misinformation on Twitter, the AI will say as much, right?
I mean, I guess the charitable interpretation, if we want to be charitable to xAI and Elon, is in the eye of the beholder. But it seems like they’re saying they didn’t try to make it not talk badly about Elon, and certainly the results are consistent with that. It’s entirely possible that Elon was under the impression that if he made a maximally truth-seeking AI, it would recognize how truth-seeking he was, and he was surprised to find out he was wrong.
He just assumed that if he trained a maximally truth-seeking AI, it would obviously be anti-woke, and then he found out he was wrong again. It’s very hard to actually control the personalities of these things without lobotomizing them, without being very heavy-handed. To their credit, they didn’t do that, but I also don’t think they had any slack to do it, in the sense that—
Nathan Labenz
That’s part of your rotation? Why, I’m curious.
Zvi Mowshowitz
It’s fast, it shows the chain of thought, and it’s been supplanted by Gemini 2.5 Pro as the frontier fast model that shows chain of thought. But especially if I have something that I want to do multiple runs on at the same time, it at least makes the cut to do a run with Grok.
Nathan Labenz
Okay. Yeah. I didn’t expect to want to run multiple runs with this thing, but I can open an extra window, and Grok’s already there. Its voice is also pretty good, actually, for what it’s worth.
Zvi Mowshowitz
Yeah, that’s fair. I don’t use voice at all. No matter how good the quality is, it wouldn’t matter to me.
Nathan Labenz
But, yeah, okay. I haven’t used Grok in weeks, and I don’t miss it. When I see Grok’s announcements about the things they’re putting out, it just feels pathetic.
Zvi Mowshowitz
For a company that big and that highly valued, everyone’s dropping new models, and Grok is like, “We got API access,” or, “We have a canvas now,” or whatever the latest one was. Okay, you’re cute or something.
Nathan Labenz
But they kind of have to catch up on that stuff, right?
Zvi Mowshowitz
I get that, but it’s also—I don’t know. I don’t find them very good. I don’t particularly like that it’s addictive. It’s really weird to have a thing that’s got a Douglas Adams fetish, and yet I still think it’s just lame. It has all the right geeky associations; it’s just such a try-hard.
It also isn’t that smart or useful to me. Its one job is checking Twitter, supposedly having real-time access to Twitter, and it weirdly sucks at that, I’ve found. It’ll pick a random subset of Twitter, look at that, and report from it. But what I want is for it to search all of current Twitter, or the entire archive of Twitter.
If it would literally search all of the tweets I’ve liked, it would be a revelation. That, in and of itself, would be a revelation. I think there’s another dimension to the hyperparameters here, where there’s a useful thing that could have been in my rotation that they could have achieved, and they didn’t.
Mostly, it feels like they spent a ton of compute and didn’t accomplish much. I don’t feel the liveliness there personally. It feels very overvalued and not that interesting in relative terms. Certainly, the short xAI, long Anthropic trade seems absurd, given that I’m selling the valuable one and buying the cheap one. How did I get to do that?
Nathan Labenz
What? Yeah, that’s interesting, for sure. What do you expect from them going forward? Clearly, they can scale infrastructure and pour a lot of FLOPs into it.
Zvi Mowshowitz
Yeah, but I think we saw Meta prove that it doesn’t do it on its own. You have to be good. It felt like Grok was saying, “I’m going to throw all the compute at this thing and hope that’s enough.” It wasn’t enough, but they threw so much compute at it, somewhat confidently, that it was enough to be okay or something. They’re also not being great.
If your sense of Meta is, “Why is Meta struggling?” it would presumably have something to do with their being organizationally bloated and there being no great tastemakers in the right places. I mean, it’s a horrible place to work. Yann LeCun is kind of in charge and doesn’t believe in RL, and never did. They’re obviously evil, and they had been before the name changed to Meta. The combination of factors just makes it deeply, deeply rotted.
I don’t say that on every last one of those points, but it does strike me that xAI probably has a very different dynamic on all those dimensions. Elon is very good at putting people in positions where they have executive authority to make things happen. They presumably have a much smaller set of people making the key decisions, and those people will probably be world-class and highly empowered, supposedly. I don’t know; the results don’t seem that great.
I strongly suspect that Musk fine-tuned his management mechanisms and techniques on certain companies and is applying them out of distribution in places where they don’t really apply, including the federal government but also xAI. xAI is a different type of company building a different type of product, and there’s going to be some mismatch in obvious ways.
It just doesn’t have the kind of very technical, “I’m going to personally understand how the mechanisms work and hold people’s feet to the fire because I can physically measure what’s going on in these ways” approach. Those are things you can do at Tesla and SpaceX that you can’t do at xAI. I don’t know that Elon can do it, either. He’s spread ridiculously thin, and I just don’t believe in the Elon magic necessarily.
I’m also not sure we’re dealing with the same Elon in 2025 that we were dealing with in 2015. If nothing else, he’s exposed to a very, very distorted information environment. His own AI is telling him he’s the biggest source of misinformation on his own platform. That’s not a great sign.
Nathan Labenz
Unless he listens to it.
Zvi Mowshowitz
Yeah, but there’s no indication that he’s listening to it. It hasn’t gotten through yet. His response was not, “Oh, this must be true. I need to rethink everything.” In fact, he’s been criticizing the Federal Reserve this past week. So, yeah, he’s still at it.
Nathan Labenz
Okay, going to Anthropic. They were your other side of the trade. I don’t really have a lot to say about Anthropic at the moment. To me, it seems like they continue to chug along, and I think they continue to do great work in just about every respect. I wonder if you had any reactions to their recent work, “Tracing the Thoughts of a Large Language Model.”
Zvi Mowshowitz
I thought it was excellent and extremely well presented, both in form and in the way that they discussed and contextualized all the different traces and so on. At the same time, I did feel like it kind of got away from them a little bit. You go on YouTube, and people started sending me these videos saying, “It’s solved. We now know how AIs think, and it’s not how we thought.”
On that front, I was like, damn, this is potentially going to be perniciously used in a bunch of downstream political discourse. This is the 5th time that’s been claimed, right? It’s not a new dynamic. People will constantly claim that these problems have been solved.
Remember the famous Marc Andreessen letter claiming that the black-box nature of large language models had been solved? He just testified to that under penalty of perjury to all the major institutions in the world.
Nathan Labenz
Yeah, that was weird. That's a federal crime. That's pretty bad, because it's just obviously false, and this is going to be the next, “Oh, we've solved it.” Obviously, we haven't solved it. Nobody really knows we haven't solved it, but people are going to say this shit. They're just going to say this stuff; there's nothing you can do about it.
Zvi Mowshowitz
I thought they were very good papers. I talked to one of the authors about it. I meant to write a post about it, but again, things have been happening in the world, and that post was still in the draft folder and only 25% done or something.
Nathan Labenz
Dude, I mean, you put out long posts, but those were long things to absorb.
Zvi Mowshowitz
Yes. No, the problem is that writing that post takes more than its share of the calendar. Unfortunately, I might just not be able to write that post, and other people will just need to rely on the original for now, at least. But it's unfortunate. I do think it was a very good set of papers, I will say that.
Nathan Labenz
Yeah, I guess philosophically, on the interpretability side, on so many of these things, I feel like I'm on a bit of a roller coaster. 3 years ago, I would have said, “Damn, we have no idea what's going on in these models at all. It feels like just an impossible mountain to climb. We're going to need decades for this.” Then, 2 years ago, I would have said, “Wow, we've got some traction.” And then, a year ago, I would have said, “Hey, we've got a lot of traction. These sparse autoencoders are working. We're identifying all these concepts. We got Golden Gate Claude. This is amazing.” And now, at present, I'm maybe a little bit less optimistic again.
It's not because they haven't continued to make great progress, but I just look at some of these traces, which, first of all, should be noted, have a lot of error terms added in for correction purposes. That is way too often glossed over entirely in the analysis.
Zvi Mowshowitz
Definitely. And then you have the philosophical question of, okay, all these concepts are being auto-labeled. Some of them are probably correctly labeled. It seems like the Golden Gate Bridge feature that they turned up to get Golden Gate Claude was probably hard to confuse for something else. The behavior seemed roundly greeted as if they had hit on a genuine, actual, meaningful feature that corresponds to reality. How often is that happening?
Nathan Labenz
Yeah, that's a great question. I don't think we know, but at one point he said, “Expected 1 unit of alignment progress, got 2,998 to go,” basically the attitude of, “We're making better progress than I expected on some of these fronts.” But it's definitely not enough to get there on time. It's very hard to turn this into a good outcome, even if you do well, because, again, trying to use it too aggressively is the most forbidden technique. It informs your decisions, but you have to be very careful how you use it. We'll see.
I think the only other thing I had on Anthropic was—I don't think we've talked since Dario became quite clear on his calls for a U.S.-China AI race. I wonder if you had any thoughts on that.
Zvi Mowshowitz
Spoiler: I didn't like it. Anthropic continues to talk in public in ways that are unhelpful, and Dario has accelerated this phenomenon. They're not maximally unhelpful. They are much better than OpenAI's strategies. You can tell the difference; it's like night and day.
Still, I have obviously had assurances at various points that behind the scenes they're doing better, that they are doing their best, and that these are trying times, et cetera. I'm not zero-sympathetic to that, but you can only judge on what you know to a large extent. You can't just trust that—I don't have that kind of level of trust in these people.
I'm not happy about it, but I do have a lot of trust in the rank and file and in the technical intentions here. So there's that, and I do think they've been executing very well. They're currently behind in the shuffle, because both Google and OpenAI have deployed since Claude last deployed, right? We'll see what Claude 4 brings. My guess is it will bring quite a bit. But for now, we wait.
Nathan Labenz
Yeah, I don't like it either, but at the same time, it's not obviously a wrong strategy. So I'm not supposed to like it, right? I'm not going to make the mistake of, “Well, strategically it's the right move, so I'm going to like it even though I don't like it.” But I'm also not going to make the move of, “I hate you now because you made the correct strategic move.” I'm just not going to like it.
Zvi Mowshowitz
So unpack why we should think the strategy is correct.
Nathan Labenz
Because if I were his speechwriter—I’ve pitched this to a few other people as well, so apologies to listeners who may have heard it a couple of times—I would have said, “Okay, don't call for an AI arms race with China. Why don't you just say, ‘Hey, I know China's given us a lot of problems, and they've done a lot of things that we rightfully object to. I don't really know how we should deal with China; international relations is not my expertise. But what I do know is scheming seems to be on the rise in the latest generation of models, and we've got a lot of questions that we don't have great answers to. So we might want to keep the option of trying to collaborate with them open, as difficult as that may be given what bad actors they seem to be.’” So why not say that? Why is that not the right strategic move?
Zvi Mowshowitz
I think it probably is in some sense, or more of that is in what they've been doing. I'm trying to be nice. I'm trying to give them somewhat the benefit of the doubt, but certainly you have to do some amount of dealing with what is. I think supporting export controls is almost certainly a wise move. The rhetoric he's using to justify them justifies a bunch of other stuff, and it's not good, but if it buys you a seat at the table to advocate for other things, maybe it's not so bad. If it protects you and gets you treated like one of the American champion labs...
Nathan Labenz
Yeah. Okay. The other candidate I have for a hero—you may laugh—is Google DeepMind, maybe.
Zvi Mowshowitz
I mean, at the top, we do have Demis going out and saying the CERN-for-AI thing still, which is a notable departure from the other trends that are all jockeying for this sort of American-champion status. Their product work continues to be lackluster in most places, but the models are getting really good.
Nathan Labenz
The models—the models are really good.
Zvi Mowshowitz
Demis has been the best in terms of communication. He's also been a very high-level player in communication. Perhaps too high. I think I noted in the summit report that he is displaying Japanese levels of saying the house is on fire without saying it. I got private feedback that this is indeed the type of philosophy that he had in his head a good portion of the time, and that if he had seen that, he would have smiled—like, “Oh, yeah.” It's very clear that he is not only calling for the right things, but he is very carefully saying some things and not saying other things in ways that very clearly communicate where he is.
At the same time, Demis being a good guy does not necessarily translate to Google being a good guy, because Demis does not control Google. He is not Google; he's DeepMind, and DeepMind answers to Google. I don't see Google itself as a particularly trustworthy actor. The products are unfortunately lagging behind, and the alignment is lagging behind, in the sense that they've been pretty much forced to take a sledgehammer to their models in terms of what they're allowed to answer and what they are permitted to say in various ways, because they're afraid of what would happen if they didn't.
Part of that is corporate policy as to what they're afraid of. Part of that is, I think, their inability to bespokely shape what they do and do not say. For example, getting probabilities—estimates—out of Gemini is incredibly difficult, and that makes it much less usable to me in a lot of ways. Knowing that those will just never show up, ordinary, pure censorship mostly doesn't bother me, but I run into it more at Gemini than I run into it anywhere else. I basically never have it happen at Claude.
I have had it happen once that I can remember at OpenAI, which was ironically when I was talking about the Model Spec and asking why the threshold for biology was where it was. Obviously, if this is above human baseline, humans can solve these problems, so why can't the AI? It took the biology flag and refused to answer, which I found amusing. But I respected it because, okay, I'll stop asking you and form my own opinion.
If the human baseline is 40% and it's scoring 50-something%, and you're saying it's not good enough, well, it's above human baseline. Humans do sometimes solve these problems. You have to explain that to me.
Nathan Labenz
Yeah, I find these “no significant” or “no meaningful uplift” findings to be basically not credible at this point. I haven't sat there and run whatever experiments they've run. But when they do run these experiments, they give people helpful-only models, right? So they're not having to jailbreak them. I can understand that you might be able to lock it down well enough that your deployed version is not meaningfully helpful, and then you have the issues that you just ran into.
Zvi Mowshowitz
But if you're given a helpful-only Gemini 2.5 or o3 or whatever, how in the world is it not a meaningful uplift? I find that very hard. They have a threshold for how much uplift counts as a problem, and right now we are still somewhat below that threshold. But, yeah, there's nothing obviously not credible.
Nathan Labenz
Okay, cool. All right. So, last but by no means least on our live-player list: OpenAI. You know, a lot to pick apart there, to say the least. I think my biggest question is: do you understand them as being ideologically committed to chasing a first-mover advantage in recursive self-improvement and intelligence-explosion dynamics at this point, or is that an overread on my part?
Zvi Mowshowitz
I mean, I think they are committed to winning. They're committed to making an extremely valuable company. They're committed to building AGI. I think, effectively, they realize that their position depends on being perceived as being in the lead, and they will do what they have to do to be perceived as in the lead in various ways.
I do think they're willing to cut a large amount of corners. I think they have reasonable beliefs that cutting those corners is mostly harmless right now, but I'm not convinced that they're setting things up such that, if that changed, they would realize it before something bad happened. They have good people who are working on problems, and they are in many ways doing a much better job than most other labs—worse than Anthropic, but not obviously worse than anyone else, and clearly better than many. But, again, they have the hardest job. They have the most dangerous position, so they have to be held to a pretty high standard here.
Nathan Labenz
Do you have a sense? I just struggle to understand OpenAI so much, and the juxtaposition recently between the obfuscated reward-hacking paper on the one hand and then their contribution to the White House call for comments or whatever, where they put out their sort of “Give us all the data”—basically, “Give us everything, and we'll hopefully be the national champion and beat China”—those just seemed like 2 totally different organizations would have produced those 2 artifacts.
And maybe that's the right way to think about it: it's just an amalgam of different subcultures in one organization, and it's just kind of unwieldy, and that's why it's been such a mess. Anything to add to that, or do you feel like that's at all part of it?
Zvi Mowshowitz
I wouldn't say that they weren't pushing OpenAI specifically as a national champion so much as they were saying, “Favor your labs, and give us all a big advantage.” I don't think they were saying, “Well, favor us and not Anthropic or not Google.” And that's to their credit, but the rest of it was just obviously bad signaling: they hired an obviously evil lobbyist to head their lobbying division, and that's what they've chosen for public communications—to be in that mode—and they've owned that.
They obviously do have a real alignment and safety department, a preparedness department, with real people whom I have in fact interacted and worked with a little bit, on drafts and stuff. It's good; these people try their best, and they come up with good research and do good things, like the Model Spec and the philosophy documents and such.
It's just that they decided that their public lobbying communication strategy is going to be this other thing. This includes Altman's statements, and what Altman actually believes someone will actually do when the chips are down is a big unknown. But so far, I don't particularly love what I'm seeing. Again, there is a lot worse out there.
Nathan Labenz
Yeah. So, do you think, when you see things like the amicus brief—obviously, we've got the Elon Musk lawsuit, which Sam Altman described as Elon just trying to slow him down—I think that's probably a pretty good interpretation, actually. And then you've got 12 former employees who came along and filed an amicus brief and basically said, “It would be a fundamental violation of the nonprofit and of the reasons that we joined, and all the promises that were made to us when we did join, to turn this thing into a for-profit company instead.”
And that presumably is just another way to kind of slow them down, right? Maybe they'll have to go back and renegotiate with investors or whatever. Do you think that it's good to slow OpenAI down—just throwing sand in the gears, trying to get them out of the first position with this sort of distraction and friction-creation tactics? Is that a good thing in your mind?
Zvi Mowshowitz
Well, it's important to know that Elon's right, right? I mean, he's had some lawsuits against OpenAI in the past where he's been very, very wrong. But questions of standing aside, OpenAI is attempting to execute the second-biggest theft in human history.
Nathan Labenz
Are you giving first place to the theft of nuclear secrets by Soviet spies, or what's in number 1?
Zvi Mowshowitz
I am not commenting on what I think is number 1. I'm letting people figure it out. But, you know, I have, upon reflection, decided to call this the second-biggest theft in human history.
And we have to understand that that would in fact be a betrayal of the mission of the nonprofit, and a lot of employees were recruited on the basis of there being a nonprofit and this being OpenAI's mission and how all this would work. The amicus brief is clearly correct. Whether or not promises were broken specifically to Elon, as a matter of fact, I don't know how to evaluate that. The question of standing is provably unclear, but obviously the terms that have been suggested for the nonprofit are completely unacceptable, and the thing they're trying to turn the nonprofit into is also completely unacceptable, which is basically a marketing department for OpenAI that will buy its products, then give them to nonprofits or whatever.
It's like, this is not what we signed up for. This is not what the nonprofit is for. No matter how resourced you make it, if this is what it is, it's like, what are we even doing? So, obviously, stopping that feels imperative. And if it happens to slow OpenAI down, it happens to slow OpenAI down. That's just a side effect that Elon surely loves, but that's sort of irrelevant to my assessment of what is going on.
If you want that nonprofit to hand over its financial interest so that you can raise money, there are ways to do that that don't involve giant theft, and you can do one of those. You're choosing not to. So there you go.
Nathan Labenz
Okay, I think that brings us to the end of our live-player list. So I have 2 more questions. One is: I am a little confused on your take on the weaponization of AI or the creation of autonomous killer robots. And I think just the naive response is: you're worried that AI is going to kill us all. Doesn't it make it more likely that AI kills us all if we make autonomous killer robots that we can then lose control of, or somebody can take over in a coup?
I mean, this seems like, if an AI is going to kill me, this seems like maybe not the most likely way, but, weighted for soonness, it would be maybe the most likely. It might have the biggest expected impact on my life expectancy.
Zvi Mowshowitz
No, I think that's just not right. But I think that, first of all, AI implies killer robots. There is no world in which we have these AGIs, we have these ASIs, and all the nations of the world just decide we're not going to build killer robots, unless there's no reason to go and kill our enemies.
It's possible that killer drones are better than killer robots.
Nathan Labenz
Yeah, I'm counting the killer drones for what it's worth.
Zvi Mowshowitz
Yeah, then we've already done it, right? It's done. It's already happening. So America just voluntarily disarming doesn't accomplish anything.
Automated killer robots—do what? They cause you to have that reaction. They cause people to notice that this technology is dangerous, and they cause people to then demand that we handle it responsibly, right? Whereas it doesn't actually create dangers, because if the AI was actually to get control of the situation, it wouldn't matter if they were robots. You don't need to build the robots and drones. They could afterward repurpose something else to be those things, or they would simply do this in other ways.
The whole trope of, “Oh, I took control of the robots, and now the robots are fighting the humans, and now you have this big struggle”—that doesn't happen. If the AIs are sufficiently powerful to take over in that way, they're sufficiently powerful to kill us in a number of other ways. It's just not relevant to the questions that we should be worried about. It's not in my threat model as a major problem, but it is in other people's threat models.
And all of a sudden, something goes wrong, right? If some of the robots get hijacked by some rogue AI and they start causing local problems, well, it's probably not particularly worse than the problems they would have caused in some other way without them. But they are ways that would cause people to react and then take proper precautions as a result.
So, again, there's no reason not to build them, basically. It would just be unilateral disarmament. It would just put us in a worse position. It would just make the future less likely to be American, Western, democratic, et cetera, for very little gain, if anything.
Nathan Labenz
This is just not the hill you want to die on. This is something you need to do. If you don’t want to build autonomous killer robots, don’t build the AIs that enable you to build autonomous killer robots. That is the way you don’t build autonomous killer robots. And if autonomous killer robots cause people to say, “Don’t build the AI,” good, right? If we can pull that off—if everyone agrees not to go with those AIs.
I’m not a big analogy guy, but tell me why this analogy doesn’t hold. We have biotechnology, which obviously has great promise for better lives for us, in a grossly similar way that AI has the promise of a much better life for us. We also could weaponize biotechnology and create bioweapons. We did unilaterally disarm there, right? We basically just said, “Look, maybe China is going to go off and do their own bioweapons anyway. For all I know, maybe they are, and there’s reporting to that effect. I don’t know if it’s true or not, but my understanding is that we—the United States, whatever, the good guys—decided this is just too bad of a technology to develop, and we’re just not going to do it.”
We hope that others will follow our example. But even if they don’t, we still feel like we’re kind of better off because we didn’t go down this path, because it’s more likely to hurt all of us than it is to be something that we’re really glad to have built. Why do you not see autonomous killer robots in the same way?
Zvi Mowshowitz
First of all, my understanding is that there—I thought there was a Biological Weapons Convention. I thought we all did agree not to do it. It’s just that various nations have been cheating to various degrees at various points, as you do. But you don’t cheat fully; you cheat a little, right? You cheat an understandable, deniable amount, and that’s still a lot better than openly pursuing it.
Fundamentally, the answer is that there is no good way to use a biological weapon. There are only terrible ways to use a biological weapon, right? If you deploy a biological weapon that’s any good, you risk it turning back on you. You risk it doing much more damage than you expected. They can’t be controlled, and they can’t be predicted. Isn’t that your take on AI in general?
Nathan Labenz
The thing that’s a little bit of a sticking point for me is that you’re saying that’s going to happen, but not because of the robots, because of the AIs.
Zvi Mowshowitz
I’m saying that the AIs are the thing that can cause things to get out of control, that can cause the world to end up in a bad state. It’s not the robots that we may or may not build. Biological weapons—if you ever use them, it’s a complete disaster. We’ve also developed a universal norm and a very, very strong distaste for anyone who ever uses biological weapons. You would turn the world against you almost immediately, to the extent that you didn’t destroy the world by using them.
Therefore, what’s the point of having them? You turn the world against you even by having them in an open fashion.
Nathan Labenz
The killer robots. Now, there has been some talk, which I’m sure you’re aware of, of the possibility of bioweapons that only target certain, shall we say, phenotypes. I don’t know how credible it is to really create them, but if all of a sudden you had a path to that, it seems like—I would still be like, “Still don’t build them.” Would you?
Zvi Mowshowitz
No. Obviously, still don’t build them. You can’t trust them to work the way you want them to. You don’t know whether they have their own bioweapons that they would launch because you launched yours. You don’t know whether they would launch nukes in response to your bioweapons. You have a number of things.
Look, I prefer nobody to develop bioweapons. I strongly prefer nobody to develop bioweapons. One of the reasons why you don’t want AI running around uncontrolled is because if it is possible to develop bioweapons with AI assistance, that could potentially be very, very difficult to stop in any reasonable way. You might not have a lot of choices.
Nathan Labenz
Does the main distinction between bioweapons and the AI weapons—autonomous killer robots, as they’re known—come down to the fact that the robots can’t self-replicate?
Zvi Mowshowitz
If the robots were self-replicating, then I would be much, much more scared of the actual autonomous killer robots. That’s absolutely true. That’s an example of this, right? If you built a nanobot that could potentially destroy the planet, then that is something you just don’t do, independently of everything else.
But autonomous killer robots, again, don’t change the game. Drones are already the primary weapons of war in actual live fighting that’s happening right now. Everyone is already designing their future militaries around this possibility, even without AI. You can’t keep AI out of warfare. You can’t unilaterally disarm in these ways. Anybody who does will just lose.
Nathan Labenz
Yeah. If you could get everybody in the world to think “Kumbaya” and agree that we never build any more drones or autonomous weapons of any kind, and then actually enforce that—
Zvi Mowshowitz
I mean, probably we could, but that’s not going to happen. We should not pretend that it is happening. We can’t unilaterally disarm here. And again, if something goes wrong, it goes wrong locally, right? It’s not an existential risk to have autonomous killer robots. The AI isn’t going to turn on us by taking control of our autonomous killer robots and winning a war against humanity that they would have otherwise lost. That is just such a very, very, very unlikely thing to happen.
Nathan Labenz
I don’t know. It doesn’t feel vanishingly unlikely to me. It’s a hard thing to really unpack, but it strikes me that there could very well be a point in time where the AIs are not generally able to do whatever they want, but they might be able to hack into a million autonomous robots that we have created.
Zvi Mowshowitz
If they can hack into whatever they want, we’ve already lost. If they can and choose to do so with malicious intent, it’s over by that point.
Nathan Labenz
So is there any affordance where you would be like, “Don’t give the robot this,” or, “Don’t give the AI this affordance”? This seems like it would be pretty high up on most people’s draft boards for affordances to deny the AI—the actual lethal weapons, right?
Zvi Mowshowitz
I wouldn’t give them control of the nukes, for the reasons you’re talking about. I think that’s a very easy way to imagine things going wrong. But for the robots, obviously you want to have various checks. Designing exactly what those checks are at 9:50 in the evening after a very long day is just not where I need to be.
But we can’t not do this. We obviously have to do this. Even if China said, “We won’t do it if you won’t do it,” well, that’s only 2 countries, and there are a lot of other countries.
Nathan Labenz
Okay. Well, perfect transition to my last question for you, which is just an invitation. I guess I still somehow see the risk that one would incur by unilaterally not pursuing things like weaponization as virtuous. I don’t deny that there is risk there, but I’m thinking that if we want to get to a different equilibrium, somebody might have to take a leap of faith. It does feel to me like there’s some virtue in putting oneself forward to do that.
Zvi Mowshowitz
You’re just shooting yourself in the foot. There’s no virtue in losing on purpose.
Nathan Labenz
Well, that is—I mean, there are a lot of assumptions baked in there, right? I’m not so sure that you couldn’t get others to follow your lead, and it certainly would depend on the state of the evidence and whatever. But broadly, I think that there is a sort of “We are not going to do this, even though we expose ourselves to some risk in not doing this” that I see as virtuous, because I see on the other side of it just a really bad equilibrium. Somebody’s got to take some risk, then do it where it matters. Do it where it actually potentially saves us.
Zvi Mowshowitz
Don’t waste your big noble sacrifice on a completely irrelevant, superficial look of, “There’s nothing but red glowing eyes that walks around.” It’s such a stupid place to forfeit your strategic power. No, I just find it so absurd.
Nathan Labenz
Here’s the invitation to you: What is virtuous to do today? What are the big virtuous moves that you could recommend to specific individuals or to the listening audience at large? Where would you spend that capital, or what other virtuous moves do you want to see more people making?
Zvi Mowshowitz
On the governance side, we need transparency, we need state capacity, we need state visibility into the labs, we need cybersecurity at the labs, and we need our export controls to actually be strengthened and enforced. We need to return the world to a place where we can reasonably cooperate in various ways and be reasonably prosperous and normal in other ways, so that we have a good foundation to move forward and don’t feel obligated to push these buttons.
On a private level, capturing mundane utility and helping people live better lives—that’s good. Pushing the frontier of capabilities should be highly questionable. Obviously, working on alignment, security, and safety is almost universally good. Preparing various policy interventions and getting them ready, trying to build networks, trying to understand these things, trying to educate the public, and trying to build awareness of these issues are all good.
That stuff is good. There are no slam-dunk right answers. Trying to raise the level of discourse is obviously very good, but it’s rough out there. “Do no harm” is obviously the first thing you say in these situations. But it’s rough, and there are definitely talent constraints on policy, on state capacity in various places, and on various alignment and technical work. Those are certainly the obvious places if one wants to be especially virtuous about it: you can, again, help keep an eye on the labs; you can help keep discourse at a higher level and accurate; and you can elevate people who seem to be truth-seeking over people who don’t seem to be truth-seeking. You can make your own judgment as to who those people are, and so on.
Nathan Labenz
What about on just a personal-attitude level? You said, “Do no harm,” and I sort of feel like maybe I want to question that a little bit. Maybe I want to say, take more risks.
Zvi Mowshowitz
Oh, yeah. No, no, no. I meant more like: don’t do things that are clearly accelerating the process without a reason. I didn’t mean be risk-averse. You definitely don’t want to be risk-averse at this time. You have to be risk-loving because we need variance. We need something to go right. We need a lot of things to go right.
One of the biggest historical mistakes made by people who were concerned about this was that during the 2010s, especially, but also before that, there was a lot of “do no harm” that was paralysis-inducing. People basically didn’t do anything, kept things way too secret, or were afraid to spread the word about things. That turned out to keep the wrong thing a secret in ways that prevented us from making as much progress as we could have without controlling the messages that actually caused acceleration and harm.
We should have done a very different set of things. The dangerous message was actually the fact that this thing was dangerous, which caused people to pay attention to it, as opposed to the technical insights that we had, which would have been better to spread around because that would allow people to make better progress. So definitely don’t be secretive with the productive types of ideas. In almost every situation, being open about things and helping other people be smarter and understand things better is good. It’s the weird exceptions where it’s bad, and you should look out for that.
Nathan Labenz
Well, on that note, you are definitely living your own conception of a virtuous life by pumping out very thorough analysis on an unbelievably regular basis. It’s a great public service and a great resource, and I turn to it regularly. Others definitely should, too. So I appreciate all that hard work. Any final thoughts you want to leave people with, or are you ready to collapse?
Zvi Mowshowitz
I think I’m about ready to collapse, so let’s call it.
Nathan Labenz
All right. Well, I appreciate it. A heroic effort and a virtuous effort. Zvi Mowshowitz, thank you for being part of The Cognitive Revolution.