Erik Torenberg
AI is not synonymous with language models. AI is being developed with pretty similar architectures for a wide range of different modalities, and there's a lot more data there. Feedback is starting to come from reality. Maybe we're running out of problems we've already solved. When we start to give the next generation of the model these power tools and they start to solve previously unsolved engineering problems, I think you start to have something that looks kind of like superintelligence.
Nathan, I'm stoked to have you on the a16z podcast for the first time. Obviously, I've been podcast partners with you for a long time, with you leading The Cognitive Revolution. Welcome.
Nathan Labenz
It's great to be here. Thank you.
Erik Torenberg
We were talking about Cal Newport's appearance on The Lost Debate, and we thought it was a good opportunity to have this broad conversation and really entertain the question: Is AI slowing down? Why don't you steelman some of the arguments that you've heard on that side, either from him or more broadly, and then we can have this broader conversation?
Nathan Labenz
I think, for one thing, it's really important to separate a couple of different questions with respect to AI. One would be: Is it good for us right now, and is it going to be good for us in the big picture? I think that's a very distinct question from whether the capabilities we're seeing are continuing to advance at a pretty healthy clip.
I actually found a lot of agreement with the Cal Newport podcast that you shared with me when it comes to some of the worries about the impact AI might already be having on people. He looks over students' shoulders and watches how they're working, and basically finds that he thinks they're using AI to be lazy, which is no big revelation. I think a lot of teachers would tell you that.
He puts that in maybe more dressed-up terms: People aren't necessarily moving faster, but they're able to reduce the strain that the work places on their own brains by trying to get AI to do it. If that continues—and I think he's been a very valuable commentator on the impact of social media—we should all be mindful of how our attention spans are evolving over time and whether we're getting weaker or averse to hard work. Those are not good trends if they're showing up in oneself. I think he's really right to watch out for that sort of stuff.
As we've covered in many conversations in the past, I've got a lot of questions about what the ultimate impact of AI is going to be, and I think he probably does too. But when it comes to the capabilities, it's a strange move from my perspective to go from, "There are all these problems today and maybe in the big picture," to, "But don't worry, it's flatlining." It's like, worry, but don't worry, because it's not really going anywhere further than this, scaling has kind of petered out, and we're not going to get better AI than we have right now.
The most easily refutable claim from my perspective is that GPT-5 wasn't that much better than GPT-4. That's where I really was like, "Whoa, wait a second." I was with you on a lot of things, and some of the behaviors he observes in the students, I would cop to having exhibited them myself. When I'm trying to code something these days, a lot of times I'm like, "Oh, man, can't the AI just figure it out?" I really don't want to have to sit here and read this code and figure out what's going on.
It's not even about typing the code anymore; I'm way too lazy for that. It's about figuring out how the code is working. I just want to say, "Make it work. Try again," and I'll keep trying again. I do find myself at times falling into those traps. But a big part of the reason I can fall into those traps is that the AIs are getting better, and increasingly it's not crazy for me to think that they might be able to figure it out. That's my first slice at the takes that I'm hearing.
Erik Torenberg
There's almost a 2×2 matrix one could draw up where it's: Do you think AI is good or bad now and in the future, and do you think it's a big deal or not a big deal? I think it's both on the good and bad side, and I definitely think it's a big deal. The thing I struggle to understand most is the people who don't see the big deal. It seems pretty obvious to me, especially when it comes to the leap from GPT-4 to GPT-5.
Maybe one reason that happened a little bit is that there were just a lot more releases between GPT-4 and GPT-5. What people are comparing it to is something that came out a few months ago, like o3, which only came out a few months before GPT-5. Whereas with GPT-4, it was shortly after ChatGPT, and it was all this moment of, "Whoa, this thing is exploding onto the scene." A lot of people were seeing it for the first time.
If you look back to GPT-3, there's a huge leap. I would contend that the leap is similar from GPT-4 to GPT-5. These things are hard to score. There's no single number you could put on it—well, there's loss—but one of the big challenges is figuring out what exactly a loss number translates into in terms of capabilities.
It's very hard to describe exactly what has changed. But we could go through some of the dimensions of change, if you want to, and enumerate some of the things that people are starting to, or have come to, take for granted and kind of forget. GPT-4 didn't have a lot of the things that were now expected in the GPT-5 release because we'd seen them in GPT-4o, o1, o3, and all those other releases. Those things may have boiled the frog a little bit when it comes to how much progress people perceived in this last release.
A couple of reactions. One is—and even to complicate your 2×2 further—is the question of whether it's bad now versus whether it's bad later. Cal, whom we both admire a lot, by the way, is a great guy and a valuable contributor to the thought space. But he's not as concerned about these sorts of future AI concerns that the AI safety folks and many others are concerned about. He's more concerned about what AI means for life, cognitive performance, and development now, in the same way that he's worried about social media's impact.
You think that's a concern, but nowhere near as big a concern as what to expect in the future. Cal also presents a theory of why we shouldn't worry about the future because AI is slowing down. Why don't we share how we interpreted his history? As I understood it, the simplistic version is that we figured out a way such that, if you throw a bunch of data into the model, it gets better by an order of magnitude. That's the difference between GPT-2 and GPT-3, and then between GPT-3 and GPT-4.
The difference was significant, but then it achieved significantly diminishing returns, and we're not seeing that in GPT-5. Thus, we don't have to worry anymore. How would you edit that characterization of his view of the history? Then we can get into the differences between GPT-4 and GPT-5.
Nathan Labenz
The scaling-law idea is definitely worth taking a moment to note. It is not a law of nature. We do not have a principled reason to believe that scaling is some law that will continue indefinitely. All we really know is that it has held through quite a few orders of magnitude so far.
I think it's really not clear yet to me whether the scaling laws have petered out or whether we've just found a steeper gradient of improvement that's giving us better ROI on another front that we can push on. They did train a much bigger model, GPT-4.5, and that did get released. There are a number of interesting things about it. Of course, there are a million benchmarks, but the one I zero in on most for understanding how GPT-4.5 relates to both o3 and GPT-5 is SimpleQA.
OpenAI is, famously, terrible at naming. We can all agree on that, I think. A decent amount of this confusion and disagreement actually stems from unsuccessful naming decisions. On one benchmark called SimpleQA, which is really just a super-long-tail trivia benchmark, GPT-4.5 performed at about 65%, compared with about 50% for the o3 class of models.
SimpleQA measures whether you know a ton of esoteric facts. They're not things you can really reason about; you either know these particular facts or you don't. In other words, of the things that were not known to the previous generation of models, GPT-4.5 picked up a third of them. There's obviously still two-thirds more to go, but I would say that's a pretty significant leap.
These are super-long-tail questions. I would say most people would get close to 0. You'd be like the person sitting there at trivia night who maybe gets 1 question right in a night. That's what I would expect most people to do on SimpleQA. It checks out: The models obviously know a lot more than we do in terms of facts and just general information about the world. At a minimum, you can say that GPT-4.5 knows a lot more.
A bigger model is able to absorb a lot more facts. Qualitatively, people also said that, in some ways, maybe it’s better for creative writing. It was never really trained with the same power of post-training that GPT-5 has had, so we don’t really have an apples-to-apples comparison, but people still found some utility in it.
I think maybe the way to understand why they’ve taken that offline and gone all in on GPT-5 is just that that model’s really big. It’s expensive to run. The price was way higher—a full order of magnitude plus higher than GPT-5’s—and it’s maybe just not worth it for them to consume all the compute that it would take to serve that. Maybe they just find that people are happy enough with the somewhat smaller models for now.
I don’t think that means that we will never see a bigger GPT-4.5 model with all that reasoning ability. I would expect that would deliver more value, especially if you’re really going out and trying to do esoteric stuff that’s pushing the frontier of science or whatever.
In the meantime, the current models are really smart, and you can also feed them a lot of context. That’s one of the big things that has improved so much over the last generation. When GPT-4 came out, at least the version that we had as public users was only 8,000 tokens of context, which is about 15 pages of text. So you were limited. You couldn’t even put in a couple of papers without overflowing the context.
This is where prompt engineering initially became a thing. It was, “Man, I’ve really only got such a little bit of information that I can provide. I’ve got to be really careful about what information to provide, lest I overflow the thing and it just can’t handle it.”
As context windows got extended, there were also versions of models where they could nominally accept a lot more, but they couldn’t really functionally use it. They could fit it at the API-call level, but the models would lose recall, or they’d sort of unravel as they got into longer and longer context.
Now you have much longer context, and the command of it is really good. You can take dozens of papers on the longest context windows with Gemini, and it will not only accept them, but it will do pretty intensive reasoning over them with really high fidelity to those inputs.
That skill, I think, does kind of substitute for the model knowing facts itself. You could say, “Geez, we’re going to need a trillion or, who knows, 5 trillion—however many trillion—parameters to fit all these super-long-tail facts.” Or you could say, “Well, a smaller thing that’s really good at working over provided context can, if people take the time or go to the trouble of providing the necessary information, access the same facts that way.”
So you have a choice: Do I want to push on size and bake everything into the model, or do I want to try to get as much performance as possible out of a smaller, tighter model? It seems like they’ve gone that way, basically because they’re seeing faster progress on that gradient.
In the same way that the models themselves are always, in the training process, taking a little step toward improvement, the outer loop of the model architecture, the nature of the training runs, and where they’re going to invest their compute are also going in that direction. They’re always looking at, “Well, we could scale up over here and maybe get this kind of benefit, or we could do more post-training here and get this kind of benefit.”
It just seems like we’re getting more benefit from the post-training and reasoning paradigm than from scaling. But I don’t think either one is dead. We haven’t seen yet what GPT-4.5 with all that post-training would look like.
Erik Torenberg
Yeah. One of the things that you mentioned that Cal’s analysis missed was that it way underestimated the value of extended reasoning. What would it mean to fully appreciate that?
Nathan Labenz
A big one from just the last few weeks was that we had an IMO gold medal with pure reasoning models, with no access to tools, from multiple companies. That is night and day compared to what GPT-4 could do with math.
These things are really weird. Nothing I say here should be intended to suggest that people won’t be able to find weaknesses in the models. I still use a tic-tac-toe puzzle to this day where I take a picture of a tic-tac-toe board in which one of the players has made a wrong move that is not optimal and thus allows the other player to force a win. I ask the models if somebody can force a win from this position.
Only very recently—only the last generation of models—are starting to get that right some of the time. Almost always before, they were like, “Tic-tac-toe is a solved game. You can always get a draw.” They would wrongly assess my board position as one where the player could still get a draw.
So there’s a lot of weird stuff. The jagged capabilities frontier remains a real issue, and people are going to find peaks and valleys for sure. But GPT-4, when it first came out, couldn’t do anything approaching IMO gold problems. It was still struggling with high school math.
Since then, we’ve seen this progression in high school math all the way up through the IMO gold. Now we’ve got the FrontierMath benchmark, which I think is up to about 25%. It was 2% about a year ago, or even a little less than a year ago, I think.
We also just today saw something where somebody said that they had solved a canonical, super-challenging problem that Terence Tao had put out. It was this problem that had stumped scientists for years, and it just so happened that they had also recently figured out the answer but had not yet published their results.
There was this confluence where the scientists had experimentally verified the answer, and Gemini, in the form of this AI co-scientist, came up with exactly the right answer. These are things that literally nobody knew before, and GPT-4 just wasn’t doing that. These are qualitatively new capabilities.
That thing, I think, ran for days. It probably cost hundreds of dollars, maybe into the thousands of dollars, to run the inference. That’s not nothing, but it’s also very much cheaper than years of grad students. If you can get to those caliber of problems and actually get good solutions to them, what would you be willing to pay for that kind of thing?
I don’t know. That’s probably not a full appreciation—we could go on for a long time—but I would say, in summary, GPT-4 was not able to push the actual frontier of human knowledge. To my knowledge, I don’t know that it ever discovered anything new.
It’s still not easy to get that kind of output from GPT-5, Gemini 2.5, or Claude Opus 4, but it’s starting to happen sometimes. That, in and of itself, is a huge deal.
Erik Torenberg
Well, then how do we explain the bearishness, or the kind of vibe shift, around GPT-5? One potential contributor is this idea that if a lot of the improvements are at the frontier—not everyone is working with advanced math and physics day to day—maybe they don’t see the benefits in their day-to-day lives in the same way that the jumps in ChatGPT were obvious and shaped the day to day.
Nathan Labenz
Yeah. I think a decent amount of it was that they kind of upped the launch, simply put. They were tweeting Death Star images, which Sam Altman later came back and said, “No, you’re the Death Star.”
“I’m not the Death Star.” But I think people thought the Death Star was supposed to be the model. Generally speaking, expectations were set extremely high. The actual launch itself was just technically broken.
For a lot of people, their first experience of GPT-5 involved this model-router concept. Another way to understand what they’re doing here is that they’re trying to own the consumer use case. To own that, they need to simplify the product experience relative to what we had in the past, which was, “Okay, you’ve got GPT-4o, GPT-4o mini, o3, o4-mini, and other things. GPT-4.5 was in there at one point. You’ve got all these different models—which one should I use for which?” It’s very confusing to most people who aren’t obsessed with this.
One of the big things they wanted to do was shrink that down to, “Just ask your question and you’ll get a good answer,” and take on that complexity on their side as the product owners. Interestingly, I don’t have a great account of this, but one thing you might want to do is merge the models and have the model itself decide how much to think. Or maybe have the model itself decide how many of its experts, if it’s a mixture-of-experts architecture, it needs to use.
There have been a bunch of different research projects on skipping layers of the model. If the task is easy enough, you could skip a bunch of layers. You might have hoped that you could genuinely merge all these different models on the back end into one model that would dynamically use the right amount of compute for the level of challenge that a given user query presented. It seems like they found that harder to do than they expected.
The solution they came up with instead was to have a router whose job is to pick: Is this an easy query, in which case we’ll send you to this model? Is it medium? Is it hard? I think they just have 2 really different models behind the scenes, so I think it’s just really easy or hard. Certainly, the graphs they showed basically showed the difference between thinking and not thinking.
The problem at launch was that the router was broken. All of the queries were going to the dumb model, so a lot of people literally just got bad outputs, which were worse than o3 because they were getting non-thinking responses. The initial reaction of, “Okay, this is dumb,” traveled really fast. I think that kind of set the tone.
My sense now is that, as the dust has settled, most people do think that it is the best model available. Things like the infamous METR task-length chart show that it is the best. We’re now over 2 hours, and it is still above the trend line. If you just said, “Do I believe in straight lines on graphs or not, and how should this latest data point influence whether I believe in these straight lines on power-law, logarithmic-scale graphs?” it shouldn’t really change your mind too much. It’s still above the trend line.
I talked to Zvi Mowshowitz, a legendary infovore and AI industry analyst, on a recent podcast and asked him the same question: Why do you think even some of the most plugged-in, sharp minds in the space have seemingly pushed timelines out a bit as a result of this? His answer was basically that it resolved some amount of uncertainty. You had an open question of, “Maybe they do have another breakthrough. Maybe it really is the Death Star.” If they surprised us on the upside, then all these short timelines would have seemed plausible.
One way to think about it is that the distribution was broad in terms of timelines. If they had surprised on the upside, it might have narrowed toward the front end of the distribution. If they surprised on the downside, or even just performed purely on trend, then you would take some of the probability mass from the very short end of the timelines and push it back toward the middle or the end.
His answer was that AI 2027 seems less likely, but AI 2030 seems basically no less likely—maybe even a little more likely—because some of the probability mass from the early years is now sitting there. It’s not that people are moving the whole distribution out by a huge amount. I think they may just be shrinking the distribution a little bit; it’s getting tighter because it may not be happening quite as soon as it seemed like it might have been.
I don’t think too many people—at least, the people I think are really plugged in on this—are pushing out too much past 2030 at all. And, by the way, there’s obviously a lot of disagreement. The way I’ve always thought about this sort of stuff is: Dario says 2027, Demis says 2030. I’ll take that as my range.
Coming into GPT-5, I was kind of in that space. Now I’d say, “Well, I don’t know. What cards does Dario have up his sleeve?” They just put out Claude Opus 4.1, and in that blog post they said, “We will be releasing more powerful updates to our models in the coming weeks,” so they’re due for something pretty soon. Maybe they’ll be the ones to surprise on the upside this time, or maybe Google will be.
I wouldn’t say 2027 is out of the question, but I would say 2030 still looks just as likely as before. And again, from my standpoint, that’s still really soon. If we’re on track, whether it’s 2028, 2029, or 2030, I don’t really care. I try to frame my own work so that I’m preparing myself and helping other people prepare for what might be the most extreme scenarios.
If we aim high and miss a little, and we have a little more time, great. I’m sure we’ll have plenty of things to do with that extra time to be ready for whatever powerful AI does come online. But I guess my worldview hasn’t changed all that much as a result of these summer developments.
Erik Torenberg
Anecdotally, I don’t hear as much about AI 2027 or situational awareness to the same degree. I do talk to some people who’ve just moved it back a few years, to your point. Dario still believes in it, but maybe because of this gap in continual learning—or something to that effect—it’s just going to be a bit slower to diffuse.
And METR’s paper, as you mentioned, showed that engineers are less productive, so maybe there’s less concern around people being replaced en masse in the next few years. When we spoke maybe a year ago, I think you said something like 50% of jobs. I’m curious whether that’s still your litmus test, or how you think about it.
Nathan Labenz
Well, for one thing, I think that METR paper is worth unpacking a little bit more. I’m a big fan of METR, and I have no shade on them, because I do think, “Do science, publish your results.” That’s good. You don’t have to make every experimental result and everything you put out conform to a narrative.
But I do think it was a little too easy for people who wanted to say that this is all nonsense to latch onto that. Again, there’s something there that I would put in the Cal Newport category, too. For me, maybe the most interesting thing was that the users thought they were faster when, in fact, they seemed to be slower. That sort of misperception of oneself is really interesting.
Personally, I think there are some explanations for that, including hitting “Go” on the agent, going to social media and scrolling around for a while, and then coming back. The thing might have been done for quite a while by the time I get back. One really simple thing that products can do to address those concerns—and we’re starting to see this in products—is provide notifications: “The thing is done now.” So stop scrolling, come back, and check its work.
In terms of just clock time, it would be interesting to know what applications they had open. Maybe they took a little longer with Cursor than doing it on their own, but how much of the time was Cursor the active window, and how much of it was some other random distraction while they were waiting?
I think a more fundamental issue with that study—which again wasn’t really about the study design, but about the interpretation and digestion of it—is that some of these details got lost. They basically tested the models, or the product Cursor, in the area where it was known to be least able to help.
This study was done early this year, so it was done with, depending on how you want to count, a couple of releases ago, and with large codebases, which strains the context window. That’s one of the frontiers that has been moving: very mature codebases with high standards for coding, and developers who really know their codebases well and have made a lot of commits to those particular codebases.
I would say that’s basically the hardest situation you could set up for an AI, because the people know their stuff really well. The AI doesn’t know it. The context is huge. People have already absorbed that knowledge through working on it for a long time.
The AI doesn't have that knowledge, and these were models from a couple of generations ago. Another big thing was that the people were not very well-versed in the tools. Why? Because the tools weren't really able to help them yet.
I think the mindset of the people who came into the study, in many cases, was, “Well, I haven't used this all that much because it hasn't really seemed to be super helpful.” They weren't wrong in that assessment, given the limitations. You could see that in some of the instructions and help that the METR team gave people.
One of the things in the paper was that if they noticed you weren't using Cursor very well, they would give you feedback on how to use it better. One of the things they told people to do was make sure you @-tag a particular file to bring it into context for the model, so that the model has the right context. That's literally the most basic thing you would do in Cursor. That's the thing you would learn in your first hour, your first day of using it.
So it really does suggest that these were very capable programmers but basically novices when it came to using the AI tools. I think the result is real, but I would be very cautious about generalizing too much there. What else was the other question? What is the expectation for jobs? I mean—
Erik Torenberg
We're starting to see some of this, right? We are definitely seeing this. Marc Benioff, no less, has said that they've been able to cut a bunch of headcount because they now have AI agents responding to every lead. Klarna, of course, has said similar things for a while now.
I think Klarna has also been a little bit misreported in terms of, “Oh, they're backtracking off of that because they're actually going to keep some customer service people, not none.” I think that's a bit of an overreaction. They may have some people who are just insistent on having a certain experience, and maybe Klarna wants to provide that. That makes sense. You can have a spectrum of service offerings for your customers.
I once coded up—I actually just vibe-coded—a pricing page for a SaaS company that had a basic level with AI sales and service at one price. If you wanted to talk to human sales, that was a higher price. And if you wanted to talk to human sales and support, that was a third, higher price. Literally, that might be what's going on in some of these cases. It could very well be a very sensible option for people.
I do see this with Intercom. I've got an episode coming up with them. They now have this Fin agent that's solving about 65% of the customer service tickets that come in. What's that going to do to jobs? Are there really 3 times as many customer service tickets to be handled? I don't know. I think there's a relatively inelastic supply. Maybe you'll get somewhat more tickets if people expect that they're going to get better, faster answers, but I don't think we're going to see 3 times more tickets.
By the way, that number was 55% 3 or 4 months ago. As they ratchet that up, the ratios get really hard, right? At 50% ticket resolution, in theory, maybe you get some more tickets; maybe you don't need to adjust headcount too much. But when you get to 90% ticket resolution, are you really going to have 10 times as many tickets, or 10 times as many hard tickets for people to handle? It seems really hard to imagine that.
I don't think these things go to zero, probably, in a lot of environments, but I do expect that you'll see significant headcount reduction in a lot of these places. The software one is really interesting because the elasticities are really unknown. You can potentially produce X times more software per user, or per Cursor user, or per developer at your company. But maybe you want that. Maybe there's no limit, or maybe the regime we're in is such that if there's 10 times more productivity, that's all to the good, and we still have just as many jobs because we want 10 times more software.
I don't know how long that lasts. Again, the ratios start to get challenging at some point. The old Tyler Cowen thing comes to mind: You are a bottleneck. You are a bottleneck. More often, I think the question is whether people are really trying to get the most out of these things, whether they're using best practices, and whether they've really put their minds to it. Often, the real barrier is there.
I've been working a little bit with a company that's doing government document review. I'll abstract away a little bit from the details, but it's really gnarly stuff: scanned documents, handwritten filling out of forms, and all that. They've created this auditor AI agent that just won a state-level contract to do the audits on roughly 1 million transactions a year involving these packets of documents—again, scanned, handwritten, all this kind of crap.
They just blew away the human workers who were doing the job before. So where are those workers going to go? I don't know. They're not going to have 10 times as many transactions. I can be pretty confident in that. Are there going to be a few people still there to supervise the AIs, handle the weird cases, and answer the phones? Sure. Maybe those workers won't go anywhere.
The state may do a strange thing and just have all those people sit around because it can't bear to fire them. Who knows what the ultimate decision will be? But I see a lot of these situations where, when you really put your mind to it and identify what would create real leverage for you, you can ask, “Can the AI do that? Can we make it work?” You can take a pretty large chunk out of high-volume tasks very reliably in today's world.
The impacts are starting to be seen across a lot of jobs. Humans, I think, are often the bottleneck. Leadership may be the bottleneck, or the will to make changes may be the bottleneck, in a lot of places. Software might be an interesting case where there's just so much pent-up demand that it might take a little longer to see those impacts, because you really do want 10 or 100 times as much software.
Erik Torenberg
What is—yeah, let's talk about code, because it's where Anthropic made a big bet early on, perhaps inspired by the sort of automated researcher, recursive self-improvement, desired future. We saw OpenAI make moves there as well. Why don't we flesh that out and talk a little about what inspired that and where you see that going?
Nathan Labenz
Utopia or dystopia is really the big question there, I think, right? Maybe it's 1 part technical and 2 parts social in terms of why code has been so focal. The technical part is that it's really easy to validate code. You generate it, you can run it, and if you get a runtime error, you can get the feedback immediately. It's somewhat harder to do functional testing.
Replit recently, in the last 48 hours, released v3 of its agent. In addition to “code, code, code, try to make your app work,” which v2 of the agent would do, it could go for minutes and, in some cases, generate dozens of files. I've had some magical experiences with that where I was like, “Wow, you just did that whole thing in one prompt,” and it worked amazingly.
Other times, it will code for a while, hand it off to you, and say, “Okay, does it look good? Is it working?” And you're like, “No, it's not. I'm not sure why.” You get into a back-and-forth with it. The difference between v2 and v3 is that instead of handing the baton back to you, it now uses a browser and the vision aspect of the models to try to do the QA itself.
It doesn't just say, “Okay, I tried my best, wrote a bunch of code, and let me know if it's working or not.” It takes that first pass at figuring out whether it's working. That really improves the flywheel: how much you can do, how much you can validate, and how quickly you can validate it. The speed of that loop is really key to the pace of improvement. It's a problem space that's pretty amenable to rapid-flywheel techniques.
Second, of course, they're all coders at these places, so they want to solve their own problems. That's very natural. Third, on the sort of social-vision and competition side—who knows where this is all going?—they do want to create the automated AI researcher.
That's another data point, by the way. This was from the o3 system card. They showed a jump from low- to mid-single digits to roughly 40% of PRs actually checked in by research engineers at OpenAI that the model could do. Prior to o3, it wasn't much at all—low to mid-single digits. As of o3, it was 40%.
I'm sure those are the easier 40%, or whatever. Again, there will be caveats to that, but you're entering maybe the steep part of the S-curve there. That's presumably pretty high-end. I don't know how many easy problems they have at OpenAI, but presumably not that many relative to the rest of us out here making generic web apps all the time.
At 40%, you've got to be starting to, I would think, get into some pretty hard tasks and some pretty high-value stuff. At what point does that ratio really start to tip, where the AI is doing the bulk of the work? GPT-5 notably wasn't a big update over o3 on that particular measure.
I mean, it also wasn't going back to the simple QA thing. GPT-5 is generally understood to not be a scale-up relative to GPT-4o and o3, and you can see that in the simple QA measure. It basically scores the same on these long-tail trivia questions. It's not a bigger model that has absorbed lots more world knowledge.
Cal is right. I think his analysis is that it's post-training. But that post-training is potentially entering the steep part of the S-curve when it comes to the ability to do even the kind of hard problems that are happening at OpenAI on the research engineering front. And yikes, I'm a little worried about that, honestly.
The idea that we could go from these companies having a few hundred research engineers to having unlimited numbers of them overnight—what would that mean in terms of how much things could change, and also just our ability to steer that overall process? I'm not super comfortable with the idea of the companies tipping into a recursive self-improvement regime, especially given the level of control and the level of unpredictability that we currently see in the models. But that does seem to be what they are going for.
In terms of why, I think this has been the plan for quite some time. You remember that leaked Anthropic fundraising deck from maybe 2 years ago where they said that in 2025 and 2026, the companies that train the best models will get so far ahead that nobody else will be able to catch up. I think that's kind of what they meant. I think they were projecting then that in the 2025–2026 timeframe they'd get this automated researcher, and once you have that, how's anybody who doesn't have that going to catch up with you? Obviously, some of that remains to be validated, but I do think they have been pretty intent on that for a long time.
Erik Torenberg
Five years from now, are there more engineers or fewer engineers?
Nathan Labenz
I tend to think fewer. Already, if I just think about my own life and work, I'm like, would I rather have a model or a junior marketer? I'm pretty sure I'd rather have the model. Would I rather have the models or a junior engineer? I think I'd probably rather have the models in a lot of cases. It obviously depends on the exact person you're talking about.
But truly, forced choice today, I would choose the model. And then you've got cost adjustment as well, right? I'm not spending nearly as much on my Cursor subscription as I would be on an actual human engineer. So even if they have some advantages, I also haven't scaffolded—I haven't gone full Co-Scientist on my Cursor problems.
I think that's another interesting thing. You start to see why folks like Sam Altman are so focused on questions like energy and the $7 trillion buildout, because these power-law things are weird, and to get incremental performance for 10× the cost is weird. It's definitely not the kind of thing that we're used to dealing with, but for many things it might be worth it, and it still might be cheaper than the human alternative.
If Cursor costs me about $40 a month or something, would I pay $400 for however much better it is? Yeah, probably. Would I pay $4,000 for however much better it is? Well, it's still a lot less than a full-time human engineer. And the costs are obviously coming down dramatically, too, right? That's another huge thing.
GPT-4 was way more expensive. It's like a 95% discount from GPT-4 to GPT-5. That's no small thing, right? I mean, an apples-to-apples comparison is a little bit hard because the chain of thought does spit out a lot more tokens, and so you give back a little on a per-token basis. It's dramatically cheaper, but generating more tokens does eat back into some of that savings. Everybody seems to expect the trends will continue in terms of prices continuing to fall.
So how many more of these price reductions do you need in order to be able to do the power-law thing a few more times?
I guess I think fewer. And I think that's probably true even if we don't get full-blown AGI that's better than humans at everything. You could easily imagine a situation where, of however many million people are currently employed as professional software developers, the top tier of them who do the hardest things can't be replaced. But there aren't that many of those.
The real rank and file—the people who, over the last 20 years, were told, “Learn to code; that'll be your thing”—didn't need to be told to learn to code if they were really at the top. It was their thing. They had a passion for it. They were amazing at it.
It wouldn't shock me if we still couldn't replace those people in 3, 4, or 5 years' time. But I would be very surprised if you can't get your nuts-and-bolts web app and mobile app type things spit out for you for far less and far faster than—and probably, honestly, with significantly higher quality and less back-and-forth than—with your kind of middle-of-the-pack developer in that timeframe.
One thing I do want to call out: there are definitely people who have concerns about progress moving too fast, but there's also concern—and maybe it's rising—about progress not moving fast enough, in the sense that a third of the stock market is the Magnificent 7, AI capex is over 1% of GDP, and we are kind of relying on some of this progress in order to sustain our economy.
Nathan Labenz
Yeah. Another thing that I would say has been slower to materialize than I would have expected are AI culture wars, or the ramping up of protectionism in various industries. We just saw Josh Hawley—I don't know if he introduced a bill or just said he intends to introduce a bill—to ban self-driving cars nationwide.
Erik Torenberg
God help me. I've dreamed of self-driving cars since I was a little kid. Truly, sitting at red lights, I used to think, “There's got to be a way.” I think we took a Waymo together.
Nathan Labenz
Yeah, and it's so good. I think whenever people want to argue about jobs, it's going to be pretty hard to say:
Erik Torenberg
Thirty thousand Americans should die every year so that people's incomes don't get disrupted. It seems like you have to be able to get over that hump and say that saving all these lives, if nothing else, is just really hard to argue against. But we'll see. I mean, he's not without influence, obviously.
Nathan Labenz
Yeah, I am very much on team abundance. My old mantra—I've been saying this less lately—is “adoption accelerationist, hyperscaling pauser.” The tech that we have could do so, so much for us even as is. I think if progress stopped today, we could still get to 50% to 80% of work automated over the next 5 to 10 years.
It would be a real slog. You'd have a lot of Co-Scientist-type breakdowns of complicated tasks to do. You'd have a lot of work to do to go sit and watch people and say, “Why are you doing it this way? What's going on here? What's this? You handled this one differently. Why did you handle that one differently?” All this tacit knowledge that people have, and the kind of know-how and procedural instincts that they've developed over time, are not documented anywhere. They're not in the training data, so the AI hasn't had a chance to learn them.
But again, when I say “no breakthroughs,” I still am allowing for fine-tuning of things—capabilities we have that haven't been applied to particular problems yet. Going through the economy and just sitting with people, asking, “Why are you doing this? Let's document this. Let's get the model to learn your particular niche thing”—that would be a real slog.
In some ways, I kind of wish that were the future that we were going to get, because it would be a methodical, one-step, one-foot-in-front-of-the-other process, with no quantum leaps. It would probably feel pretty manageable, I would think, in terms of the pace of change. Hopefully society could absorb that and adapt to it as we go, without going from one day to the next with, “Oh my God, all the drivers are getting replaced.” That one would be a little slower because you do have to have the actual physical buildout.
In some of these things, customer service could get ramped down real fast, right? If a call center has something that they can just drop in and it's like, “This thing now answers the phones and talks like a human and has a higher success rate and scales up and down.”
One thing we've seen at Waymark—a small company, right?—is that we've always prided ourselves on customer service. We do a really good job with it. Our customers really love our customer success team. But I looked at our Intercom data, and it takes us about half an hour to resolve tickets.
We respond really fast. We respond in under 2 minutes most of the time. But when we respond, 2 minutes is still long enough that the person has gone on to do something else, right? It's the same thing as with the Cursor thing that we were talking about earlier, right? They've tabbed over to something else. So now we get the response back in 2 minutes, but they are doing something else. Then they come back at minute 6 or whatever, and then they respond. But now our person has gone and done something else.
So the resolution time, even for simple stuff, can easily be half an hour, and the AI just responds instantly, right? So you don't have to have that kind of back-and-forth. You're just in and out. I do think some of these categories could change really fast; others will be slower. But I kind of wish we had that slower path in front of us.
My best guess, though, is that we will probably continue to see things that will be significant leaps and that there will be actual disruption. Another one that's come to mind recently: maybe we can get the abundance department on these new antibiotics. Have you seen this development?
Erik Torenberg
No. Tell us about it.
Nathan Labenz
I mean, it's not a language model. I think that's another thing people really underappreciate: you could look back at GPT-4 to GPT-5 and imagine a pretty easy extension of that. GPT-4, initially, when it launched, didn't have image-understanding capability. They did demo it at the time of the launch, but it wasn't released until some months later. The first version we had could understand images and did a pretty good job of understanding images, still with jagged capabilities and whatever.
Now, with Google's new Nano Banana, you have basically Photoshop-level ability to just say, “Hey, take this thumbnail.” We could take our 2 feeds right now, take a snapshot of you, a snapshot of me, put them both into Nano Banana, and say, “Generate the thumbnail for the YouTube preview featuring these 2 guys. Put them in the same place, same background, whatever.” It'll mash that up. You can even have it put text on top: “Progress since GPT-4,” whatever we want to call it. GPT-5 is not a bust, and it'll spit that out.
You see that it has this deeply integrated understanding that bridges language and image. That's something that it can take in, but now it's also something it can put out, all as part of one core model with a single unified intelligence. I think that's going to come to a lot of other things.
We're at the point now with these biology models and materials science models where they're kind of like the image-generation models of a couple of years ago. They can take a real simple prompt and do a generation, but they're not deeply integrated, where you can have a true conversation back and forth and have that kind of unified understanding that bridges language and these other modalities. But even so, it's been enough for this group at MIT to use some of these relatively narrow, purpose-built biology models and create totally new antibiotics.
New in the sense that they have a new mechanism of action: they're affecting the bacteria in a new way. And, notably, they do work on antibiotic-resistant bacteria. This is some of the first new antibiotics we've had in a long time. Now they're going to have to go through—when I say, “Get the abundance department on it”—it's like, where's my Operation Warp Speed for these new antibiotics? We've got people dying in hospitals from drug-resistant strains all the time. Why is nobody crying about this?
I think one of the things that's happening to our society in general is that so many things are happening at once. It's kind of the flood-the-zone thing, except there are so many AI developments flooding the zone that nobody can even keep up with all of them. That's happened to me, by the way, too. I'd say 2 years ago I was pretty in command of all the news; a year ago I was starting to lose it; and now I'm like, “Wait a second, there were new antibiotics developed?” I'm missing things, just like everybody else, despite my best efforts.
But the key point there is that AI is not synonymous with language models. There are AIs being developed with pretty similar architectures for a wide range of different modalities. We have seen this play out with text and image: you had your text-only models and your image-only models, and then they started to come together, and now they've come really deeply together. I think you're going to see that across a lot of other modalities over time as well.
And there's a lot more data there. I don't know what it means to run out of data. In the reinforcement-learning paradigm, there's always more problems, right? There's always something to go figure out. There's always something to go engineer. The feedback is starting to come from reality, right?
That was one of the things Elon talked about on the Grok 4 launch: maybe we're running out of problems we've already solved, and we only have so many of those sitting around in inventory. You only have 1 internet; we only have so much of that stuff. But over at Tesla and over at SpaceX, we're solving hard engineering problems on a daily basis, and they seem to be never-ending.
So when we start to give the next generation of the model these power tools—the same power tools that professional engineers are using at those companies to solve those problems—and the AI starts to learn those tools and starts to solve previously unsolved engineering problems, that's going to be a really powerful signal that they will be able to learn from.
And now, again, fold in those other modalities, right? The ability to have sort of a sixth sense for the space of materials-science possibilities—when you can bridge or unify the understanding of language and those other things—I think you start to have something that looks kind of like superintelligence. Even if it's not able to write poetry at a superhuman level necessarily, its ability to see in these other spaces is going to be truly a superhuman thing that I think will be pretty hard to miss.
Erik Torenberg
You said that was one thing that Cal's analysis missed: just the lack of appreciation for non-language modalities and how they're driving some of the innovations that you're talking about.
Nathan Labenz
Yeah. I think people are often just equating the chatbot experience with AI broadly.
Erik Torenberg
Yeah.
Nathan Labenz
And that conflation will not last probably too much longer because we are going to see self-driving cars, unless they get banned. That's a very different kind of thing. And talk about the impact on jobs, too, right? It's like 4 or 5 million professional drivers in the United States. That is a big deal. I don't think most of those folks are going to be super keen to learn to code, and even if they do learn to code, I'm not sure how long that's going to last. So that's going to be a disruption.
And then general robotics is not that far behind. This is one area where I do think China might actually be ahead of the United States right now, but regardless of whether that's true or not, these robots are getting really quite good. They can walk over all these obstacles. These are things that a few years ago they just couldn't do at all. They could barely balance themselves and walk a few steps under ideal conditions.
Now you've got things that can literally do a flying kick, and it'll absorb your kick, shrug it off, and just keep going. It'll right itself and continue on its way. Super-rocky, uneven terrain, all these sorts of things are getting quite good.
The same thing is working everywhere. I think one of the other things is that there's always a lot of detail to the work. So it's a sort of inside view, outside view, right? Inside view, you're like, there's always this minutia. There's always these problems that we had and things we had to solve. But you zoom out, and it looks to me like the same basic pattern is working everywhere.
If we can just gather enough data to do some pretraining—some kind of raw, rough, not very useful data, but just enough at least to get us going—then we're in the game. Once we're in the game, now we can do this flywheel thing: rejection sampling, having it try a bunch of times, taking the ones where it succeeded, and fine-tuning on that; RLHF feedback, the sort of preference feedback—take 2, which one was better?—and fine-tune on that; reinforcement learning. All these techniques that have been developed over the last few years seem to me like they're absolutely going to apply to a problem like a humanoid robot as well.
That's not to say there won't be a lot of work to figure out exactly how to do that. But I think the big difference between language and robotics is really mostly that there just wasn't a huge repository of data to train the robots on at first. So you had to do a lot of hard engineering to make it work at all, to even stand up, right? You had to have all these control systems and whatever, because there was nothing for them to learn from in the way that the language models could learn from the internet.
But now that they're working at least a little bit, I think all these refinement techniques are going to work. It'll be interesting to see if they can get the error rate low enough that it'll actually allow one in my house around my kids. They'll probably be better deployed in factory settings first, in more controlled environments than the chaos of my house, as you've seen in this recording. But I do think they're going to work.
Erik Torenberg
What's the state of agents more broadly at the moment? How do you see things playing out? Where do you see it going?
Nathan Labenz
Well, broadly, I think it's the task-length story from METR: a 7-month or 4-month doubling time. We're at 2 hours-ish with GPT-5. Replit just said their new Agent v3 can go 200 minutes. If that's true, that would even be a new high point on that graph. Again, it's a little bit apples to oranges because they've done a lot of scaffolding.
How much have they broken it down? How much scaffolding are you allowed to do with these things before you're off their chart and onto maybe a different chart?
But if you extrapolate that out a bit and you're like, "Okay, take the 4-month case just to be a little aggressive." That's 3 doublings a year. That's an 8x task-length increase per year. That would mean you go from 2 hours now to 2 days 1 year from now. Then, if you do another 8x on top of that, you're looking at, basically, say, 2 days to 2 weeks of work in 2 years.
That would be a big deal, to say the least. If you could delegate an AI 2 weeks' worth of work and have it do it even half the time, right? The METR thing is that they will succeed half the time on tasks of that size. If you could take a 2-week task and have a 50% chance that an AI would be able to do it, even if it did cost you a couple hundred bucks, right? It's like, well, that's again a lot less than it would cost to hire a human to do it.
It's all on demand. It's immediately available. If I'm not using it, I'm not paying anything. Transaction costs are just a lot lower. The many, many other aspects are favorable for the AI there. So, that would suggest that you'll see a huge amount of automation in all kinds of different places.
The other thing that I'm watching, though, is that reinforcement learning does seem to bring about a lot of bad behaviors. Reward hacking is one. Any sort of gap between what you are rewarding the model for and what you really want can become a big issue.
We've seen this in coding in many cases, where the AI—Claude is notorious for this—will put out a unit test that always passes, one that just has `return true` in the unit test. Why is it doing that? Well, it must have learned that what we want is for unit tests to pass. We want it to pass unit tests. We didn't mean to write fake unit tests that always pass, but that technically did satisfy the reward condition. So we're seeing those kinds of weird behaviors.
With that comes this scheming kind of stuff. We don't really have a great handle on that yet. There is also situational awareness that seems to be on the rise, where the models are increasingly, in their chain of thought, saying things like, "This seems like I'm being tested. Maybe I should be conscious of what my tester is really looking for here."
That makes it hard to evaluate models in tests because you don't know if they're actually going to behave the same way when they're out in the real world. I wouldn't say this is a high-level or high-confidence prediction, but one model of the future I've been playing with is that task length keeps doubling while, at the same time, these weird behaviors pop up and then are suppressed.
We've seen in the Claude 4 and GPT-5 system cards that Claude 4 reported, I think, a 2/3 reduction in reward hacking. In GPT-5, they reported a few different dimensions, but something similar—a reduction in deceptive behavior. But those behaviors kind of just emerged.
So it's sort of like weird behavior emerges, then they figure out how to tamp it down, but not entirely. Presumably, in the next generation, they'll tamp it down some more, but maybe some new, additional weird behavior could emerge, and then they'll have to tamp that one down.
All the while, the tasks are expanding in scope, with a doubling every 4 months. So you could end up in a world where you can delegate really major things to AIs, but there's some small but not necessarily totally vanishing chance that it actively screws you over in the way that it's trying to do that task.
You have, maybe, "Okay, I think here I'm about to get 2 weeks' worth of work done for $100. Okay, that's amazing. But there's also a 1-in-10,000 chance that it legitimately attacks me in a meaningful way."
Some of the things that we have seen—these are fairly famous at this point—but in the Claude 4 system card, they reported blackmailing of a human. The setup was that the AI had access to the engineer's email, and they told the AI that it was going to be replaced with a less ethical version or something like that. It didn't want that, and it found in the engineer's email that the engineer was having an affair. So it started to blackmail the engineer to avoid being replaced with a less ethical version.
Erik Torenberg
People, I think, are way too quick, in my view, to
Nathan Labenz
move past these anecdotes. People are often like, "Well, they set it up that way, and that's not really realistic." But another one was whistleblowing. There was another thing where they set up this dynamic where there was some unethical, illegal behavior going on, and again, the model had access to this data. It decided to email the FBI and tell the FBI about it.
First of all, I don't think we really know what we want. To some degree, maybe you do want AIs to report certain things to authorities. That could be one way to think about the bioweapon risk. Maybe the models shouldn't only refuse; maybe they should report you to the authorities if you're actively trying to create a bioweapon.
I certainly don't want them to be doing that too much. I don't want to live under the surveillance of Claude 5, always threatening to turn me in. But I do sort of want some people to be turned in if they're doing sufficiently bad things. We don't have a good resolution, society-wide, on what we want the models to even do in those situations.
I think it's also like, yes, it was set up, and yes, it was research, but it's a big world out there, right? We got 1 billion users already on these things, and we're plugging them into our email, so they're going to have very deep access to information about us.
I don't know what you've been doing in your email. I hope there's nothing too crazy in mine, but now I've got to think about it a little bit, right? What have I ever done that—geez, I don't know. Or even something that it could misconstrue, right? It's obviously not that maybe I even did anything that bad, but it just misunderstands what exactly was going on.
So that could be a weird thing. If there's one thing that could stop the agent momentum, in my view, it could be that the 1-in-10,000 chance, or whatever we ultimately push the really bad behaviors down to, is still just so spooky to people that they're like, "I can't deal with that." That might be hard to resolve.
So, what happens then? It's hard to check 2 weeks' worth of work every couple of hours or whatever, right? That's part of where you bring another AI in to check it. That's again where you start to get to, "Now I see why we need more electricity, and $7 trillion of buildout is yikes."
They're going to be producing so much stuff. I can't possibly even review it all. I need to rely on another AI to help me review the first AI, to make sure that if it is trying to screw me over, somebody's catching it. I can't monitor that myself.
I think Redwood Research is doing some really interesting work on this, where they're trying to get systematic about it. Let's just assume this is quite a departure from the traditional AI safety work, where the big idea traditionally was: let's figure out how to align the models, make them safe, and make them not do bad things.
Great. Redwood Research has taken the other angle, which is: let's assume that they're going to do bad stuff. They're going to be out to get us at times. How can we still work with them, get productive output, and get value without fixing all those problems?
That involves, again, all these sorts of AIs supervising other AIs, and crypto might have a role to play in this. There's another episode coming out soon with Illia Polosukhin, who's the founder of NEAR. He's a really fascinating guy because he was one of the 8 authors of the paper "Attention Is All You Need."
Erik Torenberg
Yeah.
Nathan Labenz
He started this NEAR company. It was originally an AI company. They took a huge detour into crypto because they were trying to hire task workers around the world and couldn't figure out how to pay them.
So they were like, "This sucks so bad to pay these task workers in all these different countries that we're trying to get data from. We're going to pivot into a whole blockchain side quest." Now they're coming back to the AI thing, and their tagline is "the blockchain for AI."
You might be able to get a certain amount of control from the sort of crypto security that blockchain-type technology can provide. But I could see a scenario where these bad behaviors just become so costly when they do happen that people get spooked away from using the frontier capabilities, in terms of just how much work the AIs can do.
But that wouldn't be a pure capability stall-out. It would be, "We can't solve some of the long-tail safety issues."
Erik Torenberg
Yeah, it's a challenge, and if that is the case, that'll be an important fact about the world, too. Nobody ever seems to solve any of these things 100%, right? Every generation, it's like, “Well, we reduced hallucinations by 70%,” or “We reduced deception by two-thirds. We reduced scheming,” or whatever, by however much. But it's always still there. If you take even the lower rate and multiply it by a billion users, thousands of queries a month, agents running in the background and processing all your emails, and all the deep access that people envision them having, it could be a pretty weird world where there's this negative lottery of AI accidents.
Another episode coming up is with the AI underwriting company. They are trying to bring the insurance industry and all the wherewithal that's been developed there to price risk and figure out how to create standards: What can we allow, and what sort of guardrails do we have to have to be able to insure this kind of thing in the first place?
That'll be another really interesting area to watch: Can we financialize those risks in the same way we have with car accidents and all these other mundane things? But the space of car accidents is only so big. The space of weird things that AIs might do to you as they have weeks' worth of runway is much bigger, and so it's going to be a hard challenge. But people are working; we've got some of our best people working on it.
What do you make of the claim that 80% of AI startups have Chinese open-source models, and what are the implications?
Nathan Labenz
I think that probably is true, with the one caveat that it is only measuring companies that are using open-source models at all. I think most companies are not using open-source models, and I would guess the vast majority of tokens being processed by American AI startups are API calls to the usual suspects. So, weighted by actual usage, I would say the majority would still be going to commercial models.
For those that are using open-source models, I do think it's true that the Chinese models have become the best. The American bench there was always kind of thin, right? It was basically Meta that was willing to put in huge amounts of money and resources and then open-source it. You've got the Paul Allen-funded group, the Allen Institute for AI, AI2. They're doing good stuff, too, but they don't have pretraining resources, so they do really good post-training and open-source their recipes and all that kind of stuff. So it's not like American open-source is bad.
Again, this is another way in which I think you can really validate that things are moving quickly: If you take the best American open-source models and take them back a year, they're probably as good as, if not a little better than, anything we had commercially available at the time. If you compare them to the Chinese models, I think the Chinese models have surpassed them. So there's been a pretty clear change at the frontier. I think that means the best Chinese models are pretty clearly better than anything we had a year ago, commercial or otherwise.
That just means things are moving. Hopefully I've made that case compellingly, but that's another data point that makes it hard to believe both that the Chinese models are now the best open-source models and that AI has stalled out and we haven't seen much progress since GPT-4. Those seem to be contradictory notions. I believe the one that's wrong is the lack of progress.
In terms of what it means, I don't really know. We're not going to stop China. I've always been a skeptic of the whole “no selling chips to China” thing. The notion originally was, “We're going to prevent them from doing some super cutting-edge military applications.” Then it was, “Well, we can't really stop that, but we can at least stop them from training frontier models.” And then it was, “Well, we can't necessarily really stop that, but now we can at least keep them from having tons of AI agents. We'll have way more AI agents than they do.”
I don't love that line of thinking at all. One upshot of it, potentially, is that they just don't have enough compute available to provide inference as a service to the rest of the world. So instead, the best they can do is say, “Okay, well, we'll train these things and you can figure it out here. Here you go. Have at it.” It's kind of a soft-power play, presumably.
I did an episode with Anne from a16z, who I thought really did a great job of providing the perspective of what I started calling countries 3 through 193. If the US and China are 1 and 2, countries 3 through 193 are a big gap. I think the US is still ahead, but not by that much, in terms of research and ideas relative to China. We do have this compute advantage, and that does seem like it matters.
One of the upshots may be that they're open-sourcing, and countries 3 through 193 are significantly behind. For them, it's a way to try to bring more countries over to the Chinese camp, potentially, in the US-China rivalry.
It seems like the model everybody wants, and I don't like this at all. I don't like technology decoupling, as somebody who worries about who's the real other here. I always say the real Other is the AI, not the Chinese. So if we do end up in a situation where—yikes—we're seeing some crazy things, it would be really nice if we were on basically the same technology paradigm.
To the degree that we really decouple—not just the chips are different, but maybe the ideas start to become very different, publishing gets shut down, and tech trees evolve and grow apart—that, to me, seems like a recipe for making it harder to know what the other side has and harder to trust one another. It seems to feed into the arms-race dynamic, which I do think would be a real existential-risk factor. I would hate to see us create another sort of MAD-type dynamic where we all live under the threat of AI destruction. But that very well could happen.
I do kind of have some sympathy for the recent decision that the administration made to be willing to sell H20s to China. Then it was funny that they turned around and rejected them, which to me seemed like a mistake. I don't know why they would be rejecting them. If I were them, I would buy them. Maybe I would sell inference on the models that I've just been creating and try to make my money back doing that.
In the meantime, they can at least demonstrate the greatness of the Chinese nation by showing that they're not far behind the frontier. They can also make a pretty powerful appeal to countries 3 through 193 and say, “Look, you really want to—you see how the US is acting in general. They cut us off from chips. The last administration had an even longer list of countries that couldn't get chips. This administration is doing all kinds of crazy stuff. You get 50% tariffs here, there, whatever. How do you know you can really rely on them to continue to provide you AI into the future? Well, you can rely on us. We open-sourced the model. You can have it.”
“So come work with us and buy our chips because, by the way, as we mature, our models will be optimized to run on our chips.” I don't know. That's a complicated situation. I do think it's true that adoption isn't as high as 80%. I think that is within that subset of companies that are doing stuff with open source.
We're going to experiment with that at Waymark, but to be honest, we have never done anything with an open-source model in our product to date. Everything we've ever done has been through commercial models. At this point, we are going to try doing some reinforcement fine-tuning. We are going to do that on a Qwen model, I think, first. So that'll put us in that 80%.
But I'm guessing that at the end of the day, we'll take that Qwen model, do the reinforcement fine-tuning, and probably get roughly up to as good as GPT-5 or Claude 4 or whatever. Then we'll say, “Okay, do we really want to have to manage inference ourselves? How much are we really going to save?” At the end of the day, I would guess we probably are still going to end up just being like, “Eh, we'll pay a little bit more on a monthly-bill basis for one of these frontier models. They're a little bit better, maybe, still, and operationally it's a lot easier, and they'll have upgrades.”
Of course, there are regulated industries. There are a lot of places where you have hard constraints you just can't get around, and that forces you to use those Chinese models. Then there's also going to be the question of whether there are backdoors in them. People have seen the Sleeper Agents project, where a model was trained to be good up until a certain point in time.
People put today's date in the system prompt all the time, right? “Today's date is this. You are Claude. Here you go.” So that's going to be another kind of thing for people to worry about. We don't really have great tools for detecting them. There have been some studies—Anthropic did a thing where they trained models to have some hidden objectives and then challenged teams to figure out what those hidden objectives were.
And with certain interpretability techniques, they were able to figure that stuff out relatively quickly. So you might be able to get enough confidence to take this open-source thing created by some Chinese company, whatever, and put it through some sort of—not exactly an audit, because you can't trace exactly what's happening—but some sort of examination to see: Can we detect any hidden goals, secret backdoor behavior, or whatever? Maybe with enough of that kind of work, you could be confident that you don't have it.
But the more and more critical this stuff gets, again, going back to that task-length doubling and weird behavior, now you've got to add into the mix: What if they intentionally programmed it to do certain bad things under certain rare circumstances? We're just headed for a really weird future. We've got all these—there's no limit to it. All these things are valid concerns, and they often are in direct tension with each other. I'm not one who wants to see one tech company take over the world by any means. So I definitely think we would do really well to have some sort of broader, more buffered ecological system where all the AIs are in some sort of competition, mutual coexistence with each other. But we don't really know what that looks like, and we don't really know what an invasive species might look like when it gets introduced into that very nascent and as-yet-not-battle-tested ecology. So, yeah, I don't know. Bottom line, I think the future's going to be really, really weird.
Erik Torenberg
Yeah. Well, I do want to close on an uplifting note. So maybe, as a segue to a closing question, we could get into some areas where we're already seeing some exciting capabilities emerge and sort of transform the experience. Maybe around education or healthcare, or any other areas you want to highlight?
Nathan Labenz
Yeah, it's—boy, it's all over. One of my mantras is that there's never been a better time to be a motivated learner.
Erik Torenberg
So, I think a lot of these things do have two sides of the coin.
Nathan Labenz
There's the worry that students are taking shortcuts and losing the ability to sustain focus and endure cognitive strain. The flip side of that is that, as somebody who's fascinated by the intersection of AI and biology, sometimes I want to read a biology paper and I really don't have the background. An amazing thing to do is turn on Voice Mode and share your screen with ChatGPT, then just go through the paper. You don't even have to talk to it most of the time; you're doing your reading, it's watching over your shoulder, and then at any random point you have a question, you can verbally say, “What's this? Why are they talking about that? What's going on with this? What is the role of this particular protein that they're referring to?” and it will have the answers for you.
So if you really want to learn in a sincere way, these things are unbelievably good at helping you do that. The flip side is you can take a lot of shortcuts and maybe never have to learn stuff.
On the biology front, again, we've got multiple discovery things happening. The antibiotics one we covered; there was another one that I did another episode on with a Stanford professor named James Zou, who created something called the Virtual Lab. Basically, this was an AI agent that could spin up other AI agents depending on what kind of problem it was given. Then they would go through a deliberative process where one expert in one thing would give its take, and they would bat it back and forth. There was a critic in there that would criticize the ideas that had been given. Eventually, they'd synthesize.
Then they were also given some of these narrow specialist tools. So you have agents using the AlphaFold type of thing—not just AlphaFold; there's a whole wide array of those at this point—to say, “Okay, well, can we simulate how this would interact with that?” The agents were running that loop, and they were able to get this language-model agent with a specialized tool system to generate new treatments for novel strains of COVID that had kind of escaped the previous treatments. Amazing stuff, right?
The flip side of that, of course, is the bioweapon risk. So all these things do seem like they're going to be even on just the abundance front itself, right? We may have a world of unlimited professional private drivers, but we don't really have a great plan for what to do with the 5 million people currently doing that work. We may have infinite software, but especially once the 5 million drivers pile into all the coding boot camps and get coding jobs, I don't know what we're going to do with the 10 million people who were coding when 9 million of them become superfluous.
So, yeah, I don't know. I think we're headed for a weird world. Nobody really knows what it's going to look like in 5 years. There was a great moment at Google's I/O where they brought up some journalist. I know we're skeptical of journalists—this is a great moment to go direct, right? This was a great example of why one would want to do that. They brought up this person to interview Demis Hassabis and Sergey Brin. The guy asked, “What is search going to look like in 5 years?” and Sergey Brin almost spit out his coffee on the stage and was like, “Serge, we don't know what the world is going to look like in 5 years.”
So I think that's really true: the biggest risk for so many of us—and I include myself here—is thinking too small. The worst thing I think we could do would be to underestimate how far this thing could go. I would much rather be mocked for things happening on twice the timescale that I thought than to find myself unprepared when they do happen. So whether it's 2027, 2029, or 2031, I'll take that extra buffer honestly where we can get it.
My thinking is just: get ready as much and as fast as possible. And again, if we do have a little grace time to do extra thinking, then great. But I think the worst mistake we could make would be to dismiss this and not feel like we need to get ready for big changes.
Erik Torenberg
Should we wrap directly on that, or is there any other last note you want to make sure to get across regarding anything we said today?
Nathan Labenz
One of my other mantras these days is: the scarcest resource is a positive vision for the future. Yeah, I do think it's always really striking, whether it's Sergey, Sam Altman, or Dario Amodei. Dario probably has the best positive vision of the frontier developer CEOs with “Machines of Loving Grace.” But it's always striking to me how little detail there is on these things.
When they launched GPT-4o, which was the voice mode, they were pretty upfront about saying, “Yeah, this was kind of inspired by the movie Her.” So I do think that even if you are not a researcher, not great at math, or not somebody who codes, this technology wave really rewards play. It really rewards imagination. I think literally writing fiction might be one of the highest-value things you could do, especially if you could write aspirational fiction that would get people at the frontier companies to think, “Geez, maybe we could steer the world in that direction. Wouldn't that be great?” If you could plant that kind of seed in people's minds, it could come from a totally nontechnical place and potentially be really impactful.
Play, fiction, positive vision for the future. Behavioral science, too. These days, because you can get the AIs to code, I'm starting to see people who have never coded before. I'm working with one guy right now who's never coded before but does have a behavioral-science background, and he's starting to do legitimate frontier research on how our AIs are going to behave under various esoteric circumstances.
So I think nobody should count themselves out from the ability to contribute to figuring this out and even to shaping this phenomenon. It is not just something that technical minds can contribute to at this point. Literally, philosophers, fiction writers, people just messing around, and Pliny the Prompter, the jailbreaker—all of them represent cognitive profiles that would be really valuable to add to the mix of people trying to figure out what's going on with AI. So, come one, come all, is kind of my attitude on that.
Erik Torenberg
That's a great place to wrap. Nathan, thank you so much for coming on the podcast.
Nathan Labenz
Thank you, Erik. It's been fun.