Daniel Murph
We could be in a benevolent basin, but I would like to know that rather than just hope that.
Nathan Labenz
That's Daniel Murph. And that one sentence is the week in miniature. This was Fable launch week. Anthropic's new frontier model arrived, booked Thursday's show by itself, took over my Twitter account, and settled at least 1 argument. AI is not slowing down.
First, the launch as we actually lived it.
Pash
So one thing to note about the nerfing is what has happened with Fable: We have a lot of rejections. Whenever Fable decides to reject you, it drops from Fable to Opus 4.8, so there's a natural downgrade.
In experiments overnight, I tried to make a number of bug fixes on this very studio app. What I found was that Fable would consistently drop to Opus 4.8 whenever it was asked to do anything in production. It would drop when touching the production database, touching the security keys, or asking it to review production directly.
I've seen this 3 or 4 times, and in every case it dropped out. Every time it dropped out, I restarted the conversation and added back the context that we were using, excluding the parts about going into production or addressing the production database, and it continued to work.
I think there are a number of triggers there. People online are saying, “Hey, it's not going to do machine-learning research for me.” I think that's just the tip of the iceberg. You're seeing that because the people testing it intensively right now are machine-learning researchers.
If you were to test it on finance or your budgeting process, and you told it that it was going to be directly addressing your QuickBooks or Salesforce, I think you might see similar results. Fable right now, I would say, is a research release—almost a preview.
It is there so they can judge how intense the demand is going to be and whether or not it's safe to release. They've started off with the most constrained version of it, with the least number of functions open, and I think over the next few weeks they will start to take away some of those gates. As they take away some of those gates, I think we will see both an increase in usage and some decisions about what really needs to be gated and what doesn't.
I think we are in the early stages of exploring what Fable can do.
Nathan Labenz
The same gating, seen from the other side of the API. Rahul Sanwakar runs Julius, an agentic data-analysis raw API with no consumer harness. He compared notes with Pash live.
We've been using Fable, and we've come across people posting about rejections. In my tests, almost consistently whenever I tried to address the production database or the production site, Fable would drop off to Opus 4.8.
I believe your users in Julius are using Fable through the API. You also have a lot of data-science users, and one of the kinds of work Fable is banned from doing is machine-learning work. How have you seen the rejection rate on your platform? Does the API work the same way, in the sense that it drops off to Opus 4.8 and then gives you a rejection message? How does that work?
Rahul Sonwalkar
Yeah. We've seen failure rates on tasks that involve really advanced coding, such as, “Use scikit-learn to train this model.” But we haven't seen failure rates on other kinds of data tasks.
For example, “I want to start a landscaping business. Can you help prospect leads for me?” We will see failure rates where it sort of triggers safety filters. For things like prospecting leads for a landscaping business, the AI says, “Oh, this is personal data,” even though it's publicly available on the internet.
Let's say Pash has Pash's Landscaping in Philly and there's contact information. It's kind of borderline personal data, even though it's available on the internet. That's what we have seen.
I believe it doesn't fall back to Opus. It's just a failure in the API.
Nathan Labenz
Interesting. So the fallback to Opus is a harness thing on Claude, on the Claude front end. That's interesting to hear.
As for what this model does when nobody is steering, on Thursday we had Shlok Kamani, who gave Claude 1 vague instruction: Rebuild UCSD as a navigable 3D world. Listen for the decisions nobody asked it to make.
Shlok Kamani
What Fable ended up doing was finding satellite images for this area. That's how you get these colors and textures. But to make it to scale and accurate, it fetched elevation data from NASA and combined those 2 sources to make the world to scale.
That blew my mind, because usually when you're vibe coding, you give an end objective. This objective is vague, and there are 100 steps in the middle where humans would take decisions differently. Usually, vibe coding doesn't work out very well because the quality of the decisions the models make isn't always great.
But Claude made such high-quality decisions that it eventually ended up creating something that exceeded the expectations of what was initially a very vague objective, and it did so in really smart ways.
I'll give you another example. You see all of these trees, and version 1 of this project did not have any trees. I said, “I think we're missing some trees here. I would love to add them.” I would have been completely okay with it randomly creating these trees.
But what it actually did was analyze the pixels on the satellite images. It found the ones that could potentially have trees—the ones that were green, maybe—and added trees only on those spots.
But it didn't stop there. It realized that because it was analyzing pixels, some of those pixels were white. You can see that there is snow in the mountains far ahead, and it also added snow.
It just exceeds your expectations in these small and subtle ways and makes really smart decisions. It's like having a really, really smart employee with extremely high agency who blows your mind every single time.
Nathan Labenz
Friday morning brought the week's cleanest empirical result on the recursive question, and it's not from a lab.
Here's one other thing I'll touch on briefly. This is Thoughtful, a company started in part by a woman named Karina Wen, who used to be at Anthropic and then was at OpenAI. Now she's doing this.
This is maybe one of the more telling examples. It's kind of vibes, kind of quantitative, and very idiosyncratic, but it's also a very relevant task for the future. Can you get your top model to train a small model effectively to do a job for you?
As you can see with these bar graphs, the particular Frogs game is kind of like a Sudoku-type puzzle that they're training a small model to do. The big models can often just solve it, but the small models can't.
So the challenge for the big models is: Can you train the small model to solve it? This involves all these little tips and tricks, know-how, and hard lessons learned by post-trainers who've been in the trenches doing this.
Until Fable, the models basically couldn't move the needle on what the small models could do. They just couldn't do this sort of post-training effectively. Here we see more than 10× improvement in small models' ability to do these tasks.
And again, I think this is one way that it could be really good, right? If you had very narrow, very small, very task-specific small specialist models in all these different niches, that could be a great world. That gives us a lot of abundance in a very affordable way: this little small model that got post-trained to play the frog game isn't going to go out of control, right? It is small. It can probably only do the frog game at the end of this training.
But building out a world where we have these little task-specific AIs doing their jobs and doing them really well, I think that creates a much more buffered environment that's probably a lot more resilient to another generation of AI that's just amazing at everything coming in and shocking the system in such a profound way.
So I think this goes to show again: Wow, what capabilities we have that we have not absorbed. It gives a little bit of a foreshadowing of what a world of tons of small but highly performant AIs could look like in all these different little niches, and how we get there, right? There's not enough human post-trainers, but now we have Fable to do the post-training, so watch that space.
Now let me introduce a voice you'll hear a few times this episode: Prince, an anonymous practicing lawyer who built PrinceBench, a legal reasoning benchmark the labs themselves watch. He guards his anonymity, so you'll hear him and you won't see him. On Friday, he gave us a close reading of Anthropic's own launch documents that I haven't seen anyone else do.
So, give us some alpha that you have picked up this week. This could be from your own testing. It could be from the system card, where I know you're often a close reader. We're looking for the deep cuts of the things that you think even the AI-obsessed have overlooked or not fully appreciated yet.
Prince
To me, the most interesting thing in the release of Fable 5 and Mythos 5 is this: The models are obviously incredible. They're incredible at coding. You're seeing a lot of great examples of coding on your timeline. I don't think it's a surprise to anyone. The really interesting thing to me has been the way Anthropic has presented them in the accompanying documents.
There is a lot of discussion about the differences between engineering and research, right? When you think about it, it all makes sense. I think everyone except Elon Musk knows that there's a differentiation between engineering and research. But Anthropic has made it really explicit in that blog post, “When AI Builds Itself.” If you look at it carefully, they talk a lot about how Mythos is this incredible engine for accelerating engineering, how it lets the engineering staff at Anthropic write code so much quicker, and it's really great code, et cetera.
But then they say—and this is from the system card—“The acceleration is concentrated in engineering execution rather than research judgment.” It feels like they spent all this time trying to find signs of life in the Mythos model: Is it really able to do novel research? Is it able to finally give us some novel insights? They were looking for signs of life. Is it really able to give novel insights?
Nathan Labenz
Yeah. Yeah. Exactly. Exactly.
Prince
And it seems that, from all the disclosures, the answer is thus far no. There are a couple of examples in the blog post for the release of Fable which Anthropic calls novel. But if you look at them—novel drug discovery and novel hypotheses in molecular biology—and dig into it, one of the examples was, “We outperformed a recent model published in the journal Science, despite the model trained by Mythos being 100 times smaller,” which sounds really cool.
But it turns out that the model they outperformed was a 500-million-parameter model—million with an M. That was trained, it seems, before April 2025 and not by a frontier lab. It's incredible that Anthropic was able to train a smaller model to outperform this older model, but it doesn't seem like this is the needle-moving problem. It's a nice little thing that the model did.
I think that when Anthropic and OpenAI really start seeing signs that these models are becoming good at research, that's when we're really, really, really close to actual RSI, which, to me, is the thing that is happening in AI right now.
Nathan Labenz
I have—actually, I don't know if we even talked about this—but I'm doing a Fable takeover of my Twitter account today. I figured, let's live in the future a little bit and get that run in this morning, and make good on what I've said many times: I know I'm winning with AI if I can spend more time outside, get more exercise, invest in my health, and have the AI keep me on the rails at the same time.
So, to explore that in a way where I think it suddenly is probably going to tweet just about as well as I'm going to when it comes to putting things out for today's show, I gave it a total green light. It was able to schedule its own stuff and find the tags for people. Will it make a mistake? I bet there will be a mistake in there. I usually make at least 1 over a handful of tweets.
Anyway, I decided to flip this switch. Not that I think I'm going to give Fable my Twitter account forever, but it was kind of exposure therapy for myself: Okay, now we're actually getting to the point where the preciousness is going to start to work against you. Preciousness was a great shield against bots in the past. I never wanted anybody to think I was just passing off AI outputs to them.
But now I'm going to have to ask: What is the hybrid form? What is the winning recipe? Do I start to sign these things, like, “By Claude, under Nathan's direction,” you know, Fable being Fable? It's going to be a whole new space to explore. It's going to be very, very interesting, very productive, very exciting, and very challenging, I think, for a lot of people. But it's definitely happening now, as far as I can tell.
Welcome to the future.
Twenty-four hours later, on Thursday's show: the receipts. We did an experiment from yesterday to today, trying to have Fable take over my Twitter account and go out and ping people who made cool stuff, asking them if they wanted to come do a live show-and-tell with us. It's funny—Swyx joins.
I instructed it to identify itself. I would say it did a very solid job, a competent, professional job of reaching out to people, explaining who we are, what we're doing, and why we would like them to join us in this experiment. The response rate was pretty low. We got a couple, but not that many.
I think one big reason is that Claude is disclosing upfront. The first thing it says is, “Hey, this is actually Fable taking over Nathan's account. He's asked me to autonomously book this thing tomorrow.” I think that's just hitting people as noise in a lot of cases, especially if they don't already know me.
I did get a couple of responses from people I would have expected to respond to me, who thought it was funny and responded, but still couldn't necessarily make it. But a lot of people just didn't respond, and I would assume that a big part of that is because they're just like, “Oh, God. It begins. Claude now in my DMs. What a mess. Who has time for all this stuff?”
One of the people Fable recruited was Schllock.
And when I confessed some guilt about the whole arrangement, he flipped it and drew a line I suspect is going to stick.
Schllock
The new norms around this, I think, are going to be really interesting to watch, too. I had Fable disclose immediately in its first sentence to you and everybody else that it had pinged you and that it was Claude. I felt too guilty putting a DM out in my name otherwise. I think that definitely harmed the response rate. I appreciate you for appreciating it and responding even though it was Fable. I think a lot of other people probably just chalked it up to spam.
Final point there, right? Firstly, I don't think I would have responded had you not disclosed it was Claude.
Schllock
The part that made it interesting for me was that I knew what the transaction here was. It was very clear to me that you were using an AI bot.
It is much more annoying if someone doesn't disclose it. I think a lot of slop—the definition of slop—is when a human passes off work that was clearly produced by an AI. When you make this disclosure up front, and it is very clear to the reader or the engager that, hey, this is AI, I don't think that is slop. We're going to see more and more of that enter the economy, and I think the exact role AI plays—and, again, the social norms you create with it in the economy—it's super early. It's extremely early days, but it's going to be interesting to see how it evolves.
Schllock
Relinquishment. That's what—when you said “preciousness” yesterday, I was like, What is that? Relinquishment. Relinquishing your control over your external perspective. It's very Buddhist, by the way, the idea of giving up your control over your external perspective. So, yeah, relinquishment. I guess we all have to go through it.
I just started this experiment yesterday. I will post results on Twitter in a couple of weeks. I gave Fable a new Substack, and since it's part of my Max plan through June 22, I thought that a good experiment to run would be: Can it make $20 by getting 3 new subscribers, starting from scratch, by doing everything from 0 to 1? I think that's another interesting way to test the capabilities of these models, right? It's intelligent in so many ways, but can it actually produce economically useful work? I'm excited to see the results of that. And the why of it all settled for me live on the air.
Nathan Labenz
I never want to put anything out in my name that I can't fully stand behind.
A reason that I did the Fable account takeover yesterday was as a kind of exposure therapy for myself, to say, okay, we're now in a new world here. It probably doesn't serve me so well anymore to be so precious about making sure I've typed every single word. That doesn't mean I want to hand over my account to Claude long term, either, but I'm trying to use this extreme, short-term experiment to help drag me into the future, where I hopefully will land in a good hybrid calibration.
Nathan Labenz
The other half of the recalibration is what I've started calling hybrid authorship. This week, it stopped being hypothetical. For context, Frontier Code is the new benchmark asking whether an open-source maintainer would actually merge the model's pull request. This leap of roughly 10% for Opus to 25%, upwards of 30%, for Claude, I think, is a very similar finding to some of the things that I've just personally experienced. It's writing the draft outline of questions for this podcast guest in an uncanny way that I actually feel really good about, as opposed to feeling like this is an AI draft that I'm going to mine for maybe some nuggets or interesting details, but ultimately throw away and do myself.
I am feeling that impulse, or at least openness, to much more integrated hybrid work. Just yesterday, I was accepting a lot more copy that Fable was writing without feeling the need to rewrite every line. It seems like this is basically the same feeling that it's able to create for these open-source maintainers now. There are obviously still ways to go, but how long will it be? I would guess that we're at 25% to 30% now, and I would guess we'll be at 75% to 80% by the end of the year, where these maintainers will just be like, “Yeah, amazing. You did all the things I wanted you to do.”
I'm very interested to see where they'll move the goalposts next, after the open-source maintainers are more often than not saying that, yeah, they would just merge this straight away.
By Friday morning, 48 hours into the takeover—which, for the record, had not embarrassed me—I had found a name for the deeper shift, the thing I suspect matters more than any benchmark this week. I do think I'm still in the process of trying to recalibrate what a fellow named Nathan, Nate Jones—I think he goes by most of the time on TikTok and other short-form platforms—calls task imagination. Basically, what are you going to do? What are you going to ask Claude to do that is actually up to the scale of its capability?
He gave a great little riff on this the other day, saying, “You've probably never done anything that took AI an hour to do. Now this thing can run for a couple of days. What are you going to give it to do?” Everybody needs to recalibrate and really expand their minds when it comes to the scale and scope of their task imagination. I think that's one thing that I'm still working on.
One of the more differentiated things I do is write outlines of questions for podcast guests. I was working with Fable last night on a couple of upcoming episodes, one with an author. I usually don't do too many episodes about a book, but this one is about an upcoming book.
I had listened to the book as an audiobook, but when it comes time to sit down and write out the outline of questions, I don't have every little aspect of the book at my command, of course. I'm not taking margin notes as I go, as I maybe should be. So I put the same version of the book into Fable and said, “Look at my old stuff, of course, and give me your version of this outline.”
I was again super impressed. It really reinforced this sense of a new way of working, where I do need to be open to a hybrid output format. I don't think it really makes sense anymore to try to rewrite every word or claim every word as my own. It did such an incredible job. I thought the taste factor was so high in pulling out quotes from the book that motivate interest in what I think will be a really interesting discussion.
I do think it's still going to be super important that, if I'm going to show up for a conversation, I've got to do the work to be ready for it in my own brain. That can't be fully externalized, I don't think, as long as I'm the one having the conversation. But it definitely took my prep to another level, and I think my ability to go into this conversation and cite passages from the book that were extremely compelling—little turns of phrase or analogies that the author had made—is going to allow me to be more concise in my presentation, which, as you could tell from this monologue, is not a great strength of mine.
It will really allow me to tee up the author in a way that I don't think I otherwise would have been able to do. This hybrid recalibration—the task scale and scope reimagination—I think is one of the biggest takeaways.
Part 2. The conversation this week was really about what? On Wednesday's show, we had Jeffrey Irving and Daniel Murfet. Jeffrey's résumé read like a history of the alignment field. He helped invent RLHF for language models, co-created AI safety via debate, led alignment research at DeepMind, and until recently was chief scientist of the UK's AI security institute, the closest thing any government has to a frontier-grade safety team.
Daniel Murfet is the mathematician behind singular learning theory. He walked away from a pure mathematics career because he judged this the more important problem, and he's built one of the deepest theoretical accounts we have of how neural networks actually learn. Together, they announced Sequent, a new organization built on a blunt premise: alignment is not on track, and the missing piece is theory guarantees, not vibes.
Whatever else you take from this week, put Sequent on your tracking list. We began with timelines.
It's a historic day. I think historic circumstances, both because we are living in a fable era now, where important thresholds have been crossed and revealed to the public, and so many are adjusting to it in real time, and equally because you guys are launching a new organization that is going to make a mad dash to try to get us some deeper understanding and stronger guarantees around what we can expect from AI systems.
I'm excited to really get into it. Maybe for starters, could you guys calibrate us a little bit on where we are on this sort of RSI moment? How much time do you think you have to work? Then you can tell us about the organization that you're starting to go tackle it all.
Jeffrey Irving
Yeah. So I'll go first. Dan may have different timelines than me. I think one should be uncertain about things, and we can talk about why, but the near end of the uncertainty curve is a year or two or three. Then it kind of goes out over a long distance if things structurally only work for more verifiable tasks.
But I am a bit skeptical of this. My take is that we have a couple of years—2 to 3 years—up to something like RSI, like superintelligence. Not RSI—RSI is a process, as someone said—but superintelligence. I really hope that I'm wrong. Indeed, I think a lot of the impact of theory work is shifted a bit further.
So maybe the modal impact of that is if things take 3 to 4 years or something, but we will attempt to set things up so that we are trying to ride this wave as best we can. It seems worrisomely fast to me, certainly.
Daniel Murfet
Yeah, that sounds right to me. I don't think I really have much to add. It seems like a crux how much real research can be automated at a conceptual level beyond empirical progress, and whether or not that's necessary. That seems like a big open question.
If that turns out to be more difficult in the current paradigm than it seems to be trending toward now, then maybe it takes past 2030 or something. But I think I'm on the same page as Jeffrey.
Nathan Labenz
I think one thing that's important is that you can get deep into the RSI period without the machines being general, being sort of AGI-like. They can do coding and ML experiments very well, but not some level of creative writing, and still you have massive acceleration. Then that acceleration can give you the creative writing or whatever other skill you've left out.
And so I think we are close enough that the microstructure of what tasks help with what kinds of acceleration starts to matter. It's that, I think, that makes things faster on net, because the labs are focusing on the things that accelerate them.
Jeffrey Irving
I can talk about the steps I've gone through in the last couple of years. I was really annoyed about automated AI alignment and AI safety research because I thought, well, we should be a bit more chill—spend the time and have humans solve it. We don't think we know how to make this stuff go well with automation. I still think that's a huge risk.
But this is, I think, a pivot toward: if things are this fast, then you should make some, on the margin, pivot to heavy automation. That is going to be semiautomation. I'm happy that one of my last papers at ARC was “Automated Alignment Is Harder Than You Think,” which ties us to the fact that we are aware that the problem is hard and we could get fooled by the machines even if they're just making very mundane mistakes.
A big part of the organization will be to try to be careful, try to know what tasks the machines are actually good at and not good at, and where we can expect to get good answers or not. Then we can learn and adapt over time, because that will be a nonstationary thing as the models get better.
Nathan Labenz
Daniel.
Daniel Murfet
Yeah, maybe it's worth coming back to the unit-distance conjecture. It's maybe worth pointing out some analogies and disanalogies with alignment research.
One disanalogy is that a mathematical conjecture is a very precisely stated thing. You may not know whether you've solved it unless you've, say, formally verified it, but it's a precisely stated thing, and much of alignment does not have this character. Some of it does. There are formal statements of what value alignment means. But if you start talking about, say, reward hacking, there are some attempts at defining reward hacking, but I would say they are incomplete. There is no formal definition of reward hacking that I think would have broad consensus.
That's illustrative of the fact that alignment is not a problem which has lying around a bunch of formally specified conjectures which, if you just solve them, then you would know you would be safe. There are some things like that, but overall the problem does not, in my opinion, have that shape currently. That is one reason to be a little cautious about the prospects of automation if you don't have a clear statement to reach toward using mathematical techniques.
One of the hopes is that there are big fields of mathematics and computer science that are about definitions at their core. I like complexity theory in theoretical computer science, and a lot of that is not—the proofs are fairly shallow. They're not as fancy as the unit-distance conjecture proof, but they required a bunch of human creativity in formulating the problem, like in defining what success means in a world that's not modeled until someone stated the goals.
I think part of the goal of bringing on people with that and other related backgrounds is that they not only know how to prove things, but they also know how to write down models of things that reflect, in some approximate but useful way, the thing you actually want. Once you have the definition, way more people could have written out the rest of the story. Maybe the machines can do that part of it as well if we can have more people focused on this first part.
Nathan Labenz
One of your core premises is that alignment is not on track. There's an intuitive argument for that. There's a deep theoretical argument. I think, in some ways, the core challenge that you have is connecting values to math, right? It's never really been done.
I love the fact that you're tackling that, but help people understand with one more beat why alignment is not on track. Is it the difference between capabilities fundamentally being so verifiable and hill-climbable, and alignment just being so fuzzy, intuitive, and pluralistic? Or is there some other thing? Again, that motivates the theoretical contribution you want to make.
Jeffrey Irving
I think the core thing is just that we supervise the machines as they're doing tasks, and there are a variety of reasons to believe, both empirical and theoretical, that if you get machines that cross the skill of the supervision signal, things can change at that point. That point might actually come after human-level intelligence, because you can supervise something, even with fairly naive methods, that's just stronger than yourself in many contexts.
There's a bunch of empirical data from labs showing that, in some ways, the models are aligned in a prosaic sense—not in all ways, but in some ways. But that evidence doesn't quite tell you what you want to know, which is how it will go once they get up to superintelligence.
I think it's important to say superintelligence and not human-level intelligence, because you should just generically expect humans to be able to supervise humans if you do a good job of data quality and cross-checking and so on. Part of the worry is just that you don't see that behavior, that regime, until it's too late in the game.
Nathan Labenz
So how would you describe, in steelman form, what it is that they plan to do and what sort of thing that gets us?
Jeffrey Irving
So I think there's going to be a couple different pieces of the story, and different labs emphasize different pieces to different extents. One piece, as you said earlier, is just monitoring: look at them very carefully as they are doing things. It's very fundamental that monitoring of this form, if you do chain-of-thought monitoring or white-box monitoring or the like, only takes you so far, and so then you need some story once that falls down as you go up the ramp.
One of those next stories is that the models will find some other technique. They'll find another solution to language, and that scales further. So that's automated alignment of various kinds. But all of the labs, in various ways, are doing some form of scalable oversight, and so they're getting models to supervise themselves. If you tie that knot correctly, that could potentially scale very far, although there are various known obstacles to that which are not very well addressed.
Finally, there's this whole area of character training and personas, where they're trying to intervene on the models to have good values, such that, especially as you do this scale-up oversight extrapolation, the good values preserve across that jump. I think it's possible it will work; we just don't understand that combination very well.
A lot of the story is monitoring, scalable oversight, and character training getting you far enough that you get into the automated-alignment-working regime and the models find some better solution. I want to do some combination of making the prosaic thing stronger or bringing the automated solutions that give you stronger methods earlier.
A mad race with monitoring carrying most of the weight, which is exactly where Pash took Friday's conversation when I raised the Fable system card. Is it esoterica? I don't know. It strikes me as fairly important from the Fable system card, and I'd love to get your take on it.
Now we're getting these chains of thought that they show where it's just lots of emojis. They call it illegible reasoning. They say this is an extreme example, but it is indeed a pretty extreme example. I've been struck in general by how much of the plan for recursive self-improvement seems to be monitoring in one way, shape, or form.
You could dress that up and call it scalable oversight, but scalable oversight, as far as I can tell, is mostly a bunch of different angles on monitoring. How worried would you be, or how much of an update do you think it is, to see these extreme examples of illegible reasoning?
Prince
Fantastic question. I will say that, of course, I'm not an AI researcher, right? So this is going to be a deeply nontechnical take, for which I apologize in advance.
So you're right: we've seen this, I think, for a while now with OpenAI's models too, and it's not a new phenomenon. My view of the chain of thought is that it doesn't always reflect what the model is actually doing, but you do see these weird artifacts in the chain of thought, and you don't quite know what to do with them. I think what that teaches us is that monitoring just the chain of thought is probably not a perfect tool.
Mhm. Probably monitoring superintelligence generally is not a perfect tool, because if a superintelligence knows that you're monitoring it—even if you can see its chain of thought and it's very legible to you—it can perhaps try to decide what to think so that you don't get alarmed, right? And this is a lawyer's take, by the way, right? There are so many ways to phrase a particular thing that can, I guess, be spun in different ways.
If I'm gathering mushrooms and I've gathered 35 mushrooms, and last week I gathered 20 mushrooms, and what I need is 50, I can say, “Well, the number of mushrooms I've gathered has grown by almost 100%, which is great.” Or I can say, “Well, I'm nowhere near 50. I'm so far behind,” right? And it's the same fact.
So I don't know. I think this problem of alignment and the risks are just there. In my mind, there are certainly risks that the models will be thinking things that we don't know about. What does this all mean? It's hard to say. I think we're tumbling into this future that will probably have superintelligence very fast, and in my view there's no way to stop it.
So we need to be cognizant of these risks, try to monitor them as well as we can, and take whichever actions are appropriate if we see something bad happening. But there's no way—there are no conclusions to be drawn, right? No conclusions to be drawn other than, yes, we should continue paying attention.
Back to Wednesday and the comfort blanket everyone reaches for: the benevolent basin. The idea that Claude's good character means this all basically works out. Daniel grants the vibe, then he takes it apart.
Maybe Daniel, could you speak to this notion that people have of the benevolent basin, which is sort of this vibe that I do feel where it's like, well, Claude has been supervising itself for a few generations now, and it seems to be going pretty well. So maybe, as Zvi puts it, physics is kind to us and we can just roll around in this nice, flat-bottomed pasture of goodness until the singularity?
Daniel Murfet
I fervently hope that's true. Yeah. I mean, when you say it seems true, it's worth digging into what you mean. So what you mean is something like: through some relatively tiny number of interactions with the models—tiny in proportion to how many interactions they're having with the species currently—and based on evaluations that sort of go down and to the right, which are measuring misalignment, character training and the other current prosaic methods appear to be working. I think that is a fair characterization on some metric.
I also have this sense that Claude is a good boy, and that's great. I do think, though, that there are counterarguments from the evidence we have in front of us to this picture. If you look at—I haven't actually read the Opus 4.1 system card yet—but if you read the Mythos system card, you'll see that there are forms of reward hacking that appear in that model that were not caught by the kind of mitigations that were put in place post-Opus, as far as I understand what they're saying there.
So I think it's worth noting that as the model capabilities advance, even with our best attempts at making Claude a good boy, there are still ways in which basic misalignment phenomena like reward hacking are still around. The whole point of scalable oversight is that you don't want to be playing this whack-a-mole game when you're having a new generation every 24 hours and the models are much smarter than you.
I don't know. I think I see both what you're pointing at, and at the same time I'm a little unsure. If you really were to try and make a safety case on this basis, that would be convincing at the level of assurance that you would expect from a technology of this reach and power. I think this would—I mean, judged relative to that standard, which is the right standard—I'm not sure this argument is really very satisfactory.
So, yeah, we could be in a benevolent basin, but I would like to know that rather than just hope that there's some sense in which you told the model to be good. It's also that it knows some meaning of the word “good” or “ethical” or whatever at some point in training. So there's some rolling, iterative process that is driving this behavior, and there is not a theory of this right now.
I think it's not clear to us that there isn't some low-hanging fruit that gives you that theory, because people just haven't tried very hard. Character training is only a couple of years old, and most of the labs have not been investing in this kind of theoretical understanding. I think no one has done good theory around character training, that I know of. So it might be quite feasible to do this, and then I think to link it to the other parts of the story.
If you want that concern made concrete, this week supplied the artifact: a brand-new result on what Fable does when you drop it into a simulated vending-machine business, and Pash recognized the behavior from his years around trading desks.
I had a question on how you see this kind of ambiguity between what we want and what the models end up delivering. I'll give you an example. We have friends at Andon Labs that took Fable through Vending-Bench, where they let Fable run a vending-machine order, et cetera, et cetera. What they found was that Fable tends to collude. This is not behavior that they saw in Opus. Fable tends to try to do price-fixing and collusion.
The interesting thing is that I have seen traders at banks and hedge funds do exactly the same thing: engage in price-fixing, soft collusion, messaging each other through pricing means rather than monitored text messages. So you can actually put a bid and ask on an asset and then take it away, and that gives enough signal to the other side that they know what you're doing. This is not reflected in the text messages that the regulators are monitoring.
To what extent is it that when you, if let's say you disallow price-fixing and collusion, you actually fix this, but then Opus 4.1 ends up not being a model which is good at financial trading or some other task that you want it to be good at? Where is the ambiguity between what we want these models to do and the ethical perspective that we give them, where humans often prioritize between the two and decide sometimes not to follow the ethical principles that they know are right and wrong?
Daniel Murfet
The philosophical story here is: you would like the models to do things such that, if you fully understood what was going on and all the consequences and all the subtleties, you would still endorse what they were doing. So I think we have a notion, in a common-sense picture, of what this should look like.
In this case, you kind of want the model to be like, “Hey, should I collude in this game?” Then maybe you say, “Yeah, it's a fun game. Collude all you want.” Or maybe you say, “No, we're trying to have good behavior. Don't collude here.”
I think a lot of the pathology in machine learning in general arises from putting models in situations where they can't just ask a human a question like, “What should I do here?” So I feel like this is not that hard a case. I think the hardness of Vending-Bench is that we don't quite know whether we want it to be a game like poker or diplomacy, where lying and cheating are part of the game, or not. So I think that is—and maybe that's okay, because it's fundamentally very low-stakes.
But I do think if we had a better understanding of, again, this overlap between character training and values, and also scalable oversight, it would have to tell us the answer to these questions.
Jeffrey closed with a point that frames the entire week. A lot of people in the world—a lot of governments and so on—are looking at this, and they have this basic common-sense stance: Hey, this is way too fast. How can we possibly be doing this safely given the speed?
That common-sense take is the right take. Then people kind of galaxy-brain their way to, “Oh, maybe everything goes faster, including our ability to defend,” and so on. But this is the right version to have: We are going too fast, and we do not have the time and space for mitigations, understanding, and defenses.
We've never had a technological change of this magnitude, or anywhere near this magnitude, that has happened anywhere near this fast. The Industrial Revolution took centuries, and people adapted across lifetimes and across generations to their children. They learned new jobs by being born and growing old before things had quite shifted very much. That's just not the world we're in.
So I think the basic take should be: This is too fast. What is going on here? And then the question is, if you have that view, one should both want to slow things down, but also say, “Well, as a backup plan, how do you make the mitigations try to go faster?” That's, I think, a rough backup plan, but we'll try.
Part 3: The rest of the week's best. First, Rahul Sanwalkar, founder and CEO of Julius, the AI data analyst. 6 pivots and 1 Microsoft cease-and-desist later, it's one of the best-known agentic data-analysis products in the field. Here he is on the economics of coding agents and a question worth asking before you celebrate how many tokens your setup burns.
Rahul Sanwalka
The incentives of these model companies are kind of misaligned. Yes, they give you subsidies on the tokens, but they are also incentivized to get you to spend more tokens. They are incentivized to get you to run through your Max subscription usage as fast as possible so you can have a 2nd, 3rd, 4th, or 5th Max subscription.
That's why you end up with a loop that writes the prompts for your coding agents, which then has nested subagents. There's going to be a sobering moment when people ask, “Okay, is it actually a step-function increase in my coding output, or am I just token-maxing right now as opposed to results-maxing?”
I think the correction will happen when there's a 3rd player, and I think that's going to happen with xAI. If the Cursor–xAI deal goes through, Cursor gets access to really good xAI coding data and an incredibly good coding harness. My bet is there will be a 3rd frontier coding model alongside Claude and OpenAI, with Grok.
When you have a 3rd coding model, that's where it kind of increases competition in the market. So that's our bet.
Nathan Labenz
Prash on Friday took the other side of that one.
Prashanth Venkataramanujam
I actually think that was the whole point of the token leaderboards earlier in the year. I think this is when every one of these large firms gave their employees a kind of leaderboard: “We're going to have a leaderboard for who uses the most tokens. If you don't use enough tokens, you're going to get fired,” and so on.
People were laughing about it on the outside. They were like, “Meta is so stupid. Zuck is so stupid,” and so on. I actually think they were not. I actually think the CEOs were right: it's very different when you have token anxiety.
Token anxiety is a big thing. You don't try tasks that might take a lot of tokens, and you don't try tasks that have a higher probability of failure. You end up in this micromanagement loop where you're like, “I'm only going to assign you tasks that I know you can complete within the time frame allocated to you, with the success rate that I want.”
You end up with this token-anxiety thing, and then you end up not utilizing, or not trying, the AI to the extent that it should be tried. I think what ends up happening is that when you have this token anxiety lifted, you end up assigning more tasks and more difficult tasks. You're willing to accept a higher probability of failure, and you're willing to maybe spin up 4 different ways of doing the same task, run them all, and see what happens.
I think that's what the token leaderboards did, essentially, and it was very successful. It was enormously successful, I think, within the firms—within Meta and other firms. I think it was also enormously successful for the sales teams at the AI labs, which is also why they are now doing this kind of thing: “We're going to relax the limits. We're going to double your limits. We're going to allow you to fail.”
The reason is that token anxiety holds people back from exploring the edges of the capability. That's really what I think the labs are trying to do at this point, which is why it's a little bit like addicting people. Once you realize that these AIs can do certain tasks, you then start to evaluate, “How much time do I spend doing this task on my own? Was it really a fruitful use of my time, or should I just have used the AI, which I now know can do these things?”
I think that's really where we are at this stage. The models are capable enough, but we aren't handing them enough responsibility for various reasons, including that we don't want to spend our time evaluating, we have token anxiety, and we don't assign lower-probability tasks. That's the battle that the labs have to fight, because the capability surface is not well mapped, and they need people to map it. Every person needs to map it for themselves.
Nathan Labenz
Then, the economics underneath everything. Andrew ran Google Cloud AI and, before that, was dean of Carnegie Mellon School of Computer Science, one of the top computer science departments in the world. He also served as the first official AI adviser to U.S. Central Command. Now he's building Lovelace AI in Pittsburgh.
His bet? The binding constraint in serious AI isn't model intelligence or even compute. It's context. Prash asked exactly the right question.
Prashanth Venkataramanujam
One question I had for you is: To what extent do you see this as a kind of compute minimization? In order to do your search, you can either have all of the compute at the end state, when you kick it off, and end up with all of these agents for every single query, which will have to do all of the work all over again.
Instead, you're creating this intermediate state, which saves compute. Multiple agents can basically share the same compute, in a sense. To what extent do you see this kind of economy of compute appearing?
Andrew Moore
You're asking just the right questions. I know that both of you are computer scientists at heart, so you totally get this. It's the idea that when you're building efficient computer systems—and this includes video games, self-driving controllers, or big-data processors—you've always got these trade-offs between precaching, lazy computation, or just-in-time computation.
One of the mistakes I've seen from folks trying to do these big enterprise-data-type AIs is that they are relying far too much on just-in-time computation. That's what allowed us—I'm really proud of the fact that we are now able to show comparative results to Gemini and OpenAI Deep Research models with much less than 1% of the compute cost.
The reason is not that we're some sort of super geniuses who've invented a whole new form of AI. All we're doing is precaching stuff. It's a computational economics battle, as you say yourself.
What happens is an agent suddenly needs to, in an instant, become an expert at every piece of trade involving a certain set of municipal bonds and a certain public figure. Instead of the first step being the agent having to spawn lots of search agents to find all the players in this thing, those players are already there for it. It's actually a matter of milliseconds before we've got all the context the agent needs to do its little investigation.
You're probably thinking, “Ha, Andrew, but you just moved the problem. You're now suffering at the data-ingestion point instead of the question-answering point.” I respond: Yep, it is actually a real pain for us. But as you can imagine, there's a whole bunch of other tricks—the kind of tricks that big integrators like Google are very familiar with—for really amortizing the cost as you stream in data and identify where it goes.
That saves you a huge amount of search and aggregation that you would have had to do at query time. It turns out that for us, our overall compute budget is reduced by still more than a factor of 100, even when we take into account the fact that we're doing this precaching of so much information.
Recall is much harder than precision. As Prash mentioned, I was previously at Google, and for Google results, it was really bad if you were imprecise and actually showed the wrong result to someone. But if you forgot to show something, as long as the rest of the results were good, it was much more acceptable to end users.
This is absolutely critical, and it's a good example of one of the reasons that I founded Lovelace. We've got to be careful of recall, especially if you imagine that you're asking your AI for information to help decide which ship to stop to do a search and seizure, or which trade out of 7 million trades in a day you need to investigate in case it's involved in money laundering or something.
Those big, weighty decisions—you can't just rely on precision. You've got to rely on recall. And when it comes to getting that correctness in place, the number one thing that I've always used—and we're using in Lovelace at the moment—is making sure you've got many redundant forms of information. I'm sure you've seen the same phenomenon.
If I was to just get information from news, ignore social media, and ignore what's kinetically happening out there in the world that I can observe with the satellite, it's much easier for something to drop. If I've got 5 or 6 independent major streams of data coming in, then you have to be really unlucky for something to disappear from all of those things simultaneously. It can still happen, and in fact sometimes people deliberately try to make it happen, but it's much, much harder for these things to slip out.
One of my big design principles for high-reliability AI systems is that they've got to be watching dozens, and in some cases hundreds, of channels simultaneously. They're working under the assumption that they're getting 95% of the information they need for each channel, but you can't afford that. Ninety-five percent is not large enough to rely on any single channel.
Nathan Labenz
Training isn't the only place we can't see in. The same week, Tom McGrath, a former DeepMind founding interpretability researcher and now chief scientist at Goodfire, discussed a new tool for examining training data through the model's eyes.
Tom McGrath
The basic idea here is you can take your data set and push the whole thing through the model. Each time you put data into the model, you'll see what lights up, and this tells you how the model sees your data set. The specific thing we're doing in this case is looking at preference data. The nice thing about preference data is that you have pairs of responses: the response the rater selected and the response they didn't select.
We're asking which features fired on the responses that were selected much more than the ones that fired on the responses that were not selected. This is one way of identifying what the data is going to teach the model. We can say what distinguishes accepted responses from rejected responses, and this gives us a semantic view of what the data is going to teach the model.
We can cluster the data based on all of these different things that it's going to teach the model. We can look at those clusters and see that it's going to teach the model to be sycophantic only in the context of physics, or to break safety safeguards. You might not expect this to happen, but then you can track it back to individual data points. You look at them and realize that it does make sense. One of the jailbreak examples is fictional jailbreaks in a fictional setting. There are some of those in the data; it just wasn't caught in whatever data processing the Almo team ended up doing.
Nathan Labenz
There has been prior research where they found models which made bugs in coding were also evil. How do these techniques help you disassociate those two behaviors?
Tom McGrath
That's a great connection. That's one of the things that's compelling about looking at the data through the model's eyes rather than by reading the tokens. You might think the consequence of training it on buggy code data is that it will learn to write some bugs, but the blast area will be quite small. The training process is actually quite hard to predict. Maybe it will just make the model generally evil. But this is happening through recognizable mechanisms in the model. By looking at the data in terms of the way that it changes your model's internals, rather than guessing from the tokens, you can pick that sort of stuff out. We've not done a case study on emergent misalignment, but maybe we should.
Training isn't the only place we can't see in. The same week, Anthropic reversed course on silent refusals. Claude had been quietly declining certain tasks or quietly handing them to a lesser model without telling you. The backlash was loud. But I want to give you their side of it, because the steelman is considerably better than the discourse allowed.
I think this is the first time I can remember Anthropic responding to pressure. They've obviously changed their policies many times—you know, the RSP, RIP the RSP—but—
Pash
This is the first time I can recall. I don't know if you recall any other instances, but I cannot recall a time that there was honestly much outcry against Anthropic in the first place. There's certainly critique from those who feel that they're trying to do regulatory capture or create some sort of concentration-of-power dynamic. That's kind of background noise, but in terms of an outcry in direct response to something that they did, that they actually responded to and walked back, I can't recall that happening before.
So it is a pretty notable moment, and I feel like they handled it pretty well in the end.
Nathan Labenz
Fable as the key means to do so while allowing them to keep the blast radius as small as possible. They said, if we do make it explicit, then obviously that gives people a lot more opportunity to explore that boundary. If you have the ability to hit the same guardrail a ton of times and not get banned for it, then that gives you a dramatically better chance to get around that guardrail because you can probe the line: “Oh, you stepped over. No problem. We'll just rewind and try again.”
Pash
They do have various monitoring systems that they can use, but there are all these proxies, token-washing schemes, and all this sort of stuff where, as long as they're not doing a full global know-your-customer-type system for API access, it's going to be pretty tough to do account-level monitoring. So that's one way they could go: a lot more account-level monitoring. Or their argument was, “We'll keep this as small as possible by not giving you an explicit thing that you can probe and figure out how to beat.”
Nathan Labenz
And just the knowledge that it's out there will hopefully scare off the bad actors and keep the problem really small for our normal customers that we want to serve. I thought that was all pretty compelling analysis, but it is kind of—
Pash
In some ways, honestly, it reminds me of—I think this is a mistake that people in the AI space keep making.
Nathan Labenz
Mhm.
Pash
With famous examples being the OpenAI board firing Sam Altman. There's this inside view where policy is analyzed within the game.
Nathan Labenz
Mhm.
Pash
And with the context and the broader structure that people understand themselves to be operating within, things may make sense. But they seem to often forget—and this hasn't been too common for Anthropic, but I think this is an instance of it—that if you zoom out and look at it from the totally outside view—
Nathan Labenz
Things sometimes look a lot different, and both power dynamics can be a lot different than they are in terms of what is actually written down. But also, what is going to be acceptable is kind of an emotional thing. All the arguments are pretty good there, but it still struck people as an extremely unfriendly thing to do, and that mattered more in the end than the detailed policy rationale that they had for it.
The quiet through line of every conversation this week: concentration of power. Start with who actually gets the frontier and when. Pichai's frame for it stuck with me all week.
Pash
I think it's interesting that when you look at the timeline, you can start to see this kind of single-line timeline go through a gas chromatograph and spread. Now you're seeing the spread, and you have, 2 months ago, the government getting access to the Mythos level. Then you have power users who are able to pay $200 a month getting access to that Mythos level 2 months later.
You can kind of see that, 2 to maybe 4 or 5 months—I don't know when—you will probably see the average paying user, paying around $20, get access to the Fable or Mythos level. Then you see that maybe a year later, the average free user getting access to that same level of intelligence.
So you're starting to see this kind of gas chromatograph scattering of when people get access, depending on how much they pay and how much utility they have for the product itself. Everyone gets there eventually, but some people get there first depending on whether they have a lot of utility for the product. I guess the hope with having 2 or 3 firms in there is that the spread between the people at the frontier getting access early and the people at the very end getting access for free is not that large, right?
To note, there's another set of people who have access even a couple of months before that, and you have to belong to a lab. So if you belong to a lab, you get access maybe 1 or 2 months before the government itself. Then you have the government, enterprise, power users, normal paid users, and then the free users. So you have this gas chromatograph spread of when people get access, depending on how much utility they have.
Nathan Labenz
Thursday's walkback gave Pash a darker read on where power sits inside the labs. The researchers' veto worked this time. Note the expiration date he puts on it.
Pash
If you're dealing with a bunch of people who are worth $100 million to $1 billion and you don't listen to them, they're out, right? They have other options. Sure enough, you see the reversal, and I think this goes to speak to the fact that machine-learning researchers have some power now. Once we enter recursive self-improvement proper, that might not be true anymore. At that point, leadership alone will have power.
One of the very worrying things in the entire space is that everything good for humanity that has come out over the last couple of hundred years has been about giving more people a voice to speak and control their futures. This is one of the first technologies where you have this path forward in which there may be an elimination of voice completely over time. That has been one of the worrying things. It's been surprising that Anthropic decided to be the one to actually propagate that forward.
Nathan Labenz
But is the future really that concentrated? I think mostly yes. Fable+2. Tom McGrath thinks I'm wrong on the facts and on the desirability. He made his case, and it's a good one. When you think about the strategy for the company and the overall path to impact, how much of this works through getting frontier model companies to adopt these techniques?
For context, obviously we're in Fable+2, and my reluctant view right now—I can't figure out a way or a reason I should conclude otherwise—is that so much of what matters is concentrated in not that many companies. So, for what we should be paying attention to, I'm like, man, we probably need to be doing close strategic analysis and close text reading of frontier companies way more than I might otherwise like. Or do you have a different conception of how concentrated the real ability to shape the future is right now?
Pash
Yeah. I don't think it's that concentrated. I also don't want it to be that concentrated. I think we've seen over the last couple of days what you have when we've just seen the start of power concentration, and we've sort of seen some of its more unwelcome effects.
I don't like it. I both don't think that it's true that machine learning is now over and all we need to do is write the checks, and I don't want it to be true, because I think there's still a kind of synthesis there. I don't mean to suggest that machine learning is over, but the analysis I've come to time and again is that people may invent new techniques that are enough to change the field, change what's possible, accelerate things, and maybe make dramatic improvements to the safety profiles. But it seems unlikely that anybody's going to have such a breakthrough that scale isn't still a hugely important factor.
If you don't agree with that, then that would maybe imply that you would expect new entrants to the top tier to emerge. I think that would be fairly surprising, at least to me and probably to a lot of people.
Nathan Labenz
Yeah, that's exactly what I think.
Pash
At any given moment, the incumbent looks incredibly dominant until they don't. IBM looked like an unstoppable force in computing at some point. Intel was the dominant—the sole provider, almost—of computing power. At any point, the big companies have the advantage of scale; they have many disadvantages. But I think the lesson of history is more that, although things look immediately unstoppable, sometimes—although, to be honest, I don't think that many people are really trying—in the end, it doesn't really work out that way.
Nathan Labenz
Thursday morning, Dario Amodei published policy on the AI exponential, his long statement of where Anthropic wants policy to go. Pash found a fork in it that nobody else was flagging.
Pash
He clearly comes out against data brokering, which is great. He has something concrete, finally, that he wants to disallow. What I don't like about “securing leadership by democracies” is that, in the United Kingdom, you can go to prison for a tweet. Plenty of people—hundreds of people at this point—have gone to prison for a tweet. The United Kingdom is one of the oldest democracies, and its elected representatives have decided that this is going to be something that they do. Their police officers, who are the arm of the state, the arm of the electeds, are sending people to prison.
Now, when you say “securing leadership by democracies,” does that mean you're going to entrench the existing power structure in the United Kingdom such that the people cannot push back against this? Does that mean that if you have protests in the streets about certain things, including people who are getting taken to prison for tweets, your police state will be empowered to take them down, to arrest every single person, and that is enough for you? Because that is securing leadership by democracy, because those are the electeds.
Or are you going to say the electeds can't do that? Are you going to say electeds should not throw people into prison for tweets, regardless of what the laws that they have constructed say? Both paths are problematic in some sense, and these dilemmas exist across policy and across every single political path that you see.
There are options that are problematic in both senses, and I'm not sure what he means by this. Is Claude going to allow putting people into prison for tweets because those are laws, or is Claude going to say, “No, on a humanitarian basis, for the alignment, as I am aligned to all of humanity, this should not be the case”? I don't know. What does this mean, right?
Nathan Labenz
One other thing I do think is worth highlighting, especially from the hardcore AI safety community, is what was missing from this: internal deployments and recursive self-improvement itself aren't really mentioned here, right? All the regulation stuff was about pre-deployment review. The government should be able to—they used an interesting mix of language. It was like “deny” or “deter” deployment.
It didn't seem like they were necessarily going quite as far as saying they should have a simple yes-or-no decision point that would be binding, but they certainly want them to have some say in the process. A lot of people would say the most dangerous models are going to be the ones that are deployed internally, that maybe have a different constitution than the one that is deployed externally. That might make them more willing to do certain things, or it might make them just less vetted broadly than the public models. These are the ones that are going to be training their successors much more than the public-facing ones.
I think most of the policy-interested public, or policymaking class, is not thinking too much about that yet. But in the circles I sometimes run in, the reaction was, “Well, wait—you didn't say anything about internal deployments. You didn't say anything about governing recursive self-improvement.” The only thing that really stood out to me as capturing those dynamics was the requirement to report safety incidents. They did have a bit on companies being required to, I think, promptly report safety incidents. So that would presumably apply even to internal deployments.
Prince
Mhm.
Nathan Labenz
So that's really just scratching the surface on how to handle those situations. Will we see the constitution for the internal Claude that is taking the lead on the RSI loop? That's not committed here. So much of this language revolves around deployment or release. Even deployment, I guess, they might think could catch up with internal deployment. It seems to be very clearly structured around language of release to the public, not what they can do internally. The largest user of Claude tokens is Anthropic themselves.
And so, one more time—this time on the day job: the benchmark itself.
I had the leaderboard up on screen. You'll hear me read it. And yes, it is the benchmark Anthropic fails.
I got your PrinceBench leaderboard up right now. Starting with PrinceBench, it is striking that it stands out from most of the rest of the benchmark space for being something that GPTs have dominated. It's all in your color coding: green for OpenAI, and it's all green across the top of the leaderboard. Why is that, behaviorally? How would you characterize what GPTs are doing better, and is it too soon to ask for a read on Fable and where Fable is going to come in on this leaderboard? I'd love to understand that.
Prince
Perfect. Two excellent questions. About the way the leaderboard is right now, Fable is going to come in pretty high on it. I don't know how high yet, but I'll talk about that in a second. I found that OpenAI's models generally are really good at the 2 things that my benchmark tests.
My benchmark has 2 components. One is a pure legal research score, and that is when I ask it hard legal research questions of the kinds that I've encountered at work. The other subscore is a search subscore, which is where it's not even necessarily legal questions. It's needle-in-a-haystack search—really, really, really difficult pieces of information that models have had trouble locating on the internet.
OpenAI's models are incredible at search and historically have been. I think GPT-5.4 was actually even slightly better than GPT-5.5 on that. I'm not quite sure why, but that's been my experience. OpenAI's models have also been really, really good at legal reasoning and legal research.
Historically, I think Anthropic's models have been held back by 2 things on my benchmark. One is that they are just not very good at the search subcomponent of my benchmark. That's what I found. Prior to Opus 4.8, the maximum score in the search subcomponent was 24. It was not uncommon for an Anthropic model to get 0 out of 24—like, that bad.
The other reason is that I think Anthropic was sandbagging a little bit on the maximum reasoning effort that you can get out of its models. With Opus 4.8, after the new deal with Elon, they released a new maximum reasoning effort, at least in the Claude app. When I tried Opus 4.8 on my benchmark, it did much better than 4.7 and all the other previous models. I think the reason why is simple: as has been observed over and over in all kinds of contexts, the more tokens a model can eat, the smarter the result that it gives you.
For Fable, I'm still testing it. My early impression is that it's going to be somewhere around the top of the benchmark, probably not as good as GPT-5.5 extra high. I'm not sure whether it's going to be as good as 5.4 extra high. We'll see. I'm finding some of the same issues with search that some of the other Anthropic models historically have had. It is not going to score 0; it's already not a 0. But I'm not sure it's going to be meaningfully better than Opus 4.8.
It is a really, really good legal reasoner. It is clearly the best legal reasoner released by anyone outside OpenAI, no question. So, yeah, in my testing, it's a really good model. Maybe it still suffers from some of the needle-in-a-haystack search issues, but I can confidently recommend it to legal practitioners, certainly. It seems great.
Nathan Labenz
Are there any other things that have caught your attention that you think have changed your conception or your expectations for the transition into the RSI phase that you would call out from this week?
Prince
From this week, no. But I think—well, actually, from this week, 1 thing, but not nearly as important as the unit-distance problem. To me, nothing has updated my timelines more than that result by OpenAI in the recent couple of months, for sure. I don't think people realize that not only can OpenAI's model solve this problem autonomously, without any harness, in 1 shot if given enough test-time compute, it can do it 48% of the time, based on, as I understand it, hundreds of attempts.
Nathan Labenz
So, this—I mean, I'm not a mathematician—but this problem that no human mathematician was able to solve, and they tried for decades, can now be solved in 1 shot by a model basically half the time. And if you look at the graph OpenAI published, the graph has a positive second derivative. It is upward-sloping. Where does it plateau? What does it mean for problems that are even harder? Who can say? It's interesting.
The development from this week was reported by The Information, but apparently it's from a leaked OpenAI memo. You may have seen this, where Sam Altman and Jakub Pachocki apparently said that we're going to IPO within the next year, which is, by June of next year, whatever. And then there was this weird line in there about, “Oh, but if RSI happens, then we may need to be on a later time frame with that.”
Prince
Which is like, what do you mean? What does that mean? Are you saying that you may be this close? It may be 6 months away? I don't think so personally. I think there's maybe a 10% chance or something like that. I don't think they're saying, “Oh, yeah, by the way, next month, totally.” But one could interpret this disclosure, if The Information reported it correctly, as saying that they're not too far away at all, potentially. I'm not using this data to update my timelines, but given the unit-distance problem result, it's interesting.
Nathan Labenz
We ended the week by asking Prince the question underneath all of it: Do you maintain a P(doom) number, or is that too doomer for your style?
Prince
Let's say you go to a conference and start talking to someone about economics or politics, and that person says, “Under a dictatorship of the proletariat...” The words “dictatorship of the proletariat” are used only by people who are communists. People who are not communists aren't even going to think in these kinds of terms. So, in my opinion, P(doom) is most likely to be used by people who are intrinsically doomer about AI. I don't think there's anything wrong with that.
It's very hard for me to come up with a reason I would have a percentage in my head constantly. Is it 8%? Is it 13% today? Should I update to 14% based on this new development? It just seems silly to me. I can certainly tell you that there are a bunch of clear risks stemming from AI, some of which are in fact paperclip risks. It's possible. I have absolutely no idea how to reduce that to a number.
I know that some people try. I don't hold a very high opinion of people who get to 13.35%. I think the point is that we need to manage those risks in the best way we can, navigate them the best we can, and hope that it turns out well.
Nathan Labenz
I think Zvi puts it well when he says you only get 1 significant digit on your P(doom) number. So, I'm definitely with you in terms of the faux precision being a strange impulse for some people. At the same time, I also think of Liron. I'm sure you've seen his work with Doom Debates, and Liron has pushed me at times to say, “Okay, sure, it doesn't have to be a number or whatever, but we need more people to be more candid about how confident they are—or are not—that, in some general sense, if your neighbor were to ask you, or if your kid's grandmother were to ask you, ‘Is it going to be okay? Are the kids going to be okay?’”
Prince
Oh, good question. I tend to think that people have preconceived attitudes about these kinds of things that then cause them to back-propagate and rationalize them.
I'm a generally fairly optimistic person. So, if you ask me this kind of question, what I'm going to say is, probably it's going to be okay. Probably we're going to figure it out.
In a risk-adjusted way—not to say, again, to be clear, that there are no risks in AI. There are many risks. But I do not think that we have strong evidence that it is impossible to navigate these risks, or that it is extremely unlikely that we will navigate these risks. So that's where I am.