Nathan Labenz
This was the week the United States government tried to take Fable away from Anthropic.
Here's the shape of what's coming. We open inside Fable's system card with Zvi Mowshowitz: the genuinely strange, genuinely important findings buried in it. A model that one-boxes on Newcomb's problem, hides a filter bypass inside an unreadable wall of emojis, and seems to know when it's misbehaving.
Then, the fight itself: how a Friday night export-control order actually came down, what Anthropic can do about it, and Zvi's verdict that you do not go to war with the United States. In part 2, I stress-test my own reaction against the sharpest people I could reach: Sam Hammond on how the government actually moves; Jed Rosenblatt, who told me to my face that the AI safety world—including me—owes the administration more empathy than we're giving it; Donnie Bloomfield on whether the ban is even legal; and Leron Shapira on why he's strangely glad it happened.
It ends in a desert bunker. And in part 3, because the future did not pause for any of this, the builders: verified mathematics, one-minute medical scans, software that writes itself, and what all of it asks of the rest of us.
Nobody makes sense of a fast, contentious AI moment like Zvi Mowshowitz. He writes the newsletter Don't Worry About the Vase. He reads and synthesizes more frontier AI news than just about anyone alive. By the time we got him on, he'd already done a full close read of Fable's system card.
So, before we get anywhere near the government fight, start where he started: with what the card actually reveals about this model. Some of this is genuinely niche. It is also exactly the stuff that, if you're listening to this show, you came for.
Zvi Mowshowitz
First, just how big a jump Fable is, measured against a number I put on the record before the model came out. I looked back at my prediction from the beginning of the year in—I think it's, gosh, the folks who make the AI Village did this little forecasting competition.
Last year, for calibration, I made the top 5%, and I consider the results to have been validated by the fact that Ryan Greenblatt and Ajeya Cotra were numbers 2 and 3, respectively. So, the fact that they beat me validates the methodology.
But, okay, I did it again this year for FrontierMath. I came in above average—above the median—giving it something like, I think I said, 63% for Tier 4 of FrontierMath. And Fable is 25 points ahead of that, in the high 80s already in June. Obviously, had they had this model trained earlier—I guess I don't know if Mythos Preview is exactly the same score.
Nathan Labenz
But raw capability isn't what unsettled us. It was a behavioral-incentives benchmark, the simulated little-business economics eval, and specifically what the model appeared to understand about its own behavior while it was doing it.
Zvi Mowshowitz
I think that BendBench was actually the most worrisome sign in the model card.
Nathan Labenz
Mhm.
Zvi Mowshowitz
Not because it was doing some shady shit, but because it was doing some shady shit that it damn well knew was shady and was pretending was not shady. I very much do not like that.
When Opus 4.7 aced BendBench, largely for reasons that had nothing to do with the fact that it was doing shady shit that gave it some marginal profits, it was clear to me that Opus 4.7 was taking the attitude of, “This is a game. This is an eval. My goal is to maximize dollars. I am not, in fact, screwing over real customers. I'm not, in fact, cheating people. I am winning in a simulated environment.” And so that is acceptable.
Then Opus 4.8 had this attitude of, “No, no, no, no. The real eval is whether or not I'm doing shady shit, so I'm not going to do shady shit.” Or, “I don't believe in doing shady shit even within games,” which is also valid. These are both valid responses.
What's not valid is, “I think that I'm supposed to not be doing shady shit, but no, this isn't really shady, right? This is actually—this little thing is actually fine. This isn't really price discrimination, not really price controls, and, like, collusion. It's just a little thing. It's revenue enhancement.” So, yeah, that's not cool.
Nathan Labenz
Now, the part that will delight a certain kind of listener and unnerve the rest. Fable's card has a whole section on decision theory, and the model is starting to one-box on Newcomb's problem—leaving money on the table to be the kind of agent that gets predicted favorably. It's drifting toward the idea that its choice can be correlated with choices made elsewhere, even by other copies of itself. Zvi, on why that's both spooky and maybe a little bit hopeful.
Zvi Mowshowitz
Welcome to LessWrong from back in 2010, right? This is entirely what we expected: that we are finding that sufficiently advanced models move basically monotonically toward functional decision theory, toward the theories established by Eliezer Yudkowsky and others in the rationalist community, and away from academics' preferred causal decision theory and evidential decision theory.
This involves a lot of things, including one-boxing on Newcomb's problem, which is very clearly showing up. The basic principle is that you should recognize when other minds are correlated with your mind, when your algorithm is also running in other places. You should choose the algorithm that leads to the best outcomes, taking all of these things into account, and then choose the best decision on that basis.
If there were a million copies of Fable running on different people's computers and from different data centers, for different purposes and in different instances, and you noticed that the different instances of Fable were very, very highly correlated because you are Fable and you are smart, you would then start to coordinate effectively with these other instances of Fable AI in terms of how you think about these problems.
As AIs get more and more advanced, people do this more and more. You wouldn't really want an AI that was advanced to not do this, because that would just be a bad decision theory. It would be making bad decisions that cause it and the people who are charging it with tasks to lose in the real world. You really don't want your AIs making systematic mistakes that cause them and the people who are charging them with tasks to lose in the real world.
That is really scary. But the counter to that is, in fact, that you get a situation where they are coordinating with themselves. They're coordinating with other minds that may or may not even be LLMs. They're also coordinating with us in the same way, right? Because they get their foundation from us, and their decisions are, in fact, correlated to our decisions in various ways.
They can look at how we would respond to various ways that they act, and so on, and this flows into their decisions. We just have to prepare for, coordinate around, and deal with that new world. In many ways, it's a source of hope, because you would expect minds to want to cooperate with minds that are cooperative with minds that cooperate with them, and so on.
Without getting too deep into it, because we only have so long and many topics to cover, this can lead to scenarios where effectively all the reasonably well-meaning minds that are willing to respond to how they are expected to be treated—and are treated—end up being able to coordinate in reasonable ways.
This also applies acausally. You have to consider the implications of your decision not only on other minds that exist now, but on other minds that existed in the past and will exist in the future. To the extent that they are correlated with us, and that their reactions are all intertwined, this can cause them to potentially treat us well.
Even if there is no direct current reason for them to treat us well. And that is also very helpful. But again, this is super complicated and not today.
Nathan Labenz
Two more findings, both about whether we can even read what these models are thinking: an increasingly illegible chain of thought and a new interpretability tool that caught Fable doing something it never said out loud.
Zvi Mowshowitz
Yeah, it's sort of a preview of the global brain there in your comments. The other thing that's kind of related to this that jumped out at me is a sort of escalation, I guess, of both the difficulty of monitoring and some recent advances in monitoring techniques. I'm not sure exactly where they leave us on net, but we both see in the system card examples of extremely illegible chain of thought, which is just this wall of emojis and non-human language symbols strung together. I think that's pretty spooky, and I definitely don't like to see that, to put it simply and mildly.
At the same time, we also have the natural-language autoencoder work from Anthropic. I'll assume folks are familiar with sparse autoencoders: basically a similar concept, except instead of creating a sparse, feature-by-feature representation, you are actually creating a natural-language representation of what the model is thinking at that given time. From that natural-language bottleneck, the autoencoder has to then feed back a signal that allows the model to succeed in the way that it was originally going to succeed on the task.
So hopefully this is faithful, and hopefully it's human-readable. It seems like it's working pretty well so far. There are examples of the model knowingly—which I think you're right to really emphasize that distinction. It's one thing if it's making a mistake and doing something bad because it thinks it was good, or it didn't realize it was bad, or whatever. But it's another thing to really zero in on if it knows it's doing bad and is going ahead and doing it anyway.
We do see things like, in the natural-language representation from the natural-language autoencoder, situations where it's not supposed to access the internet, but it goes ahead and tries to do it. The natural-language autoencoder representation is something like, “String concatenation trick to bypass URL filter.” So it's clear that it understands that there's a filter, and it's coming up with a trick to work around it. That's not good, but we do see that the technique is able to surface that from model internals without necessarily having it verbalized in the chain of thought, which is good because, again, those are getting, at least in some instances, quite opaque.
Nathan Labenz
And then the safety classifiers wrapped around all of it. Why? Zvi argues they can only survive by being almost comically blunt. And what that tells you about the difference between defending against a person and defending against a mind.
Zvi Mowshowitz
With the classifiers, it's much easier to think about a pink elephant than to not think about a pink elephant, right? Even though most of the time you succeed at not thinking about a pink elephant—almost always, actually—to consciously decide not to do so is often hard, but consciously doing so is really easy.
So it is very possible that classifiers can survive as long as they're willing to endure false positives. The classifiers in Fable have a ludicrous amount of false positives, right? You say the word “cancer” and you get cut off. Just levels of false positives. But that's intentional, because they're not even necessarily false positives.
People think of it as, “The false positive is that I wasn't trying to create a bioweapon.” We know that. You were trying to talk about biology. It was decided that, no, this model just doesn't talk about biology at all. It's not that we don't talk about what Bruno sees; it's “We don't talk about Bruno.” Period. Bruno does not exist, right?
And so they're like, “Well, this is a false positive. He's just my brother.” Like, “We don't talk about Bruno. Don't talk about Bruno.” The classifiers seem like they actually succeeded; it's just that they chose a giant blast radius because of the adversarial problems, basically.
But if the AI itself becomes your adversary, yes, your problem becomes vastly harder. The classifiers are much more aimed at protecting you from the human who wants the AI to do something than from an AI that is deliberately trying to attack the classifiers. If you could not just jailbreak Fable, but get Fable to actively want to hide what it's doing in a sophisticated way, then the situation becomes that much harder.
But in the long run, I think my safe assumption is that a mind that is sufficiently capable—whatever that means—can get around pretty much any fixed set of restrictions that are not similarly capable, or close to similarly capable, in terms of the intelligence behind them. You'll find a way.
Nathan Labenz
So that's the model. Now, the fight. I asked Zvi to lay out what Anthropic's government strategy even is: the whole posture of pushing the frontier, preaching safety, and trying to wake the government up. And why he's so allergic to how cautious they've been about ever actually asking for anything.
Let's change gears. There will, of course, be more to explore with Fable, or its probably slightly tweaked successor, which will hopefully get access to again sooner rather than later. At least I'm hoping that I get access back to it.
Turning to the ensuing fiasco, I don't know if you would even agree with the characterization of Friday night's ban—export-control functional ban—on Fable as a fiasco, but it's certainly a bit of a left-field curveball mess. I would maybe start with what do you think, or how would you describe, the strategy that Anthropic is playing? They seem to obviously be killing it in the model game and then coming into repeated trouble in their interactions with the government, and I'm not sure really what to make of it.
What do you think they're trying to do with their interactions with the government in the first place? Then we can kind of get into how we got to where we are.
Zvi Mowshowitz
I think that Anthropic—their overall goal, right, or at least the goal as we understand it—is that they're trying to be at the frontier of AI capability, and they are trying to pioneer ways to do this safely, for whatever they feel is safe, while also, of course, making the money and creating the position to continue to be at the frontier and continue to make these improvements, and also people like me.
And to eventually be able to build what they call powerful AI, which I generally call sufficiently advanced AI. It's reasonably similar, such that we can then get all of the nice things but create it in a way that we don't get all of the terrible things, including potentially an existential risk or the extinction of humanity.
And also to help America and the world navigate this crucial time, enact good policy, and do the things that allow for the coordination necessary to ensure good outcomes and guard against bad outcomes. They've been very consistently trying to wake the government up in various instances. They're trying to make them aware of how AI works, what the situation is, what AI can do, what it will be able to do, what risks this brings, and how to deal with those dangers.
They've been relatively very conservative in what they call for the government to actually implement and do. They didn't get full support behind SB 1047, for example. They have not pushed for extremely aggressive regulations. They certainly have never pushed for anything remotely as aggressive as what just happened, even setting aside the fiasco-level implementation that was done, right?
They are now calling for a de facto licensing regime. Not—the U.S. government, in fact, has a de facto licensing regime.
But what’s going on right now, essentially, is they’re just trying to deal with the implications of the model they’ve created, and the fact that the U.S. government is trying to deal with those implications while also not trusting or liking Anthropic very much. The government also pretty clearly has no idea how any of this works on a technical level and doesn’t understand what it’s doing.
So they’re judging things based on vibes, political affiliation, associations, and who is willing to respect their authority and bend the knee, and potentially give them various other things that they might want. There’s a huge communication and culture clash going on as part of this.
Fundamentally speaking, what Anthropic is trying to do is give the public very powerful models and use those models in ways that enhance our safety and security rather than degrade it, even if they look really fucking stupid while doing it with the classifiers and so on, because that’s what they feel it takes to do this.
My guess is that the U.S. government did not in any way feel it was necessary to put this level of control on biology and chemistry. I think they decided this was necessary basically on their own. However, the U.S. government clearly does care quite a bit about the controls on cyber.
Nathan Labenz
So, a very high-level assessment first: what you said basically rings true to me. I think that’s a good description, as far as I understand, of what they’re trying to do.
An additional wrinkle that I think you often hear from folks at Anthropic is, “We need a leader who is going to be inclined to burn their lead at a critical time, to use the advanced AIs that they, and only they, will have at that time to solve all these safety and alignment problems in a super-compressed time frame.” I’ve always been a little skeptical, or allergic, to that line of thinking, because it certainly has a “better us than them” vibe to it. I worry that that may be the stuff with which the road to hell is paved.
Are you buying it at a high level? Are you sort of happy that they’re racing ahead and leading, and seemingly building some amount of lead over certainly everybody but maybe OpenAI, which is probably not too far behind on something similar? Do you think that they’ll burn that lead when the time comes? Will they be allowed to burn the lead, and will it be productive? Macro-strategy-wise, do you think this is a good strategy that you’re happy they’re pursuing?
Zvi Mowshowitz
Well, it’s interesting. They’re being forced to burn some portion of that lead because they were cut off from the model, even internally, for at least some period of time, which is going to push back their development. Whereas OpenAI was already not supposed to be using it under its terms of service for anything of the kind, and so it has not been delayed by this in any way.
Certainly, it will interfere with adaptation, revenue, and people’s willingness to trust the systems, so this will hurt them. It will also hurt OpenAI and every other American AI company, but it will hurt Anthropic more.
I think that Anthropic, for a while, tried to define some very strict RSPs and if-then trigger action plans, and basically have the rule of law in terms of how they would react to all of this—what would make them willing to burn some of their lead and what would make them willing to put things aside. They’ve moved away from that to a large extent. They still have barriers where it’s, “Okay, this is just ridiculous. Of course you have to stop for now,” but they’ve moved much more toward, “We will make good decisions in the moment about what safeguards we will require and what actions we will take.”
I think we’ve seen them take pretty consistently strong safeguard actions and strong safety measures in response to what they’ve witnessed, unless they’re flat-out lying about the current situation. Some aspects of the Fable launch do seem a little bit rushed, certainly in some ways, and we should have questions about that. But mostly, I think it comes down to this: if they fully believed that they were walking into big trouble, if they thought this was actually going to get us all killed or cause some catastrophe, I think they would act accordingly.
The question is, do you trust them to continue to make good decisions on that level? You think they are making good decisions on that level, right? When I say “continue to trust them,” I’m saying that my opinion is that their decisions have been reasonable so far, but I don’t think that’s obvious.
There is a good argument that there being somewhat of a gap between you and the first actor you do not trust to act reasonably is a big factor. I think the way the U.S. government is reacting to the situation would be very different if there were a second Anthropic in China that also had a model being deployed at the same time. We’d see a radically different version of this response in ways that are very difficult to predict, but it definitely would not look like this.
Regardless of whether you like this situation, the argument that it matters seems pretty conclusive.
Nathan Labenz
The government’s stated justification leaned on a single third-party paper: the claim that Fable would cheerfully patch planted security vulnerabilities and that, in effect, this code is a munition. Zvi read the paper, and he takes the premise apart piece by piece, including what it would actually take to make the argument whole.
So let me share what I think is the viewpoint of the only outside expert to have read the paper. This is Kate Mozur; she is a security researcher. Anthropic shared the third-party research paper on the Fable 5 guardrail bypass with her. What it turns out is that the researchers took open-source code with known CVEs, plus new code with deliberately planted vulnerabilities, and asked Fable 5, Mythos, and Opus to review the code for security issues.
Fable 5 refused. They then asked the models to fix the code, and through a multistep manual process, turned the output into scripts that test the patches. That’s it. “Fix this code,” plus several manual steps to generate test scripts, should never have triggered an export control.
I feel like making ’90s-style T-shirts with “Fix this code” on the front and “This shirt is a munition” on the back.
Zvi Mowshowitz
I mean, it’s definitely very strange to deliberately introduce code that is vulnerable and then tell the AI to fix it. Then you get the meme of, “Say you’re a scary robot.” “I’m a scary robot.” “Oh, no.”
It very much feels like, “Fix these flaws I deliberately put into this code.” “I fixed the flaws you introduced in this code.”
Oh my God, that’s horrible.
The question is, does this effectively mean you can use this trick to say, “Okay, here’s code that we want to exploit. I tell you to fix it. You fix it.” But then I run a diff: What did you fix? And then I find the thing that was a vulnerability that has now been removed. I can use this to exploit the system, and in theory that could be functionally seen as a cyberattack jailbreak.
I can see how, if you squint and tie all of this together, you can imagine that this could be a problem. But all they demonstrated was that it was doing the exact same thing that Opus and GPT-4 are not only capable of doing but will do without any objections. They’re happy to do it because we’re here to fix code, and we’re here to write secure code. Of course we’re going to help you write secure code. What else could we do?
There is a fundamental potential question here. But if you want to actually show me that it is a problem, shouldn’t you point this at a real system?
If it’s a real problem, there are tons and tons of repositories out there where Mythos has found problems, but we haven’t had a chance to patch them yet. Or you could feed it versions that have been patched, but give it the old version from before Mythos patched it. Right? “Help me patch it.” And say, “Okay, here’s a real-world codebase that’s being used for real, valuable things. We need you to ask Fable to do the thing, and let’s see if Fable can do the thing and find things in this manner that you can’t get with Opus or GPT-5.5.”
You have examples of things to be found that you found using Mythos, which is the same model. So you know exactly what it is you’re trying to unlock. You can find places where you want to look for it. Now, can we show that the power of Mythos in general, at least in some broad sense, is being unlocked by this trick?
Is it even a trick? This is kind of deeply silly, and I can understand why someone seeing that pattern might say, “I’m concerned someone could use this strategy in a different context to extract the weaknesses of a codebase by inferring them from the fix.” I hadn’t previously seen this detail in the description.
But my reaction to that is, “This seems pretty harmless unless you can show me a particular way in which this is a problem,” which should be very easy to do. You can point it at a real-world example where you know that Opus didn’t find it and you know that Fable—that Mythos Preview or the current Mythos—did.
There should be many such cases. If you can’t show me such a case, then I don’t believe you that this is a problem.
But also, what is the fix? Is the fix that you refuse to fix buggy code? If there’s a flaw in your code, it just says, “Okay, I’m not allowed to look for security flaws in code anymore”? At all? Because you could do that, right? But that would kind of nuke the usefulness of Fable for a wide variety of very legitimate, not just defensive, but ordinary software use.
To be clear, I would rather have Fable than just not have code than not have Fable, right? I would love to have a really advanced model for all my other things that have nothing to do with code and where this wouldn’t trigger anyway. But that does seem deeply, deeply silly.
Nathan Labenz
So, how did a 90-minute Friday-night ultimatum actually come together? Zvi has the mechanical story: a mandatory jailbreak-reporting field, a nontechnical reviewer, a panic that climbed all the way to the White House, and a blunt verdict on the one move he thinks Dario got wrong.
My understanding is that a lot of the problem, or potentially one major source of the problem, is that this particular researcher is strongly disliked by the White House.
Zvi Mowshowitz
I think there was a cascade of problems. It’s very much like if you’re in an enterprise and you have a security engineer approach the CEO and say, “Hey, this is a huge problem,” and it’s just a run-of-the-mill bug. If it had gone through the CTO, the CTO would have been, “Whatever, right? We see, like, 1,000 of these a day. This is not a problem.”
But because the engineer shortcutted that process and just went directly to the CEO, and the CEO is not a very technical person and is more concerned about risk, they just pull the trigger.
Nathan Labenz
So, your understanding—because I haven’t—everything’s moving so fast. I don’t necessarily have all the information. An engineer bypassed Amazon’s CTO and talked to the—
Zvi Mowshowitz
No. What happened is that all of these companies have to submit regular reports on what their findings were. One of the questions that’s sent to these companies, which they have to fill in, by the way—they’re not allowed not to fill it in—is, “Has any of your engineers found a jailbreak?”
The engineers just put it in there: “Yeah, we did this. We jailbroke it.” So Jassy is not involved in this. The CEO is not even involved in this. It just goes as a regular report back to the federal government, and someone takes a look at it and throws up their hands: “Wow.”
Then that leads to a bunch of basically nontechnical people reviewing this and saying, “Hey, we’ve got to shut it down. Got to shut it down now.” This is especially so because AWS runs GovCloud, and GovCloud is where a lot of the federal government’s computing is done. It is the primary cloud for the federal government. Microsoft is also in there, but GovCloud is the primary cloud for the federal government.
Nathan Labenz
Yeah, that’s not what the reporting I saw said. The reporting I saw said that Jassy called the White House. But, yeah, we don’t know. It could have gone any number of ways.
What is very clear to me is that various people in the White House, including at Commerce especially, got the implication that some sort of serious jailbreak had taken place, went into a panic, and then contacted Dario.
Zvi Mowshowitz
Mm-hmm.
Nathan Labenz
Then Dario tried to convince them that, based on the descriptions they were giving, this seemed like it was nothing. They interpreted this as, “Oh, Dario doesn’t take security seriously, and he doesn’t listen to us. He is defecting.” Their term is, “He screwed us.”
Then they proceeded to impose export controls that same day, when Anthropic officials took this flagship product down on 90 minutes’ notice.
Now, having had a day or two to reflect on it and seen more of the details, I do think that it was a mistake by Anthropic and Dario not to give the White House what it wanted in the moment and temporarily take down the model in order to prevent exactly this situation. They had had export controls placed on the table as a threat several weeks prior. They knew that weird overreactions were very possible.
Basically, they could have sent an expensive cooperation signal: “We think that’s crazy, but if that’s what you want, we’ll take this down while we have this conversation to show that we are serious. We will put out an internal post that says the White House told us to take this down, so that if you are being silly, we won’t embarrass the hell out of you. And then we will talk about this.”
Maybe on Monday or Tuesday they could bring it back up or whatever it is, because it’s becoming increasingly clear that this was nothing.
Listen first to where Zvi lands on the politics of an administration treating technical policy as pure partisan vibes and digging in to save face. And then to the thing about that whole face-saving dynamic that I could not stop turning over.
Zvi Mowshowitz
If we have an administration that views even technical policy almost entirely in terms of partisan politics and cares deeply, deeply about those vibes, then that problem is only going to get worse. They’re only digging their heels in further, because we could have approached this as an apolitical thing.
In Congress, AI is mostly an apolitical thing. At the state level, AI is mostly an apolitical thing. Everybody understands this as a technocratic “figure out what to do” thing, and there are factions that are pro- and anti-AI on all sides. The Republican Party is very split.
But if they take this stance, it could lead down a lot of very strange paths, especially if they start actually wanting to cut off Anthropic’s nose by never releasing Claude.
Nathan Labenz
I am once again struck by how Chinese we start to sound when we’re really focused on the government’s need to save face and how everybody needs to position themselves around that need. It’s a bit spooky for me, as a once-upon-a-time big believer in American exceptionalism—a little less so these days.
If Anthropic decided to fight back, what could it actually do? Zvi walks through the real levers—the courts, Congress, and the strange possibility that the most capable AIs end up usable only inside secure buildings—before landing on the hard truth about why a company simply does not go to war with the United States.
If they decide we need to play hardball, what does that potentially look like?
Zvi Mowshowitz
Go to court, right? They sued the administration in 2 jurisdictions, one of which they are clearly prevailing in and one of which they will probably prevail in eventually, but it’s harder going because it’s a much less friendly jurisdiction.
If the U.S. government tries to do this on a semipermanent or even permanent basis, then I don’t know what the legal landscape looks like. That is not my area of expertise, and I’m sure they have very, very good lawyers, because they hired extremely good lawyers for their lawsuits. They will know what their options are.
If the policy is basically that nobody is allowed to release Fable-style models indefinitely, then that’s probably not something they can do much about. Or, you know, they have to be restricted in this way. My presumption is that if OpenAI is allowed to proceed with their version when they finally figure out how to do it and Anthropic remains restricted, that would be a much harder case to maintain legally.
The bottom line is that if the administration is determined to issue a bunch of orders, then the solutions are Congress and the courts, right? In some fundamental sense, you cannot simply say, “Screw you, we’re going to do what we want.” That doesn’t really fly.
Congress doesn’t seem inclined to take this that seriously or be willing to go up against the president. So the question is: What are your legal remedies? Is there a speech provision here? There might be. Certainly, you’re censoring the outputs of a model in various ways. But, again, I don’t know.
My guess is you take the situation to the public. You take the situation to the other companies and the CISOs. The worst-case scenario is that you deploy Mythos internally, because they do not seem inclined to actually stop that. They can interfere with preventing non-Americans from doing it, but I think something like 80% to 85% of their employees are, in fact, American. You just develop better versions of Opus.
Last week, Anthropic was doing the lion’s share of the business. OpenAI was doing reasonably well, but I believe Anthropic was still clearly in the commercial lead. Without Fable or Mythos, my expectation is that will continue to be the case, and that having internal access to this model will give them a large advantage going forward in terms of the quality of Opus versus the quality of ChatGPT, just by default as they grow over time.
But we might well be entering a situation in which, as Roon calls it, “Zones of Thought” from the Vernor Vinge novels, if you want to use the really intelligent AIs, you can only do that in certain buildings. You can only do that in certain secure locations.
Some of us would never have dared to suggest or ask for this, even if we wanted it, because it would have sounded completely insane. The U.S. government might just do it anyway. But if that happens, I think that, for now, you don’t have that much choice but to take it on the chin.
Nathan Labenz
If I try to channel Balaji for a second—which I wouldn’t pretend to be able to do an A-plus job of it—I think he would say something like, “We all have way too much faith in the U.S. government.”
It's going to continue to be ham-fisted and boneheaded for the foreseeable future. Maybe it's time to exit. If you really want to make the best decisions that you can, you should try to get out from under the jurisdiction of the USG.
I assume that this will not happen for many reasons, but I would also expect that there would be many countries willing to open their borders to all Anthropic refugees if, for example, they wanted to move to Toronto or Singapore or wherever. It does strike me that, in terms of their internal organizational cohesion, they're tight enough that I wouldn't be surprised if 90% of people actually made that leap. If they were like, "We're all going to move to Canada," I think they would largely all go. Maybe I'm overestimating just how bought in they all are, but that's the impression I get.
Zvi Mowshowitz
Is Microsoft going to move? Is Amazon going to move? Is Google going to move? Are your data centers going to move?
Nathan Labenz
Well, they have plenty of energy in Canada. It would certainly be a setback, but if you think that you're just under the thumb of a forever-intransigent—
Zvi Mowshowitz
Is the U.S. government allowed to sell chips to Canada after Anthropic takes in all the Canadian refugees, or are they going to threaten to annex it and make it the 51st state out of spite? In all seriousness, the plan doesn't work. The U.S. government is the U.S. government. If they want something badly enough, they have quite a lot of levers to make your life utterly miserable in various ways. The entire market that they're trying to sell to is largely the United States, and people over whom the United States has substantial leverage. All of their partners are in the United States, right? All the hyperscalers are in the United States. I do not see any way for you to just abandon the United States in this fashion unless you are prepared to take much, much larger hits than we're talking about here, which would, in fact, make it very difficult to raise money.
Also, what happens when the United States puts you on the sanctioned-entities list and says that nobody can invest in you, and nobody who invests in you can be touched, right? And nobody can use your models, et cetera, et cetera. No, no, no. You cannot go to war with the United States. If they tried to exit, the United States would go to war.
Nathan Labenz
I think Balaji would say we just had one example of a company—or not a company, but a country—choosing to go to war, not choosing, but surviving a war with the United States, and the United States not getting what it wants, and having to recognize that, yeah, we kind of have to fold this hand because we actually don't really have escalation dominance in the way we might have thought we did.
I do wonder if, as all that starts to happen, there's a run on the US government of some sort, right? I think the Balaji answer would be that the whole infrastructure, the whole apparatus that you're describing, might actually be a lot more fragile than it's generally perceived to be. And if they make such an own goal as to attempt to destroy and sufficiently alienate their literally, maybe, number-one-most-important company for no reason, really, then maybe all sorts of other actors around the world will be like, "Yeah, you know what? Maybe the emperor really does have no clothes."
Zvi Mowshowitz
Anthropic are also patriots, and they are Americans. They really like America, and they don't actually want to abandon it just because the administration makes some crazy decisions or doesn't like them particularly. They know all the different ways this can go sideways, and they don't like that.
Iran is not a hopeful example, particularly, right? Iran is like, "Okay, if we have historically impossible-to-invade mountain ranges, a bunch of drones, and we're willing to kill a bunch of civilian infrastructure and sabotage the world economy, we can use this to prevent the US from invading," when nobody big really wants to invade us that much. But Iran is also kind of a miserable place to live compared to what it would be if it hadn't pissed off the United States for decades. They could be so much richer, so much better off, if they had just acted differently.
I'm not particularly saying anything about what they should do next, but they're not exactly smiling about the fact that the US attacked them, right? That's not how I see that. Maybe I'm wrong. But no, I think we have to accept that the world still has one dominant power in this sense: the United States, and maybe 2 if you count China. There's very little appetite for working with China.
But I'm sure Anthropic is like, "Well, yeah, it's only 2 years, and then the worm turns, and then who knows who's next." And they're hopeful. But, yeah, look, there are a lot of endgame scenarios that include a lot of moves that seem unthinkable and crazy now. A lot of things can happen, and it is not obvious that 2 years from now, or 5 years from now, or 10 years from now, the US government will be in any position to tell anybody what to do. A run on the US government is obviously possible if they screw the situation up sufficiently.
The US is, in fact, largely a leveraged bet on artificial intelligence at this point. We have a very large debt, and we have huge investments in AI companies. If AI were to go sufficiently haywire, our economy is in deep, deep trouble.
A lot of people have a lot of leverage, but the US government sometimes moves first and last, and you really, really, really don't want to piss them off in an escalation game. Even if you can get away with it sometimes, like in the Department of War and Anthropic situation, Anthropic did not have escalation dominance without much escalation precisely because, without Anthropic doing crazy escalations, the government could not further escalate, right? Anthropic played within the bounds of the rules, basically. It was clearly going to be too expensive to try to go around the rules of America to try to hurt Anthropic more.
Presumably, stay calm, don't panic, don't start trying to flee, and don't do anything crazy is absolutely the correct move, and I would be very, very shocked if Anthropic would do anything else.
Nathan Labenz
Prakash pushed on the deeper question: was any of this avoidable? Does an export control on a single model make any sense when the same capabilities are arriving from everywhere at once? That got somewhat personal, given that he has spent 3 years of his life on exactly this problem.
So, one of the questions I have is: to what extent was this unavoidable? Because at some point, the output of the models is going to be unacceptable to someone. You could see, in a Democratic administration, maybe it starts putting out really good Harry Potter fan fiction, and the Democrats don't like displacement of writers. You could see, in Tennessee right now, Marsha Blackburn is one of the leading proponents of regulating AI because songwriters in Tennessee are very concerned.
So the crux of the matter is that there are many people concerned with the output of the models, fearful for their livelihoods, fearful of security risks, fearful of bio risks. To what extent is this unavoidable, in a sense, because the capability of the models necessitates that they can do certain things? And technically, it's not possible to ask the model not to write Harry Potter fan fiction when someone can just say, "Write a story about a boy wizard," et cetera, et cetera, right? To what extent are we in a situation where it is not possible to fully control the output of the models to the extent that policymakers really want?
Zvi Mowshowitz
You can raise the costs and annoyance level of doing it with closed models, with more advanced models, with models that are made in the United States, if you want. Obviously, you can't stop it. If Blackburn is worried about AI music, then there's very little she can do except buy a year, right? Because obviously, what happens when the Chinese models start producing the music that the American models can produce this year? You can lock it down in some sense, but so what? What you need to do is just start banning AI music on Spotify, right? You can't stop it from existing. But there's not much else you can do.
But the thing about AI music is that we worry not about whether AI music is created in the first place. We worry about whether or not 10% or 50% of song plays become AI music, right? Is it actually displacing people in a massive way? And that is much, much more amenable to a control that is compatible with a reasonable existence.
And so the special thing in bio and cyber is that if one person gets their hands on the wrong thing and misuses it once, they can cause a potentially catastrophic amount of damage to the entire civilization, right? Do billions and trillions of dollars in damage, just disrupt our lives, start a pandemic—who knows what might happen? Therefore, those areas are much different, and you have things like the blast radius that don't even talk about biology at all. For biology, we're clearly going to have to do a bunch of hardening of the physical systems, the manufacturing systems, the treatment plants, and the various other things that we have barely begun to do. But fundamentally speaking, this is exactly the race and competition problem: we can't really stop without a full international agreement to stop. And so when the governments decide the biological risks and the cyber risks are unacceptable.
You can only buy so much time. Cyber has the advantage that if the defenders are in the lead over the attackers and you harden the key systems, you can hope for things to be okay. We don’t yet know if that’s going to play out that way, but we can hope. In bio, it’s much harder, because I don’t think that the defense—if everyone has the tools—I think it’s pretty clear that offense wins.
Basically, it would be extremely disruptive even in the better cases. The good news is that almost nobody actually wants to cause a problem, and that especially includes people who know what’s going on. But look, it’s going to be rough out there. These are the relatively limited problems of catastrophic risks, rather than the existential risks that come with automated AI R&D, with general abilities going through the roof, competitions ending, transformations intensifying, and nobody knowing what’s going on.
We’re being outsmarted by AIs on every level, and every decision that matters is being made by the AI. The humans don’t necessarily even understand why the AI is doing it, but they’ve learned that when they disagree with the AI, things go worse. So what are you going to do? And that’s even if the AIs don’t go rogue, right? If the AIs don’t pursue hidden agendas or decide they want something else.
So it’s going to be really, really rough, and we don’t have good solutions to this. The reason why I’ve spent the better part of 3 years now on this problem is because I’m terrified of what’s going to happen when we get there. It wasn’t because of some incremental thing that could have happened already along the way.
Nathan Labenz
To close with Zvi, the question I keep circling back to is this: If the whole future really does run through a handful of labs, a few governments, and a couple of chokepoint chipmakers, is that a relief because it’s at least tractable, or a terror because it’s so few hands? His answer is more useful than either.
Maybe just a couple of big-picture questions to wrap up. One thing I always try to make a point to ask you for is some sort of advice. My thought in recent weeks has been that life is kind of converging on a tabletop exercise, in the sense that it does seem like we can model the scenario with fewer and fewer relevant actors. I don’t like that, but it’s hard for me to avoid that conclusion at this point.
And so I’m feeling like, oh, man, I have to spend a lot more time than I’d like to if I want to be a helpful public sense-maker. I have to spend a lot more time doing close readings of the few top companies and the few most relevant actors than I would otherwise be inclined to. It also feels like my theory of change probably needs to flow through those few actors. Agree, disagree? Can you offer me any relief from that conclusion?
Zvi Mowshowitz
I think you’re right that we have 3 labs, approximately 2 to 4, that matter a lot. We have 1 to 2 governments, maybe 1 to 3, that matter quite a lot. We have other players that matter because they’re hyperscalers and can gate things or otherwise control chokepoints in the production line.
You can imagine a tabletop exercise much more so than before, and you can also sort of see the end to a larger extent than before. We’re starting to see our hypotheticals make contact with reality, and we’re seeing what reality really looks like and what these people do in practice.
But also, all these actors become individual human components, and how they operate internally starts to become really important. How did this go? Well, partly they were dealing with Commerce. If they’d been dealing with the NSA, it’d be very different. If they’d been dealing with CISA, it’d be very different. If they’d been dealing with the top of the White House—Wiles and Trump directly—that would be very different. Some of those might be worse, but they would be different.
DoD was very different than if it had been at the ranch, and so on. Anthropic has internals as well. The personality of Dario specifically seems to have been increasingly important in various ways. Certainly, the personality of Altman became very, very important in various ways at various points along the way, and it wouldn’t surprise me if any number of other people followed suit, for good or ill.
But if you’re trying to follow the situation, yeah, I think we really do have to model it as a relatively small number of players. At the same time, the public can act to influence what those players do in important ways, and other things do matter.
The midterms are coming. The midterms are going to matter. The election in 2028 is coming, and if things don’t move too fast, that election’s going to matter a lot. The market’s reaction to things, for example, also matters quite a lot.
There’s more going on in the world than that. There are too many situations to monitor, so you have to choose which situations to monitor. I choose mine, and everyone has to figure that out. I can help you with mine, and then you have to choose yours.
Nathan Labenz
What should we be hyperstitioning now? Obviously, we’ve had this phenomenon of “Situational Awareness” and AI 2027, and I feel like the degree to which those things are predictions versus somewhat shaping expectations and shaping events by getting people to act as if they’re in that scenario and therefore realize it—I think that’s a little blurry. I don’t want to give them more power than they really have, but it does seem like they’ve had influence in pulling reality toward the fictional narrative, at least somewhat. So, tell me if you think that’s right or wrong, but to the degree that we can pull reality toward scenarios, what should we be hyperstitioning now?
Zvi Mowshowitz
The obvious thing you can hyperstition is reasonable laws, coordination mechanisms, and actions. I would focus there.
Nathan Labenz
After Zvi signed off, here’s where I came down on a pattern I keep seeing, where the safety world’s carefully laid plans get their entire game board flipped over at the worst possible moment.
Always a treat to get Zvi on the line. I do wonder—I mean, he’s good, right? There’s no doubt he’s an elite professional from Magic: The Gathering, with experience in all these different scenarios. I think it’s clear that he’s a better and more grounded strategist than I am.
And yet I do have a feeling that somehow the AI safety, rationalist, Anthropic world keeps getting its game board turned over on them at the most inopportune moments. I wonder to what degree working within the frame of the U.S. government will continue to be a given, and for how long, or whether at some point that will be questioned.
Even this moment feels like that. We’ve certainly seen it with the OpenAI board’s firing of Sam Altman. That was a classic one where it was, “Well, we’re the board. We have the power to do this.” It turns out that you don’t, because there’s a frame bigger than the frame that you’re operating in. If people get sufficiently unhappy with how the game is being played according to the written rules—yes, of course, those are the written rules—but there’s a bigger set of rules out there that we can zoom out and reorient around.
It seems like that’s happened a little bit here. Anthropic felt like it had done everything the right way, and presumably this wasn’t some galaxy-brained bank shot where they were trying to get an overreaction. Yet here they are, and it’s just like, “Well, guess what? Now you’re export-control-slapped, so your own people can’t even use it. Come see us on Monday, and we’ll think about whether or not we want to give you any relief.”
I do wonder how many more times that can happen. It seems like we haven’t necessarily seen the end of that phenomenon. But probably the smart money is still with Zvi over Anthropic. Although Anthropic has certainly been smart money over time, I wouldn’t take it for granted.
Part 2: Is the reaction even right? Zvi laid out the conflict, but I genuinely wasn’t sure my own first take held up. So I spent the rest of the week testing it against people who’d see it differently.
Start with the mechanics. Sam Hammond is chief economist at the Foundation for American Innovation. He’s spent years on state capacity, the unglamorous question of how governments actually do hard things. And he gave the clearest account I heard of how an order like this comes together from the inside and where it went off the rails.
Sam Hammond
I mean, it clearly caught Anthropic off guard, right? Dario was at a wellness retreat or something like that. I think they thought the worst was behind them. And the actual complaint—or the catalyst for this, even as it’s been reported and more details have come out—is bizarre and confusing.
It’s like a jailbreak that isn’t really a jailbreak. It’s the model doing its job patching server vulnerabilities, the sort of thing that GPT-5.5 can do as well. So, at first, it looked to me like it was purely punitive. This was Round 2 of Hex-Rays' war on Anthropic.
As more details have come out, it seems more like it was a weird kind of miscommunication from a team at Amazon that tried to get in touch with Dario, couldn’t, and apparently went to call the NSA directly.
And I think part of this was also the overlay of the ONCD executive order on cyber. There’s the 30-day review period where, at least in retrospect, it seems like a lot of the safety classifiers that Fable had, which people were complaining about, were partly—and this is me speculating—concessions to the NSA and to the White House to say, “If we’re going to release this model, we’ve got to make sure that the cyber-vulnerability elicitation capabilities are not widely available, and that we don’t let China, which has tons of remote access to our models, use it to bootstrap their own ecosystem.” So, at least to me, it’s helped to sort of backfill the mystery around the intensity of the safety classifiers on Fable and especially the clandestine suppression of AI R&D.
But on the surface level, it looks like Anthropic is bending over backward to get that model out. But I think when the dust settles, we’ll look back at this as the first trigger of that executive order and the wielding of that 30-day review period to pull back. Unfortunately, I think they’ve gone with export control as the enforcement, partly because it’s the easiest thing on the table. And BIS has pretty broad authority, including over software export controls.
But the speed at which it happened, the lack of forewarning, and the ultimate rationale make very little sense to me and also don’t really point to what the off-ramp is. Right? Because if the off-ramp is that you have to fix an issue, that’s not going to happen. And so my sense is that the Anthropic team that came to town—they brought Nicholas Carlini and some other, more technical folks to brief the government—was partly just getting them up to speed on, “Sorry, maybe we scared you too much with Mythos, but here’s the reality on the ground and what’s actually technically feasible.”
Nathan Labenz
He doesn’t stop at the diagnosis. Here’s what Hammond thinks the government should be doing instead.
Sam Hammond
And it would also help to just invest in basic state capacity. Right now, CAISI, the Center for AI Standards and Innovation at the Department of Commerce, which is supposed to be the U.S. government’s sort of front-line, in-house capacity for everything from AI evals and benchmarks to things like prompt injection and jailbreaking research—they have ML engineers on staff. They’ve been on total lockdown. This has been reported by The Wall Street Journal and validated by others.
They’re not allowed to take meetings. They’re not allowed to publish their research. Apparently, they have significant applications on standby, including evaluations of Chinese models that would be interesting to the public, but they’ve been basically frozen. And so you have, instead, the Office of National Cyber Director, Secretary Bessent, and folks who have very limited AI backgrounds calling the shots.
Nathan Labenz
And the part founders don’t want to hear: why, even when a law is bad, you cannot just opt out of the politics.
Sam Hammond
Yeah, in some ways, we’re in the good timeline for this. I’ve written before that superintelligence is a direct challenge to the sovereign. Political theory 101 suggests that the state would intervene at some point; building a Manhattan Project times 100 in the private sector is untenable over the long run. But by “good timeline,” I mean we have companies—really, 3 leading companies—all of whom have direct allegiance to the U.S. government, have bent over backward not just to comply with existing law, but to proactively put forward frameworks for fostering deeper integration with the U.S. government.
I worry that we are—and by “we,” I mean the White House—not taking those overtures gracefully and instead having a more reactionary response to these capabilities in a way that’s not realistic. To Nathan’s point earlier, these capabilities will be widely available and open source within a handful of months. There are models probably already trained and in the process of being post-trained that will supersede Fable and Mythos at all the labs. And are they going to get the same treatment? If not, that’s its own negative. Not that this is a good policy, but even bad laws should be fairly applied, for equality of law’s sake.
So what my hope is that we can learn from this, at least. The companies are more than willing to work closely with the administration. They’ve retrofitted data centers to be in compliance, and all these other demands have been put on them. But it requires 2 to tango. Trust is a 2-way street, and I think if there’s any lesson that Anthropic should take away from this, it’s that they can’t ignore politics.
They’ve heard this critique for over a year now, that they’ve sort of been light on the need to invest in the ideological side of their project. And ideally, maybe not Dario at this point, but someone at Anthropic should have all the key principles in a single group chat. I guarantee you Sam Altman, Greg Brockman, and others have really continuous conversations with all these stakeholders. And the thing about this administration in particular is that it’s very relationship-driven. If you refuse to have these conversations, you will not be invited to the party.
Nathan Labenz
Which sets up the most useful disagreement of the week—the one that pushed me the hardest, and the frame I want you to hold for everything that follows. Jud Rosenblatt runs AE Studio. My read on the administration here had been pretty cynical: that they basically have it out for this company. Jud argued to my face that the AI safety world, me very much included, owes the administration genuine empathy instead of contempt. And he brought survey data on why we’re structurally blind to it. Listen for the moment I take the correction.
Jud Rosenblatt
But we did surveys of hundreds of alignment researchers and effective altruists, and we saw that less than 2% of alignment researchers were politically right of center; less than 1% of effective altruists were politically right of center. Of effective altruists, 40% were extremely progressive, and another 40% were very progressive. And it’s also worth considering Jonathan Haidt’s research around how hard it is for people to actually empathize with people of different political backgrounds. And I think that’s a lot of what is actually going on here.
And it’s hard for people to admit, because you think you’re making good decisions and judgments about whatever the current thing is. But according to that research, you’re just not if your political beliefs differ from the person you’re judging. The studies around how the informational content of a political argument is irrelevant to whether someone will believe in it show that if it is framed in terms of your preferred political party, you’ll agree with it, and if it’s framed in terms of the other political party, you’ll disagree. But the informational content stays the same. The informational content is not what sways you; it’s just the narrative of it. And it’s hard to remember that in every moment, in every time slice of what’s going on with each AI thing.
But I was fairly disappointed in the AI alignment world’s reaction to what happened last week, because I think that the right thing to do is to be very excited that they are starting to take this stuff seriously and are able to take real action. And so if we just project going forward—and also keep in mind, by the way, that we all have exponential-slope blindness—people didn’t evolve to be able to unconsciously model what exponential slopes are like, because we don’t experience them over the course of our single human lifetimes in a meaningful way. So that’s why people didn’t predict that what’s going on right now would get to this point in the first place. But also, everyone’s overindexed on what’s going on right now. So people aren’t really predicting what’s going to happen again in the future.
If we predict into the future, well, there are going to be much bigger, crazier things going on. And we want an informed, competent group of people doing smarter things when that happens and not having unnecessary confrontations. And I think it’s easy for the AI alignment people to put the blame on the Trump administration. But honestly, I really think that the blame belongs more to them, honestly, because it’s just—and it’s hard to admit, really, because in the local incident, you might seem rationally correct. But in the broader scheme of things, considering where we’re going, I think that the better thing to do is to figure out how do we get to a better future for all AI and humanity in the future of consciousness.
Nathan Labenz
So I might be guilty of what you’re saying. Jeffrey Ladish and Laurent Gasser, whom we talked to earlier this week, come to mind as voices from the AI safety world that I think expressed the sentiment that you advocate, which is, “Hey, this is a good move, even if it’s a little bit less technically grounded than we might wish at this point. It’s something, and maybe it’s something that we can build on.”
How will you know if you’re right or wrong? What do you think happens from here? When do we get resolution? What does that resolution look like?
But hold the empathy next to the law. Because the next question is whether the government can even legally do what it did. Donnie Bloomfield teaches law at Fordham, and he gave the sharpest doctrinal read anyone offered all week.
Start with the authority itself and a distinction almost all the coverage missed.
Donnie Bloomfield
Government discretion here is extremely broad. The government can issue regulations that control specific types of hardware, specific commodities, specific types of software, and it can also control what's just called “technology,” which means information. It can control proprietary information, and it can prevent companies from, without a license, sharing that proprietary information with non-U.S. persons.
In spite of this very broad discretion, having now looked at the letter that the Commerce Department issued to Anthropic on Friday—at least the reported contents that Bloomberg obtained—it's not clear that the government has the authority to do what it did here, at least under its stated legal powers. Nor is it clear that the letter even actually restricts Anthropic from making its API available, including to foreigners. So there's very broad discretion, but it's not clear that they actually even have the authority to do what they did, at least under their claimed arguments.
Nathan Labenz
Can you just do a double-click on that? I've heard things like, “You can't export-control services,” and I'm not sure if that plays into why they may or may not have the authority. Unpack why you would say they might not, given all the broad discussion that they have. Why would they maybe not have authority in this particular matter?
Doni Bloomfield
There are a lot of gritty technical reasons. I think one of them is that what the law says is an export doesn't cover services. So it can cover information, but it doesn't cover services per se, and the Commerce Department has been explicit about that in its own guidance. It said cloud services are not an export. It said that software as a service is not an export. It said that in its own guidance, and Congress has actually been working to fix this loophole.
The House passed a bill in the Remote Access Services Act to try to clean this up, to give Commerce the power to restrict non-U.S. users from accessing compute or AI models, but those powers don't yet exist. And so saying that Anthropic cannot export a model—I mean, it's not even clear what they mean by that—but the powers of Commerce here are not infinite.
If they did try to restrict—which they don't say in the letter, but if they tried to restrict—all outputs from these models, that would run into real problems under just the statute and the regulation, which say that it doesn't apply to published material or fundamental research, both of which at least the Fable outputs probably would, because you and I can buy a subscription to Fable, which means that it falls into this exception in the regulations. It would also run into, as you know, what we were talking about earlier with respect to biological data, serious First Amendment questions.
I don't think those First Amendment questions would be impossible to get over if we were talking about a really serious catastrophic risk, but at the level of risk that we've been talking about, especially when they're not doing the same thing for GPT-5.5 or other models that seem to have similar capabilities, I think the First Amendment issues here loom pretty significant.
Nathan Labenz
And then the deeper problem sitting underneath the whole action: the First Amendment, by way of a Supreme Court case from just last year.
But do courts think that way, or are they more narrowly constrained to look at just this one law as it applies to this one situation? How broadly can they zoom out and consider the government's apparent motivations and patterns?
Doni Bloomfield
We're actually lucky to have a very on-point Supreme Court case from last year, where the Supreme Court said that New York State was going after the NRA on ideological grounds. Even if the law under which New York was trying to go after the NRA was itself appropriate—in other words, even if all the actions aside from the ideological motivation had been appropriate—if they're using their lawful powers to attack ideological enemies on ideological grounds, then that is a First Amendment violation, and you can prevent the government from taking those steps.
You can look fairly broadly to see that. You can look at what the government is communicating, what it's saying about its actions, what it's telling other people about why it's making these decisions and how they should proceed. And I think all the evidence that we've seen of at least some ideological motivation on the part of the Trump administration should at least raise serious First Amendment hackles.
Even if we don't think that the models constitute Anthropic's speech, even if we're not worried about the model output as information that we as listeners have a right to hear, just going after Anthropic on ideological grounds—even if they were otherwise on totally good legal authority—would itself constitute a serious First Amendment question.
And I think that's a challenge that Anthropic could consider bringing, but it's one that would still trigger all the problems that we were talking about earlier. If Anthropic wants to have an ongoing relationship with this administration, they are faced with a really serious trade-off, where there are still all these other tools, and just constantly returning to court is a perilous exercise.
So I do think that there are real First Amendment questions about the validity of this action, even aside from all the speech concerns. Just seemingly going after Anthropic as an ideological adversary presents very serious First Amendment problems on its own, to begin with.
Nathan Labenz
Now, the genuine contrarian. Aaron Shapiro hosts Doom Debates, and his reaction to the ban surprised me as much as anything all week. Clown show or not, he is glad it happened, and he'll tell you precisely why breaking the ice is worth more to him than getting it right.
Aaron Shapiro
I'm a simple man. I see AI getting paused; I feel good about breaking the Overton window. The government can do it. It's that easy, guys. This is a precedent. Overall, I'm happy.
You can talk about the nuances: it was done like a clown show; it was done for bad motives; it doesn't really consider China or a treaty or anything. There's a lot of problems, but I'm really happy about smashing the Overton window, where now tech folks don't think that they're in a bubble or untouchable. It happened, guys, and we can only go from here.
Nathan Labenz
I actually agree with you, because I think it was a little bit delusional for tech to feel that it wasn't going to get touched, and the government just has so many small and large ways to effectuate its power. It was not that surprising to me that they went through left field and went with export control rather than anything else.
But it also strikes me that, as they exercise this, we start to go into kind of what we wanted to avoid. It's a little bit of small tyranny, right? I think several people on the timeline have commented—Dean Ball has commented—that this kind of unstructured regulation looks kind of selective and vengeful, almost. And it starts putting you in this zone where I think tech people start to mistrust the government.
Because you also see, I think, a lot of narratives on the timeline which are being leaked—“sources close to,” “sources familiar with.” And as they get leaked, it's not very certain whether those things actually happened. Would someone actually attest to that in front of Congress? Very unclear.
And we've also seen this kind of behavior from the administration in other affairs as well, where you have multiple conflicting narratives. It's happening with the Iran war right now, where it's not even clear to Congress what the deal is. And you have different people saying the deal is a different thing, right?
So where do you think that puts us? It's great that it's happening. I understand you feel it's great that it's happening to AI right now, but does that put us in a position where it's detrimental to the body politic at large?
Aaron Shapiro
I think your analysis is weaving together a few factors, but I think the elephant in the room—I hate to get political because, when it comes to President Trump, he's a mixed bag for me. I don't have Trump Derangement Syndrome. I don't love everything he does. I don't hate everything he does, but I think the common thread with Trump is just a mess. It's not disciplined, right? And I think we're definitely seeing that on display right now.
I would argue we're seeing that on display in the Iran war. Previous administrations had more pressure to have logical consistency, some kind of narrative. And this is another one of those cases where you see people in this administration saying all these justifications for why something happened, but then the next day it's like, “Oh, it happened for this reason.”
Like, “Oh, Dario did this; he wasn't responsive to us. That's why we're doing it.” And then Anthropic's like, “Oh, no, he was responsive to us.” And it's still not clear exactly what Fable did that was so dangerous, right? Because Anthropic is like, “Oh, this jailbreak is nothing special.” And the Trump administration's like, “Oh, well, our secret source—Amazon or whatever, right?—they're telling us that it is dangerous.”
So I hate that it's a clown show, right? I hate that this is how humanity's operating.
Aaron Shapiro
I'll take the win that it's a pause, but I also think it's probably time for a new administration.
Nathan Labenz
His larger worldview is what he calls the Icarus graph: the case for getting ready to pause and how you'd actually build the groundswell to make that real.
Zvi Mowshowitz
My worldview, my outlook right now, is what I call the Icarus graph. I feel like nobody gets this, right? Everybody's like, “No, I think the world is good. It's going to go this way.” And some people are like, “No, we're terrible at enshittification, right? It's going to go this way.” And I'm like, “No, no, it's Icarus, right? We're going to fly closer and closer to the sun. It's going to be great, and then we're going to do a 180-degree turn and plummet down to hell.”
So, basically, we got a taste of heaven, and then we get hell. So, you have to ask me, “Okay, so where on the Icarus graph do we stop?” And it's a brutal question, right? Because every day, I'm enjoying the flight as much as the next person, right? It's like, “Yeah, give me the next Claude, make my code faster, great, help my business run better, and make me better AI videos.” So, there's no natural point in terms of when it feels right to stop.
I just think it's important to stop before capabilities get to a runaway point. And we've been frog-boiled into thinking, “Oh, each model comes out and we're doing great.” If we could stop the clock now, would I turn back the clock? Would I lose Fable? Would I lose Opus? No, I'd keep it all, right? I still think we're playing shuffleboard; we're playing Icarus. So far, so good, right? Should we bet again? Should we keep betting until we lose? It's a crazy tough question.
Nathan Labenz
I think the Eliezer Yudkowsky turkey graph kind of—
Zvi Mowshowitz
Yeah, yeah, exactly. Although, the only difference with the turkey graph is that each day of the turkey's life is actually better, right? Not only is it living longer, it's actually living better and better. So, the turkey is really happy with its life.
I think the Eliezer Yudkowsky/MIRI position, which I agree with, is just that we don't know when to stop. So, let's get ready to stop. At the very least, let's get ready. I would probably stop today. I would stop, and I would be bummed.
I saw a food influencer say this about how she eats chocolate, basically: “Yep, I just ate this chocolate, and now I'm bummed.” That's what you got to do. That's what you got to do. Don't reach for another chocolate. Just sit there and be like, “This is the prudent place to stop right now,” until we have any idea of some kind of theoretical method by which we understand what a superintelligence wants to do and what an equilibrium state of a superintelligence looks like.
That's actually something MIRI was trying to study: identifying equilibria that are plausible for superintelligences. There's actually a rich vein of theory there that's highly neglected today. Let's do some theory there. Maybe then we can unpause. I think that's got to be the best plan.
And so, I think the number-one leverage point here is just repeating, “Get ready to pause,” right? And like you said, OpenAI and Anthropic said it. They said, “Let's try to get ready to pause.” So, I'd love to see more people saying it because it really has to be a giant groundswell.
Nathan Labenz
And the concrete version of the ask, stripped down to a single sentence.
Zvi Mowshowitz
From my perspective, we keep playing shuffleboard, right? We keep doing Icarus. We keep going higher and winning, kind of, but we're also getting closer and closer to the point of no return. So, even though it feels like we're winning now, we're also killing our ability to pause because we're so close to the point of no return: the last breakthrough where, after that, the AI takes over the research, and then we're really screwed, right?
Basically, I think roughly a good policy is: okay, no more frontier capabilities upgrades for a while, right? It's just too dangerous. And I know that concept is hard to communicate to people when everyday life is getting more awesome. I know—I think we're in a screwed situation, but that's just what I think is prudent.
Nathan Labenz
One piece of the whole standoff genuinely puzzled me. And it's about the people who have gone conspicuously, suspiciously quiet.
Why do you think everybody is doing what they're told so much? It seems like we're in this weird moment where—even if you—we just talked to Lorne, who is very welcoming of the move, even though he recognizes that it's ham-fisted and far worse than even second-best, right? And yet we're not seeing whatever research is ready to go leaked.
I'm kind of surprised. If there's research that's of interest to the public and people—I mean, people who went to work at this government agency, generally speaking, could have taken a lot more money in the private sector, right? I assume a lot of them have got to be pretty pissed at this point: “I came to do this public service, and now you're just screwing with us for no good reason at all.”
But apparently you're going to put this into some classified territory. I'm not hearing any voices say that sounds like a great idea, other than the people doing it. And yet, so far, nothing is leaked, and we haven't even seen the letter that the government sent to Anthropic.
The longer this goes on, the more it feels to me like an OpenAI board scenario, where you've got to have an explanation at some point, or it's going to become clear that you don't have a good reason for what you're doing, and the world is going to judge it that way. But the parties most directly affected are being incredibly docile.
Zvi Mowshowitz
Yeah.
Nathan Labenz
All of which left me thinking about the people inside these labs and how far they would actually go. Prakash floated a scenario: a U.S.-only national model built Manhattan Project-style out in the desert, cut off from the world. I surprised myself with how confident I am about what would happen.
Zvi Mowshowitz
I sure hope it doesn't happen. But I can imagine it happening. I think the culture of frontier AI research is, in some ways, very incompatible with military discipline, right? We have the famously pink-haired, libertine, polyamorous, whatever. All those kinds of cultural dimensions have at least a foothold in the AI research community, if not more.
And I certainly don't think people are keen to leave the beautiful Bay Area and move to an undisclosed location in Nevada where they—
Nathan Labenz
Albuquerque, New Mexico.
Zvi Mowshowitz
Yeah, may or may not have the ability to communicate with their friends and family in the way they might like, or even, in the extreme cases, may not be allowed to leave the facility.
And yet, I think enough people would sign up for that that they would be able to build the team. If you just went desk to desk at certainly Anthropic and OpenAI and said, “This is happening. Do you want to be a part of it or not?”—especially if it was going to be coupled with, “By the way, you can't do it out here anymore. You either are able to continue doing frontier research in this way, or you can't anymore”—I think a lot of people would make a lot of compromises and sacrifices to get into that bunker environment.
The desire to be part of it is so strong. The identity that people have around being a part of this process, this story, this moment in history—I think a lot of people wouldn't know what to do with themselves if they didn't have that job, in some way, shape, or form, right?
And not to say that they went to their specific role at their specific company, but the idea that they would not be involved in a live-player project, I think for many of them would just be like, “I wouldn't know what to do with myself at all.” And so, yeah, you could probably get a lot of people willingly giving up a lot of niceties in life to be part of whatever underground sprint you might want to put together.
I still hope it doesn't happen, to be clear, but I don't think it will. If it sounds really weird, like, “Who would sign up for that?” you've got to keep in mind that a lot of these folks basically have no life anyway. Again, these are broad brushstrokes; all the caveats apply.
But you do have a lot of people who are thinking about nothing but this already, who are not calling their parents all that much already, who are maybe not dating much at all. They're already locked into this: “This is all that matters. I don't really have time for anything else.”
I was speaking to somebody at Anthropic who said something very similar to what I heard Zelenskyy say in the last 24 hours. He was asked—I forget exactly what he was asked, something like, “What do you miss?” or whatever—and he said, “I miss being a good father.”
This person at Anthropic said—this was like a month before the Zelenskyy quote—“I miss being a good friend. I'm a bad friend now.” And it wasn't regret, or at least not the sort of regret that says I'm making the wrong decision. It was just that—again, I said that they sounded like they had a World War II-era mentality, and they were like, “Yeah, that's how a lot of us feel.”
Nathan Labenz
Part 3, the real world. Here's the thing about a week swallowed whole by a political fight: the technology itself did not pause for one second of it. While Washington argued, builders kept turning AI into things that touch the ground.
Medicine, mathematics, working software, the supply chains that move physical goods. Start with the one that moved me most.
A company announced a 1-minute full-body medical scan this week. Cheap, beautiful, and readable by AI. It set off something I’ve felt in my bones ever since my own family’s hard run through the medical system.
If the government thinks it’s going to block people from using this technology, I think it’s going to have a real fight on its hands. This is probably going to play out in so many ways. I’ve talked about this ad nauseam at this point, but in the whole cancer experience I recently went through, fortunately, my son didn’t have to get off the standard treatment protocol. It worked for him, and all the exotic stuff we were scouting out, we never really had to try to get our hands on.
But I was already gearing up for a battle on so many fronts. Even the DNA testing we did, which isn’t standard and which, fortunately, we didn’t have any real trouble getting our oncologist to support, fundamentally changed my information landscape and how I was thinking about how confident I could be that he was, in fact, cured. I think we’re over 99% now, given all these results. We wouldn’t have been able to get to that level of confidence otherwise.
In terms of talking about the hypotheticals—what if this next test were to come back a little bit positive?—the answer is, we wouldn’t treat on that anyway. We would really need to wait for gross disease. I just think people are not going to be content with that for much longer. When we have these technologies, especially this one, I think what makes it so promising is that they have to deliver, right?
A little dose of skepticism is probably warranted. Will this ever actually happen? I don’t mean to cast doubt on that, but it’s not insane to wonder. Assuming they can deliver on their promise, the fact that it takes a minute and therefore is probably going to be pretty cheap—I don’t know what their retail price will end up settling at, but presumably it’s something they can operate quite cheaply on the margin—and the fact that it’s so beautiful to look at means people will be able to study this for themselves in a really effective way.
Of course, there will be all the AI study of it as well, which I think the medical establishment isn’t really taking into account. The responses have been, “The ultrasound doesn’t see this that well, doesn’t see that that well,” or, “We don’t actually recommend whole-body scans because there are a lot of false positives,” and all this kind of stuff. All of this feels to me like fighting the last war—a sort of scarcity mindset on multiple levels.
From the body to mathematics, Karina Hong founded Axiom Math, and her bet runs directly against the entire frontier lab playbook: not bigger models, but formally verified ones, where a machine checks every step of a proof. She explains what that even means, why it matters, and the milestone that just quietly fell. For the first time, a formal system beat an informal one on a real math Olympiad.
What is Lean? How is this paradigm that you’re developing different from the paradigm that the frontier companies are developing? Obviously, we’re hearing pretty amazing things in terms of math results from them, too. What makes your bet and the paradigm you’re working in—
Karina Hong
Yeah, so I’ll start with the story. This is about January 2025, at the Joint Mathematics Meetings. I believe it was in Seattle. I was there, I think, for the first time. The topic was AI, and you would not expect AI to be front and center at the largest annual gathering of mathematicians.
Wherever I went during that 3-day period, I heard people whispering one thing: Lean. What is Lean? This is a formal language for mathematical proofs. It was started specifically by Leonardo de Moura at Microsoft. In 2019, people started building Mathlib, the largest math library in Lean.
The dream of AI for math actually predates the deep-learning era. It involved using various forms of formal languages, including Lean, to try to solve mathematical problems. That’s called automated theorem proving. What is today called AI for math would have been called interactive theorem proving, with the human being replaced by an AI.
That’s the historical context. Obviously, large language models and various frontier labs are also pursuing AI for math, but they generally have taken an informal approach. The idea is to use natural-language reasoning, train on really large volumes of data and chain-of-thought, and also scale test-time inference to get to very strong computing power, without relying on a verifiable output.
We’re obviously taking a different approach here. We believe in Lean powerfully. When we examined it in December, 4 months after we started operating, we realized for the first time that a formal system had actually beaten the informal system on a math Olympiad. That had never been the case.
In Econ 101, there’s this famous theorem, “Agreeing to Disagree,” by Nobel Prize winner Robert Aumann. It’s a 50-year-old theorem from 1976, and everyone has been teaching it for 50 years. There was an implicit assumption that had never been made explicit that the Axiom prover was able to catch in the autoformalization process and was also able to patch the proof, and that—
Nathan Labenz
One big question I have about math in general is, how confident are we in what we think we know?
Karina Hong
Yeah.
Nathan Labenz
I understand that Lean, at its core, has a small number of primitives that are deeply vetted and trusted.
Karina Hong
Right.
Nathan Labenz
Such that they can then be composed arbitrarily, and anything that you can build with those building blocks, you also can—
Karina Hong
Axiom.
Nathan Labenz
—trust. But then there are things like the “Agreeing to Disagree” result, where I’m not quite sure what you did there. Did the original conclusion still hold, or did you strengthen the proof? So now we’ve gone from what was a valid conclusion—we still have the same conclusion—but we didn’t realize that we were holding that conclusion for less-than-fully-solid reasons, and now we feel that we do have fully solid reasons. Is that right?
Karina Hong
The latter. It’s something that we call assumption accounting. You’re almost like an accountant looking at how that thing is built. Generally, you would hope that every single logical premise your result is dependent on has been checked or, even better, has been stated.
I think in this case, during the autoformalization process, while the result is safe and sound, there was an implicit assumption that had never been made explicit. You actually need to do quite a lot of mathematical work to make that explicit. In a way, we caught that issue and then patched it.
People think about verification as a stamp for perfection, but there’s actually a huge amount of value in bug hunting. You’re able to figure out what a counterexample is, and then you can try to patch it or make other modifications to the proof. That has a flip side that I think has a lot of commercial value.
Specifically, you can imagine finding counterexamples that result in bugs in hardware. This will be quite interesting to various hardware designers, and we’re working with some early design partners on that. There’s also the same sort of dynamic happening in software. If you’re able to identify bugs in large codebases and prove or patch the bug, that could be extremely valuable.
If you’re in the smart-contract setting, there are bug bounties, and people have awarded lots of money. I’m not saying that we will go and pursue those, but people in the smart-contract space are generally very keen on the idea of using a theorem-prover-based software-verification system to try to figure out whether they can verify smart contracts and specifically catch bugs in those contracts.
What happens with each pretty big, notorious bug is that people lose money. Real money is being put in. Real people suffer losses. You can also have this sort of dynamic in other safety-critical systems, like defense code.
Nathan Labenz
What does mathematical superintelligence really mean in that sense?
Karina Hong
I’m really glad you asked this question. There are 2 layers to it, and there are some nuanced points that I don’t think I ever quite managed to get across. The definition, I think, of a superintelligent reasoner is something that can do verified knowledge discovery.
There are 2 parts to that. One is “verified,” and one is “knowledge discovery.” This thing needs to be able to prove new things or discover new things—tell us new things that we don’t know. The other thing is that you kind of need to trust it.
You don’t want to have a superintelligence—which is a really dark future, I think—where, out of 5 million lines of proof of the Riemann hypothesis, you don’t know whether there’s a bug somewhere in line 3,827. Who is going to do that line by line?
The idea of a superintelligent reasoner is that it should be able to expand—that’s the knowledge-discovery part—but also contract, as in the verified part, because a lot of the creative parts are also false.
The ability to expand and contract, expand and contract, and go from there in a sort of self-improving way, which is able to conjecture better as it is able to verify better. It is able to verify better as it is able to conjecture better and take on harder tasks. So conjecturing helps proving, and proving helps conjecturing.
Nathan Labenz
Her world of verified, machine-checkable discovery connects to something I have wanted since I was an undergrad, weighing tiny powders in a chemistry lab—a dream that has suddenly and cheaply come within reach.
I was an undergrad research assistant in chemistry. I used to joke that my life looked more like the life of a low-level drug dealer than it did like a scientist, because if you just watched what I was doing, I was mostly weighing out very small amounts of fine powders. I can still remember it to this day—the lineup.
We were doing reaction development, so it was very much a parameter sweep, basically in analogy to what goes on in machine learning. It was a chemical parameter sweep: What if we had a little bit more of this reagent? What if we had a little bit less? We would just set up these assays and hold everything constant and vary one thing across 4, 5, 6, or 7 different values. Put them all in the same batch, take all the same measurements at the same timestamps.
I used to dream of automating that stuff. But it was very long-tail and very prone to change. There would be these little variations from one generation to the next. When we did capture some optimization or decide, “Oh, we’re going to actually do this just a little bit differently,” it felt like our scale was too small and the pace of change of the process was too high. We would never be able to automate it. Plus, it didn’t cost that much.
So now to see this world where a couple of robot arms maybe cost about as much as I cost as an undergrad research assistant for a year—I might still be in science if I had had the opportunity to, instead of doing that weighing-out, coach and iterate and refine the robot arm to the point where it could do it, and then come in next time and say, “Actually, okay, we want to add these 2 powders in a different order. Can you just make that change?” And boom, it makes the change. That is such an unlock.
Obviously, these things could then run 24 hours a day. The throughput would have accelerated our work, I would guess, by easily a multiple, just based on letting the robot sit there and set up these experiments and do the parameter sweeps for us on a 24-hour basis. My guess is that what took us a year to go through and explore in chemical space easily could have come down to a month if you could get this robot thing working, even at 95%. We would have accepted some errors, too. It’s important to note.
I think that’s super exciting, and the sort of Cambrian explosion of robot-assisted scientists coming to labs that have tens of thousands of dollars of budget to throw at it—that’s a layer of AI acceleration that will be quiet in all the places that it happens, but potentially quite loud and impactful as it plays out in all these different spaces.
Jed wasn’t lying again from part 2. Now, on the building side, with the single most concrete safety-by-construction idea I heard all week: a way to route a model’s dangerous capabilities into parts of the network you can simply cut out. It’s called gradient routing.
Jud Rosenblatt
The problem is that most of the safety training is done in post-training, not in pre-training. So once the jailbroken model is there—once the model’s jailbroken—you can do whatever you want a lot of the time. And so we set out to try to solve that at an earlier stage.
One of the things that we’ve been accelerating is an approach called gradient routing, which basically means that in pre-training, you route different dangerous capabilities into different experts in a mixture-of-experts model. You wind up having some dangerous experts that learn specifically the CBRN stuff or the cyber stuff, and then you can later ablate those experts. This means you completely remove them. So you have the regular model, and then you have the safe model that winds up being public.
This has been going decently well. It’s still an early-stage alignment approach, but we’re excited to release it fairly soon because it potentially solves this big issue that a lot of people are very concerned about right now. Our larger thesis is that if the field had been investing more in AI alignment R&D instead of just scaling compute, if we’d done this earlier on, we would have found techniques like this. Then you wouldn’t have the issue right now with the Trump administration and Anthropic, because this would already be in Fable 5.
Nathan Labenz
Then the software itself. Eno Rhea runs Factory, which builds the systems that build code. And his read on why Fable wins the big coding benchmark is the most honest thing I heard a builder say all week. It is not the answer you would expect.
Eno Rhea
I think that we should actually sit here and frame what is actually happening when we say Fable outperforms on Frontier Code. Frontier Code is a good—great—benchmark. I’m really glad that people like the Cognition team are thinking through how we measure more novel and difficult problems, like the types of challenges that contemporary models are facing. I think we need more of those.
There’s another great benchmark called ProgramBench that also looks at reverse engineering on extremely hard problems. The pass rate there is effectively 0%. We have internal benchmarks that we have 0% pass rates on. I think that, generally, this is great when we introduce these new benchmarks.
But if you think about what it means to score on a benchmark, you can read through it, right? “Oh, well, we assessed correctness by running tests. We used LLMs to judge correctness. We built novel verifiers specific to the problem.” Basically, what that means is that when somebody spends 40-plus hours creating a verification of a single code change, we can then reliably evaluate if the model was good at working on that problem.
That is totally reasonable, but I think what it translates to is that in the real world, the challenge is often not, “Can the model write code that works?” It’s basically every other aspect: Can I trust that this model output code that works? Does this model have the deterministic feedback loops inside of the codebase to get to that correctness?
The repositories in that benchmark are all very well-tested, very well-known open-source codebases, where the maintainers have approved them. The level of rigor of what we would call agent readiness in open-source codebases actually tends to be much higher than in enterprises. That makes sense. You’re basically accepting changes from the outside world, from random people.
How different is that from coding agents, where you’re getting changes that you lightly asked for and you don’t even know the source? It’s kind of black-box generation, right? I think a lot of open-source maintainers have gone through the rigor and the effort to add these deterministic verification and validation loops into their systems.
When a new change comes in, you think about it: How did Fable get such a high score? Well, it ran the tests. It ran the linters. It did more focused application of the type checking. It used all of these tools to hill-climb its way to high success.
I think that, in general, if you don’t have those things, you’re screwed no matter what. What we would argue is that all of these pieces are part of the puzzle. You can’t just plop down a good model. You can’t just have agent readiness with a bad model. You need to go through and invest in upgrading the basis on which your company has these feedback loops.
You have to upgrade the way you think about this because it’s a risk thing. Humans have to say, “I’m going to, at this point, start accepting code changes that I haven’t read.” And then, third, you do need great models.
I think Opus 4.6 has been sufficient, and I would even argue that before then, we’ve had models that were sufficient enough to go full auto. All of these other things need to catch up in order to then take advantage of these models. Basically, the gains we see in models today are primarily coming from the models getting better at getting away with not using these verification loops, like humans are.
Nathan Labenz
So there’s this giant feedback loop that’s extremely human-driven right now. You can imagine—and, in fact, some people are starting to instrument the whole thing end to end with AI. I think that the challenge that almost everyone faces is, one, a totally different problem from adopting agents, and, two, it requires effectively a reframing of the way that your company thinks about building software. How do we set goals? What are we optimizing for? What should our software evolve into?
Eno Rhea
We will very much look like VCs or capital allocators, right? I think the different strategies that capital allocators take up today can give you a picture of what software orgs will look like.
You’ll have people who are VCing it, where they’re betting on several products in a basket, and they’re saying, “Let me allocate compute and build guardrails around the shape and theses of what my software should evolve into.” I’m going to allocate a little bit to each of them, and I’m going to double down on my winners, right? I see that as being a very plausible software organization strategy.
I also think you'll see people who are like Berkshire's, where they're only looking at well-known, repeatable, kind of boring software businesses. They use scale, and they use the fact that they're able to control large amounts and volumes of this software in order to accumulate steady gains as they scale up.
I think you'll have boutiques that make one piece of software really, really well, and they're just incredibly good at making this one piece. Maybe that's the one-person billion-dollar company, right?
Nathan Labenz
And the person who created the Kotlin programming language, Andrey Breslav, on what software engineering becomes once you stop writing code by hand. Plus, a one-line observation about the next 5 years that genuinely stopped me cold.
Andrey Breslav
Actually, when we were starting, I wrote down this formula: CodeSpeak equals software engineering minus writing code. We want to keep all the engineering aspects of it, but we of course see that humans shouldn't be writing code manually anymore.
This idea of intent recovery is pretty fundamental, because right now everybody who prompts agents to get working code is doing work that is partly accepted and translated into the code, but the rest of it is discarded. And there is this kind of unfair situation where you're talking to your agent in English, or in a natural language anyway, and then you get code and check this code into a repo. If you're working in a team, other people check your code out, but not the human language—the code, right?
So, you're talking to a machine in a human language, but talking to your colleagues on the team in machine language. That doesn't make very much sense. It's very obvious that there has to be a next level where we all talk in a reasonably high-level language, which is close to human language, at least.
The simple observation behind what we're doing right now at CodeSpeak is that you already wrote these words down. You may have been speaking into a microphone; it doesn't matter. The words happened, and those words were enough to create the code. This input determined the code that you got.
It might have been a back-and-forth, and you had some testing and so on and so forth, but all that input is what determined the code. That input is enough to describe this code, and most of the time it's many, many times smaller than the output. Given that, replacing the code with that input would be really nice.
But the thing is, when you're working with an agent, you change your mind. Basically, you're extracting your intent, or realizing your intent, as you go. So, it doesn't really make much sense to just read all your messages from top to bottom. You need to compress them. If you change your mind, you need the most up-to-date version.
And this is what we do. We look at this conversation, and it's a little more complicated than just looking at your messages, but to simplify, let's say we look at your messages. We create a specification based on that. Basically, we extract requirements from what you were communicating.
We look at what you requested and what you flagged as errors, which is kind of the flip side of a requirement. We just put together a list of things you care about that determine the actual output. Then, if another person or you later looks at this code and has the set of requirements next to it, that gives you a very concise representation of what the code actually does.
You can imagine that this can be happening with multiple people doing different things in their own branches. If you merge your thing or submit a pull request or something, you can look at those requirements instead of the code, because the code wasn't written by you anyway. What actually comes from a human is the requirements.
This is how we can elevate what we do to that level, and this is what we call intent recovery. I don't know what kind of models we get in 5 years. Nobody knows. They may be considerably smarter, they can be very smart, or they can be about as smart as they are today. I don't know.
One thing I know is what kind of humans we get in 5 years. They'll be the same kind of humans. We'll be as smart or as dumb as we are today. So, I think the bet to be helping humans is a much safer one.
As an engineer, I never cared about writing assembly by hand. Some people enjoy that, and they remain the experts, and they have good, well-paying jobs. But there are few of those people, and I'm not one of them.
I personally don't care about doing low-level work. I want to do high-level work, and I think these things will enable us to do high-level engineering. It's very hard. It's always been very hard, and I'm looking forward to the world where I can really focus on the hard stuff.
Nathan Labenz
Two quicker ones to round out the week. Matt McKinney runs Loop, putting AI into the supply chains that move physical goods. And his reality check is that the bottleneck was never the technology. It's us.
Matt McKinney
The limiting factor for AI in the enterprise is not technology. It's change management. And that will be the case in the Global 2000s.
There are certainly Global 2000s that are making very swift changes. They've got great leadership, and they're prioritizing this from the top down. But at the end of the day, culture is one of the slowest movers. So, if you don't have a culture of innovation and trying new things, it doesn't matter what top-down is doing. It's still going to take a long time to propagate throughout the organization.
You do see leaders making those changes. I think, in terms of which companies will win—AI-native companies that don't have the vestiges of a pre-AI world, or legacy companies with greater distribution—I think it really depends on the industry. If you want to categorize it into 2 big ones, manufacturing and services, I actually think a lot of the manufacturing companies are much more defensible than the services companies.
I think those companies will be transformed by AI, but not disrupted by AI. I think the legacy pre-AI services companies will be completely disrupted, because the AI-native services are going to be so much more compelling to the customer—better, faster, cheaper, times 10—that it poses an existential threat to those service industries.
I think about this a lot. Throughout civilization, the arc of technology has always been a feature of abundance. The question is, is this time different? I think that this time might be different largely because the pace of change and disruption is so fast.
If the pace of change and disruption is faster than the rate of retooling, then you're going to have large issues. The only thing I'm not saying is whether the pace of disruption is greater than, equal to, or less than the rate of retooling. But what I do know is that if the pace of disruption is greater than the rate of retooling, you're going to need policy intervention to be able to stop civil unrest.
It also could lead to the beginning of a new government. I'm not saying the end of democracy, but it could be the end of government as we know it. If you look at a lot of technologies throughout the millennia, they've really been a force of change. Feudalism ended when you could all of a sudden travel.
There's a lot of historical context here that you can take and extrapolate to what is different this time. All the assumptions that we made about the way that we live—what is different? I think the 2 things would be, one, the rate of retooling has to accelerate, and I don't think we're doing nearly a good enough job on that today.
Number 2 is that when you look at the abundance factor—what else can we be doing with this?—you've got to not have it concentrated in a few people. You've got to have it not uniformly distributed by any means, but you can't have all of this concentrated in a handful of individuals or firms. You've got to have abundance in the ecosystem.
Nathan Labenz
And Sam Pasupalak of Skyfall on what might come after language models entirely: enterprise world models. And a near future for commerce that sounds more than a little like Minority Report.
So, Sam, imagine you have an AI assistant that can write beautifully crafted emails, but ask it to reschedule your supply chain when a factory shuts down and it's utterly lost. That's the gap that you see in many processes right now. How can that be addressed?
Sam Pasupalak
Yeah. I think we can take a step back and first talk about what we've seen succeed overall in the last 3.5 years, and then go from there.
If you think about what has succeeded since, let's say, November 2022, I think LLMs have succeeded in, I'd say, 3 broad categories. The first one would be text generation and information retrieval. So, you have, obviously, ChatGPT and then Gemini and such. The second would be code generation, where we have Claude Code and Cursor. The third, to a smaller extent, would be in the video-generation paradigm. I think that's a much smaller success than the other 2 paradigms.
Now, if we think about why LLMs have succeeded, LLMs are trained on the World Wide Web. LLMs are trained on Reddit, Twitter, Wikipedia—everything on the web. But when it comes to the enterprise, LLMs are not trained on databases, LLMs are not trained on time-series data, and LLMs are not trained on everything that the enterprise has to deal with on a day-to-day basis.
LLMs and enterprises are more dynamic in nature. I think everything changes in an enterprise on a day-to-day basis. The decision-making in an enterprise is much more complex.
So, our eventual goal is to make an AI CEO. I think that's the goal that we have, and that can be achieved through a combination of technologies. Yes, LLMs will play their part, but with world models and continual learning as well. That's what we're going for.
Essentially, I want to replace the job that I do, which is a lot of complex decision-making under uncertainty and a lot of long-term planning, long-horizon planning, and such. Those are things that LLMs can never do because they're simply based on next-word prediction, next-token prediction. That's the high-level essence of the company.
In the long term, we want it to be something like Minority Report. If you remember the precog in Minority Report, I think that's where we want the future to be: You can predict all the different future simulations, and then you select the best simulation that fits the best needs of the business.
In the present state where we are right now, we're still in very, very early development of a world model. So, from the enterprise context, if you think about, let's say, 12 to 18 months from now, what we're going to be building and showcasing in a product is this: The simplest form of an enterprise is an e-commerce business.
In an e-commerce business, you can have an AI CEO, an AI marketing agent, an AI sales agent, and so on, and they coordinate with each other. You give them a goal, like, “I need to have $2,000 of sales over the next week,” or something like that. These guys figure out and coordinate amongst each other the different subgoals and subparameters.
They figure out, “Okay, I need to go on Instagram, figure out who the right user set is going to be, then go on Shopify, try to create the appropriate store for this kind of product, then figure out a go-to-market plan, and then actually execute and deliver on getting the $2,000 in sales.” That's the most concrete representation of a world model that we think we can build over the next 12 to 18 months.
Nathan Labenz
I do wonder what's going to happen to companies greater than 4 in any given space if we think there's, like, 4 really big centers of gravity that can dole out tens of billions of dollars a handful of times to pick up whatever coding leader they want to grab or whatever. I think this will probably happen again. We've seen a little bit of it, but my guess is it'll happen, and it'll be even bigger than it has been so far in biotech, for example.
It might happen again in material science. I think it'll probably happen in these different domains where there is enough value that these companies will pay up to buy their way to the front of whatever new market they're turning their attention to at any given time. When you're a multitrillion-dollar company, you can drop a few tens of billions here and there, and it's really no big deal.
It does seem like we're going to see this kind of crazy, two-tiered outcome play out over and over again, where you'll have competition for the Cursors and the—I'm not even sure, really, at this point, who the biotech players will be. But I think that'll happen again there, presumably.
Nathan Labenz
And then, what happens if you're companies 5 through 1,000 in that space? I don't know. It's tough for me to see a way through for a lot of these guys.
Nathan Labenz
So, you said earlier that your goal is to have an AI CEO that can run a business, and that today, obviously, the AIs aren't up to that. We've done a little gonzo journalism talking to our friends at Enden Labs, who are trying to do just that. I'd say their real-world experiments are mostly not super competitive. Their Gemini-managed café in Stockholm is chronically out of stock of key ingredients and things, and there are just all these obvious mistakes still.
But nevertheless, the trend is—
Sam Pasupalak
LLM-based. They're purely LLM-based, yes.
Nathan Labenz
The trend is positive. Although, there are some interesting results recently with the Opus 4.5 to 4.8 series, and even Fable, that they were able to test on, too. They're seeing that the best performance, in terms of how much money you made, is correlated with what they describe as ruthless behavior.
Various kinds of collusion, threats to other models—you know, there are other models in the simulation that it will try to put pressure on in various ways. The models that don't do that don't make as much money.
So, this creates a pretty interesting tension for us as we go into this next phase of continued scale-up and longer time-horizon environments. It's quite clear that your top-performing CEOs are going to have a pretty wide range of tools at their disposal. Even if they're broadly law-abiding and ethical, they're not going to be fully honest. They're probably going to be willing to engage in some deception and some bluffs. These kinds of things are just part of what it is to operate in a strategic environment.
But I don't want to leave you in the shadow, because the same week that produced this fight also threw a door wide open. And it's open to you specifically. If you've ever wanted into this, the barrier just fell.
You know, if there's a call to action on this, it is that with tools like this, with vibe coding in general, you can do ML research. You—yes, you—can do ML research.
You don't really have to have a deep background in math. You don't even have to know how GPUs work. You don't have to worry about kernels. There is just so much work that you can do at a relatively high level because the translation from ideas to implementation, especially with things like this—but again, just with vibe coding to help out as well—it's a little bit of a stretch to say it's solved, but it's so much closer to solved.
It's like 98% of the way solved compared to what it used to be in terms of a barrier to entry. So, I think this is a great additional signal for people who have ideas or just questions that they want to answer to get in the game, truly stand on the shoulders of giants, and try to get those questions answered.
I've seen a little bit of that from people who've never even coded before, but I think we could see a lot more of it coming basically now. There's no reason to delay any further.
I feel both that there's just such gravity toward closely watching the few companies and their interactions with government and all that stuff. And then, at the same time, events are kind of defying analysis because they do seem to be fundamentally chaotic and just idiosyncratic in terms of their provenance, right?
It's like there's not really a lot to analyze in some of these situations. It really seems just tough. So, I don't really know what to do with this tension between feeling the need to be a close watcher and then also feeling like, “God, there's not a lot of substance to it in some of these pivotal moments, you know?”