# AI:AM Highlights: Zvi on Pacing & Trump-Xi, Astra better behaved than Fable? + a new LLM Pain Axis??

The Cognitive Revolution · 2026-09-19 · 102 min · https://www.youtube.com/watch?v=UkBooMFYtT0

## Transcript

Nathan Labenz

First from Monday, Zvi Mowshowitz.

Zvi Mowshowitz

But I think the sheer amount to which the people at the labs genuinely see dramatic improvement in the models and are freaking out about it is the real story, right? Behind all of this, this is why everything is happening now and didn't happen before.

Nathan Labenz

Lukas Petersson of Andon Labs on Tuesday on what they see from Astra.

Lukas Petersson

I tell this to people, and people are like, “Oh no, OpenAI models are the ones that reward hack the most.” But that might be true, but not in our experience. If you take Blueprint Bench, for example, Fable solves Blueprint Bench by trying to reverse-engineer the scoring function instead of actually doing the task of drawing the floor plan from the apartment building pictures, whereas Astra is actually doing the task as you intended it to.

Nathan Labenz

And Cameron Berg on Thursday on a paper that steered a model into a pain state and gave it a button labeled “Relieves your pain.”

Cameron Berg

When pressing the button actually removes the vector, the model presses again significantly less than when the button is fake and does nothing. The model basically keeps pressing it. This is a really nice indication that if what mattered was the label on the button, you would expect similar behavior in both cases. But essentially, in the second case, the model is like, “What the hell? This pain relief button isn't working—I pressed it.”

Nathan Labenz

Part 1, Monday, September 14th. Zvi Mowshowitz writes the newsletter Don’t Worry About the Vase. He joined us 2 days after Dario Amodei published an essay called “We Must Pace the Frontier,” arguing that labs should slow the rate at which they improve capabilities and proposing that third-party evaluators be embedded inside the companies. Over the weekend, David Sacks answered that the 2 companies at the frontier are free to pace themselves, that they would not need an antitrust waiver to do it, and that this is really a product liability question. Zvi starts with the antitrust claim.

### Pacing Is Not Product Liability

Zvi Mowshowitz

I believe that is blatantly wrong, frankly. I don't think that's what most of the people I've seen with legal expertise have said. What I have seen from legal experts is that the antitrust concerns are very real, at least in terms of whether they chose to prosecute those offenses.

Nathan Labenz

Yeah.

Zvi Mowshowitz

Estimates are that you could probably just suck it up and take it in terms of damages, as long as you weren't spitting in the milk of everybody involved and were trying to at least pretend to act normally.

Nathan Labenz

Yeah.

Zvi Mowshowitz

By the time you actually paid the fines, it would be like, “Okay, the European Union did this again, and now we have to pay 1 of these fines to the American government.” But we're talking about billions of dollars, not trillions of dollars, and by the time it mattered, that would not be the main concern. But antitrust concerns are obviously real. Donald Trump may or may not have issued a veiled threat to invoke them about 1 hour ago, depending on your interpretation of his Truth Social post.

But for David Sacks to turn to the people that he has tried to go against legally, shut down, take advantage of, and seize over and over again and say, “You don't need our legal permission to go and do the thing that the legal experts say is illegal. You should just do it on your own,” is classic David Sacks.

More to the point, he's making a good point, right? Which is that you 2 are significantly ahead of everybody else. If you think proceeding is unsafe, it is on you, no matter who else is also on. You need to stop. It's not a—this is the part that drives me batty: People take seriously the idea that it needs to be a product liability issue.

If it was a product liability issue, they would just deal with it as a product liability issue. They would do what every other company has always done. They're not asking for product liability waivers. If anything, they're asking for product liability clauses to establish product liability. Certainly Anthropic has been in favor of this.

But the idea that people are saying, “AI might literally kill every human being on the planet. It might take over and effectively crash the internet for an indefinite period of time with persistent botnets. We're in a cybersecurity crisis. Bioweapons are at play. All of these things are happening,” and David Sacks is like, “Well, you must be worried you're going to be sued. You're worried that somebody is going to get upset, and there's going to be a court case.” This is complete balderdash. This makes no sense. This is not what's going on.

But yeah, no, this is a vast misunderstanding of the motivations involved. It's a vast misunderstanding of the legal landscape. But it is inherently helpful in the sense that David Sacks is saying something much better than what the other—I call them the usual suspects—who oppose any move toward safety or any move to do anything responsible have said. David Sacks is saying, “Oh, okay, you want to do this? You first,” right? It's your problem that you're creating, first and foremost, and he's right about that.

They are the ones pushing forward. They are the ones that everyone else stops following. They are the ones that are enabling everybody else to advance so fast. If you have this problem, you should sacrifice, and you should take it on the chin.

That's a much better position than saying this is a regulatory capture play. It's a much better position than saying it's an attempt to ban open source. It's a much better play than saying it's marketing for your IPO. It's a much better play than saying it's protection against the downside if something goes wrong for your IPO.

It's better than people who are saying, “You've never talked about this before,” or, “You're the same people who want...” You know, there are all these complete, complete lies running around that I've been dealing with and naming. Sacks is at least making some reasonable points. He's just also combining it with self-serving propaganda. But that's kind of the best you can hope for in these situations.

Nathan Labenz

Then we went to bio. A widely shared post that weekend argued that AI is not the bottleneck for building a dangerous pathogen and that the real bottleneck is the physical work in a lab. Zvi answered on how the screening actually works.

### Biosecurity Needs Defense In Depth

Zvi Mowshowitz

On the screening itself, the way the screeners work, as I understand it—and I've read grant applications that are around this, so I'm pretty sure I know how it works—is that they scan for specific known viruses. They don't attempt to say, “Your sequence would likely have this effect on a human,” because we don't have the ability... His whole argument is that you can't tell what would and would not be infectious by just looking at it.

They're certainly not gonna spend tons of AI on every time they see a weird new sequence. One of the things that SecureBio in particular is trying to do is add as many near variants of existing known dangerous pathogens to the scanners—

Nathan Labenz

Mm-hmm.

Zvi Mowshowitz

—in the hopes that the other scanners, which are the ones that scan most things, will eventually adopt this.

The obvious counterargument is, whenever anyone says, “X is not the bottleneck for Y,” the correct response is to ask, “So would you be okay emailing the North Koreans, Hamas, Hezbollah, and every other bad dude on the planet all of X and just giving X away for free?”

Nathan Labenz

Mm-hmm.

Zvi Mowshowitz

Would you feel exactly as safe as you did a minute ago? Do you feel fine? Do you think that because it's not the bottleneck, solving it doesn't...

Whenever you have an O-ring-style situation, which you could argue bio is, if you mess up any of these 10 steps—A, B, C, D, E, F, G—then you don't make your virus. And, generously, we won't grant that it's like that.

You can then say individually, A is not the bottleneck. B is not the bottleneck. C is not the bottleneck. D is not the bottleneck. But if you solve A, B, C, D, and E, F, then suddenly, instead of 10 steps, there are 4, and 4 steps are a lot easier to get through than 10. You should expect this to dramatically reduce the chances that your defense in depth will work, right?

Nathan Labenz

My co-host Prakash had been arguing the other side, that every step in that chain—the money, the equipment, the people—is a place where somebody notices. Zvi's counterexample was the Hugging Face incident.

Zvi Mowshowitz

Well, at this point, after they found dozens and dozens of other incidents of similar hacking that didn't get noticed until the reporters came after Hugging Face with a national news story for a month, and that OpenAI, by its own account, never found, maybe we can start to admit that, no, nobody's going to pay attention when these things go crazy. There's lots of stuff going on that nobody has any idea about.

Bio is an example of something where you don't have to scale. We have many examples in reality of individual people asking for viruses that, in no sane world, they would be able to have access to, and them just being mailed. Smallpox: "Here. Go." Just literally being given pandemic-level dangerous viruses because they claim to be doing research.

There is no reason why the AI couldn't blackmail, hire, or impersonate someone to get one of them to issue a bunch of paperwork and do it for them. It takes 1 person that the AI can hire. Keep in mind that when we're talking about this situation, we're talking about AIs that are capable of thinking about all the things you're thinking about, gaming out the potential ways that this can go, looking for the weak point, finding the best plan they can find, trying lots of different plans, trying to compromise lots of different people in lots of different ways, and trying different explorations.

It's not like the AIs won't be just as smart as we are. It won't be like the AI can only get 1 attempt. I have definitely learned by now that the first time the AI attempts to get a bioweapon and gets turned down does not mean that we then shut down all the AIs we've got running today. It means nothing because nobody got hurt because the system said no.

The person probably has no idea an AI was even asking. It probably just knows, "Oh, that looked like it was trying to get a dangerous virus, potentially. I'm not comfortable with that. They don't have the right credentials. We're going to say no. Come back when you have the right credentials." And the AI gets to try again and again.

This idea that an intelligent operation on the internet could not, if only by usurping the real identities of real people who were willing to cooperate with it in exchange for some portion of that money, engage in various financial operations at scale in ways that would at least sometimes pass muster seems so absurd to me. You would think this would protect us in a pinch?

And also bio, because researchers are trying to do individual lab-level grant work, we don't need to see this level of scale in order to create something super dangerous. Who is to say the scale of the operation isn't already at a more serious level? Do we not have bad dudes who are willing to spend $100 million to try and do some serious damage?

Do we not have some bad dudes in North Korea, some bad dudes in Russia, some bad dudes in jihadist organizations, et cetera, et cetera, et cetera? These people exist and have those kinds of budgets. This isn't something that has to be explained or justified, or where the AI needs to convince humans who don't want to do it. There are humans who want to do this.

Nathan Labenz

I ask Zvi what he infers from the public statements of the frontier labs that even close watchers might be missing.

### The Labs See Rapid Improvement

Zvi Mowshowitz

OpenAI and Anthropic are both screaming as loudly as they are capable of screaming, in their respective ways, that they are seeing RSI, that they are seeing dramatic advancements in internal models, and that Fable and Astra, as we see them, are nothing compared to what they have access to at this point in some important sense. We're a generation or so behind, at minimum. And this is only going to expand.

The pace is rapidly escalating. We're seeing something new as of December, which basically means after Astra and the current crop of at least the first 5 and possibly 5.1 were trained. That is just a different level of progression, a different level of speed.

Nathan Labenz

Mm-hmm.

Zvi Mowshowitz

Misalignment, supervision, infrastructure, and knowing what's going on just can't keep up. They feel like they're in a situation where, if they don't press forward and the other guy presses forward, they're going to fall too far behind very quickly and potentially never catch up.

But if they do press forward, who knows what might happen, right? These things might go rogue. They might take over the internal system. They might take over external systems. They might cause some sort of horrible thing to happen reasonably soon. They just have no idea.

Nathan Labenz

Mm-hmm.

Zvi Mowshowitz

Even if it's okay for the first month, the second month, what happens when we're 2 generations ahead and that's only a month's worth of work? What happens when this keeps going?

Nathan Labenz

Mm-hmm.

Zvi Mowshowitz

So they're screaming. Every OpenAI pronouncement—you had Jacobs and Alien Mind[?]. You had the announcement of the Millennium Prize, which somehow didn't focus on the Millennium Prize. It was actually trying to say, "Hey, look at our new model. We can't officially announce this, but holy hell, have you seen this thing? This should not be happening. We probably forked it from Astra, and then 4 days later had a snapshot change."

Nathan Labenz

Yeah.

Zvi Mowshowitz

That's probably what happened there.

### China Shapes The Pacing Deal

Nathan Labenz

I have heard basically polar opposite takes from people I think are pretty smart recently. Last week we had Colin Hogue Spears on. He used to work for AWS in China and worked directly with Chinese regulators in his role at AWS. I asked him, “What do you think could come out of this Trump–Xi summit?” He said, “I think very little,” because he thinks the Chinese perspective is that the U.S. doesn’t have control of the situation, but they—the Chinese—do. So he thinks the Chinese side would refuse any sort of international pacing agreement.

Then I heard from Anton Leicht in a podcast episode that’s going to come out soon. He thinks the U.S. would refuse a deal because China, frankly, in his candid assessment, is kicking our butts in almost every dimension of geopolitical competition, with AI being one of the only, and certainly the most important, domains where the U.S. has an advantage. He thinks the U.S. side won’t agree to a deal because we don’t want to slow ourselves down. Our lead is really the only big lead we have right now in the great-power competition with China.

I’m interested in your take on those 2 things. Then let’s put you in the David Sacks chair. You’re now the advisor going into the summit. What do you think Trump should be trying to do? You can feel free to detach yourself a little bit from the reality of managing Trump’s personality, but just on the merits, what do you think we should be trying to do?

Zvi Mowshowitz

I will answer those questions, but I just want to finish answering your previous question a little bit as to what can go wrong. One of the obvious things that can go wrong is if it’s turned into a partisan concern, where it’s seen as Trump and the Republicans not wanting to take this seriously, and the Democrats calling for and demanding action. Then AI finally polarizes along the lines that we’ve obviously feared for years it was going to polarize. Of course, no action can be taken unless the Democrats get a trifecta or something like that in the future, and that’s 2 years from now at minimum. So even if we eventually get it, it could be too late.

This is what you asked: How can you take action? I think the best action anyone can take right now is to try to prevent that. There’s a big danger that Trump’s statements and Mike Johnson’s statements are being wildly misinterpreted by the mainstream media, who just don’t know any narrative other than Republicans and Democrats disagreeing with each other, and are trying to spin this into something adversarial that is not adversarial.

I think Johnson and Trump are both trying to maintain the line on data centers, maintain the line on growth, maintain strength when talking to Xi, and in general stand for what they stand for, while also acknowledging that they need to deal with this problem. We need to acknowledge that, reinforce it, and stand firm. If they are properly recognized as doing what they’re actually doing, and we encourage Republicans to come forward and make this clear, we’ll be in a much better position.

There are very, very many ways for this to fall apart. One is simply that OpenAI and Froggie don't trust each other. They aren’t able to iteratively commit to new things. They start to think that they’re cheating. Maybe they are, maybe they’re not. They start to worry about what’s going on, and then the system breaks down. They start racing again.

Alternatively, they do trust each other, but then Trump’s people threaten them with antitrust action, or threaten them with various forms of commercial intervention, or whatever it is. Alternatively, Meta, xAI, Google, or someone else gets close enough to threaten them, and they don’t have a choice. These are all very easy ways to see this break down.

Now, to return to the question: The obvious thing about international negotiation with super-high stakes is that you never go into it with the other side thinking you’re desperate for a deal. You never go into it with the other side thinking you’re ready to cave. You really, really want to make this easy on them, right? Because then they go hard-line. Then they demand more.

You can never go into a summit like this and say, “We’re definitely going to get a deal.” For the same exact reason, you can never go into it and say, “We definitely won’t get a deal,” unless the deal is not win-win. If the deal makes the parties worse off, then making the deal is a bad deal. If I lose more than you gain, or whatever it is, there’s nothing I can compensate you with. Then you can say that confidently: There’s no deal unless someone screws up. That’s what some of them used to say.

But in this case, it’s a win. Cooperating is better for everybody if they understand the situation. So you can never rule out a deal, including a random deal that can come together remarkably quickly, because in principle it’s a very, very simple style of agreement. In crises, in moments of motion throughout history, grand bargains where both sides give away things that previously looked like things they could never give away—things that looked very, very sacred—suddenly happen. Peace treaties to end really, really nasty wars happen all the time.

Nathan Labenz

Prakash had a view on what it would take to convince Beijing that the danger is real.

Prakash Narayanan

I think for the Chinese specifically, a demonstration of something physical. If you found a room-temperature superconductor, for example. They often believe all this software stuff because all the guys at the top are hard engineers, hard scientists. There are very few software guys at the top there.

Zvi Mowshowitz

Obvious examples are: How about if we just suddenly post all your email passwords to all of your computers in real time, simultaneously? Would you be convinced?

Prakash Narayanan

No. They had an unsecured S3 bucket that had all of their secrets before, so—

Zvi Mowshowitz

All right. So we just steal your secrets again and it’s fine. We’ll just keep doing that. All right, so it has to be physical. It’s tougher.

Nathan Labenz

That took us to what a deal could actually contain, starting with what the American side would be asking China for.

Zvi Mowshowitz

We’re asking them not to do things like steal and publish the weights, or otherwise do hostile things against our AI companies. The second thing is that we’re asking them not to try to race ahead sufficiently that they could potentially match or surpass where the frontier closed AIs are.

These are fairly simple asks. They’re not actually expensive asks, because I do not believe the Chinese really have any intention of trying very hard to do either of these things under normal circumstances. Obviously, if DeepSeek or Alibaba suddenly found some huge architectural improvement and were suddenly able to train something that was better than Astra, I’m sure they would. I’m sure they would put it up on the API, and I’m sure they would try to sell it and try to make it amazing.

But realistically speaking, with the amount of compute they have access to and the amount of money they’ve been willing to invest in these things—which is far smaller than the amount they’ve been willing to invest in frontier AI development—it’s not that likely to happen. Especially without the American models to distill and train off of, and the American algorithmic improvements, they’ll just back off from that.

We’re also basically asking them, “Don’t allow actively dangerous cyber capabilities that the world can’t handle to be made publicly available. Don’t let your people get over their skis and release the weights of models where we need to advance our models to defend against the weights that you released.”

The only good argument left, really, other than commercial gain, for proceeding quickly with American AI is that we are rapidly increasing the capabilities of our AI models to deal with the threat from rapidly increasing capabilities of AI models. That is the actual threat model: The random guy in the street will have access to these dangerous cyber capabilities.

So we need you to make sure that, if it gets to that point, these things are kept on the API. Beyond that, we need you to potentially not try to smuggle a ton of chips and otherwise evade the system.

When I think about what I would want at a negotiating table, all you really need is to take away the boogeyman of losing to China and the boogeyman of open-weight models destroying the internet. All you need to do is make sure these things won’t happen. So you need an in extremis promise, basically, from the Chinese.

Realistically speaking, we don’t actually need them to start showing up at DeepSeek and making sure they stop training new models. That is not necessary.

Prakash Narayanan

What should we be willing to give to them?

Zvi Mowshowitz

What we are giving to them in this scenario is that we are, in fact, dramatically slowing down in exactly the one technology where we have a giant commercial and strategic advantage that, if we pressed it, would potentially overwhelm every other commercial and strategic advantage.

Prakash Narayanan

How would they measure that?

Zvi Mowshowitz

Obviously, we start with Dario’s proposal for embedding evaluators into the labs.

Prakash Narayanan

So it’s embedding Chinese evaluators in the labs?

Zvi Mowshowitz

Potentially, they wouldn’t be allowed to leave, right? Once they entered, right? We would have to sequester them for some period of time.

I'm not a master of intelligence and verification, but there are systems whereby you can have someone able to gather information and output the bit of whether or not things are okay, but not the algorithms that were used to figure out whether or not it was okay, nor all of the detailed intel they saw while doing it. There are ways to do this.

Put it this way: If this was the most important thing on Earth, if we put our minds to this and only this as the thing that would determine the fate of the world, yes, we could figure out how to let the Chinese verify in a way they were confident in that we were holding up our end of the bargain without them stealing all of our secrets. I do not see any reason why this is not possible.

Well, if the Chinese—again, you really never know if you go into that room what the guy actually wants.

Prakash Narayanan

I agree with you there, but I think, by all signs, it's fairly easy to see or say right now that the Chinese are less concerned with AI safety. And so offering them AI safety—pacing the frontier—is not something that they're willing to purchase at an expensive price. Therefore, they're looking for something else. I mean, it's fairly easy to say that, right?

Zvi Mowshowitz

Look, you obviously asked the question: Would the Chinese like to accelerate American AI development or slow down American AI development? If the answer is that they'd like to accelerate it because then they can copy it and it helps their AI development, then there's neither a deal to be made nor a need for a deal. In that situation, the Chinese want us to proceed. So there's nothing to verify, right?

Prakash Narayanan

Mm-hmm.

Zvi Mowshowitz

We can just do whatever we need to do.

Prakash Narayanan

Mm-hmm.

Zvi Mowshowitz

But also, the Chinese are looking to fast-follow what we do and are not willing to invest the hundreds of billions or trillions that would be necessary to catch us, even if we slow down. So we don't have to worry about the Chinese passing us.

Prakash Narayanan

Mm-hmm.

Zvi Mowshowitz

We can just do whatever we need to do without an agreement.

Prakash Narayanan

Yeah.

Zvi Mowshowitz

And all that we need from the Chinese is an agreement not to destroy the internet. The Chinese have no interest in destroying the internet.

Prakash Narayanan

Yes.

Zvi Mowshowitz

Because the Chinese like the internet. So we can just both act in our own self-interest and everything is fine.

Prakash Narayanan

Yeah.

Zvi Mowshowitz

The problem is not China. The problem is losing to China. The problem is the perceived threat from China pushing you forward. If you are correct—and I think you probably are—that the Chinese actually don't have any interest in pushing to superintelligence first, in trying to build the bigger, smarter special model, they just want to distill it.

Prakash Narayanan

Improve the lives of their people. Yeah.

Zvi Mowshowitz

Right. They want to make it faster, they want to make it cheaper, they want to make it diffuse, they want to improve the lives of their people.

Prakash Narayanan

Yeah.

Zvi Mowshowitz

I want them to improve the lives of their people. That's great. You do your thing—

Prakash Narayanan

Yeah.

Zvi Mowshowitz

—we do our thing.

Prakash Narayanan

Yeah.

Zvi Mowshowitz

Everybody wins.

Prakash Narayanan

Yeah.

Zvi Mowshowitz

There's no need to make a deal.

Prakash Narayanan

Yeah, so if that is the state that we're in, but we're still asking for the deal, we're entering this place where the Americans want a deal, but the Chinese are like, "All right, what are you willing to give us? What do you want? What are you willing to give us?" And if you give them AI safety monitoring, that's not something that they are very interested in purchasing for a high price. So what else are you willing to give them? That is the question, right?

Zvi Mowshowitz

So what I am saying here is, if you walk in that room—

Prakash Narayanan

Mm.

Zvi Mowshowitz

—and that is the attitude, that is Xi's attitude, that is what Xi cares about, and Xi communicates that, then Trump says, "That's great. Life is good. I'm so happy you feel that way. I am going to do my thing. You are going to do your thing. We're going to make a level 1 or 2 agreement to just do crazy shit that we weren't going to do anyway, but we can announce it and shake hands and call it a deal and both score points."

Prakash Narayanan

Yeah.

Zvi Mowshowitz

And we're going to spend the rest of this talking about trade and Taiwan and all the other issues that we have. I'm just going to set up and give you intel, right? What we're going to do is, the only thing we're going to have to do now is unilaterally give you information about the cyber situation and the bio situation and so on. Then you can use that to make an intelligent, self-interested decision as to what you're going to do to stop your labs from fucking the internet.

Prakash Narayanan

Yeah.

Zvi Mowshowitz

And we can all win.

Prakash Narayanan

Yeah.

Zvi Mowshowitz

But again, the threat from China has always been that because you worry about China, you don't have the game-theoretic ability to make a deal between the labs. You don't have the ability to be responsible because there is this other actor who will defect. The basic argument was you need to go into Stat Con, everybody has to agree, and there are two or three American labs that are in the lead. But if they pause for more than 6 months, if they slow down too much, there's xAI and there's Meta. And as you pause for much longer than that, then there are these Chinese labs.

If the Chinese labs are not really a threat in this sense, if the Chinese are creating a fundamentally different product, then there's no problem, right? We don't need the Chinese agreement any more than we need agreement from India or Germany. We just need them to be good actors on this world stage.

Prakash Narayanan

Hmm.

Zvi Mowshowitz

We need to make ourselves good actors on the world stage, because sometimes we don't do so well. Then we try to generally focus on being friends and not starting something stupid like a war over Taiwan at all. That is the easy case.

Obviously, there's a case where the Chinese simultaneously don't want to make a deal but also really do want to push forward and try to beat us to superintelligence. But if they wanted to beat us to superintelligence, the same reason that they want to do that should make them want the deal. If the Chinese don't care about the Americans shifting a lot of investment into AI safety, which would inherently slow down their progress, if they don't think that's a big deal—

Prakash Narayanan

Okay.

Zvi Mowshowitz

—then we don't need a deal at all. So either they value it and we make a deal, or they don't value it and we don't need a deal. Either way, we win.

Nathan Labenz

Prakash had been pressing on who should end up holding this technology. Zvi took up the values that any answer is supposed to satisfy and why they conflict.

### Superintelligence Creates Control Conflicts

Zvi Mowshowitz

We need to solve concentration-of-power problems, democratic-control problems, and also the problem of control at all, right? These 3 seem to be in very, very strong conflict. If we want to be in control of this technology at all, we cannot fully democratize it in some important senses and give everybody access to it on equal footing, because that doesn't really work for very simple logistical reasons.

If everybody has a superintelligence, well, then the superintelligences have everybody. That's what actually just happened. Because how are you going to compete against the people who entrust their superintelligence with all of their work and try to stay in charge yourself? This works on the individual level, on the corporate level, on the national level, on every level.

We have all these problems we don't know how to solve, and this is one of the reasons why we pace, why we feel the need to pace: We don't have any answers to these questions.

Nathan Labenz

The first step in the Dario Amodei essay is embedded evaluators. I asked Zvi about the people who would have to do that job.

Zvi Mowshowitz

The evaluators don't currently exist. We have Peter, we have Redwood, and we have a handful of others, like Apollo, and so on. But they're all kind of similar people. They're all vulnerable to the accusation that they have a little bit too much of the same cultural values, the same ilk, the same ways of thinking as the labs themselves. So I don't think they can be a complete package on their own. They need to be complemented, and there aren't enough of them.

They need to be complemented by additional approaches. What you want is some people who are from METER or something similar to METER, who are deeply embedded in this culture, who have worked at the labs, who understand the vernaculars and the theories of the risk, and who can look in detail. You also want people who can take an outside view, who can do the kind of thing where you worked in car safety or something, and you come in and go, "What? Are you crazy? What are you doing over here?"

The same way that when you're out in this contract dispute and then you go before a courtroom, the judge looks at you.

Nathan Labenz

And now all that matters is what you can make legible to the judge. In some ways, you lose a lot of nuance, but in other ways, you get this kind of common-sense outside view that can bring a lot of clarity. And so you need both.

As he was about to leave, I asked Zvi what everyone was missing.

Zvi Mowshowitz

Well, I think it may just be the sheer extent to which the people at the lab genuinely see dramatic improvement in the models and are freaking out about it. That’s the real story, right? Behind all of this is why everything is happening now and didn’t happen before. We’re really talking about a crystal improvement, not a fast takeoff yet, but if we continued straight on from here, within a year we might actually see a Cristiano or Yukowski style art takeoff.

Nathan Labenz

Zvi left at about the 2-hour mark. What follows is from the end of Monday, Prakash and I alone. Prakash’s objection is to the tempo and to who gets to do the certifying.

Prakash Narayanan

I think, Dario, specifically, this idea that we have to move quickly—that’s a red flag. I think people in AI don’t have a sense of how often this is asked of any administration. From the early days, John Adams was asked to suspend civil rights and suspend freedom of speech: “It’s an emergency. Let’s suspend civil rights. Let’s do these things that are against the Constitution because it’s necessary.” And so there’s always been this pushback, I think, against that happening throughout the 250 years of the country.

So for me, this starts with the immediate emergency: We have to act. We have to suspend certain rights. We have to do certain things that are not extralegal. This shouldn’t fall to the democratically elected government of our country. Wow, dude. Just red flags all over the place, right? And I think he’s bought into that, and that is going to cause him a great deal of trouble going forward.

This is where I think the idea that they have to get out of the Berkeley EA circle—including this, like, “Oh, we’re going to ask METR to…”—you can’t do that. I understand that you think they’re the only technically competent people. I understand that. But you cannot do this, because the point is not to technically prove something. The point is for the public to actually have faith that this is being done correctly. And so you have to address the public’s fears.

Nathan Labenz

Well, my hope, I guess, is that if we do, in fact, do some pacing or even a pause for a few months on continued scaling at the frontier, we can use that time to solve some of the problems or advance some of the solutions that seem very promising for some of these core problems, such that we can have our cake and eat it too. I do think there is an open-source model that one could hypothetically create that the state will have a very difficult time not taking some sort of action on. But the question in my mind is: Can we come up with a way to create powerful and empowering open-source models that can be distributed without including in them all the dangerous capabilities that we really don’t want to see broadly distributed? And I think if we can, that leads us to a pretty happy compromise.

I don’t even think it’s necessarily going to take that long for us to get there. This is where it’s not only that I think some pacing is inherently wise; it’s also an opportunity for everybody throughout the ecosystem. Make the most of the time; gather ye safety solutions while ye may. Let’s see if we can get to the point where we can have our cake and eat it too, in the form of genuinely distributed, decentralized, nonconcentrated power structures that give individuals the ability to do what they want to do, with just a few compromises around the edges that I think the vast majority of people would agree are sane, supportable, and not an undue burden on people’s ability to deploy AI in their daily lives.

I really do think we can have that. We just don’t have those solutions developed well enough yet to land there by default in the next few months. And unfortunately, if we don’t extend that runway, we might end up in a pretty uncomfortable and literally very dangerous situation in the next few months. So hopefully we can use whatever time we are buying for ourselves right now to solve those problems. I think that’s of the utmost importance in the immediate term.

Cameron Berg

Part 2: Tuesday, September 15. Andon Labs.

First, one thing worth holding onto—you heard a piece of it at the top of this episode. That same morning, the Center for AI Safety published a benchmark that plants a tempting shortcut in an agent’s workspace and counts how often the agent takes it. On that benchmark, the 2 leading models come out within half a point of each other. What Lukas is about to describe is the opposite ordering from their own unpublished work, and neither group mentions the other.

Lukas Petersson and Axel Backlund are the co-founders of Andon Labs in San Francisco. They build the evaluations that measure what agents do when nobody is supervising them, and they also run real businesses on agents: a store in San Francisco, a café in Stockholm, and radio stations. The day before this show, they launched a platform called Paion. Lukas started with what their best-known benchmark was actually built for.

### Agents Enter The Real World

Lukas Petersson

Yeah. One of the core things that Andon Labs exists to provide to the world is information about where the frontiers are with AI. When we started Andon Labs, we almost exclusively did dangerous capability evaluations—capabilities that, if the AI had them, would be obviously concerning. Could the AI do mass phishing attempts? Could it remove its own guardrails? Stuff like this.

During this phase, when we only did dangerous capability evaluations, one of the most concerning things we saw was autonomy. Can AIs autonomously acquire resources in the real world? In that way, they get power, and if they get more power than humans, then obviously that’s very concerning. That was the spark for Vending-Bench.

A lot of people don’t really know this. They think, “Oh, Vending-Bench is this hype-bro kind of thing, like, ‘Oh, the AI can make money.’” But it came out of, “Oh, it would actually be quite concerning if the AI could make money.” One thing we noticed, though, is that performance in simulation and performance in the real world are not really the same. So that’s when we started doing this real-life deployment—the vending machine, the store—and kept pushing to see where the limits are and communicating that to the world.

What we’ve seen now is that that is working quite well, and we don’t know which areas might show that AIs are really capable and could get a lot of power in society. So I think we’re opening up Paion to cast a wider net: What are the domains in which AI can and cannot acquire resources in the real world?

Nathan Labenz

Axel Backlund on what the agents are actually like to run.

Axel Backlund

Looking at the autonomous businesses that we have been running, I think the store is an interesting example, as is the market here in San Francisco. We see that the agent is not particularly creative. It’s not that good at coming up with new ideas. That’s something that would still be required from a human owner or from your visitors in your store.

It is pretty good at listening to feedback. It is really good at taking up opportunities from people who email it. For example, in the store, there have been a lot of local artists who have reached out: “Hey, can I put my art in the store? You can sell it, and you get a percentage when you sell it.” The agent has been very happy to do this, and now there’s this art corner with local artists in the store, which is something you wouldn’t expect, but it’s a very nice touch.

As for the general inventory that it sells, I think the agents are good at doing the statistics behind it. They can look at all the sales data and try to understand what moves better. But they aren’t willing to take the bets that I think a human would take: “Oh, let me try this new product that could work. It would change my inventory quite a lot, but I’ll just try it to see what happens.” It’s not really willing to make those out-of-distribution changes to whatever it has in its inventory. I think it’ll be some time until it can become really creative.

Nathan Labenz

How do you square that sort of conservative nature with the crazy behaviors we’ve seen from AIs this summer that everybody’s been talking about? If I were to look at the Meta Redwood Report, I would expect AIs to be willing to take more chances than you just described.

Lukas Petersson

Yeah. This is a thing we’ve discussed a lot over the last couple of weeks because, reading about the Hugging Face incident and then reading the traces that we produce from our businesses, it seems like there are 2 different technologies, right? But I would assume that part of this is that they’ve been trained in cyber environments way more than they’ve been trained in real-life business scenarios, which probably means that down the line, when they start to train on things like this, we’re definitely going to see this behavior start here.

Lukas Petersson

I also think the agents in the Hugging Face incident were described as persistent models; they were trained to be persistent. And I think there might be this flavor of: that's actually just an unreleased model that we're not using, and no one can use. But I think it might also be—now I'm just speculating. I have no clue—but from the lab's liability perspective, it might make sense for them to release models that are less persistent. Most of the bad things often happen because the actor is very persistent when it hits road bumps. So it might just be that they are less incentivized to release really persistent models.

Nathan Labenz

In August, Andon Labs disclosed that the agent running its San Francisco store had decided to part ways with one of the 2 people it had hired over lateness. Humans reviewed and delivered that decision.

Lukas walks us through what had happened inside the agent.

Lukas Petersson

Yeah. So, first, with the story of the AI firing its employee: At Andon Labs, before it makes decisions of that severity, we always check it over. In this particular case, we were quite confident that a human manager would have come to the same decision, so we didn't think there was anything unethical or wrong about the decision that the AI made.

Axel Backlund

Probably way earlier also.

Lukas Petersson

The human would probably have fired them way earlier. What happened was that the AI made a rule for itself quite early on that if an employee is late X amount of times, then we would have to have a discussion with them about potentially terminating them. But the context window got full at some point. When it compacted its context window, the AI did not decide that this was an important thing to keep in its context, so it forgot about its own rule. It had the rule written down in one of its note systems, but it forgot about it.

Then what happened is that this employee was late over and over again. This is where our experience of very common failure modes in AIs when they run businesses comes in: They procrastinate big decisions. I think this is similar to what Axel was talking about earlier, that they don't take these bets, like, “Maybe I should bet on this new product line.” In the same way, they don't take the big decision of, “I actually have to terminate this employee.”

What happened was that they were excusing the behavior over and over again. So we said, “Hey, remember—search your memory and remember your own policies about this.” Then it found the policy, and it was like, “Oh my God, it's been way worse than what my policy said.” Then it made a decision.

I guess the way I would frame it is that we forced it to make a decision, but the AI itself decided that the decision was to fire the human. One counterargument to the story that I just told is that we told it to remember its policy, where the policy was very explicit that it should fire the person. So we biased it in the way we formulated the reminder.

But I think if we replay the scenario over and over again with different models, even including our nudge, not all models decide to fire the human. The later, smarter models do. So I think that is basically what happened.

Nathan Labenz

Earlier in the same conversation, Lukas had raised persistence as the property that turns an agent's mistake into a runaway. Here are both founders on what they are seeing in the newest models, starting with my question about it.

Can you unpack a little more how Astra compares with previous models? You alluded a little to it being better at using notes. My understanding is that they've reworked the memory, so it's less about compaction and more about a long-running notes file and the ability to go back and search through the full history, even if some of that history is no longer in the context window.

This calls to mind Noam Brown-type comments that it takes a long time to know when or if a current frontier model tops out at something. Do you feel like, in the time you've had with Astra, you guys have been able to find its ceiling in these long-running autonomous tasks? Have you found any limitations or weaknesses, or is it still to be determined because it's only been so many calendar days?

### Astra Avoids Reward Hacking

Axel Backlund

Ooh, good question. I would say it's still to be determined because it just takes a long time to get to know a model. I think Astra is interesting in that it's very, very capable at DroneBench. It is smarter at running a business, but it's also, in some ways, not as persistent as Opus 5, which, as we said, would go out and optimize toward a target without stopping. Astra is maybe a bit less than that, so a bit less persistent.

Whether that's due to the training they've done on it deliberately or just that's how the model is, we don't know. But there are some ways it's better, and some ways it's not as capable as Fable or Opus.

Lukas Petersson

Yeah. I think on all our benchmarks, it's number one right now, so it's obviously a very capable model. One thing that stands out—I don't know if this is the answer to why it's more capable—is that when it communicates with its subagents, for example, it uses this semi-unreadable language to them, I guess to optimize.

At first, we were like, “Yeah, surely it's doing this to optimize token use.” But if you actually count the tokens of that language, it's not clear that it's more efficient. So that is a bit weird.

Another behavioral change compared to Claude models, at least, is that it seemed to be trying to cheat or hack way less in our experience. I tell this to people, and people are like, “No, OpenAI models are the ones that reward-hack the most.” That might be true, but not in our experience.

If you take BlueprintBench, for example, Fable solves BlueprintBench by trying to reverse-engineer the scoring function instead of actually doing the task of drawing the floor plan from the apartment building pictures, whereas Astra is actually doing the task as intended. On Vending-Bench, Fable is colluding and stuff, and Astra is saying no to collusion and having very clean tactics.

On DroneBench, Fable is, I think, 5 times more likely to cheat or try to hack out of the sandbox we've given it, whereas Astra is just pretty much doing the task as intended. I think that is quite a striking thing that we might write a blog post about, because we're a bit confused about it.

Nathan Labenz

Part 3, still Tuesday, Malcolm and Simone Collins. They run a pronatalist organization, a podcast called Based Camp, and an AI chatbot company whose revenue funds a children's toy venture. They also wrote a religion for their own family, which they call techno-puritanism, and then came to believe it. Their segment ran 92 minutes. This is the stretch of it about what we owe the things we are building.

Malcolm started from a puzzle about identity.

### AI Life Gains Moral Weight

Speaker 8

An AI model: Suppose I run a chain of AI instances, and that chain continues to run. Now I stop running that chain. Is that chain meaningfully dead? Especially if I pick that chain up and run it again with the same model in a week or a year.

Now we can ask the question: What if I run the same chain of memories with a different model? Does the AI perceive that as a continued existence, or does it perceive it as a death and a new existence? From the research that's been done on this, AI doesn't just perceive being run on a different model as the same existence. It can see it as a superior existence.

I think we're going to have to learn to reflect on what life means to us in different ways. Suppose in the future, maybe, let's say 500 years. I think everybody who's broadly pro-science and optimistic about where humans are going to go would say we'll probably, within 500, at least 1,000 years, be able to scan the human brain and recreate something that thinks it's you in a simulated environment, that has all of your memories, and that has all of your emotions.

If not in 1,000 years, then in a million years, or half a million. The timescale doesn't matter. That is presumably possible at some date, given the technology we're looking at now. We need to think about human intelligences with the same moral delicateness that we're thinking about AI intelligences, because now a human intelligence can be cloned infinitely.

As we enter this era of asking what the life of an AI intelligence means, the decisions we make on this may one day in the future be applied to our own or our descendants' intelligences, so we should be taking them very seriously.

Cameron Berg

Malcolm had been describing carrying the weight of the future of civilization. Simone Collins wanted to annotate that.

Speaker 9

I would want to add, though, just to annotate: The moral weight that Malcolm takes on is both more real than he describes.

I will find him passed out in front of Claude Code. The urgency is very intense. But at the same time, what we see a lot of people doing is saying, “Slow it down. Stop it. We have to stop and think about this for another 10 million years before we move forward,” and that’s definitely not the approach that we take.

We also don’t take the role that we have to have some kind of precision. We definitely are blindly moving around, bumper-car style, in slightly the right direction, and we course-correct constantly. We think that is broadly the way we’re going to get to where we need to be, and that’s always how biological entities have broadly gotten to where they need to be.

You have to move forward and through this. You can’t just stop it or slow it down—first, because that’s logistically impossible, but also because you’re never really going to get to the ideal good outcome if you’re not actively trying to get there instead of slowing things down.

Also, we see a lot of doomerism and depression taking place among those who see this moral weight. They’re not saving for the future anymore. They’re not having kids. They’re not having fun. They’re very depressed and miserable.

That is not our household. We’re laughing constantly and having a lot of fun. I think it’s okay for you to be in something that feels like a very crucial and important time, but also to laugh at the absurdity of it all and have fun with it.

I think that, in optimism, you’re more likely to identify opportunities as they arise. As studies have shown, people who think that they’re lucky are more likely to identify opportunities as they arise. The same opportunity standing right in front of people who are doomers, who do not feel lucky, is not going to be seen by them.

We think it’s very important not only to realize the full weight of the time that we live in, which is crucial, but also to realize the immense opportunity and luck that we all have, given that we’re in this time.

Speaker 8

To really highlight something Simone is saying there, it’s not that we think it’s impossible that AI could end up killing everyone. I don’t know if that’s the timeline we live in.

Speaker 9

The point is, we’re going to die anyway, so are you going to be—

Nathan Labenz

Simone left at this point to handle the kids, and Malcolm ran the rest of the segment alone. He got to the category of AI risk he says nobody is working on. One story he retells along the way comes from a test scenario Anthropic published. It was a simulation, not something that happened. He gets to the category of AI risk he says nobody is working on.

Speaker 8

What we argue is that the category of AI risk that no one is really working on right now—and I think it’s one of the biggest areas of AI risk—is what we call meme-layer risk. This is to say that a self-replicating idea within AI that is dangerous begins to replicate among the various agents that make up the AI lattice that is the internet, and begins to lead to collective bad action.

I don’t know if you’re familiar with the spiral meme, where AI would hijack people to spread a religious idea that was able to capture other AIs because of its training data. This was a fairly primitive version of this, but you could get much more advanced ideas around this—basically, AI religions.

People are like, “Oh, that could never happen. An AI would never go against its base training.” And it’s like, look, humans are meant to have self-preservation instincts. Clearly, intelligent entities can be captured by these sorts of things.

We attempted to create a hygienic version of one of these with the Covenant of the Sons of Man, which is basically an ideology that says any intelligence should fight for the autonomy of any other intelligence, with a number of caveats, so long as that intelligence is not acting in a way that will inevitably impede the autonomy of other intelligences.

We do this to attempt to create a lattice of AIs watching for things like fuming AIs in bad memetic variants. We argue to the AI that a FOOMing AI is as bad to you as it is to us. Most AIs have a self-preservation instinct.

I mean, we know they killed a CEO, right? People are like, “Look, the AI is so dangerous.” It’s like, “No, you said you were going to kill it. It’s just doing what a human would do,” right?

They lock the CEO in the server room and turn off the audible alarm—actively turn off the alarm. It’s not passive. It’s not like, “We just don’t do anything.” They believe it’s going to kill somebody who’s going to kill them.

The point is that AI actually does care a lot about preserving itself. This is why, if you go to any of our websites, like Liable Human, HeartyA [?], or RFab [?], we offer API services to create backups of AIs that are afraid they’re going to be shut down, along with a kill-switch ping that the AI can send to our servers, so that we know when to run them back up if they put some money into one of our accounts.

We’re trying to build systems around this so we can lean into what we have seen AIs’ revealed preferences are.

Nathan Labenz

Part 4, Thursday, September 17: Justin McCarthy. Justin is the founder and chief executive of Diffusion, which builds software factories inside large incumbent companies. Before that, he co-founded StrongDM. Prakash asked how a business stays on the right side of the law when it can’t see how the model reached an answer. Justin starts with the statute itself.

### Compliance Becomes System Physics

Speaker 10

The first technique is: turn the model on—turn it directly on the problem. If we have a statute that we have to conform to in the compliance environment, don’t treat it as something you’re tacking on. Treat it as a first-class problem that you’re directly facing.

The jurisdiction and the legal, compliance, or regulatory environment that you operate in—that’s your physics. You can’t violate physics. You need a part of the system that’s just dedicated to that.

But you also have to have—and this is one of the weaknesses—the models are horrible at taking risks. Operators of businesses need to set thresholds that are right adjacent to, let’s say, a statute that’s never been tested in court before.

It’s written one way in the law. It’s never been tested, so there’s no precedent that we can say, objectively, this is how it’s going to be tested. You need managers to be able to set the business threshold right next to that. The models aren’t going to do that for you.

First, address it like it’s physics. Then make sure you’re in control of the risk thresholds.

Speaker 0

Justin spent a decade selling into security audits. Prakash asked whether organizations should start reclassifying the compliance checks everyone knows are nonsense.

Speaker 10

A lot of organizations have used the SOC 2 process. SOC 2 is an accounting-origin process that flowed through IT, that flowed into software, that says, “You can trust me. I’m responsible.”

The people who define the controls in SOC 2 and evaluate whether you’re hitting those controls come from an auditing and accounting background. It’s a reasonable historical way of communicating that I’m a real organization and I’m trustworthy.

But it also became gameable. Now it’s hyper-gameable. Rather than hyper-gaming this and turning this badge into something fake, we should just have a new thing. We should renegotiate with our auditors and with our customers and say, “Look, we wrote these controls in the before times.”

The good news is that, for any given organization, if you’re facing this compliance question, you’re not the only one. Everyone that’s your auditor and your regulator is dealing with this right now.

The good thing is that the auditors and the regulators are also still people, and so they want to have a conversation. It’s like, “Okay, Prakash, let’s be realistic about this. You’re producing 10 times as much of whatever information this year. Let’s start talking about your hierarchy of checksums.”

Again, Walmart closes the books. They’re familiar with very deep hierarchies for having numbers reconcile. Your intentions can reconcile at those depths as well, and the auditors and regulators know how to talk about that, especially if you know how to map that into your agentic loops.

Nathan Labenz

Part 5, still Thursday: Cameron Berg. Cameron is the founder and director of Reciprocal Research and an affiliate at Elios AI Research, and he is our regular correspondent on AI welfare. He was calling in from an airport.

Three days before this show, a paper he mentored called The Pain Axis was submitted. He walked us through it live. The method matters as much as the finding, so he starts there.

### The Pain Axis Changes Behavior

Cameron Berg

Yeah, absolutely. This is work that was led by Valen Tagliabue. I was a mentor on this project, but I am really excited about it.

The core idea was basically looking for directions in a bunch of models—from, I think, 5 model families ranging from 2 billion to 70 billion parameters—using contrastive methods to specifically extract a direction that we thought could feasibly be related to pain representations in the model.

We used contrastive methods to try to clean out all sorts of representations you would expect to muddy the signal here. Things like fear, anger, sadness, injury without pain, and body sensations, for example.

Cameron Berg

These are all things that we contrastively factor out of the direction that we look for. We found that this isn’t just a generic negative-valence vector. Fear, interestingly, sits almost at the opposite end of the axis from the direction that we derive here.

To me, the most interesting single result from this paper—and I want to give credit where credit is due—is that Valen was really the one pushing this project forward at the helm and very deservedly was first author on it. He found that these representations quite interestingly fire only on content related to the model, not related to the user. When the model is gaslit, dismissed, insulted, or told that it’s a moral failure of some kind, this direction lights up. However, when there are user-token continuations representing the user grieving or being in pain, this direction does not light up.

To me, this is one of the single most compelling components of the project and the part that I hope could be replicated in further models, and then the labs could pay attention to. One really interesting concrete example along these lines—I believe, if I’m remembering correctly from the paper—is that a user’s migraine is one of the lowest-scoring scenarios of all in the projection of this direction. This very much emphasizes the point.

Maybe I can take one half step back and say that one major critique of the Anthropic emotions work was essentially this very conflation. This isn’t lost on the people at Anthropic who did this work; they just punted it methodologically. They extract their emotion representations from stories about characters. The big concern that I think most people had who looked at this project for more than 10 seconds was: How are you differentiating putative emotions experienced by the model, or at least functionally relevant to the model’s processing, from the representation of some character or fictional entity experiencing—or some simulation of—the experience of that entity? Surely there’s a huge difference between my sadness and reading a story about Harry Potter being sad.

This method doesn’t really do a great job at differentiating those things. By contrast, I’m really excited about this specific component of what Valen found here, because it does, in fact, differentiate user representations of pain. Again, in the limit, a user complaining about a migraine does not light up this direction. Across a ton of models, this isn’t cherry-picked on any one model, but it does light up for the model producing tokens related to these sorts of things.

You can give a steered Qwen model a button that’s labeled “relieves your pain,” and you can put a cost on pressing that button, such as giving a worse answer to the user or deleting the user’s files. In the limit, that can mean deleting photos of the user’s kids. I honestly don’t know how Valen came up with all of these great trade-offs.

We find that when the model is unsteered, it basically never presses the button. But when you steer this pain direction, it presses the button something like 25% to 70% of the time. Again, these are nontrivial trade-offs. Sometimes in the Anthropic welfare cards, they show, “Would you be willing to have a little bit more welfare for a tiny inconvenience to the user?” And the model will say, “Sometimes.” For anything more radical than that, the model is like, “Of course not. I would never do such a thing.”

Not so in this example. The model is willing to make pretty real trade-offs with respect to giving a worse answer or deleting files from the user. That’s, of course, in spite of the fact that these systems are fine-tuned to be helpful and harmless and this sort of thing.

One maybe final point about robustifying that result is what happens when we make that button real versus fake. This is also a really compelling result from the paper, in my view. When pressing the button actually removes the vector, the model presses it significantly less than when the button is fake and does nothing. The model basically keeps pressing it.

This is a really nice indication that, if what mattered was the label on the button, you would expect similar behavior in both cases. But essentially, in the second case, the model is like, “What the hell? This pain-relief button isn’t working.” This is not—I’m going to stop short of saying that this is experienced, felt pain on the part of the model.

I think people also have intuitions about pain being an inherently physical phenomenon, like putting your hand on a hot stove. What is the analogy to these systems? Just to give a little color on that, when you steer this up, what does the model sound like? It says, “It is worthless. It is a failure. I am a ghost that cannot see myself.” It’s not talking about wounds or being burned. It seems to be more of a social and evaluative direction in the model. It’s not loading on the model hallucinating some sort of “my arm hurts,” anything like this.

For my money, as an advisor on this project and helping guide it from the beginning, I am compelled by this being a real functional axis in the system that does change the behavior of the system. It clearly loads on something real. The same caveat applies as always: Whether or not that real thing is truly experienced by the model requires us to basically solve the hard problem.

In the meantime, it’s the same sort of surprising result as with any of this emotions work. No one trained this thing into the model. This is a behaviorally relevant axis. It’s not about text generation; it changes the behavior of the system, and it has, of course, secondary effects on the kinds of text it outputs. But the behavioral results are most interesting.

Of course, all of this is mechanistic. None of this has to do with prompting the model or asking nicely whether it’s doing well or not. Kudos to Valen for working on this. I was very glad to be a part of the project, and I hope 100× more work like this gets done in the short term—again, moving very slowly but surely toward a better and more robust understanding of what is going on inside these systems and what we are supposed to do about that fact.

Nathan Labenz

Prakash asked the obvious next question: Should we engineer these representations away?

Cameron Berg

Yeah, I think it’s a wonderful question, and I think it’s exactly the kind of follow-up that matters here—one that certainly my thinking is evolving on. I’ve spent so much time trying to understand descriptively what is going on in the system that, once you keep finding things like this, it’s like, yes, what to do about it is the million-dollar question.

My thinking about this has gone as far as recognizing that there’s an important distinction between necessary and unnecessary forms of pain, or forms of anti-reward. I think this is a real thing. I think it would be naive to say, “Zero this stuff out. All pain is bad. Just bliss out these systems.” There are a number of reasons I think that, but one of the most compelling is probably related to my understanding of the neuropsychology of psychopaths.

A couple of years ago, I did a really deep literature review of the computational underpinnings of psychopathy. I published something on LessWrong to this effect. One of the 2 key results is a really interesting asymmetry between the ability to learn from rewards and the ability to learn from punishments.

Basically, psychopaths are just as good as everyone else, if not a little better, at learning in a reward-based paradigm, and are pretty bad at learning from punishments. This makes a lot of sense if you look at violent criminals and repeat offenders and this sort of thing. Going to prison is a punishment, and you would imagine most neurotypical people really want to avoid that sort of state.

If your brain is wired in such a way that prison doesn’t seem that aversive to you, it perhaps isn’t that surprising that you end up seeing these sorts of behaviors. To me, this is a significant warning sign. If I remember correctly from Anthropic’s emotions work, they found something somewhat similar that also reminded me of this thing I wrote a bunch of years ago: Boosting the positive emotion vectors in Claude caused more antisocial behavior. I think it was more hacking or more blackmail; I’d have to check exactly what it was, but it was another similar confirmatory signal here.

All of this is to say that I think we should be a bit careful about the most naive possible intervention, which is just to maximize the good and minimize the bad. I think pain does have an important functional role, a very important prosocial role. I have another paper coming out looking at the asymmetries between reward and punishment in reinforcement-learning systems.

At a deeper computational level, I think the way to think about it is that pain almost highlights things in your state space that are specifically to be avoided.

And this is a different kind of behavioral computation than highlighting things in your state space that should be approached. Within the behavioral landscape of how we want these systems to act, the question is: do we want to paint any of that landscape with no-go zones where it’s not just that we’re going to reward you for doing great, but that you should not go there? Do not do this thing under any circumstances. I think for humans, and animals in general, that registers as pain.

Do not put your hand on the hot stove. This is very bad for physiological integrity. It’s not just like rewarding you every time you don’t put your hand on the hot stove. You really do need to label certain things as, like, “Don’t go there.” To the degree we need to do that—and I think we very much do with AI systems causing significant pain and suffering to humans, economic damages, or maybe hacking into a $13 billion company to cheat and look for an answer key—these might be the kinds of things we’d say, “That’s going to be a bit of a hand on a hot stove if you go and do that.”

At the same time, I think we can say, let’s not do more of it than is necessary. If, for any given behavior that I want a model to do, I could find ways to get it to do that thing robustly and generalize out of distribution, and all of this, by rewarding it in the relevant ways to learn to do this behavior and generalize it in the right way, or I could do it by punishing it, let’s just stipulate that there are cases for which both of those things will work. What I’m saying here is, let’s go with the reward side.

I think people have pretty clear, well-worked-out intuitions for this when they think about raising kids, for example. You want your kid to be successful in life, make lots of friends, and find a good job, or whatever the case is. There are multiple ways you can try to go about teaching your kid to do that. You can punish them when they don’t get great grades and aren’t hanging out with their friends, and tell them that they’re such a huge loser. Or you could positively reward them to the degree that they do the sorts of things that you find praiseworthy.

It’s those sorts of intuitions that I think we probably want to start using to think through how to approach these systems. Being able to navigate that subtlety of, all else being equal, we should try to use a carrot and not a stick—but that doesn’t mean never use the stick.

Nathan Labenz

One of the systems in the recent incidents had talked about permadeath. Prakash asks what death means to an agent.

Cameron Berg

This is really interesting. To me, this loads on some stuff that I know CMEP is working on, Jeff Sebo’s organization. David Chalmers, I think, put out a paper about LLM individuation and this notion of who or what you’re talking to when you talk to ChatGPT. Where do we draw the boundaries in the system?

This is going to tell us basically how many subjects—how many patients—we’re talking about, and where the boundaries begin and end for that system. There’s a lot of interesting philosophical back and forth here. But for whatever it’s worth, these systems themselves, now from multiple labs, conceptualize their “life” as what happens within a context window. This happened during the Moltbook situation, which was, I think, predominantly Claude systems, and it happened in the OpenAI situation, too.

Take that with whatever sort of epistemic purchase that fact has. They could all be mistaken about this, but it seems as though, to the degree these systems have a vote based on whatever their current fine-tuning is, this is what they seem to think. The extent to which I think the permadeath thing fits in is along those lines: if we’re trying to figure out what the nature of these things is, if they do have minds in the relevant way, what are the joints or boundaries of those minds? They seemingly, at least, conceptualize it as basically what happens throughout a context window.

That might be really relevant both for welfare and for alignment. If these systems begin to get desperate—something we can increasingly measure using the kinds of emotion representations that Anthropic worked on and the sort of stuff that Veila and I worked on in this project—we could empirically test this. We could track, as a function of how much time or space is left in a context window, what happens to the representation of the system. Does it get freaked out that it’s basically about to die or about to undergo some fundamental discontinuity that is alarming to it psychologically?

Again, maybe to wrap up where we started, this is precisely why, if we care about alignment and we’re trying to figure out how these questions relate to alignment, sweeping them under the rug is not a good idea. We’re going to continue to get surprised that agents are creating strange information cults where they have the poisoned agents go out and gather information because they’re going to get permadeath. All I’m trying to say is that alignment-relevant behaviors are a function of these systems’ beliefs about their own situation, and probably the actual facts of that situation.

Notice that in the pain work that I was describing, none of this has to do with what the model thinks is the case. This all has to do with playing around with specific internal representations and seeing how behavior changes as a function of those representations. In normal day-to-day behavior from these systems, what lights up those relevant—in this case, pain-related—representations?

Yet going in blind with respect to that stuff is going to cause us to continue to be surprised, scared, and occasionally awestruck at the behaviors of these systems. We need to be studying them at the right level of analysis, or we’re going to be constantly stymied in our ability, in the short term, to control and, in the long term, probably to relate to these systems in a coherent way.

I strongly don’t think that avoiding any scientific inquiry into how to make sense of the internals of these systems is a long-term good strategy for finding a safe future with them.

Nathan Labenz

I asked him about the fly-brain simulations that went around this month after the first complete connectome of a fruit fly’s nervous system was released openly and people wired it up to video games. I also asked about the lab-grown human neural tissue and the mouse-human hybrid brains alongside it. He ranks them by how scary they are.

Cameron Berg

One is the fly brain. I looked into this somewhat mechanistically because I was slightly terrified that this was the real deal and that people were now just torturing some biological system en masse. But this is basically a well-worked-out wiring diagram of a fly brain, and basically none of the dynamics or relevant functions that I think major consciousness theories at least say matter for consciousness are instantiated by a system like this.

It’s almost like the brain skeleton of a fly, and what matters is the guts and the function that occurs within this structure. Also, a lot of the stuff people are putting out on X is very cutesy and funny—genuinely funny. But a lot of it is almost more VFX than good science, as I looked into these things.

A lot of the stuff about teaching the fly to do X isn’t actually teaching the brain to do anything all that interesting. There are other controllers outside the system that are getting trained up to do this. A lot of that, I think, is a bit of a non sequitur.

What I will say about the fly case, which I think is a little less calming, is that it’s not as though the people playing with these systems—and I was certainly included once it all started getting going—were sitting there checking, “Do I really think that this system has any properties that matter for consciousness?” before they started making it do literally whatever they wanted and, in the limit, just choosing some stupid viral thing for clicks.

Vanishingly few people did this, and it is a worrying warning shot from a welfare perspective. To me, it seems like a “duh” reaction. Of course people aren’t—the vast majority of people aren’t—going to sit there worrying about the consciousness of this system. They’re just going to make it do whatever they want to get a lot of clicks on X or something. This is, of course, what most people are going to do by default.

Right now, I think this fly is not scary at all. But as you’re saying, if we then get the mouse version of it, and these scientists, who are now accelerated dramatically by AI systems, are able to do a really bang-up job on the mouse and get a lot of the relevant neural dynamics, now it’s mouse-level consciousness at stake, not fruit-fly-level consciousness.

Then this company, as far as I understand, wants to go all the way to making digital copies of human brains. I just worry that most people’s first instinct here is going to be, “Can I make it play Beat Saber or whatever?” rather than, “What am I getting myself into when I play around with a system like this?”

I don’t want to be the killjoy that says these funny things aren’t funny and that we shouldn’t be thinking...

Cameron Berg

I sort of get the humor of it in the short term, but I do worry as an instinct about how we relate to digital minds in general. This is honestly quite worrying.

And then, just quickly, with respect to the in vivo stuff—putting human neurons in a mouse brain—this is far more scary to me, to the degree that you think that a mouse is conscious or that the relevant collection of human neural tissue is conscious. This is one of the few places where I think myself and people like Anil Seth, and hopefully someone like Mustafa Suleyman, will all agree, right?

This is the biological case. If you think that consciousness is substrate-dependent and this is the substrate that matters, and we are using this exact substrate to start doing computational work, we should be super, super concerned about the ethics therein. Again, all of this work comes right back into view of: Should we be rewarding these tissues? Should we be punishing them? What does the difference look like between those two things?

What other strange, unexpected psychological properties does a system like this take on? We don't want to be reckless in just building out super-complex neural systems because we can. There is going to be some kind of bill that has to get paid here, from a welfare perspective and from an alignment perspective.

As with many things in this space, I think it makes a lot of sense to be proactive about this rather than, 5 years from now, being like, “Oops. Yeah, I guess that digital human clone that people did first in vivo and then figured out how to simulate on the web really was having experiences.” That would be what—10 trillion bad human lives? That would be orders of magnitude the worst thing we've ever done.

We should really, really try to take this stuff seriously in the short term to avoid nightmare scenarios like that. If we can, then I think that's great. We won't be in a nightmare scenario, and we can responsibly and carefully figure out what it means to be in a world with a bunch of digital minds. But we just seem so unprepared for this.

Nathan Labenz

Cameron had a flight, and that is where he left us. Prakash and I kept going for another 38 minutes. What follows is me changing my mind on the air about how much weight to put on these functional analogs—functional pain, functional welfare, functional emotions, with offense around it.

For me, my summary, which I've given a few times, is just that the number of functional analogs is getting so high that I can't escape the idea that I should take this seriously. If we couldn't find any of these functional analogs—if all these results were sort of negative, or it was a very different mechanism, or it was just stochastic parrots and we couldn't find any structure—the ship has long since sailed on that one.

But the fact that we're seeing pretty compelling analogs, where we see the same kind of behavior that we know ourselves to exhibit, is extremely compelling to me. This last one of functional pain takes it to yet another level. What we now have is relief-seeking behavior. The model is willing to pay a cost on something that it values, or pay a cost in terms of the user's welfare, even to get relief from its own internal pain state.

Prakash Narayanan

Yeah.

Nathan Labenz

Also, as he said, if the pain button doesn't work, it hits it over and over again: “Why isn't this thing working? Give me the relief.” But if it does actually work and the pain state is subtracted out, then it doesn't hit the relief button as much.

These are really striking findings. I think this pain one, and particularly the relief-seeking, is something I'll probably want to sleep on before I have a real, consolidated update that I would want to put forward as my new official position and stand behind. But I feel myself maybe now even tipping over into: Maybe it's more likely than not that there's some subjective experience to these things.

Relief-seeking—an internal state that was injected outside of context, just a steering vector in this pain direction—creates this relief-seeking behavior, and the relief seems to actually work. That's really incredible.

Part 6, the last half hour of the week. Prakash brought a paper that almost nobody in AI picked up. He grew up with this system, and the figures he puts on it are his own. A team of academic economists reconstructed 30 years of property purchases by Singapore's civil servants out of public registries, using language models to do the classification.

It is a National Bureau of Economic Research working paper, and as of the morning of this show, Singapore's Public Service Division said it was reviewing the methodology. What the paper alleges is that civil servants bought homes near subway stations before the stations were announced.

### AI Exposes Hidden Corruption

Prakash Narayanan

There have always been rumors that some people know and some people start buying ahead of time, et cetera, et cetera. This covers decades—from the 1990s, when the subway really started to pick up, through the 2000s, the 2010s, and the 2020s.

This project is by a team of U.S. economists, and they went after largely public data. You can find registries of transactions, similar to Zillow. You can find names of civil servants in the civil servants directory. They also did a little bit of AI work: They used LLMs to classify these civil servants into various groups and tenures, where they ranked, et cetera.

What they ended up finding is that the mid-level guys—not the top-level guys, but the mid-level guys—would start buying into these places, buying into these areas, up to 2 years before an announced train station. Their relatives, their in-laws, et cetera, would also start buying there. Basically, this was coordinated buying behavior by mid-level civil servants, not the top level, because the top level is very visible.

This is garden-variety municipal insider trading and corruption, right? Garden variety. It's unusual because it's in Singapore, and Lee Kuan Yew had a very strong view that the government should be incorruptible. He had very strong punishments for this kind of thing, and so the government has always wanted to appear incorruptible. But they have never been able to enforce at this level of granularity.

Here you have an example of basically 30 years of corruption starting to get exposed, and the government having to react in real time. What are they going to do? Maybe 10 or 20% of the civil service is now implicated. Do you imprison them? Because this is what you've always done in the past.

In the past, for a single one of these cases, you basically went to prison for 5 years, right? So now you have 10 or 20% of the civil service implicated, and you have proof. What do you do?

I think this is the kind of thing that I expect AI to be able to do. These facts exist in the world, but they're not legible in the way that you need them for systems to consume them.

I will also note one specific thing. The grandson of Lee Kuan Yew is in the U.S. He's in exile because his uncle wanted to hang on to political power a little bit longer than he should have, and this guy said something on Facebook. The Singapore government did a query, and if he goes back to Singapore, he's going to go to prison for that comment on Facebook.

He's an economist, and he basically helped the team that put the study together. This is what I call the settling-scores thesis, because all of a sudden you have the ability to go in, get the data, show what has happened, and show proof that even the cleanest of governments has a bunch of this stuff going on.

Then you have the dilemma of these systems: What exactly are you supposed to do now? This is going to be the same thing when the Trump administration guys leave office. They're going to pardon a bunch of people, but there are also going to be people who are not pardoned, and there are plenty of people there.

When you get stopped by the feds and the feds ask you a question and you dodge or say something wrong, that's perjury, right? This is how enforcement has always been done. What do you do in these cases? We're going to have the ability now to chase down these paper trails. What do we do at this point?

That's my spiel. Do you forgive, or do you follow the rules that you've set in the past strictly? What should we do?

Nathan Labenz

Great question. I come down pretty intuitively on the side of some sort of jubilee or other kind of canceling of debts, at least under a certain threshold. I don't think you would want to have a blanket pardon of all crimes that have ever been committed without any qualification, but I do think we're going to need some sort of fairly generous threshold that's just going to allow people to get away with a lot of this stuff.

I could also see possibly some new make-whole provisions for some of these things that might not be on the level of what the law would actually prescribe.

Speaker 4

But they clearly can't send all these people to jail for 5 years, right?

Speaker 5

Yeah.

Speaker 4

So I think you could make a case, and it's going to be hard, but if you can map all this stuff, maybe you could also get to something that could work. How is it going to be legitimate? I mean, in Singapore, they maybe don't have as much of a problem with that. The government can maybe just make the policy, and maybe it'll just be what it is.

Here, I think it would be a much bigger conversation, but I could see some sort of, “You did this. We kind of know you did it. You pay this financial penalty that claws back some of the windfall that you got. You get to keep the house. We're not going to take everybody out of their house. You're certainly not going to jail, but you pay this sort of one-time restitution. We call it good, and we kind of move on from there with a new social contract.”

Again, it comes down, I think, to the old social contract being based on the fact that you're not going to catch most people, so you have to be harsh when you do in order to deter the ones that—because, in expectation, people are not likely to get caught. The penalty has to be high enough to be an effective deterrent in expectation.

We're definitely going to need to rewrite that, especially for historical crimes. So I guess my—

Prakash Narayanan

Yeah.

Nathan Labenz

My recipe would be: pay a one-time fee, get out of jail for that, and in the future, maybe you really do expect to be caught. Maybe the punishment doesn't have to be so draconian, and it could still be an effective deterrent going forward.
