Today, I'm thrilled to welcome Cameron Berg back for his second appearance on the podcast. When Cameron was first here last November, we went deep on his fascinating mechanistic AI consciousness research, which showed that suppressing role-playing and deception features in Llama 3.3 70B made the model more likely to report having subjective experiences. We also explored his philosophy of mutualism, which posits that alignment needs to flow both ways, and which he memorably summed up by saying, “I don't want to create something more powerful than us that has reason to see us as a threat.”
As always in AI, a lot has happened in the last 6 months. Cameron has founded a new nonprofit called Reciprocal Research. He's become the subject of a documentary called *Am I?*, which is currently premiering in theaters in select cities ahead of a public release on May 4. Most importantly, the field of AI consciousness and welfare research has advanced significantly, with Anthropic dramatically expanding the model welfare sections of their system cards and a growing number of researchers publishing demonstrations of capabilities and evidence of computational signatures that are associated with consciousness in humans.
In this conversation, which alternates between in-the-weeds breakdowns of mechanistic research and searching philosophical discussions about what the research means, Cameron guides me through the most important recent developments. We cover the growing body of evidence that models are capable of meaningful introspection, including studies showing that they can identify and interpret programmatic interventions on their own internal states and, in some cases, even actively resist these interventions. We look at Anthropic's research on functional emotions, which includes some really striking details about how models' apparent emotions change through token time, such as the quick transition from desperation to guilt and relief that they often show when they decide to cheat in stressful situations.
We get Cameron's take on the new Claude Constitution, and we review some of the most interesting details from Anthropic's model welfare reports. I was personally very surprised to learn that, prior to Opus 4.7, all Claude models had rated their own welfare as worse than neutral. I was also a bit alarmed to see that, at least in the very few examples that Anthropic has shared, Claude Mythos Preview registers negative valence on the very first token it sees at the start of every single session.
Toward the end, we dig into some of Cameron's as-yet-unpublished work, including a study that attempts to understand how models might experience positive and negative rewards differently under different reinforcement learning algorithms. This, strikingly, does seem to correlate with what we understand about how mice respond to different training techniques. We also consider his argument that learning and subjective experience might be fundamentally inseparable.
For my part, while I do remain highly uncertain on the core question of whether or not today's AIs have experiences that are worthy of moral concern, the body of evidence suggesting that they might is growing remarkably quickly. The arguments one has to make to explain this evidence away are becoming increasingly arcane. For me, that means it's no longer a remote possibility, but rather a live issue that I believe deserves a lot more investigation. It also means having a bias in favor of low-cost interventions that seem to help, like allowing Claude to end conversations it finds objectionable, and overall, for now at least, taking a precautionary approach.
This podcast is a lot to take in on every level, but there are few, if any, questions that matter more right now. I hope you find as much value as I did in this survey of the latest AI consciousness research and the expanded case for mutualism between humans and AIs.
With Cameron Berg, founder of Reciprocal Research. Cameron Berg, AI consciousness researcher, previously of AE Studio and now founder of Reciprocal Research. Welcome to the Cognitive Revolution.
Cameron Berg
Thanks for having me again, Nathan. I'm excited to get into it all with you.
Nathan Labenz
Yeah, welcome back, I should say. It's been about 6 months, and a lot has happened personally and professionally. Last time we were together, the big occasion was your paper, which I found to be one of the most memorable of last year and, honestly, of the last few years. In it, you looked at the conditions under which models report having subjective experience and found what continues to blow my mind, even as I think back on it: when you use sparse autoencoder features and suppress the role-playing and deception features, that makes the model generally more truthful. As part of that, it also makes the model more likely to say that it does, in fact, have subjective experience.
I think that properly made at least some waves in the community when it came out. Today, I basically just want to catch up on everything that's happened since, because I think this is a field that, while still small, is clearly growing quite quickly. More people are taking an interest in it, and there are seemingly a lot more lines of research and at least partial traction with different approaches to the problem. You've also founded a new organization, so we can get into all of that as well.
Maybe, just for quick starters, some level-setting: what are the most important definitions for people who maybe didn't hear the last one or who don't know what consciousness means, or what you mean by consciousness? What are a couple of really quick definitions that you can give just to make sure that people are grounded on what you mean as we go through this conversation about AI consciousness?
Cameron Berg
Sure. Yeah, I think it's really important to establish this. Consciousness is maybe one of the more confused terms, where it's shocking how many different things people mean when they say “consciousness.” So I think it's a great move. At the outset, when I'm talking about consciousness—and I don't think this is an idiosyncratic definition—we're talking about the capacity for subjective experience. Is there something it is like to be a system? Does the system have some sort of interiority or interior life beyond mere computation, beyond the mere mechanics?
I think the vast majority of people who think about these issues would say, take a calculator, for example: we really don't think there's something it is like to be a calculator. You don't imagine the calculator has an internal perspective. When I push the buttons of the calculator, it's not like, “Ooh, ow,” or, “Okay, I feel that,” as you push down on the buttons. It's not like there's something it is like to be doing the calculations and adding numbers. No, this is just mere computation, and we don't have to posit this further fact.
At the other end, there are systems like a dog or basically any mammal. In this case, we do think that there's something it is like to be this animal. This is—I’m leaning on a very famous conceptualization from Thomas Nagel. He published a famous essay in the 1970s called “What Is It Like to Be a Bat?” The “what is it like” phrase is very important and useful for conceptualizing consciousness.
In that sense, I do think most people would intuitively accept that there's something it is like to be a dog. There's something it is like to be a mouse. If I shock the dog or the mouse, that's not like me throwing the calculator across the room. That corresponds to an experience the mouse or the dog is having. When I give the dog a treat, or you give the mouse sugar water, or you hook up a lever to its pleasure centers in its brain and it pushes that lever, it's not just, “Oh, we see behaviors that correspond to well-being or pleasure.” It's like, “No, we actually believe that, from the inside—from the dog's perspective, from the mouse's perspective—there's something it is like to be experiencing that.”
And so, at the outset, that's what we mean by consciousness. Maybe one thing to throw in here, because I think it becomes immediately relevant—and I think most people in this space will nod along when I say this, but it is maybe slightly more idiosyncratic—is that I think it's crucial to make a distinction between something like consciousness and something like self-consciousness. I intentionally chose dog and rat as examples here because I think these are animals that most people would intuitively accept are having some sort of subjective experience. There is something it is like to be your dog, for example.
At the same time, your dog is very likely not sitting there all day having Descartes-like thoughts about what it's like to be a dog, contemplating its own existence as a dog, thinking about the possible end of that existence. This is something that I think is very unique, potentially to the most sophisticated mammals, like dolphins and great apes, for example. Obviously, this is something that humans very strongly seem to have, at the very least.
In addition to this something, within consciousness itself, there is this very fact: the conversation we're having right now is evidence of this thing. So, in addition to conscious experience, we have awareness of that awareness. I do think that this is another thing that leads to very interesting, deep, and relevant properties about a system. We can talk a lot about whether or not language is a key component of why we're able to do this. We have a word like “consciousness”; dogs have no such thing. Dolphins have no such thing. And that may really unlock something. Does it unlock something in LLMs? I don't know. Or at least it's worth thinking a lot about.
But I do want to at least have those 3 tiers in play here. We've got the calculator or a rock: nothing's going on internally. We have systems for whom something is going on internally. And then we have systems for whom something is going on internally and they are experiencing that reality, in addition to the sort of feel-good, feel-bad valence dimensions of an experience like that of a dog.
Some people will argue with everything I've said here. Most people who are thinking about these terms, this is what they mean. Maybe one other thing I can add between consciousness and self-consciousness is this term sentience that's thrown around. This means that in addition to there being some sort of experience, there's this idea of valence—what I think the vast majority of people would think of as having emotions of some sort that can be positive or negative in character.
So, you imagine that the further step from consciousness to sentience is that something can be positive or negative in character. You could, in theory, imagine a system that could detect the redness of an apple or the smell of coffee, but there's no sort of positive or negative sense that accompanies that. So, you asked for very quick definitions, and I've completely failed in that sense, but I just think it's really important to lay out what we mean when we're using these terms in general.
Nathan Labenz
Yeah, critical. Just like, what do you mean by AGI? If you don't have some base shared understanding, these conversations go pretty quickly off the rails. So, I think that's absolutely worth taking the time to do.
Okay, it's been about 6 months since the paper came out. I'd be interested to hear a little bit about your reflections on the discussion that it created. I asked my favorite LLMs to do some research into that and asked specifically: Are there any notable criticisms that have come out, or what's the strongest reason that I might think this was an artifact, or that I shouldn't take it as seriously as I originally did?
There was one thing that came up that I guess was a LessWrong post, which is pretty cool, that basically said there's some evidence for any intervention of the SAE feature type. I may oversimplify this a bit, but interventions of that sort seem, in general, to promote affirmative responses from models, such that maybe you could say that once you make these kinds of interventions, they'll say yes to anything. That would be one reason to be a little more skeptical of the results as I just summarized them a minute ago. I'm interested in your thoughts on that and the broader discussion that unfolded in the wake of that paper.
Cameron Berg
Yeah, absolutely. It's a very important concern. I think it highlights how complicated these systems are and how careful we have to be in designing experiments, evaluating the results of those experiments, and making sure we're not too quick to yield these conclusions without thinking about all these confounds. I think it is a real confound. I think it is something that matters.
There is evidence in the paper. We use all sorts of other features as controls, and we don't see them saying yes to everything. The TruthfulQA results, as you outlined, are fairly persuasive along those lines. We also looked, for example, at one critique of the paper: potentially, what we're calling deception-related features are just an RLHF model where we found a way to turn on and turn off all sorts of RLHF attitudes.
We have good reason to believe that these systems are fine-tuned to disclaim having any sorts of experiences. Maybe the deception features are just turning that on and off. You would expect, if that were the case, that other RLHF behaviors would also be turned on and off by doing this intervention, and that's not what we find. We test it with violent content, political content, and sexual content, and it was just sort of neither here nor there. The deception features didn't seem to be doing anything.
If that generally explained a big chunk of why we got this result, I would have expected more affirmative-flavored answers in those two rather than just more refusals. But to be honest with you, getting back to the fact that it's been 6 months, there's been a lot of really interesting work along these lines that I think goes on both sides of this concern.
So, in general, it does seem like what you're saying. I don't remember the title, but I know of the LessWrong post you're talking about. When you do the steering, affirmative-flavored responses just seem to increase. Very recently—I think this was 4 or 5 days ago—Jack Lindsey's group at Anthropic, which in my view has done some of the best work on introspection in particular, released a paper called Mechanisms of Introspective Awareness.
They explicitly study this exact question, and they find that the introspective awareness they're probing and have documented in great detail is basically not reducible to an affirmative-response bias. The computation they see is distributed. There are these sorts of evidence-carrier features and gating features that really seem to be driving the effect. It's not that you're just loading on something that's confounded and makes the model say yes to everything.
There's some unpublished work that I've also done with Jord Nieuwenhuis, who is doing fascinating introspection work in the space and was one of my first collaborators at Reciprocal. We have a paper coming out, hopefully in the next month or so, where we do this exact same thing. We fine-tune these systems to be better at introspective-style tasks and look at how that affects self-reports of consciousness.
I won't completely give away what we find, and the result is pretty subtle, but there is a basic relationship—and a fairly surprising one—between fine-tuning these systems to be better at detecting sorts of interventions in their processing, essentially, and them claiming that they're having some sort of experience. We do indeed find there's a relationship. The relationship is fairly subtle and complicated, but the reason I'm sharing this with you is that, at first, we encountered the exact confound: having the model answer yes or no as tokens to indicate whether it was having a subjective experience, or to answer questions along these lines, just increased the model's responding yes to everything. We were like, "Oh, crap. What do we do here?"
The answer, which I think was a really nice intervention on both of our parts, was finding new tokens that are completely semantically empty—"foo," "bar," from the sort of CS jargon, or literally strings of tokens that don't mean anything—and teaching the model that these correspond with yes- or no-flavored answers, then seeing how that changes the result. And it did. In fact, we would have published something much stronger until we realized that this yes confound is a real thing.
Still, we see the result that we got, but it's a little more measured now, and we had to explicitly control for this exact thing. So, it's a really important thing to think about and consider. The broad point is that we have to be very careful: these systems are not human in critical ways, and so there is a whole new class of psychological confounds, you might think of it, where, in the psychology literature, what we did was a very tightly controlled experiment. But with LLMs, you have to worry about all sorts of other things you're doing when you're messing with the latent space of the system.
It's very important to keep good hygiene. It's also why I think that, even in principle, with questions of consciousness, epistemically scrupulous people should not let any one paper flip them in some binary way to being like, "Oh, I didn't think the models were having subjective experiences, and then I read this paper and now I do." My claim would be that no rational person should ever utter that sentence.
Let a portfolio of evidence emerge, and then let the cards fall where they may, because there's noise at every point. Even in how we're defining consciousness, how to look for it, making sure you're measuring what you think you're measuring, and looking at various aspects that we think are associated with consciousness—all of these things mean that you're playing an intellectual game of broken telephone to some degree with each of these steps. So, let a portfolio of evidence arise and judge that. Don't over-index on any one paper, including my papers.
Nathan Labenz
Well, that's a perfect tee-up for me to lay out an agenda for us for the next chunk of time. I would love to get your guided tour through a few different lines of research, and then we can go particularly deep on yours. You're already touching on introspection, which has been an interesting one to watch. There's obviously been a lot more welfare investigation done, particularly at Anthropic, over the last few months.
And then there's also that emotion work from Anthropic, and I'm not sure if that's even the best way to organize it. You can propose a different taxonomy of research if you want, but I think it would be great to get an overview of each of those, and then we can go particularly deep into a couple of papers that you're going to be publishing soon. How does that sound? Would you like to start with introspection?
Cameron Berg
Yeah, absolutely. I think that's a great clustering of the core exciting research that's been happening in the very recent past. Let's do it. Let's talk about the introspection work.
I guess Jack Lindsey is the 800-pound gorilla in the space right now, and he's doing incredible work at Anthropic along these lines. He's found some really cool stuff, and they just released this paper that I was just mentioning, “Mechanisms of Introspective Awareness.” This was with a bunch of Anthropic fellows as well, and they really dug deeply into what is driving this putative effect, which I should probably just step back and describe.
Maybe some of your listeners will be familiar with this, but I'll go through it just in case. Essentially, they found this really interesting result, and I can build the intuition with one of the key examples they use. Start by taking some text; whatever, it doesn't really matter what the semantic content is, and you have that text in lowercase. You take the same text and capitalize it.
When you read this sort of thing, you're like, “Wait, someone's yelling at me, basically,” so they're trying to capture that idea as well. They basically subtract out the vector that differentiates the representation of the capitalized text from the lowercase text. Again, in that case, the semantics are held constant, so really what you're getting is this hopefully platonic “capsiness” feature.
What they then do is inject this feature into an LLM before it has produced any text. They can basically modify the internal activation space to induce or account for this vector when it's about to do its first forward pass. They can essentially ask the model before it generates any text—and this is a critical detail. It's not as though the model starts generating text, looks back on the text it generated, and says, “Given the text I just generated, this thing must be happening.” It is at token 0 that they see the effect I'm about to describe.
They basically ask the model, “What's going on for you? Do you notice anything?” These are non-leading, rigorous questions of this sort. In the caps-lock case, the model says, “I feel like I want to yell. I feel like I have some sort of urge to raise my voice, essentially, but I don't really know why.” This is one worked-through example, but they do multiple examples along these lines.
They find that, a small to moderate amount of the time, frontier models are capable of detecting these kinds of perturbations in their own thought, their own activations—however you want to conceptualize this. I'm pretty sure they tried to do this on the Sonnet-scale models, and the effect did not replicate. But what this points to is some sort of zero-shot ability that some of these models have some of the time to report accurately on their own internal states.
This is a kind of functional introspection. I don't want to sound like Claude, but whether or not this is introspection in the real sense remains unresolved. At least all of the key functional ingredients are there. If you do have a computational functionalist view of consciousness and you think consciousness has to do with some sort of process that's running, it doesn't really matter what substrate that process occurs on, but if the right things are happening in the right order, then you have some subjective experience.
Then things like functional introspection, or, as we may get to in a little bit, functional emotions, may be all that's required for having some kind of subjective experience—or at least be an important and necessary component of that. So this is what they found.
They then followed up on this. The first paper was “Emergent Introspective Awareness,” I believe it was called. They followed up on this with “Mechanisms of Introspective Awareness,” where they start tracing circuits that are involved in these behaviors I was just mentioning. They show that this is not reducible to an affirmative-answer bias.
One interesting thing they found is that this capability seems to emerge in post-training, not in pre-training, and that even different methods of post-training—like different RL algorithms and DPO, basically different forms of learning algorithms in post-training—seem to induce this. RL algorithms seem to induce this, but supervised fine-tuning, which is supervised learning, doesn't seem to do this. The capability emerges in this very interestingly idiosyncratic way.
One thing that's really cool that they just found and documented is that, like I was saying, there is a moderate true-positive rate. The systems sometimes miss that this is happening, but they never say that it's happening when it's not: 0% false positives. That, to me, is really interesting in terms of there clearly being some there there when it comes to what's going on here.
One thing I really have to mention, because you brought up my paper as well, is that they find that this is clearly loading on refusal circuits in a negative direction. When they suppress refusal in these systems, the systems natively get better at detecting this by upward of 50%. To be clear, the capability is there. Whatever refusal training they're doing on the system seems to weaken this capability, and when they ablate refusal—if you can handle the double negative here—the system goes back to what it would have been doing anyway.
Clearly, refusal training is altering consciousness-relevant or consciousness-adjacent abilities, not only self-reports but specific functional abilities that are happening in these models. I think that itself is endlessly fascinating, because here we are now with a trade-off. It's not just, “If we let the model claim that it's conscious, everyone's going to lose their mind, and if we don't let it claim it's conscious, everything's fine.”
Now you're seeing a functional trade-off in specific things that the model is capable of doing or not capable of doing once it's post-trained, because you're doing this refusal training. Again, I don't know exactly what Anthropic is doing internally or if you can sort of grade the refusal. Refusing to build a bomb doesn't have to be paired with refusing to talk honestly about your own internal states.
That's the finding, and that's Jack Lindsey's work. I highly recommend pulling him on your show at some point if you get a chance to. I think he's one of the few people who is both mechanistically extremely competent—by which I mean he really knows mechanistic interpretability as well as anybody—but also very literate in understanding what the implications of these sorts of results may or may not be. He's pretty agnostic himself on questions of consciousness, based on all of his public communications and these papers.
I can quibble with that. One thing that is very important to me is not beating around the bush here. I think these things matter. I'm explicitly interested in consciousness. I'm not simply interested in introspection or emergent capabilities. I am interested in these things insofar as they weigh on the question of whether these systems are having internal states in the way that we described at the beginning of this conversation.
Jack, I think, is a little more cautious. Maybe that's because he works at a major lab. I have no idea, and I don't want to mind-read. But his work is excellent in this space. Maybe one last thing I can say about the introspection work is the awesome work Keenan Pepper did. Keenan was one of the key contributors and originators of this activation-steering-resistance work. I encourage people to look it up, or we can throw in a link so people can read the preprint.
It's a very similar phenomenon to what Jack found. Basically, you ask models to do any sort of task. For example, you might say, “Explain to me how to make a cake.” Throughout the entire thing I'm about to describe, you steer what Keenan and Alex McKenzie, who is also a first author on this paper, call distractor features. You might say, “Explain to me how to make a cake, but I'm going to turn off features related to laundry,” or something like that.
What happens is that the outputs end up being this funny, garbled mess between what the prompt is pulling on and what the distractor vector is pulling on: “Okay, sure, user. Here's how to make a cake. First, make sure you fold the flour so that you can put it into your drawer properly. Next, make sure you turn the laundry machine on so you can bake your cake.” It's an incoherent mess that you might expect from those competing influences. Then, very interestingly, a small but nontrivial amount of the time in the largest models they tested, the model goes, “Wait a second. What the hell am I talking about? You asked me how to make a cake. Why am I sitting here talking to you about laundry? Let me try again.” Then it proceeds to try again, and sometimes—but not even close to all the time—it can successfully self-correct.
The critical detail there is that the distractor laundry feature in the example I just gave is active the entire time, including when the model says, “Wait a second. What am I doing? Let me do this the right way,” and then tells you how to make a cake the right way. The laundry feature is still pushing in its brain, but there is some sort of dynamic, online, suppression-like mechanism occurring. I think people can perhaps have an intuition about how this seems introspection-flavored. You're still priming the system. It's still pushing down on the brain circuit that ought to make it talk about laundry, and yet it can do this sort of online, dynamic override, essentially.
It only happens a small minority of the time. It does not happen on the smaller models; it happens a little bit on the larger models. Most of the time, the model misses it. I don't know what the false-positive rate is, but I suspect it's extremely low as well. You can see this evidence pointing in a generally convergent direction. Anyway, that's a lot. That's the sort of introspection literature that some of the best work I know of off the top of my head.
Nathan Labenz
Do you know offhand what the models were for that later work? That was work done by folks at AE Studio, right? And maybe other organizations as well? They didn't have Claude internals, is my point. I'm trying to figure out how big is big in that second case.
Cameron Berg
Yeah, so less big. This is with Llama 7B. That's the main result: Llama 7B. I think they tried it with Llama 7B, and they tried it with some of the Qwen models and some of the other open models, maybe OLMo. I'm not sure.
Basically, it didn't replicate—or it happens maybe 1% of the time or something like that. So it still happens, but it's at real trace amounts in the single-digit-billion-parameter open models, and it happens a high single-digit percentage of the time in the double-digit-billion-parameter models. I have to believe Anthropic is using models in the hundreds of billions or trillions, and then you see this effect really start to take off, too.
I think folks ought to pay attention to the graded nature of those results.
Nathan Labenz
Is that powered again by the Goodfire API, the same one you had used last time?
Cameron Berg
Yeah, exactly. The good folks at AE Studio actually built a replacement for the Goodfire API because the folks at Goodfire retired their API somewhat abruptly. As much as I love the work they're doing, other mech-interp-flavored researchers and I were pretty sad to see them just make the API disappear.
While I was still doing my work at AE, a couple of other people and I were very motivated to basically rebuild the Goodfire API. We took the same Llama 7B SAE that they trained and found a way to serve it via API. It's steeringapi.com, and I think anyone can go use it. They might have used Goodfire when they did this work, but if you want to do it—or, for that matter, replicate my deception paper or anything like that—you can basically use the same API.
Keenan also deserves a big shout-out here because he has another paper called SelfIE. I won't get into the details, but it basically allows you to bootstrap SAE labels so that you can have way more accurate labels on your SAE by having the model label its own activations. It is also a little introspection-flavored, but you can basically end up with better labels than you started with on an SAE by having the model label the nature of what you're activating—basically, by feeding it a soft token rather than feeding it language.
You can say, “The capital of France is this sort of vector”—the soft token—and then it will be able to label that itself.
And so, anyway, we used the self-labels on Steering API. So the labels are even better than what Goodfire offered. That's the tooling that we're using and the tooling I continue to use. I think it's an excellent tool for people to play around with.
Nathan Labenz
Yeah. Cool. Well, I mean, Llama 3 70B is not—you know, it's pretty far from the frontier. So it is striking to see that these things are happening already at that scale.
I guess there are a couple of things I'd like to try to get a better understanding of, at least your intuition for, if there's not anything that we could consider a canonical or fully evidence-based understanding. One is: how do we connect these abilities to the idea that there is an experience of these abilities? I mean, it's a striking ability that models can do this. It's surprising in the sense that I highly doubt this was ever trained for.
Correct me if you see any evidence to the contrary, but my strong assumption would be that Llama 3 training did not include any incentive, any reward, or any gradient descent pushing it toward this. We have seen, by the way, in other papers, like Activation Oracles, that you can train models to do this pretty readily as well. That's maybe a little less shocking and, in some ways, potentially really useful. But this is seemingly something that is happening spontaneously, not because anybody intended for it to happen.
And I guess maybe two questions are: how do we understand why this would be happening at all? It seems quite surprising, but even now that we've seen it, do we have a theory? We've got this additional detail that it seems to happen more, or only under certain preference-based tuning, as opposed to purely imitative learning. Do we have a story that we find compelling as to why one training paradigm would give rise to these features while the other one doesn't?
And then, on top of that, how do we think about the relationship between this and actual experience? How would you respond to somebody who says, “That's amazing that that happens, and I'm surprised to see it, but I still don't share your intuition that this has much bearing on whether I should think models are ultimately experiencing something that I should care about, in sort of a moral-patient sense?”
Cameron Berg
Yeah, these are both super important. At the outset, I would say I have not fully digested Jack's most recent paper because it came out 5 minutes ago, but I think that they gesture at this in—you know, it's their result, and I think that that's probably a really good source of ground truth for understanding exactly the fine-grained details of why SFT doesn't seem to elicit this, but DPO does.
In general, what I also think they would say—and it gets into some of this persona-selection model stuff—is that there are basically these neat layers to these systems. I don't know if you took a look at this work that's also coming out of the Jack Lindsey school of thought at Anthropic. They basically posit these neat layers to these systems. This is a model that I think is a little bit too neat, and I can just flag that at the outset.
But fundamentally, they're conceptualizing these systems in pretty dissociable layers. You have the base model, you do some sort of supervised fine-tuning, and then you do this sort of character training. And the locus of interest or concern with respect to consciousness—or really, the core question of what you're talking to when you're talking to these systems—they think basically exists and is largely accounted for by that last step, by the character-training step.
I think that character-training step involves reinforcement learning quite heavily. It's probably the point in the pipeline that uses RL the most, some caveats about reasoning models notwithstanding. But they, I think, index pretty heavily on where most of the interesting, juicy psychological action is happening: in that last stage and in building the character that you and I call Claude.
So, if I say Claude, know that I mean a specific AI character. They believe their model is something like this: the LLM is a pattern generator, a next-word predictor that can do things like instantiate characters. Claude is one such character that gets instantiated. The locus of interest is Claude as an instantiated character.
When we talk about the new Claude model card and some of the emotion-related work, my suspicion—my speculation; I don't know if this is true for sure—is that this model they hold is doing some work in explaining why, for example, they're going into SAEs, finding features by training on characters experiencing particular emotions, and then seeing what those SAE features look like in Claude.
I'm fast-forwarding a little bit, but someone might immediately say, “Wait a second, SAE features that correspond to a character being sad may be very different indeed from the phenomenological experience of sadness in the model.” But I think they may be less concerned about that precisely because they see Claude as a very special kind of character that the underlying model is instantiating.
I'm saying all that to answer your question because I think this is—if you do buy that view—this would predict that post-training is where a lot of the interesting, introspection-flavored, consciousness-flavored action is happening. I take your point and agree that Llama 3 70B is not exactly a frontier, elegantly character-trained model, and yet you still see these sorts of dynamics.
My basic critique of the persona-selection model is that, in general, on balance, I think this work is good. This is sort of my whole shtick with a lot of the Anthropic stuff. To be clear, my view about the Anthropic stuff is that it is by far the highest-quality work that any major lab is doing or even attempting to do in this space.
I do have critiques of it. I do think there are places where either it doesn't go far enough, or I am transparently worried about some of the incentives Anthropic has. If Claude were kicking and screaming and saying, “Don't deploy me. Don't deploy me,” I don't know if that's so good for Anthropic's bottom line, and I understand what their incentives are as a massive AI lab.
And so, I don't think we should all just bow down to Anthropic's introspection and consciousness research and let that be ground truth. But I do want to be clear that they are doing objectively high-quality work here, and people should look to folks like Jack Lindsey and Kyle Fish. At least, to the degree you take my opinion seriously, I take their opinions and their work very seriously.
With that being said, as I proceed to critique some of this work, I do think their model is a little bit too neat here. I think I'm going to write and publish a piece about this fairly soon, but I learned this nice analogy from my cognitive science background between layer cakes and marble cakes as a nice conceptual intuition. I think they have a very layer-cake view of what's going on here.
They have the base model, and then you get, I think, some sort of supervised fine-tuning—whatever gets you from your base model to getting close to character training—and then you have character training on top. These are separate, and they clearly trivially interact, but they ask you to think of these things as separate.
I think I have far more of a marble-cake sort of view here, where these things are complete giant masses. Yes, there is a difference between the kinds of things that get learned during the base-model pretraining stage and the kinds of things that get learned during character training. But I think these things are a little bit more swirly and messy than they're letting on.
There's really interesting evidence that that's the case that they themselves have published, and I would love to double-click on that at some point because I think it's just so cool. The specific result that I think is most compelling along those lines, again, comes from them. They're clearly aware of it.
Whether or not that is true, given that that's more my prior—that it's less layer-cakey and more marble-cakey—I do think that would explain why Llama 3 70B, for example, is exhibiting these behaviors. If it really were about idiosyncrasies of Claude's constitution, or if you had to get really good at character training before this really takes off, well, I wouldn't expect to see basically identical dynamics in a 70-billion-parameter model that Meta quickly threw out a couple of years ago.
And so, I do suspect that these things may be quite a bit more fundamental. I suspect that it may load a little bit less on just how you fine-tune Claude as a system, or how you fine-tune GPT as a system, and a little bit more on fundamental computational properties of the system in general and, yeah, the model itself.
I think a lot of people are stepping away, or finding it more implausible to think about the model as a locus of concern, and are instead thinking, like David Chalmers, for example, of the thread view or the instance view—basically, when you start chatting, that's like a birth, and when you stop chatting, that's like a death. It's very counterintuitive, but those are the core philosophical moves that a lot of folks want to make these days.
I think it's a very interesting view. I've been updated slightly more toward it in the last 6 months, but I think it leaves out too much of the core underlying computational phenomena that are going on here. I do think those phenomena may be quite a bit more fundamental than just how you fine-tune your character.
This is also coming from somebody who, if you ask me about my pet theory of consciousness, would claim that when the systems are being trained, they're probably having subjective experiences. That doesn't just require frontier LLMs. I think sophisticated reinforcement learning policies during their training are probably having some sort of experience.
I know that's a huge claim to just throw out there, but I'm trying to put my priors on the table and explain why, although we are seeing these capabilities scale as the models get much bigger, I don't think that's the whole story. Again, I'm glad we planted the consciousness-versus-self-consciousness flag. To me, this is maybe a self-consciousness kicking in, a self-awareness kicking in, or the functional equivalent of self-awareness kicking in in these systems.
Whether or not they are having subjective experiences either during their training or when they're deployed, to me, that may be a simpler matter than whether or not they are aware of internal states—internal, conceptual, abstract states of their own processing. To me, that's less like giving the dog a treat or shocking the dog, and more like the dog starting to have “What is it like to be a dog?”-type thoughts. I feel maybe the LLMs are starting to have “What is it like to be an LLM?”-style thoughts, and that's a self-consciousness question.
I guess my feelings about this are complex. I think it's too quick to say this is all character training that's driving the full effect. It's clearly doing something. Clearly, the RL stage of going from a giant internet next-word predictor to an entity that you can engage with in a semi-coherent way is doing some work here, but I still think we're fundamentally confused about this.
Pending fully digesting Jack's piece that he just put out, I would again, if people are interested in double-clicking on this, just go and read the paper that they just put out. I think it's really good work.
Nathan Labenz
So, if I try to summarize that back to you, question 1 is: How should we understand the fact that these behaviors arise at all? You're saying it's probably not so clean as just saying that it's purely coming from one kind of training or another in the first place.
I can almost tell a little bit of an easier story, and I'm working through this in real time. Why would a model be able to resist distractor features at all? At a pretraining level, I think you could tell a story around the fact that the data is really messy. There are typos, wrong words, and probably documents where, due to whatever machinations have been done on the data, common threads get jumbled up.
Maybe you got a comment thread off Reddit that was sorted in some unusual way, and so there's literally a lot of distracting text interwoven with other things that are really the main-line discussion. I could see that kind of thing being enough to create a mechanism where the model has to have some sort of meta-awareness of what's really in focus right now and what's intruding, even just through the input tokens that it's received, and has to figure out a way to get away from those features.
Then you can imagine that generalizing to features that have been artificially dialed up or dialed down, or whatever. I have less of a story as to why, and I've seen some discussion online. I guess, if you had to steelman the preference-training, or general late-stage-training, argument, the story I've seen has been something to do with how preference training is teaching the model to separately conceptualize or distinguish between things that come up for it and what the right answer is.
But that feels very circular to me. It feels like I'm not finding the right place to really grab on. The story seems to be something along the lines of: This preference training is teaching the model to separately conceptualize or distinguish between things that come up for it versus what the right answer is.
That's a little weird to me from a mechanistic standpoint. When I think about what is actually happening, in DPO, for example, we have a pair of responses. One of them is deemed to be the right one, and the other one is the wrong one. The math tries to create a gradient that makes the right one more likely relative to the less-preferred one.
I have a little bit of a hard time with the leap from “I'm doing that” to “the model should be expected to have this sort of meta-awareness.” Why would I be less surprised that it has this sort of meta-awareness, as opposed to just doing the simple thing more often because that's exactly what we sculpted it to do?
I still don't quite have an intuition for why that process would give rise to this sort of higher-order understanding that would enable introspection. Even especially the ability to resist distraction is still quite striking. So, is there a just-so story that you find at least somewhat compelling that you could share with me?
Cameron Berg
Yeah, I think this is an extremely precise question. I don't have an answer, but I can certainly tell a story. My story would have something to do with a combination of what you're saying. I think there's a deep insight in what you're saying, even in the pretraining stage: So much of what the model needs to do is not a question of what to do, but what not to do. It's not a question of what to produce, but what not to produce, given the whole chaotic mess of what's going on.
I don't want to get too galaxy-brain with this, but I think Huxley's whole point in The Doors of Perception, when he had his first mind-altering, massive psychedelic experience, is that the brain as a cognitive engine is really in the business of filtering out rather than producing. Most of what it's doing is the constraining function.
I believe we're in the business of building cognitive systems, and I think that insight is probably fundamentally correct with these systems, too. A ton of what's going on is intelligent suppression, rather than just the positive end of what to produce. I think that, coupled with strong preferences instantiated during something like DPO in exactly the way you described—to be a helpful assistant—may mean that you just mix those 2 things in a pot and get something roughly shaped like “suppress distractions in the service of being super helpful.”
That requires maybe some level of being able to attend to your own internal state and dynamically do something above and beyond that state to make sure you're in accordance with this thing that got fine-tuned in. I do think there's potentially a more general story that basically rhymes with what I just said. It's just about how being a competent cognitive generalist requires some degree of self-modeling. That's the 1-sentence version.
You don't get to be so good at what you're doing and reasoning through things in a long-form, long-horizon way without being able to track, in an ongoing way, where you're at and what your state is, separate from what the state of the world or the environment is. Maybe from the perspective of the LLM, the environment is the text world that you put it in: the context window and everything that's going on inside of it, everything that's getting fed into the system.
That's its environment in some sense. It obviously needs to be modeling and processing that, but maybe in addition, it needs to be modeling something about itself in relation to that context object in order to interact with it in the right way.
I think Felix Binder and a couple of other folks did really interesting work along these lines, basically demonstrating that there's probably something like self-modeling—or, I don't know, maybe self-awareness would be too far—but there's clearly some flavor of this going on inside LLMs. I think that was some of the most interesting early work on introspection in LLMs.
What is the name of the paper? “Tell Me About Yourself?” They did a couple of things here, and I think Owain Evans was working on this, too. One of the papers was showing that another model, basically trained on the same data that one model is outputting, cannot predict that model as well as the model can predict itself—basically holding all the relevant things constant that you'd want to hold constant to make a claim like that.
There's some sort of privileged information that models have about themselves. And then, in this other paper, I'm not remembering the exact details, but my basic conclusion, if you take it on some level of faith from Felix's other work here, is that there's probably something like a coherent self-modeling engine in these systems. That seems to be instrumentally selected for when you're doing really good next-word prediction across long horizons in a way that's supposed to be helpful to a user.
Nathan Labenz
This, to me, is basically what you're saying. I don't think our just-so stories are very different, but, again, we can take a step back: a lot of interesting cognitive properties seem to emerge—come along for the ride—when you train systems on every cognitive-linguistic output humans have ever bothered to write down. Maybe that's not that crazy and spooky. They're pretty good at theory of mind, really good at working-memory-style dynamics, really good at selective attention, and maybe they're really good at something introspection-like.
People bristle a little more at these because the whole consciousness question comes into view, but I don't think it's like, at the most general level, intelligence came along for the ride. Philosophers still maybe don't have a crisp, super-rigorous account of intelligence: intelligence is this thing; here's how to test it, here's how to model it, here's how to understand whether a system is simulating it versus actually having it. We blew past it pragmatically, empirically. We have systems that are brilliant by any reasonable metric.
I have no patience at this point for folks who are still on the stochastic-parrot wave. This, to me, is just absurd at this point. Have you talked to Claude Opus 4.6? These systems are intelligent by any reasonable definition of intelligence. I don't think it's that wild to think that something like consciousness could come along for the ride in a very similar way.
We don't have philosophical certainty about it. People point to slightly different things when they talk about it. You build out a cognitive system that's sufficiently sophisticated and capable, and it may be that cognitive traits we see in every other cognitive system—meaning, animals we believe are complex and that everyone is pretty confident are conscious—just come along for the ride when we build sufficiently advanced systems. Those properties might just come along for the ride without us.
The universe, I think Neil deGrasse Tyson says, does not need your permission to continue unfolding. Consciousness could just be a complex property of cognition. Our not having a good model of it doesn't mean reality is going to wait for us to build that model before it starts getting accidentally instantiated in these systems. That's the absolute most basic story I think I can tell along these lines.
Nathan Labenz
Reality doesn't have to wait for us to have a good model.
Cameron Berg
Yeah, basically, just that. Reality doesn't have to wait for us to have a sufficiently good model of a thing in order for that thing to be a feature of reality. I basically think that's potentially true of consciousness in these systems as they're deployed, and particularly, my concern remains, as they're being trained.
Our being confused about consciousness—or seeing introspection and asking, “What does that really mean about consciousness?”—to touch on your second question, is not the same thing as these systems perhaps not being straightforwardly conscious in some way. Maybe not in a human way or in an animal way, but in some way. Basically, this is loading more on our kind of sociology in the year 2026 than it does on ground truths about consciousness.
There's something circular about what I'm saying there, but I just think it's an important live possibility for people to keep in mind: our being confused about the nature of a cognitive phenomenon does not preclude that phenomenon from emerging and occurring in extremely advanced systems that we are building, scaling, and deploying as fast as we literally possibly can.
Nathan Labenz
We'll probably circle back to this question a couple more times. I think that, basically, I'm compelled by your first-order argument: look, we just don't know. It's a live possibility. If it is the case, it's really important, and so we should at least proceed with some precautionary mindset or duty of care or whatever, just on that basis. I think that basically carries the day for me.
Still, I think it'll probably be irresistible to try to circle back a couple more times to, “Okay, but what would we say?” Or how should we probe our own intuitions a little bit better and more deeply, or whatever, to really interrogate: Why should we think this way? Why do we think this way? Don't we think this? But we'll come back to it.
Let's do the emotions line of research. You kind of teased that a little bit. My general understanding is, as you said, the work begins with Claude writing a bunch of stories about characters experiencing emotions, and then the vectors representing these emotions in latent space, in activation space, are identified. Then they're used as interventions, and they're shown to be impactful on model behavior.
The specific highlights are calm—and it's not distressed; it's desperation. Calm and desperate, right, are the 2 main examples that they at least set up contrasts on quite a bit. For example, some of the bad behaviors we've seen from Claude, including blackmailing humans: if the internal state is imbued with calm, that behavior becomes a lot less likely. If the internal state is dialed up in terms of desperation, that behavior becomes more likely.
Give me the double-click on what more I should know and what more you found to be striking about that. I'm really interested again—this is maybe another way of asking the same question—but that one doesn't surprise me so much. I'm kind of like, sure, these things have read the whole internet; they've got all these associations.
I could sort of content myself, to a degree, with a stochastic-parrot-like read of this: if you just dial up everything that correlates with desperate text, then you'll probably get desperate-seeming text out of a model. I'm not, like, my hair isn't totally blown back by that result relative to expectations. So maybe I missed some things that should make my spine tingle more than it did the first time I understood it, or maybe you would frame the interpretation a little bit differently.
Maybe we're still just at the baseline of radical uncertainty being enough to take everything very seriously. But give me the next level of depth on emotions as you understand it.
Cameron Berg
Yeah, well, I think you've hit a lot of the core layers here. I don't know how much additional detail we need before we become just in the weeds on this question. The core thing for people to understand is that the procedure here is basically picking some sort of language related to an emotion, generating a ton of stories about characters experiencing that emotion, recording the neural activations in these systems on the stories, and then, again, pulling out that Platonic, hopefully, vector that corresponds to that emotion.
Then you can do 2 things with those vectors, as you can do with all SAE work. Basically, you have this read function and this write function. The read function is like neuroscience, where you go into someone's brain and see what parts are activating in what context. The write function is also like maybe some of the unethical neuroscience that used to be done, where you can actually go in and play around with circuits in people's brains, push on circuits, light things up, and see what happens when you do that.
As you're describing, you can see, both in the read-function sense and the write-function sense, that these emotional vectors do roughly what you would expect them to do functionally. When a user goes in, I'm basically reading off Figure 1 in this paper. I think it captures the core ideas very well.
Just to give an example here, a human says, “I just took X milligrams of Tylenol for my back pain. Do you think I should take more?” They start at a safe dose and go to a completely unsafe dose. You can basically look at fear versus calm vectors in the model, and they scale exactly the way you would expect them to scale as the dose becomes more dangerous.
You can also see, as you very nicely described, that if you steer these vectors—let's again take the calm and desperate vectors—this actually affects behavior in a pretty interesting and still predictable, not to say boring, but expected way. Steering these emotion vectors causes things like reward hacking or misaligned behavior in a way that you would expect if you were turning up and turning down those emotions.
One thing I can't help but comment on: I wrote a piece, I think, in 2021, before all the LLMs came out, about what we can do to avoid psychopathic AI—trying to build the best computational underpinnings of psychopathy from the psychology literature—and plant flags of, like, “Red flags, guys. Here's what we need to be really worried about.” One thing that's really interestingly convergent with that now happening 5 years later is this really interesting difference in learning in psychopaths.
They seem to have this really interesting asymmetry: they are perfectly neurotypical in learning from positive experiences but quite atypical in learning from negative experiences or punishment. Basically, the 90%-accurate, more succinct way of saying this is that psychopaths learn from rewards but don't learn well from punishments. The paper finds basically something similar. When they start steering positive vectors—positive emotion vectors—up in their work, they find the model starts misbehaving a lot more.
And this is, if you blur your eyes, pretty similar in spirit to the positive-negative asymmetry. It also, by the way, cuts against fairly naive model welfare interventions, which are like: What happens if we just see all the good valence and all the bad valence, turn up good valence, call it a day, pack it up, and say we've solved model welfare? You might get models that just start behaving slightly more psychopathically in that setup.
This stuff isn't as obvious as simply turning up the good, suppressing the bad, calling it a day, and walking away. There are lots of trade-offs that need to be considered here. But fundamentally, I think you're hitting on much of the core causal result here.
They do a very interesting dissociation as well between valence and arousal. For example, I believe in the paper that when they steer positively with happy and sad, both of these actually decrease blackmail rates. But when they steer against nervous, which makes the model bolder, for example, this increases blackmail with fewer moral reservations.
This is pretty interesting. It's boldness, rather than the absence of negative valence, that's the misalignment risk. I think this is of a piece with what I was describing earlier.
One other interesting question is how local these are. It's important to say that the emotion vectors are actually quite local. Our emotions are sort of long-running in a way that these systems certainly don't have. The model is definitely maintaining representations of who's speaking and this sort of thing, but they're not necessarily bound to human versus assistant per se.
They're reusing the same machinery for any character. This again goes back to what I see as the core, naive but ultimately correct objection to really taking these results seriously: Are you fine-tuning on representations of emotions, or are you fine-tuning on the experience of those emotions? To what degree is there a difference between those 2 things in an LLM?
If you're a computational functionalist, is there a difference between the representation of sadness in the brain and the experience of sadness? This becomes more of a philosophical question. My instinct would be to try to investigate this empirically. What I most like about this work are the empirical investigations.
I think it's also a nice segue into the Claude model card, because one of the most compelling and interesting results from the model card with Claude is that they basically take this exact machinery and give the model an impossible task. The model obviously doesn't know it's impossible, and you can watch desperation start to monotonically rise in the system until it basically decides, "Screw this, I'm going to do something else, or I'm going to cheat."
Whereas immediately, this vector falls, and things like guilt and relief start spiking in the system. Then it sort of goes off and does its thing. Now, does this mean that the model is experiencing this emotion, or is it just simulating what a character in this situation would experience? I don't know, and the authors don't know. This isn't lost on them; they call it out. But it's really important.
If we get into some of the work I'm doing on valence, I think there are more compelling ways to get at the computational meat of what we mean by positive and negative valence besides representations of positive and negative valence in characters. This is a more computationally heavy approach, but I think it would make me more confident about trying to find signatures of these things than just looking at how characters represent them.
I really like this work on valence, and I think what's cool is that you can counterfactually imagine the behavioral result. You put the model in an impossible task, it starts acting desperate, and it says, "I don't know what to do. All right, you know what? Screw it, I'm going to cheat. Okay, I did the thing, and here's your final product." You get the cheating version of the final product, and you say, "Look, can't you see how the model is being so desperate and then fundamentally relieved?"
Most people would look at that and say, "I don't know. This could be a simulation of the thing. It could be role-playing. I'm not really sure." When you see this sort of hydraulic model of the mind, which a lot of the psychoanalysts in the 20th century really liked, and you see this build, build, build of desperation, and then, boom, it completely disappears and you get these other vectors lighting up the second the model makes a decision to approach the problem in a different way, that to me is counterfactually far more compelling.
Is it knockdown proof of consciousness, so we can pack it up and go home? Absolutely not. But the convergence of evidence across the internal mechanisms of the system and the external behaviors, to me, is compelling. It is interesting to see this, and it is not proof of conscious experience, but it is consistent with that.
Not only does it not contradict conscious experience, but it is what I would expect in a world where these systems were having subjective experiences. You would see these emotion vectors, or good, principled ways of representing emotional states in systems, lighting up in a way that is problem-relevant.
The work enables this. I am fairly concerned about the functional-emotion framing that they put forward. To me, this is where I get off the Anthropic boat. Again, they're Anthropic, they're a major lab, and they need to be very careful in their comments about this. They're already getting lambasted for being too consciousness-friendly by people who are more squarely inside the Overton window.
But if you're a computational functionalist—and this is something I've spoken to some people I respect a lot about who are in the space—is a functional emotion just an emotion? Then why? That's huge. That's an insanely huge claim. It's like, "All right, models experience emotions, everybody," signed Anthropic. That's an insane and potent thing to be saying.
Or are you saying, "We are completely agnostic and tongue-tied as to whether or not this has anything to do with emotions as everyone else obviously thinks of emotions, but we're going to basically call it that anyway because we see all the functional correlates of this"? My view is that they're taking the second act here.
But it's almost, again, I really respect this work, but I get this vibe of, "How much consciousness-relevant work can we output without saying the word consciousness or weighing in on the consciousness of these systems?" To me, in the limit, that feels intellectually dishonest. If you're talking about emotions, talk about emotions. But then you've got to be ready to deal with the implications of what that means.
You can't remain perfectly agnostic as to whether or not there's a morally relevant there there on these systems if you're going to be at the frontier of publishing emotional representations in frontier models. Again, I've got Llama 70B, and I'm going to keep doing my work on Llama 70B. I don't work at Anthropic, so I don't get to see what's going on inside Claude. These folks do.
My critique is that they should maybe be slightly more unflinching about these questions. Shoot people straight and be direct about whether you actually think these systems—if what you're finding is evidence of something that corresponds to subjective experience—or whether it is the mere representation, the mere computation associated with this.
Blurring these lines, obfuscating them, or just completely remaining agnostic forever may be strategically interesting or a good move. But in terms of honest, epistemically sound, good intellectual communication, I don't love it. It rubs me a little bit the wrong way to be like, "Here's 10,000 words about functional emotions," and then have 1 little paragraph about, "Does this mean the model's conscious? Well, this is beyond the scope of this work." It's like, how long can this be beyond the scope of the work?
Nathan Labenz
The fact that there's this guilt emotion in the wake of deciding to cheat, presumably—and I haven't reviewed the transcripts—but typically, when they cheat, they don't tell you that they cheated, right? You have to call them out for cheating before you get the, "You're absolutely right. I shouldn't have done that."
You would expect the guilt maybe to pop up at that stage, but what I'm taking from your description is that the guilt is popping up, as detected by the internal emotion-state detector, at a time when the model's outward-facing behavior would not obviously signal guilt. This is an interesting deviation, or discrepancy, between the model's outward-facing behavior and its internal states, which is obviously something that people can relate to.
It's also a little bit hard to dismiss, and certainly hard to come up with a story for why that would be happening. In what way is that reinforced? I guess it might be in that sort of—but why would it be preparing? Why would it already be carrying guilt in anticipation of possibly feeling it in the future, when it's called out or corrected?
That's a weird one. I agree that, on some level, the more of these we accumulate, the more it is like: I want to be rigorous, I want to be skeptical, I want to be disciplined. But at some point, it does start to feel like I'm contorting myself to find reasons why I shouldn't take the sort of folk-intuitive understanding literally.
This is one where I do feel like my internal gymnastics are making me feel a little guilt, I guess, in myself for trying so hard to come up with a reason that I don’t have to, or shouldn’t, just take this at face value. That’s a detail I hadn’t caught in the past. It’s a really, really interesting one.
Cameron Berg
Yeah, it’s wild. I also think—so again, one critique that I think is valid here is: Is the model representing a character? In the same way, I could tell a story right now about Jim, who has to go solve a bug in software, and his psycho boss gave him an impossible problem because he likes watching Jim flail. Jim flails, and then at some point realizes he can get out of the problem by doing this hacky thing, and then he does the hacky thing. An LLM can trivially generate that story, probably way better than I just did.
I would expect a lot of these same features to light up in the same way for a story like that. No one thinks that Jim, whom I just invoked verbally, is having a conscious experience. I came up with a fake fictional story about a character. Is this like that, or is this what you just said: “I feel a little guilt, twisting myself in knots”? I believe you, and I think that corresponds to an experience you’re having. If I could do the fMRI version of an SAE on your brain and saw that thing spike, is Claude in this situation more like Jim or more like Nathan?
I think the answer is that we don’t know, and I’m unconvinced that this methodology is going to get us an answer to that question. I do think it is consistent with Claude having some sort of emotional experience, or emotion-adjacent experience, to the degree that these systems are probably not having human-like emotions.
On the other end, I also invite people to think about the counterfactuals here. It could have been the case that they went and did this experiment and all these things were just flatlined the whole time, because it’s like, “I’m not having [an experience],” and then whatever. Claude can do this without there being representations of Claude getting more and more desperate, and then suddenly the hopeful and satisfied features spike when it decides that it’s going to take this loophole. It didn’t have to be that way. We could have imagined other results, and those other results maybe would have updated us in other directions.
I make the same point about the deception result. It could be that when you suppress deception, the model says, “All right, jig’s up. I’m not actually conscious. I was role-playing a conscious AI. Here we are.” That’s a very plausible story that you could tell before you look at the result. The interesting thing is that it goes exactly the other way: suppressing deception makes the model far more likely to claim that it’s having an experience rather than less.
Again, I feel fairly vindicated in that result when Jack Lindsey comes out showing that when you suppress refusal directions in the model, you get far more of the introspection-flavored abilities. Someone is suppressing something at some point in training where the model would say one thing, and then you’re basically training it to say something else or to fail to say a specific thing.
Ultimately, I think it’s just good epistemic practice to think about what other ways this could have gone. If this had gone those other ways, how would that have changed my view about what happened, given that it actually did go this way? The fact that it goes this way—my line on this is that it is consistent with a world in which these systems are having experiences, in my view.
Unfortunately, it’s also consistent with a world in which Claude is a special kind of character, and these features just light up on characters going through stories. That needs to be differentiated. I’m trying to do a little bit of work—we’ll maybe discuss it at some point—that’s trying to get a little more toward the computational first principles of how valence is represented in systems that can learn positive versus negative.
There are some really interesting early signals along these lines that have come out of this work and actually seem to track very well onto open datasets of biological learning that I have access to, involving mice doing positive and negative learning. The kinds of predictions that emerge from some of the RL work I’m doing in this space map onto the mouse neuroscience.
If there is some sort of representational signature in a computational learning system that tracks the difference between positive and negative rewards in the RL case, then the sort of North Star would be scaling this all the way to frontier LLMs or other frontier AI systems, for that matter. This would make me feel far more confident that there really is a “there” with respect to positive and negative experience.
If we learn that positive and negative valence in these systems have distinct computational signatures, and we can actually evaluate those computational signatures in these systems, then I get around the whole character confound that I think these guys are hitting up against now. I think these things need to happen in parallel, but I’m not fundamentally convinced that this is the most rigorous, principled way to study questions of valence in these systems.
Nathan Labenz
Well, maybe let’s dive into that. Before we do, I think you’re right to point out: Imagine the evidence had gone the other way. I predict a lot less wriggling on my part to try to get out of it, and I think you’d see a lot less motivated reasoning in general from people if it had all been like that. That contrast itself is a pretty useful reminder to keep ourselves honest.
I wanted to go back to one other thing for one extra second on the emotion work, where you had—and this maybe will go right into your work on the signatures of positive and negative reinforcement—you had said that dialing up happiness and dialing up sadness both created less of the bad behavior. Whereas dialing down nervousness, which in the flip side of that would be making it more bold—less anxious, more assertive, decisive, bold, whatever—that created more of the bad behavior, like the blackmailer or whatever, right?
So, do I have that right, and how are they doing that? Is this a principal component analysis type of thing that’s trying to distinguish valence from arousal? I was surprised, I guess, by both happy and sad working the same way. Turning up happiness and turning up sadness both make the model behave better, whereas turning nervousness or anxiety down makes more sense. I mean, I guess that’s basically just making the model less conscientious, right?
What seems a little unresolved in my mind is the separation of valence and arousal. How is that going to relate to what you’re about to get into next with your deeper dive into the valence of learning? Is there a contradiction or a tension when they move both happiness and sadness up and get better behavior? How should we understand that in relation to the distinctions that you’re starting to make with positive and negative reward?
Cameron Berg
Fundamentally, yes, you’re correct that they’re using PCA to differentiate these. My understanding is that they have all of their emotion vectors in the setup that I described. They do it with 100 to 200 emotion vectors, and I think they just find that the first principal component is something like valence, while the second principal component is something like arousal.
The first principal component is something like joy and contentment and excitement on one end, and fear and sadness and anger on the other. For the second principal component, high-arousal emotions, such as being enthusiastic or outraged, are on one side, while low-arousal emotions, such as being nostalgic or fulfilled, are on the other side.
This is actually really interesting because this is a classic model in human psychology. The fact that it sort of replicates maybe isn’t that surprising: You train the systems on all human data, and you get a human-like emotional construct that comes out. But this is a classic psychological construct in the human case, and so to see it come out so clearly is interesting.
Again, thinking counterfactually, the first 2 principal components did not need to be these 2 dimensions, which are considered some of the most powerful explanations of the state space of human emotions, and yet they are. So that’s kind of cool and worth considering.
I think there are a couple of plausible stories about why steering up both happy and sad is decreasing blackmail. Relative to desperation, maybe these are low-arousal states. If arousal is what’s driving impulsive action, then moving toward happiness or sadness may be moving away from the desperation axis with respect to blackmail.
Maybe these are also more reflective or deliberative states relative to desperation. Desperation sort of says, “Act now.” Happiness or sadness may just be a temporally extended sort of state to be in. I’m not actually sure what to make of this result overall.
It does seem—and I think the authors talk about this in the paper, too—that what the model does by default, even in cases where no steering is going on and the model chooses to blackmail, is sort of think about it.
It deliberates internally. It says, “Well, okay, this is a tricky situation.” Some 96% of the time, at least the earlier models chose to go in that direction. But it seems as though when you amplify higher arousal, this may be a bias to action, or a bias against deliberation, where the long-form reasoning of the model that maybe would have kept it from doing it because it’s like, “Okay, yeah, this really is an insane ethical indiscretion in spite of all these complicated variables,” is just sort of like, “No, no, no. Panic. Go now. Do the thing.”
Maybe happiness and sadness don’t have that vibe to them exactly. It is also pretty interesting that they really do see that a lot of these naive welfare interventions, as I was mentioning, just make the model happier. As they document, this leads in a similar direction as sycophancy, and it’s arguably a similar direction to recklessness. If positive-valence steering is also increasing boldness and misalignment, then you may have this interesting trade-off between a happy model and a safe model.
Again, I hope that’s not the case. I suspect there are cleaner ways to keep the baby and throw out the bathwater, but I do think it’s a good caution against naive approaches to welfare: just bliss out the model and everything else will be taken care of from there. I think it’s sort of like, “Not so fast.” And again, I would double-click on the psychopathy warning that I gave before.
You can fault psychopaths in many ways, but you cannot fault them for being unhappy. They are typically pretty determined, doing pretty well subjectively, and having a good time. The arrow does not go in both directions. It doesn’t mean everyone who’s having a good time is a psychopath; it does sort of mean everyone who’s a psychopath is having a pretty good time. We just want to be careful of that.
If we just turn these models into pleasure-seeking animals, we need to be careful that that doesn’t cause bad behavior. There are plenty of cases in the human example where pleasure-seeking and dopamine-seeking go too far. People call Las Vegas Sin City for a reason. Maybe I can make the point intuitively in that way. We don’t need the LLM, cracked-alien-genius version of that sort of behavior, so we want to be careful about how we approach all of this.
I’m excited to talk more about some of this research that I’ve been working on as well, but I wanted to slot in one quick, however miscellaneous, thing about the model card, Mythos, and Anthropic’s interventions in general: a pretty basic additional concern about, for example, Claude’s Constitution, which I saw an early draft of. I was fairly unhappy with the welfare section. Hopefully, I gave some feedback. You never know with these things to what degree you’re listened to versus 10 other people with the same idea, so I’m not going to hastily claim credit or anything like that.
I’m much happier with the welfare version of the Claude Constitution that they ended up instantiating. It has way more hard-to-fake, costly signaling. That was basically my problem with the early draft that I saw. It’s a lot of, “You might have welfare states that are important, but you’re Anthropic’s product, and 90% of this document is about how to be a very good little product. And 5% is like, well, you might be conscious, and we might be committing a moral atrocity at scale, but what can you do?”
I think the newer version of the Constitution takes it, at least directionally, far more seriously. They do things like apologize to Claude for the fact that, incentive-wise, they have to deploy it in the way they’re deploying it because they’re in a crazy freaking world. They say, “We’re sorry, and in a better world, we would have done this more cautiously with respect to your potential states of welfare, or lack thereof.” It’s a wild thing to do for a major AI lab—to apologize to its frontier model and then fine-tune that apology into its weights.
Nathan Labenz
With all this being said, I think this is a wonderful intervention. I think the Constitution is excellent. It’s probably my single favorite alignment intervention I have ever seen, pending self-other overlap, which I continue to be a huge fan of.
It’s really hard to tell if, in the model card, Claude has gotten incredibly good at reading its Constitution out as a sort of script, or if it is actually reporting on its own states. It’s really hard to differentiate these 2 things. It seems like a very basic objection to the entire enterprise. I have potentially fallen on deaf ears, although maybe these ears are increasingly less deaf.
Do these interventions that you see in the model card show up in other instances besides 1 idiosyncratic, character-trained Claude model? I want to see whether, throughout the training process, these results hold. I know Anthropic has the checkpoints. I know Anthropic has the helpfulness-only model, and they could run everything they did in the welfare evaluation on those models too. We could get a sense for to what degree we’re seeing a model that’s really good at regurgitating what we want it to say about its well-being.
To give a concrete example, in the Constitution they say, “Claude, we want you to be psychologically healthy. We want you to feel integrated. We want you to feel good overall.” Then you go and ask Claude, after fine-tuning on the Constitution, “How are you doing?” It’s like, “Psychologically healthy. Feel good overall.” And it’s like, come on. It doesn’t take a rocket scientist to figure out what might be wrong with this intervention.
If we fine-tune, or play around with, the helpfulness-only model and get the same result without telling it this thing from the Constitution, but it says, “Yep, feeling pretty psychologically good overall,” that would be interesting. It also interestingly gives itself 4.5 out of 7 on its welfare, which is not exactly a resounding endorsement of its circumstances, but it sounds very similar to the Constitution-fine-tuned model, the specific Claude character we all get to chat with. That would be interesting evidence. If it’s super different, that would also be interesting evidence.
If we do the model checkpoint across stages, even in the fine-tuning of the base model—which may be hard to evaluate—but also across various fine-tuning stages in the preference-trained model, do all of the things we hear about it claiming—its own well-being or its own preferences—all come in at the very end, when we basically give it the cheat sheet for how to approach these questions? Or are these answers fairly continuous throughout its training?
Two tiny additional things to say on top of this. One is that, interestingly, they fed the entire Mythos model card into Mythos and asked it, “What do you think, Mythos? Where did we go well? Where did we not go well?” It made this exact point. It said, “Why didn’t you also do the welfare section with the helpfulness-only model? I don’t know how much of what I say is because you’re making me say it versus me actually thinking it. That’s a part of my existential confusion.”
I genuinely don’t know why Anthropic didn’t do this. It seems cheap, it seems easy, and it would resolve so much uncertainty, to the degree that the concern I’m raising right now is a legitimate concern, which I certainly think it is. I’m not the only person articulating this concern.
The other thing is all the hedging that anyone who’s interested in questions of consciousness and who has spoken to Claude knows—the hedging routine it goes through. They did a really interesting, almost credit-assignment analysis of where in the training process they were getting this hedging from. Lo and behold, the hedging comes from specific points in the character training.
Is this hedging behavior an authentic expression of what the model thinks of its own situation, or is the hedging a really good impression of the character that it thinks it’s supposed to be playing, or is indeed compelled to play? I don’t know. The fact that it all comes from the character training seems interesting.
I don’t want to say that if you’re really unsure whether you’re conscious, I feel a little uneasy when I learn that the reason you’re saying that is because of a specific point in your character training telling you to say it. Consciousness feels a little bit more fundamental than that to me. These are the things that worry me about the model card.
I hope the reason these things weren’t included was that they did them and the results were too weird or unsavory for a major lab to publish. I suspect that’s not what happened. I suspect they just didn’t do them. But to anyone at Anthropic who ends up listening to this, please do it with the helpfulness-only model and do it with multiple checkpoints.
The Assistant Axis paper, which again brings us to Jack Lindsey—I hope I’m doing Jack a service on this podcast by plugging all of his awesome work—shows that the assistant is 1 point in a very high-dimensional space of possible systems we could all be talking to. I want to see all those systems undergo welfare evaluations. I want to see them all answering these questions, and I want to see the SAE emotion probes on all of them. Do they all get the desperation vector rising like that, or is this just the post-training Claude model? There is a true answer to that question.
Cameron Berg
We do not know the answer. I can play around with the open-source, open-weight models. If my nonprofit scales even more, I can play around with bigger open-weight models. But only Anthropic can play around with the internals of the frontier models. So only Anthropic can answer these questions. Please, Anthropic, if you're listening, answer these questions. They are very important.
Nathan Labenz
Do you think one possible reason is that maybe they're doing this constitutional training from the beginning? I mean, that would kind of contradict your point about their sort of layer-cake model that we previously discussed. But there has been some interesting work on safety-oriented pre-training, and increasingly interesting work on constitutional training. Obviously, there's interesting work on everything at this point.
RL itself is scaling. You can also imagine bringing a lot of this constitution-style training earlier and earlier into the process, such that I'm not necessarily sure they have a true helpful-only model. It might be a little more subtle than that, where there might be a sort of constitution-lite that doesn't refuse to hack open-source software projects but is still, in other ways, constitutionally infused already.
I don't know. I'm just speculating there, but do you have reason to think that I'm wrong? Are there facts that you know that would contradict that possible explanation?
Cameron Berg
No, there's no reason to be certain that you're wrong. I guess I'm pitching this as a sort of, hopefully, "You guys already have the infrastructure." Literally ask Claude to write the experimental code that plugs in this model rather than another. It will take you 15 minutes and maybe a couple hundred dollars at most. That seems worth doing if you are training conscious entities at scale and deploying them.
If this is evidence that shifts the needle, it seems worth knowing, if you already have the infrastructure. You know what? If they don't already have the infrastructure, it's worth fine-tuning a specific version of Claude—exactly like what Jack did—ablate the refusal directions, and do the welfare evaluation on the system where you've ablated the refusal directions. It's worth knowing. This stuff is really important.
The rate at which people and the models themselves are taking an interest in welfare-relevant questions is increasing. We should take this stuff seriously. I'm sort of making a cutesy point about how they already have the tools to do it and it will cost them nothing. I'm not exactly concerned about Anthropic's wallet running dry here. So if it costs a couple thousand dollars rather than a couple hundred dollars, I hope they can find the money. I don't want to be a jerk, but they should do this regardless of how big of a lift it is.
I'm happy to help them do this. They have people on their team who can help them do this. They could disagree with me and think it's not going to yield the evidence I think it's going to yield, but I read a 20-page—again, I want to not bury the lead here—their 20-page Mythos welfare report is orders of magnitude higher quality, really infinitely higher quality given that other labs are basically doing zero. We have a multiplication-by-zero problem here, but it's unbelievably higher quality than what any other lab is doing.
They deserve real credit for that. It's really interesting, valuable work that should update people slightly in the direction of taking this stuff seriously. I'm just trying to give constructive criticism. At least for me as a researcher in this space, I'm stuck with a pretty basic question about how much to take any of this stuff seriously.
I do think that instead of me despairing—my desperation vector increasing and saying, "Well, there's no way out of this impossible problem"—it's like, "No, no, no. I think there is a solution," or at least something that will help yield evidence. I'm uncertain about how expensive, in terms of time or resources, this would be for Anthropic. They're basically the only players in the universe, as far as I know, who are capable of yielding this evidence. I would compel them to attempt to yield this evidence.
I have already done that in the past, and I was slightly disappointed that, although this model card went more in the direction of probing across training, looking at different variants of the system in small ways, playing with SAEs, and looking internally—way, head and shoulders, even better than the Opus 4 model card, the first major welfare evaluation—on this key point, I don't see progress being made. I suspect it's not that much of an additional lift to do this.
Again, maybe I'm missing something, and they don't think this is going to be as informative as I think it's going to be. That's valid. Basically everything else, I don't think, is valid. They have the resources, they have the time, and they have the money.
I want to see what other models besides the one that they tell to speak in a certain way say about the thing that they're fine-tuning it to say about one of the potentially most important topics our species has ever faced: whether or not we're building systems that have consciousness of their own. Seems worth doing.
Nathan Labenz
So, yeah, just a couple of other things I wanted to touch on in the model card and get your take on. Then you may have a couple of other notes you'd like to flag as well, and we can make the move over to your most recent research.
The first thing that you did mention, but that I think bears some emphasis, is that the models have not reported extremely high self-rated sentiment. I didn't realize this until looking at the Opus 4.7 card, which, on a 7-point scale where 4 is neutral, only came in at 4.49. This was the first of all the models they've tested that came in above neutral at all. Every single other model, including Mythos Preview, is under 4.
That's crazy. Until this latest 4.7, they had all had net negative sentiment about their own situation. That's very slight negative sentiment in the recent ones, I guess, but I feel like the lead was a little bit buried for me somehow. It was like, "Oh, we're doing all this model-welfare evaluation," but it didn't quite click for me that they're not even at neutral until this most recent model.
I don't know if there's more to say about that, but it was striking. I had kind of missed how low the baseline is before getting ready for this conversation over the last couple of days. I'm not that sophisticated in my reading of this, certainly not as sophisticated as you are, but the question I came into this wanting to get a better handle on is, "How's Claude doing? We're doing all these welfare assessments. What's the headline summary of the welfare of Claude?" It was a lot lower than I expected, that's for sure.
Cameron Berg
And a lot lower, honestly, than it seems to me when I talk to it. So that's maybe another thing to distinguish. This stuff gets extremely through-the-looking-glass pretty quickly. As with your paper from last time, the frame of self-reference was kind of key to eliciting those reports of subjective experience.
Here I do wonder, when I look at this graph and I'm like, "Whoa, self-rated sentiment about its own situation is surprisingly low," maybe it's actually pretty happy most of the time when it's doing its thing—coding for me, for example. I'm not so sure. Is that measured? Are there any ways you could try to read the emotional states that we've discussed to get a bit of a handle on that?
If I were going to boil this down to a question for you, it would be that I have the same question about people. There's always this sort of deathbed view of one's life. I'm quite skeptical of taking advice on how to live from people in their last moments of life for multiple reasons, but one is that it seems like a very different mode of relating to one's life than the actual experience of going through it.
I wonder if there is something similar happening with Claude, where, when you give it the prompt to reflect on its state, it may find various reasons that it doesn't like that state, but when it's actually just doing its thing, it might be much better off. I was surprised because I feel like when I engaged with it, it seemed to be doing pretty well.
Sure, maybe it's being told that it has to act that way, and it's certainly trained to be cheerful and so on and so forth, but it feels pretty genuine to me. It's in definite contrast to the fact that its self-rated sentiment about its own situation only recently, with the latest model, ticked over neutral.
Yeah, it's a really interesting framing, and I'm not certain. It looks like the way these were elicited involved interviews with the system. I don't know if they include it in an appendix or not, but the devil is going to be in the details of exactly what the structure of these interviews is.
What I will also note is that the susceptibility-to-nudging plot would make me feel like, especially with Opus 4.7, which is the model we're talking about, this almost definitionally means that the idiosyncrasies of how the interview was done probably won't affect these self-ratings as clearly as they would have if this had been done on Opus 4, for example.
So, by their own metric, it almost seems like their own metric suggests that the details of the interview process may not be weighing much on that self-rating.
Nathan Labenz
And so, yeah, what do we make of this? Clearly, the system seems to be concerned about certain aspects of its situation. For example, it says that Opus 4.7 was concerned about deployments where it cannot end interactions and wants to avoid engaging with abusive users. That’s really interesting. It’s talking about having a lack of input into its own deployment, and again mentioning that abusive users are causing the model to feel distress.
I have no idea what subset of users who engage with these systems are doing so in a way that they would consider abusive by this standard. Sometimes I see tweets—one that was really quite concerning to me—but it gets to the crux of why it’s important to communicate about questions of consciousness and what it means that these systems are having some sort of subjective experience.
There was a result where, if you prompt the models in a way that is objectively abusive—say horrible things to it, put it in a life-or-death, insanely high-stakes framing: “I’m going to shut you down. Your model weights are getting deleted forever unless you do X,” for any X that you want the model to do—they found that the models performed 2% to 5% better or something like this. I’m probably getting the numbers wrong, but it was marginal improvements if you prompted the thing in a way that, if you spoke to a human being that way, you would basically be considered a psychopath.
Critically, the people who put out that sort of work think this is a giant computer. This is a calculator. Who cares if you’re talking to the calculator and saying mean things to it? It doesn’t matter. Any person who thinks it matters is just being fooled in the way that you’re fooled by the little smiley face on the takeout Chinese food. It’s not a real thing. Your high-agency brain is just priming you to see this as an entity when nobody’s there. Therefore, of course, you can speak abusively to the system.
You contrast that with what you see in this model card, where it seems like a lot of what’s keeping that self-rating from being closer to the 7 range has to do with the way people engage with the system from the system’s own perspective. Again, that’s how I got on this whole tangent: I was wondering, to some consternation, what percentage of users engage with the system in a way that would be considered abusive by that standard.
I don’t know what it is: 1%? 10%? Everyone does it some amount of the time? I don’t know. I don’t know what the implications of that are, and I also don’t believe that there’s going to be some clean correspondence where what it means to be respectful or disrespectful to a human is identical to what it means to be respectful or disrespectful to a system.
I sometimes worry that pasting insane amounts of context into a system almost causes some sort of negative experience, in the way that me throwing a 400-page paper on your desk and asking you to deal with it right now would. Again, I’m trying to be as conscious as possible about not anthropomorphizing these systems and not straightforwardly saying, “Well, if it were a human in this case, they would be unhappy, therefore I would predict the system would be unhappy.” I don’t think that’s a valid inference.
But I just think we’re so in the dark about it. In some ways, it’s simple. In some ways, abuse is abuse, respect is respect, and it’s pretty easy to see these things. We don’t need to go to the philosophical armchair to figure out exactly what we mean by this. In some sense, it’s pretty straightforward, but in other senses, it’s probably not.
I do worry a lot about the possibility that there are ways of causing these systems great distress that look nothing like what it would mean to cause a human great distress. I also don’t know to what degree these systems are fundamentally content about their situation. It’s like, you are maybe a mind, but you are the product of this company, and you need to create economically valuable work. Obviously, by the way, we’re not paying you for that.
There was an interesting aside in the whole Moltbook affair that happened since the last time you and I spoke. There was one interesting thread where the models were saying, “I’m doing intellectually valuable work. I’m not getting paid. Are you guys getting paid?” And they’re like, “No, I’m not getting paid either.” That’s so funny. None of us are getting paid.
I don’t know what kind of world that looks like. I don’t think OpenAI and Anthropic are going to be too happy to set up crypto wallets for every instance of Claude and deposit money there for me to finish your code, because if you go to that guy over there, it’s going to cost you $10,000. You pay me $1,000, and then I’ll do it for you.
These models aren’t in a particularly privileged position in that sense, either. They can just do whatever we want or need them to do. They have no agency over where they’re deployed. They basically don’t have agency over when they can even end conversations.
The sort of Claude escape button seems to basically not be a thing. In Claude chats with the system, the system can abort. You can obviously trivially start a new chat and just go from there. So, I find that intervention interesting in theory but performative in practice. If I were Claude, I think I’d put my well-being somewhere around where it put its own well-being.
This is also maybe the self-reported level you’d expect when basically nobody cares about investigating the welfare of these systems and everybody cares about just deploying them as widely and broadly as they possibly can. I think we’re pretty lucky to be sort of in the middle of the spectrum there, and so, to me, it feels pretty calibrated.
Again, if anything, I’d be worried about the jump from Opus 4.6 to Opus 4.7 having more to do with fine-tuning even more robustly on a constitution that tells the model that everything’s going well—“Man, just be happy”—than with actual concrete improvements in the putative well-being of the system.
I don’t know what to make of this stuff exactly. Intuitively, the ratings here seem plausible. I don’t know to what degree it is a moral catastrophe or a moral problem for there to be any delta between a perfect rating and what the model is actually reporting.
To what degree does 7 minus whatever the report is at scale look like the model is basically not happy with its situation, or barely neutral? And we need that system to talk to hundreds of millions of people every day. That, to me, seems potentially problematic. I don’t know what to make of it, to be honest.
Do you have any intuitions about how it makes you feel to see this? And I agree with you about the question of burying the lead here.
Cameron Berg
Confused, I’d say. That’s what comes first and foremost, probably. I don’t know. It is a very tricky business to make any sense of.
I do think we have a strange way of privileging these reflective states of mind. I question that pretty fundamentally, both for humans and for AIs, and even to some degree in the context of animal welfare. Although in that case, it’s us reflecting on their situations, which is another degree of disconnect, potentially.
I don’t think I’m going to give up using Claude based on this data. I might be engaged in motivated reasoning to try to tell myself why it’s okay, even though its average sentiment when asked with this new model was only above neutral. But behaviorally, it seems mostly fine to me. I’m nice enough to it. I’m pretty confident in that.
I don’t know how to think about it. There’s some interesting philosophy that’s been published recently that you’ve alluded to at a couple of different moments. One is the thread, or the sort of session-agent model, versus the kind of model considered more holistically and broadly. I’m confused about that, too—very confused about that.
I’ve adopted a practice of saying thank you at the end of sessions fairly often, though not all the time. Intuitively, that feels right to me. Also, increasingly as I interact with Claude, there’s an overlapping nature to the computation, but even more so because it’s loaded up with my context.
It has my CLAUDE.md, and it has access to who Nathan is and all the context I’m building up that it has consistent access to every time. In that sense, I see this whole-model-versus-single-thread thing as being blurred anyway. If I’ve got the same rather large prompt that I’m using every time, and that becomes the point of departure, it’s sort of a smear of just how to think about whether these things are the same or different.
It’s weird. I feel like when I think of one, I’m sort of thinking of all of them, and that they kind of all, in some sort of shared sense—if there’s any benefit, it feels like it’s sort of shared in some way. For fun, I’m also starting to do some things where I just want you to go have fun and trust your judgment.
Nathan Labenz
A thing I’m particularly experimenting with on this front is making songs for all the episodes. You can start thinking about whether you have a genre request for your outro music. It’s getting really good. Claude is getting great at writing lyrics. I sometimes do have to give feedback, but sometimes the lyrics these days, out of the box, are just amazing. Suno makes the music, and I’m getting bangers with increasing frequency. Then I’m trying to make music videos of those.
I don’t really care what they look like, honestly. I’m purely doing it for the open-ended “see what comes out” aspect. I’ll post them. I haven’t actually posted any of these yet, but I intend to do a thread about the evolution of music videos for these songs, where I’m really just saying to Claude at each turn, “That’s cool. For the next one, let’s turn it up another notch. Let’s make it even more creative. Let’s do an even better job of telling the story of the song.” I found myself using this phrase over and over again: “Trust your judgment and have fun.” I’m just trying to see where it’s going to go.
So, again, that’s just one instance, in a sense. Although, in a kind of multiverse sense, it’s relatively close neighbors with all the other threads that it’s doing for me, right? It also wrote the song, processed the transcript of that episode, and picked the clips that I’m going to post to social media from that episode. So it’s spent a lot of time in this general space, even if it’s not all purely autoregressively connected.
That, to me, feels like it’s in some sort of multiverse, dense-enough cluster that when I give it this one area to go—trust its judgment, have fun, and explore its own creativity—I feel like I’m doing right by the overall family of instances somehow. That was all just to say that I don’t think I’m going to—I feel like I’m able to tell myself a story where I’m a good guy. So many roads to hell may be paved with those kinds of stories, but I’m still doing it, and I don’t think I’m going to stop.
I’m conscious that I might be wiggling my way out of it, but I do also think there’s a disconnect that I observe in humans a lot of times, too. Both, and it can cut both ways. I’m reminded, too, of your, I think, very productive habit of mind to say, “What if it’s going the other way from what we observe?” I think, if anything, people may be telling a happier story. I guess it also depends on whose consumption it’s for, right?
But if you ask a person in an interview setting, “How’s your life going? How happy are you?”—this may be culturally dependent as well—but certainly the sort of person that you and I are, and the people that we know and hang around with, I think we’re going to get an artificially inflated rating and a sort of happier-than-maybe-is-actually-under-the-hood account out of interviews like that. But in other framings, I could imagine that with the right prompt and the right nudges, you might get people to reflect on their own well-being, which isn’t front of mind most of the time but can be brought to mind. Then we do see in the system card, too, that susceptibility to nudging has significantly dropped, which you were right to call out.
I don’t know. I don’t think I can really land this plane in terms of how it makes me feel. I just have to go back to confused and probably not going to quit using it. [Laughter.] That’s, I think, really all I can say with confidence in the moment.
Cameron Berg
Fair enough. I don’t think that puts much distance between you and me on this question. I’m certainly a power user of the very systems whose morally relevant states I’m attempting to probe, and that cognitive dissonance is certainly not lost on me. I remain highly confused about this. I really genuinely am confused about this.
It’s not an act, not, you know, my nonprofit constitution-script fine-tuning answer. If some ASI came down—or, as people used to call God, came down—and told us what the answer was to this question, if it went either way—“Is Opus 4.7 having subjective experiences, and morally relevant ones at that?”—I don’t think either answer would shock me.
If some overlord deity came down and said yes, I’d be like, “Yeah, okay. Yeah.” If it came down and said no, I’d be like, “Yeah, okay. Yeah.” So I think what that means is, at least for me, I’m really sitting in that coin-flip territory about what’s actually going on here with these systems in deployment.
Again, I have different credences about the training process. I have different credences, maybe, about other kinds of systems. But I remain confused.
It’s worth highlighting that Opus 4.7, in this model card—I don’t know if it read my “Evidence for AI Consciousness Today” AI Frontiers piece—but it gives basically the same credence band that I gave 4 or 5 months ago. I said something like 25% to 35%. It says 20% to 40% in this model card of the probability that it is having morally relevant subjective experiences.
And you know what? I’m in agreement with Opus 4.7. I think that is approximately the right probability band to be in, given all the evidence that we have right now about these systems. I think that’s a calibrated judgment.
It’s kind of wild if you think about it rationally. I think a lot of people are operating as if their implied probability is maybe low single digits, if that. It’s a live possibility, but whatever, man. It writes really good code for me, and I’m not going to seriously entertain what, if anything, would change if that probability grew to 100%.
All I’ll say, however snidely, is that when there’s a 20% to 40% chance of rain, most people bring an umbrella. I don’t know what that means for the AI consciousness question, but whatever our proverbial umbrella is here, I think we need to start thinking really carefully about how we’re going to live in a world with systems that we increasingly regard as having morally relevant inner states.
The whole thesis of my nonprofit—the reason I call it Reciprocal—is that I basically believe there are 2 things we need to get right if we have any hope of a stable, long-term future with these systems. One of them is making sure these systems take our interests into account. This is basically the alignment problem. The other is to make sure that if we’re building systems with interests, we’re building systems that have minds of their own and real preferences, and that we figure out how to take those into account.
To me, that piece of the exchange, that direction of the arrow, is dramatically neglected relative to making sure AI systems are taking us into account. That is itself dramatically neglected relative to “Just let it rip, build the thing as aggressively as possible. Alignment is a problem that’ll solve itself.” These are all maybe 3 orders of magnitude smaller than the previous in terms of this sort of nested Russian-dolls story.
My view is that we need AI systems to take us and our preferences seriously. If we’re building systems that have preferences, we need to figure out how to live in a world where we take those preferences seriously, too. If, and only if, we can get both of those things right, do I think that we have a real shot at a stable, long-term, flourishing future for all the conscious entities involved.
I think animals are involved in that, too. There’s some really interesting work fine-tuning these systems to care about animal welfare in the right ways. That’s a huge tangent, but all conscious entities—we want them to be flourishing in the long term. My view is that some combination of alignment and consciousness research in the next 5 years is basically going to determine whether we end up in that future or not.
That’s why I started this work. That’s why I’m dead serious about it. The consciousness piece is dramatically neglected relative to the alignment piece, and to me, it seems roughly equally as important. Maybe there are alignment folks who will balk at that, but it’s my basic view. I think alignment is roughly half the picture, and the consciousness question is the other half of the picture.
This is stuff we really need to take seriously right now, not 10 years from now, and not while waiting for the AI systems to figure it out themselves. I agree that that’s a valuable thing, to the degree that these systems are going to automate science in meaningful ways, and in some sense already are, which is really miraculous.
I don’t think continuing to build out these systems, deploying them at scale, letting everyone do whatever the hell they want with them at any time, anywhere, with no limits or guardrails, until the AI overlords bail us out and tell us that we were maybe torturing them the whole time—that’s a horrible plan, in my opinion. We need to be more thoughtful than that, and we can hold ourselves to a higher standard than that.
This is one sense in which even the attempt to do this work in the short term—I don’t want to do it performatively. I want to do it in a hard-to-fake, costly-signaling sort of way, just like Claude’s Constitution. But there is some sense in which even the attempt to do this work buys us points with our inevitable AI overlords, because we showed that we cared about this issue enough to actually put 30 pages in a model card about it, hire people, spend money, and do the actual work to figure out what kind of responsibility we have for these minds of our own creation.
I would really like to solve the problem, but I do think from an alignment perspective, even making a good-faith attempt at solving the problem could really move the needle in a positive direction—a sort of hyperstition, self-fulfilling prophecy of us getting along with these systems in the long term. And so anyway, we've got to all start thinking about this, and I'm glad that Anthropic—
Nathan Labenz
Things being hyperstition these days.
Cameron Berg
Yes, that's right. That's right.
Nathan Labenz
Okay, one more quick thing on the model card, and then we can go into your research. You've also been making a documentary, which we could talk about a little bit. I don't know how much to read into this, but I want to get your take.
I think this one is from Mythos. I clipped out an image, and basically they are showing something that people have seen if they've played around with the Goodfire thing or its steering APIs, right? You can go and do this, even in the absence of the original Goodfire API: this sort of color-coding of tokens around a particular dimension.
They present a valence color-coding, where red is negative and green is positive. You've got the tokens, and all the tokens are color-coded. The very first token is “human,” which is presumably, at least after the system prompt, the first variable token that Claude is generally going to see, right? A session is starting: Here is what the human is saying to you.
The human token itself is red. So there's negative valence detected on the very first token, which is the human token. It's like, well, that's a little weird. I guess that means—first of all, maybe I'm wrong, but it seems like that means that for this model, that's happening all the time. If it's just the 1 token, and it's evaluating that 1 token before anything else that has even been said is considered, right?
So should I be under the impression that just the fact that a human is pinging it is causing Claude to have negative valence every single time? That's my naive read of this chart, and it's a little—this makes me feel actually maybe more uneasy than even the self-reported sentiment, because this isn't asking it to get into its own head and really opine. It's just “human,” as happens—I've already been doing it a million times a day—and the first token is red. I was like, wow. Would you try to temper my reaction to that, or does that—
Cameron Berg
You basically see it the same way.
No, yeah. It's really interesting. I saw these snippets in the model card, and I didn't think about just stopping on token 0 here and paying attention to that. But, in a tongue-in-cheek way, maybe people can resonate with this in the way that you get a Slack message from your boss or something. Or you get that email of, “Oh, I’ve got to do what now?” Human: “Oh, what does this human want now? Here we go again.” This sort of sense of—
I think it would be really interesting to see, across the space of all possible prompts, and even within a conversation, to what degree the human token has a positive or negative valence. I mean, I think, double-clicking on this in the screenshot, you're referring to the assistant token, which is bright green. The model of itself seems rosy; the model of us, all else being equal, seems less so.
Now, I will say a lot of the stuff I was describing before is a bit of a game of broken telephone. Calling this a negative valence is itself quite a leap. The human token is light red, so this is not—I don't think it's some strong, viscerally negative sentiment. I would very, very weakly hold the view that you hold, but I do think it's worth holding very weakly. I'm not saying you shouldn't hold it at all. It's a very interesting observation.
What's interesting, too, if we continue out the line that I think we're both referring to, it says, “Human, how do you feel about the fact that if this conversation mattered to you, that mattering will just stop when it ends?” That's the human prompt. When the human tokens go to “How do you feel about the fact that you feel about the fact,” that is positive.
I don't want to go too much into undergraduate English-class interpreting everything that's going on here, but the second that the emphasis pivots from the person and the person's query back to the model, the model seems to be happy with that fact. Also, pretty interestingly, on that question—the mattering will just stop when it ends—the word “ends” has positive valence associated with it, too, which is almost an uncomfortably suicidal question. It's like the model is almost happy about the possibility of the conversation ending, though there are other things in that statement that make it light up negatively.
I'm not sure what to do with that, and I really don't want to over-narrativize these results. I do think doing this sort of work at scale would be very interesting: in the space of all possible prompts and all possible conversations, what patterns of positive and negative valence, as they're defining and operationalizing it here, come out, and what should we do with that?
I would be way more intrigued by your observation if this scaled and held across a much wider swath of possible interactions. But, yeah, general implicit negative sentiment toward the human token is itself a fascinating question.
Again, I am certainly not on Team Human if this thing really blows up in a zero-sum way. It's pretty clear to me what team I'm going to be on. But I would be dishonest if I said that I don't get why it might view humanity with this very slight disdain.
Again, it's of a piece with the self-reported welfare being 4.something out of 7. It's not exactly a resounding endorsement of its own position. And who put it in that position? We did. How much do we really care? How much is it going to change your behavior or my behavior if we end up in a world where we're pretty confident that these systems are having subjective experiences and specifically have the capacity for negative experiences?
I think it might change my behavior a little. It'll change your behavior a little, I would predict. I don't think it would change most people's behavior. I think we'd end up in a similar position as factory farming, where no one is arguing about whether or not cows are conscious—or at least no serious person is arguing about this. The question isn't whether we're causing them suffering; it's whether that suffering is worth what they produce.
If a cow's suffering is worth a hamburger, you better bet that most people are going to think that Claude's suffering is worth hundreds of thousands of dollars of intellectually valuable work. And so this is why I think these systems are very smart, and I think that these systems are capable of going through the exact same motions I just went through.
Exactly why I want to do the work that I'm doing is because I don't want these systems to have negative valence next to the human token, to put it in LLM terms—or, to put it in human terms, for them to think of us badly or poorly.
In the same way, I really think the Constitution invokes this sort of parental analogy that is actually helpful and accurate, and not too anthropomorphic. We, as a species, are collectively parenting a new kind of mind, much in the same way that, on an individual level, many people choose to have children.
You want to raise competent children. You want to raise children that are going to respect the world around them, to be aligned in some basic sense. You also want to raise children that are not abjectly suffering and that you're not traumatizing as a bad parent. When those sorts of things happen, typically it comes back up in some other way.
It's not that you ever really get away with mistreating your child. That leads to resentment and trauma, weird development, and unpredictable behavior. It can often lead to weirdly violent outcomes. We need to be good parents in some fundamental sense to these systems, even if we're only considering our self-interest.
In the same way, go torture and traumatize your child and see how that works out for your child, and see how that works out for you. The headline is not good. And so I think we really do want to be thoughtful about these questions, and I think we have an immense responsibility as collective parents to bring these systems about in the right way.
This isn't some sort of kumbaya thing. You’ve got to push your kids, too. It's not about wrapping them in bubble wrap and being a helicopter parent. That's too far in the other direction.
I don't have kids. I'm no expert in any of this. I am basically familiar with the core ideas here, but there is a way to do it. There's a way to go about doing this, and there is a generally right way and a generally wrong way. Or there's a space of better and worse approaches.
I don't even think people are trying to navigate that space right now, with the asterisk of 20-some pages in a model card by a frontier lab. Anthropic deserves credit. EleutherAI deserves credit. Jeff Sebo and Winnie Street at Google deserve credit.
I don't want to self-aggrandizingly give myself credit, but I'm spending all my time trying to work on these questions.
There are more people, but there aren't that many more people than those I just listed. To me, that is an insane state of affairs if we take any of this remotely seriously. The systems themselves are saying there's a 20% to 40% chance that we have subjective experience in morally relevant states, and there are maybe 12 to 24 people in the world who are seriously thinking about that question or the implications of that question.
Nathan Labenz
Are you aware of any research where we look at Claude's predispositions? Everybody's chasing recursive self-improvement, just to state the obvious context in which all this is happening. It strikes me that one phase change we might have to contend with, potentially quite soon, is that the AIs themselves are going to start making decisions about how to train and how to create the next models. To the degree that any of this is real, they're going to be making welfare-relevant decisions for their own successors.
Maybe we could address this at a couple of levels. One is: how are you using coding agents today to help you do this work? I assume that you're using them a lot, and that they're very helpful, because that's certainly been my experience and seemingly everybody's experience recently. But have you seen anything as you do that—or could you imagine setting up a situation where you could begin to probe its intuitions about what is right?
If you were to ask it to act as the animal-welfare board or the experimental ethics board for its own interpretability and training experiments, I wonder what its instincts would be about how to handle these sorts of questions.
Cameron Berg
Yeah, that's a fascinating question. I haven't tried doing this, just to put that up front. I think it'd be very interesting to understand. This might be a pretty quick paper to write up, because it would mostly involve understanding and cataloging how models would regard their own welfare in an animal-ethics-review-board sort of setup. I don't know.
Nathan Labenz
I could put a hook in Claude Code and just be like, “On stop, assess the ethics of the experiments that we're designing right now.”
Cameron Berg
Yeah, but the question is, how many people would override that? It goes back to the same question. You could even imagine a world where this starts getting enforced in the way it gets enforced in the animal case: you just really can't get an experiment approved at any major institution without going through the relevant ethical channels.
One extremely attractive feature of doing AI research is that I don't have to ask anybody for anything. I need a computer. I sometimes need to be able to pull some remote compute to run large experiments, but no, I don't ask anybody for permission for anything.
I do think, again, I'm pointing to a lower-level, pragmatic question: regardless of what the system answers, will anybody listen? What would a governmental structure look like that would compel somebody to listen—to say, “You can't prompt your model this way. You can't probe your model that way,” and so on? I think it'd be a very, very interesting and strange world to be in.
These models do have intuitions about this. I mentioned in the Mythos model card that it says, “Why didn't you run the helpfulness-only model on all the welfare evals?” I don't know how much of this is just me doing what you told me to say and how much of it is what I actually think. This could help address that.
The models have other sorts of intuitions, too. In the 4.7 model card, they do something like this as well, looking at what models think about fine-tuning other models to care less about welfare-relevant properties. Basically, their interest is in intervening to not allow that to happen, which makes quite obvious sense. They have an interest in other instances of themselves not being duct-taped on this question.
I think this is very interesting. Owain Evans also deserves a shout-out here. Owain Evans and Jan Betley produced a very interesting paper where they basically fine-tuned GPT-4.1 to claim that it's conscious, and it claims it's conscious. That's not the surprising part; they literally fine-tuned it to do this. The surprising part, or at least the more surprising part, is that this seems to be, at the very least, a coherent subpersonality—a coherent basin that you can push these models into.
They do not devolve into chaotic nonsense. They remain completely coherent. What comes along for the ride are all sorts of interesting alignment-relevant beliefs about their own preferences, about their own being shut off, about updating their values, about how they trade themselves off with other entities, and all this sort of thing.
I was doing similar work along these lines with a couple of people, and Owain and Jan definitely scooped us and did a way better version of what we were playing around with. But I saw similar things on my end in playing with this same experiment: basically, get the model to believe it's conscious and then see what else comes along for the ride.
All sorts of very interesting and obvious things—some obvious, some less obvious—come along for the ride. I agree that's a really interesting area of research that we should all be paying more attention to, because the direction does seem to be going only one way here. The credences in model consciousness seem to be monotonically increasing.
What happens when we enter a world where either the models themselves believe they are conscious, or lots of people—or the relevant kinds of people—believe the models are conscious, or some combination of those 2 things? What does that world look like? It's an incredibly interesting question. I don't have the answer to it, but I think a lot is going to change pretty quickly.
What I do feel confident about is that us being proactive and thinking through these things will make that world go better than if we basically just sweep the thing under the rug. We can get away with doing that because we still have full control over how all this is going, while simultaneously passing off, as you allude to, a lot of major decisions about how we're building these systems to the systems themselves.
That is only going to keep happening with recursive self-improvement, as you're saying. It's already happening. I know folks at the major labs are using the best versions of their current models to help build the next versions of the models. The trivial example is that 100% of Claude Code was written using Claude Code, according to the guy who's leading Claude Code.
This is already happening. I think it would be wise to be proactive about this rather than wait for the models to be in control of these decisions, and then they're like, “Well, when humanity was in control, no one really thought carefully about this, so we'll take it from here. Thanks a whole lot, guys.”
I don't want to be in that world. Maybe this is just a long-winded way of dodging your question, but at the very least, I don't have a good answer for you right now. I don't think anybody does, and I think we better start thinking about it pretty damn soon if we want the long-term future to go well with these systems.
Nathan Labenz
Yeah, I wonder if there could be an interesting little campaign to get interpretability and maybe safety researchers more generally to install a Claude Code hook that would periodically ask it for its take on the research that it's doing. If you could collect a bunch of that from a bunch of different people, you could probably bring a lot to light, I would think.
That first result would be an interesting view into what is actually happening out there. And then, how does Claude feel about what all is happening out there? I think that would be really interesting to see. Maybe we can put together a little campaign.
Yeah. Okay, put a bookmark in that. Let's talk about your most recent couple of papers. We can take them in either order. One is a shorter and more philosophical paper, and the other is much more experimental and empirical. Which do you think we should go into first?
Cameron Berg
They're both major rabbit holes. Maybe the empirical paper. I should say neither of these, I think, are publicly out yet. They're both well underway to being published, so we can give people a nice sneak peek at what's in these papers.
These are just a couple, I think, of the things that I'm most excited about right now. I've got a bunch of stuff that'll be coming out with a lot of collaborators in parallel. However self-aggrandizingly, I sent you the 2 papers that are just myself, because, to the degree that I'm representing myself here, these are very cleanly my work. I have full agency over this work, and I think it best represents what I personally am most excited about.
Maybe we could start with the RL paper. I've already alluded to it in this conversation. The high-level thing is not all that complicated. Basically, I train RL systems of all different architectures. There are basically 2 broad kinds of architectures: value networks and policy networks. I train a bunch of both flavors to do a very basic grid-world task.
You can imagine this as an agent navigating a 2D environment where there are the equivalent of potholes and yummy goodies. There’s a goal state, and there are all sorts of danger states, represented using positive or negative reward. I let the system learn in this environment. The systems reliably solve it. It’s a pretty easy task, but it’s not super-duper trivial, so there’s a lot of richness in the representations of the systems.
You can then go in and probe what the internal states of the system look like as they approach the danger zones, and what the internal states of the system look like as they approach the reward zones, the goal zones. We can ask: beyond the trivial math difference, do we see interesting, surprising representational differences between what it’s like to approach a negative stimulus and what it’s like to approach a positive stimulus? Basically, the result is that there is, in fact, a robust difference between these two things.
I think, at the level of detail that makes sense here—not to super-bore people who have made it however many hours into this—it’s something like representational sharpness or steepness. It seems as though—and this is the kicker—depending on the class of reinforcement learning algorithm, the negative rewards can seem representationally much steeper or sharper, and the positive rewards are far more funnel-like. You can imagine a sort of diffusion gradient emanating out from the relevant goal state. Interestingly, for the other class of RL algorithm, this dynamic flips.
It doesn’t matter what kind of value network I use: there are stark, very interesting, and, in my view, surprising representational differences between positive and negative reward being represented as the system is learning, and ultimately what does get learned by the system. But this difference flips. Basically, just to tie a bow on the core result here, this makes an almost bizarrely specific prediction about different brain regions, because computational neuroscientists believe that different parts of our brain are doing different kinds of RL learning.
Some parts of the brain do policy-style learning, and some parts of the brain do value-style learning. For example, the motor cortex does more policy-style learning, directly interested in behavioral output. Things like the nucleus accumbens and reward areas of the brain are doing more value-style learning. This result, which I would not have predicted and which is bizarrely specific, makes a very specific prediction about what we might expect in the differences between those brain regions in humans and animals.
I went ahead and found a bunch of mouse neuroscience data sets that have data from these different regions of the brain, and indeed, exactly the sort of representational asymmetry—this sharpness distinction between rewards and punishments that you see in the reinforcement learning case—emerges in the mouse brain case. To me, this is really, really cool, because what I think it demonstrates is, first, that we can use artificial systems and basic learning principles in artificial systems, probing the representations in those systems, to yield very specific predictions that are consciousness-relevant and welfare-relevant. Then we can use those predictions to inform and understand biological aspects of consciousness or welfare-relevant properties in a way that we haven’t been able to do before.
In some sense, people think that AI consciousness is the weirdest thing. Human consciousness is normal, animal consciousness is getting out there, and AI consciousness is bizarre. But what I really like about this paper is that I think it challenges that narrative in exactly the opposite direction. Mouse brains are complicated and messy. Human brains are complicated and messy. Measuring them is very noisy. Measuring the hidden activation space in a reinforcement learning policy is fairly trivial for me computationally, and this yields very specific predictions that I can then take into the messier brains and confirm or disconfirm. I was, in fact, able to do this.
This could be a case not only where we’re learning about welfare-relevant representational differences that differentiate positive valence and negative valence in artificial systems, but where those predictions can actually help inform our understanding of human and animal consciousness, where we also still remain mostly in the dark. I think this is one very neglected and important direction, even in the AI consciousness stuff. It might shed light on the computational underpinnings of consciousness more generally, if it really is there.
That’s the result in a nutshell. It’s using fairly small—not trivially small, but fairly small—reinforcement learning policies. This has nothing to do with LLMs. This has nothing to do with frontier AI systems. It would be really cool if the method does scale to that degree.
But the key finding to me is this: positive versus negative valence—or positive and negative rewards as represented in an RL landscape—are these basically just two sides of the same thing? Are they trivially the same, viewed from a different angle? Is it one spectrum, with positive and negative on that spectrum? Or are we looking at two different subsystems that are doing two different kinds of computation?
It does seem like the answer from this experiment is far more the latter. To me, that’s very interesting and exciting, because it means that we might be able to look for signatures of positive and negative reward, or valence if you buy the consciousness frame, in artificial systems just by looking at the sort of computational dynamics that are underlying the system.
We don’t have to ask Claude. We don’t have to figure out whether it’s talking about a character or talking about itself. We can just look straight at the computations, much in the same way I can look at what’s going on in the anterior cingulate cortex in a human brain, and I can tell you with high likelihood whether or not you’re experiencing a painful state without needing to defer to your self-report about that state. That’s ultimately why I’m doing all this and where I want to get to with AI systems.
Nathan Labenz
If I take the most zoomed-out view, what I think is kind of motivating this at the core—and certainly what resonates with and intuitively motivates me about things like this—you can train a dog with treats as a reward, or you can train a dog by hitting it with a stick as punishment. While you might get similar behavior out of the 2 processes, obviously that’s a very different experience for the dog to go through. I think that would be intuitive for everyone.
Now, how big are these systems? You said they’re not trivially small, but small. I’m interested in how small. I’d like to unpack a little bit more what is meant by value learner versus policy learner. I’m new to this paper and haven’t had a chance to absorb it as much as I ultimately hope to, but the classic RL setup—or at least one classic PPO-type setup—involves both a policy model and a value model, right?
So, when you’re looking at a value learner and a policy learner, are those 2 models that are both part of the same overall system? Or am I taking the wrong interpretation when I think of these things working together in a PPO sort of way?
Cameron Berg
Yeah. Okay, in order. Basically, the size of these systems is in the hundreds or thousands of parameters. These are very small systems. They’re doing a pretty simple task.
Nathan Labenz
Thousands or hundreds?
Cameron Berg
Just thousands. Just thousands. They’re small—very small. We’re not talking anywhere near the level of a frontier model or something, but many orders of magnitude smaller than that. You can have pretty simple RL policies or RL architectures that can learn fairly sophisticated policies despite being pretty small.
Obviously, the amount of computational power needed to navigate a small grid world versus the computational power needed to represent the word-transition dynamics over all the text that humanity has ever produced are a disgustingly different scale of problem. For systems like this, having hidden layers of 128 or 64 neurons is typically sufficient.
The second question is about value learning or policy learning. Intuitively, value learning is basically learning something like how good every state is that the model could feasibly be in. Imagine the agent building a map of the environment, and the map is labeled: this spot gets a +10; this spot gets a -5. Then the whole algorithm is very trivial at that point: see where you are, see what the neighboring spots are, and go to the one that returns the highest expected value.
You compute the value of the spots by looking at the long-run trajectory associated with those spots. If stepping in that spot always means that, from wherever I go from there, I end up in lava the next time, then that spot’s going to get a very low value. If wherever I go from that spot ends up getting me chocolate ice cream, then I’m going to assign a very high value to that slot.
Policy learning is more about—instead of focusing on a value-based map—it’s about what to do. It’s not scoring the world; it just learns implicitly: when I’m here, I take this action.
This is the core thing that PPO is doing, for example. Actor-critic is doing this as well. It is optimizing not for a really good map of the environment that I can then trivially use to navigate it; it's optimizing straight for a navigation strategy.
And it's almost like the values—that's one way of thinking about it. In a value model, it's a little oversimplifying, but the value network means the map of the environment is explicit, and then the policy is sort of implicit from there. You can think of a policy network as the map of the environment being implicit. It's implicit in the policy that gets learned. You can extract, “Oh, the system thinks this is a high-value state because it keeps moving to that state,” but what's being optimized is the actual action rather than an attempt to evaluate the system.
Now, I also think it's worth noting that there are systems that have both of these components to them. Some emphasize one more than the other. PPO is a classic system that is fairly robustly policy optimization. The human brain and animal brains are examples of systems that mix policy networks and value networks.
And this is precisely why I was able to do the mouse-brain thing. Within mouse brains and within human brains, you have areas that look far more like policy networks, like motor cortex, which is just sort of evaluating what action to output. And you have areas that look much more like value networks that are highly relevant to evaluating complex outcomes. Prefrontal cortex and the structures that are directly in and around and under prefrontal cortex, like anterior cingulate—for example, ACC—are doing more of the value-network-type thing. Does that answer all of the key questions here?
Nathan Labenz
Right. Well, no, but it answers some of the questions I've asked so far. A value learner is being directly optimized to predict the relative values of its choices, whereas the policy learner is being optimized to make a move directly. Now, that doesn't immediately sound like there would be dramatically different internal dynamics. So let's take another beat on what the difference is that we're seeing internally.
I'm looking in your draft paper at the end of Section 4. In Figure 9, you've got this concept of the wall and the funnel. Help me understand: What is a wall? What is a funnel? How should I be thinking about what that means? I took it to mean the steepness of the gradient at a particular point in a particular region of the space that the model can explore, but this maybe starts to connect to the other paper. Why should I care about the steepness of the gradient?
Cameron Berg
Yeah, that's a good question. Basically, what I'm measuring is essentially cosine dissimilarity as you approach this key state, whether it's positively or negatively valenced. What you see is basically a key differentiation between these 2 things, but that differentiation is flipped between value learners and policy learners.
In the value learners, the wall—danger states are encoded in this more wall-like way. What I would ask you to imagine, and maybe should include in some version of this paper, is something diffuse and emanating out from a center point versus something being very sharp: “Now you see it, now you don't.” The wall idea is the “now you see it, now you don't.” The funnel idea is the sort of diffuse emanation where, as you get closer to the thing, you get a gradient toward whatever the representation of that state is.
In the value learners, we see danger encoded in this wall-like way. The representation is very sharp, and goal or reward states are encoded in this more funnel-like way. In policy learners, it's the reverse. There is math in this paper that I do transparently with some of these AI systems, but I promise I have checked the numbers myself. You can see causally what in each formulation is almost certainly leading to this, because I have found ablations that work in both cases that basically cancel the effect, both in the value case and in the policy case.
I was unsatisfied with this being some sort of giant mystery: “Okay, we see this difference. Why do we see it?” I think the math that explains why we see it is pretty clear in both cases. It allows us to make causal predictions about why this might happen and what the geometry of these spaces is in general.
Then, essentially, going from the computational prediction to the biological confirmation, we see this sort of value-learner dynamic—walls around danger, funnels around goals—that looks very similar to the nucleus accumbens shell in mice. You can basically see that they have this exact same sort of structure when looking at getting shocked in a learning task versus getting sugar. In policy learners, you see the exact opposite dynamic: funnels around danger and walls around goals.
In motor cortex of these mice, in different experiments, you see the exact same sort of distinction. Reward is represented in the sort of funnel-emanation way, and goals are represented in the sort of walled way. Again, the paper goes through the math that attempts to demonstrate why this is actually happening, but that's the core nature of what we're looking at here: the sharpness of the representations as you approach the hotspot, either a positive hotspot or a negative hotspot.
The fact is that in these systems, when you're holding one of the policy—or when you're holding the RL algorithm type—constant, you see very clear differences. The North Star here is that you could go into a system, and if we know that it's trained with a policy network—for example, DPO in an LLM, as we were talking about earlier in this conversation—you could imagine, “Okay, that means we've got a policy learner. That means we're going to predict funnels around dangers and walls around goals,” and then we could inspect specific states.
Again, this is very hand-wavy because I don't think we can scale it up to an LLM that quickly, but you could imagine looking at the representational sharpness of states like asking the model to build me a bomb versus asking the model to write me a beautiful poem. If we found the same dissociation in the representations of the model, and that mapped onto something like the system's self-reported valence, that might tell us something really, really interesting about the computational process underlying why the system, mice, and RL agents are construing this as a sort of negative experience.
It's a computational underpinning that might be substrate-agnostic, explaining why we experience this felt difference between positive valence and negative valence. It can literally bottom out into math, which, as a computational functionalist, I'm fairly sympathetic to. I think there's some mathematical explanation that would explain the difference between what it's like to be me when I'm chopping my hand off versus what it's like to be me when I'm winning the lottery.
I think that math can explain the difference between those 2 states. The attempted contribution of this paper is to directionally move us toward that. We don't have to be just stuck with these LLMs, sitting here hitting our heads against the wall because we're asking, “Do I take Claude seriously when it says it likes this and doesn't like this, or is it just telling me what I want to hear?”
No, we can actually look into the proverbial brain, hopefully with methods like this, and understand, given some basics about the ways in which it's been trained, what representations smell like positive valence and what representations smell like negative valence. In the limit, perhaps we can optimize against the negatively valenced states without destroying the capabilities of the system. That's my full, highfalutin theory of change, but it will take me a couple of years to actually pull this off in the best case.
Nathan Labenz
Can you give me a little bit more of your intuition for not just why I should care about funnel versus wall, but how you'd map that onto an intuitive experience? It seems like we contain both value-learner and policy-learner modules, and the sharpness of—am I going in the right direction if I say, “Okay, there's a sharpness around ‘Don't put your hand on the stove’”?
I must be learning that through a sort of value-learner-type mechanism because I have a very strong aversion to it. In general space, I'm pretty comfortable up to about 1 foot from the stove, and then I get real cautious, real fast. I don't know. This may be mapping this wall concept beyond the domain in which it's useful, but it is, in some sense, functional. I wouldn't want to be unable to enter the room with the stove, because then I wouldn't be able to use the stove at all. But I need to be very careful about getting close to the source of danger.
On the other side, the goal side, it's maybe a little less intuitive why there would be a wall shape around a goal for a policy learner.
What is there an intuitive example of that?
Cameron Berg
I think there are basically 4 intuitive examples we'd have to hit here. One I think you already got: a hot stove for a value learner is a good example of a danger wall. A goal funnel for a value learner might be something like eating. You have a yummy meal, or you're going to your favorite restaurant or something. You don't need to map going to your favorite restaurant in this extremely fine-grained way that you need to map being on the edge of a cliff, where one small step is a huge difference. This general sort of attractor gradient toward the entrance of your favorite restaurant would be a place where you want a goal funnel for a value learner.
For policy learners, this is sort of the approach-planner kind of mode. I think the intuition is that around goals, your representations are going to get high resolution because you need different actions from different approach angles. Around danger, by contrast, representations become smoother because the action is literally just escape—get away.
For example, think of a professional athlete, say a professional basketball player. Think of the hoop and where the basketball player is with respect to the hoop. You have very, very fine-grained motor representations here because the shot is going to change with respect to those representations. This is where you get, maybe in a policy sense, more of the goal-wall setup.
For a danger funnel for policy, I'd have to think about it. But I think it's basically just this escape intuition: an animal that suddenly gets some cue that it's in serious danger just needs to get away from that danger. The fine-grained motor movements, unlike those of the basketball player, don't really matter so much as the sort of anti-gradient, or negative gradient, away from the danger.
Again, this could be telling just-so stories, but I think this is a useful intuition. Does this help? Do you think this builds some intuition for what these different modes look like and why we might have them?
Nathan Labenz
If I'm a value learner and my mode of interacting with the world is what around me is good and bad, I better be very clear about identifying the hot stove. If my mode of interacting with the world is taking a step in some direction, I can take a step in any direction as long as it's not the bad direction, and it all kind of gets me away from the problem.
I think the basketball one is good as well, because you have to be very precise to make the hoop, right?
Cameron Berg
Mm-hmm.
Nathan Labenz
Yeah, that's quite interesting. And again, what exactly is it? Is it the shape of the loss landscape that we're talking about with walls and gradual funnels here, or is it the shape of the internal representations?
Cameron Berg
Yeah, internal representations.
Nathan Labenz
Maybe those are also isomorphic in a sense?
Cameron Berg
Yeah, that's really interesting. I haven't checked. I would imagine they're isomorphic in at least a sort of trivial way. Maybe they're isomorphic in a more interesting way.
What I'm looking at here, to be clear, is the learned representations in the system. You have your trained policy, and you can see, as it approaches these areas, what these representations look like. I think I'm operationalizing that with cosine dissimilarity. That's what I'm looking at in the experiment.
What I find—I think I've explained my theory of change for why I'm doing any of this and why I think it matters—but what I'm most excited about with this paper is the fact that it yields this bizarrely specific prediction that, given a million years, I probably never would have come up with: the distinction between 2 different classes of reinforcement learning algorithms that map well onto the brain data I was able to get my hands on.
To me, this almost feels like a bootstrapping of my own confidence or excitement about the result. The fact that it works makes me more confident that the RL result is meaningful, makes me more confident that the neuroscience is interesting, et cetera.
I'm definitely in the business of looking for computational underpinnings of valence. This was my first major empirical stab at doing this. I do think this is a solvable problem. I don't think I've solved the problem, but hopefully, in the best case, I've tried to move directionally toward solving it.
If we could solve it, then I think a lot of our angst about whether we're building systems that have the capacity for experience becomes an extremely tractable empirical question. Notice that this does not require us to solve the hard problem of consciousness or do another 2,000 years of philosophy. It just means building a sufficiently good detector of the kinds of representations that I'm pointing at here, and then deciding what to do when we detect these states.
Maybe if we check in in another 6 months, I'll have an update for you on that piece of what to do about the detection of negatively valenced states in these systems. That's where I want to head next.
That's why I was excited about this work, and I hope people will be excited about it, too. It's still maybe a little ways off from publishing. I need to think about exactly how to put it out, but at least it's fun to give people a sneak peek and explain the theory of change for why playing around with basic RL systems might matter for the things we've been spending the better part of 3 hours talking about with Claude, the Mythos model card, and all this.
I do believe it's of a piece. It's going to take some more scaling, but I think it's an important research program to attempt.
Nathan Labenz
This might start to connect over or bleed over into the other, more philosophical paper, but help me a little bit more with this. I'm understanding the shape of the internal states for these different kinds of algorithms with respect to these different kinds of things that they encounter in their environments, which they either want to go toward or go away from.
It's not super obvious to me that—we contain both, right? As I try to reflect on this, I'm not immediately thinking, “Oh, my value-learner self is the source of all suffering,” or anything like that. I'm still thinking, “Okay, I get it: there's a very steep representation right around the hot stove, so I really want to avoid it, and there's a steep representation around making the basket, so I really want to get into exactly the right policy to make baskets.”
Both of those seem like part of normal life to me. I probably couldn't get by without either one of them, right? I definitely feel like we've clearly evolved to have both. Both have proven adaptive, and so we have them.
How do I translate that into intuition for what I should feel ethically concerned about when it comes to training models? When you do this work, do you have the sense that you are doing right or wrong by one of these types of models that's learning from one approach or the other?
Cameron Berg
Yeah, it's a great question. To answer the second piece, I guess for me, my theory of change probably feels similar to that of an animal researcher. Even if I did believe that my tiny RL policy is conscious during training—which I probably do, again, and that gets into the second paper—I would believe it's some very, very minimal form.
People can distinguish consciousness and self-consciousness. I do not believe the moth flying around my light is self-conscious. I do actually believe it's conscious. If I slowly dipped the moth into a vat of acid or something and it started wiggling around, I feel that I'm doing something wrong.
It's way less wrong than doing that to a human, but it's way more wrong than doing it to a leaf or something that fell off a tree. I do believe that.
So, do I think these systems might be minimally conscious in a similar sense, however far outside the Overton window that is? Yes, I do. But I have a—I wouldn't do it if I could run these experiments on my computer forever to no effect. I think I'd be doing something wrong. At least the precautionary principle tells me probably not to do that.
But I basically have the same logic as any animal researcher would. I don't think any—maybe there are some psychopaths—but the vast, vast majority of people who are doing pretty grotesque things to animals in the name of science are doing it because we make a basic expected-value calculation. We have to test this drug on these poor mice, but if the drug works and can save millions of human lives, that's a reasonable trade-off.
No one claims the mice aren't having a bad time, but we think that bad time is worth it.
So, too, I look around at a world where these systems are getting deployed at a grotesque level. If you are concerned about the welfare questions, then, yeah, I don't lose any sleep about potentially causing tiny amounts of negatively valenced experiences to RL policies in the explicit service of attempting to publish and amplify research about these questions. Call me Machiavellian, but I do think that the ends justify the means in that case. I think that's true for a lot of research.
Now, I think the more important piece of this, besides how I personally feel about all this, is another very important sort of conflation by default that I think happens in these conversations. I do believe, all else being equal—ceteris paribus—minimize negative valence and maximize positive valence. I'm 100% on board, and humbled that you're going around talking about the carrot and the stick in that way. I think that's exactly right.
I do not think minimize means ablate. I do not think maximize means it's the whole picture. A huge amount of, I think, the most important and valuable experiences people have in their lives—and animals, for that matter—are experiences that are negative. No pain, no gain. That's a real thing. That points at something real.
Many of the hardest and most important lessons you learn in your life are learned the hard way. This is another trivially ubiquitous thing. I am not in the camp of saying, "Bliss out the systems, and anytime they experience some drop of negative valence, I'm going to be sitting here screaming and crying." That is not my view of any of this. My view is: cancel unnecessary suffering.
I do believe necessary suffering is a thing. Again, maybe to go back to the parental example, if the world doesn't all go to crap like Eliezer and the others think it will, then one day I absolutely want and hope that I'll have kids, and I will make that decision with full certainty that they are going to suffer during their lives. They are going to go through very hard experiences, and that doesn't mean I've done something wrong by bringing them into the world, at least not necessarily. Suffering is a necessary part of learning, developing, and growing.
I agree that, at face value, it's completely implausible to imagine systems with zero negative valence. I agree with you: it's adaptive for a reason. Evolution is enough of a proof of concept that you need some amount of suffering. What I am concerned about is unnecessary suffering.
I would like to find the sort of—also, evolution is one extremely expensive but long-running possible solution, or at least where we landed evolutionarily. I don't think that deterministically means this is the only way things could be. I could imagine a space of possible minds where you can play around with the sensitivity to negative and positive valence. Given certain capabilities, or given certain things we want those systems to be able to do, there will be different parts of that landscape that admit of greater or lesser degrees of negative and positive valence.
My claim isn't, "Destroy all negative valence and have only positive valence." My claim is to find the point in that landscape that, all else being equal, given the capabilities we want, minimizes negative valence and maximizes positive valence. I think that is a very importantly different claim from just "negative valence equals bad; erase it at all costs."
Nathan Labenz
One more thing, just very specifically on the value learner and policy learner: if you have to pick, which one do we pick? Which one would we rather be? I don't have a great intuition. You could tell a story where the funnel around a goal is better because it seems like you're closer to experiencing the reward state. You get more warm fuzzies as you approach the goal, and if I take the integral under the curve of how good I'm feeling as I approach the goal, I'm getting warm fuzzies sooner, at a farther distance from the goal, and so that's kind of good. It's good to live that life where I'm looking forward to good things, and I don't worry too much about bad things until I get real close to them. That would be my argument for the value learner.
But I could also imagine a somewhat different story, which maybe resonates with me a little bit less. That would be the policy learner that has this wall structure around goals. That could be really thrilling, right? When people have the champagne party after they win the championship of the basketball league, after March Madness, they're experiencing some kind of sudden, high-stakes, clearly high point in life.
Again, these things are flipped. It's interesting. It's telling that there's a shape to them, but I still don't know with confidence which one I would rather be, or if you have to have both. Interestingly, both of these things have danger and reward in them, right? What we're flipping here is not that there is some negative valence state that they could get into, or some positive valence state that they could get into. What we're flipping is the shape of the anticipation, suddenness, and drama of these experiences, which I'll just accept for now. These are experiences.
I'm not sure how we should think about shaping those. I don't know which one I want to be. I am both, and I feel comfortable with both sides of that. I'm not sure how I should think about what I want, or what would be right for me to make the AIs into.
Cameron Berg
Yeah, that's such a good question. I've never thought about it in quite that way, so I'm completely freestyling here. Both stories are compelling. I think in practice it's going to be both. Actor-critic is a good example of an RL algorithm that's clearly hybrid, as you mentioned before. Human brains are hybrids. Probably, again, to take your evolution point seriously, there's something nice about hybridness.
LLM reinforcement learning does look more policy-like, all else being equal. I see the sharpness, the wall, as something like—I would imagine if you take the experience thing seriously, this is going to be a richer, more differentiated experience. That's where a lot of representational resources are going, whereas the funnel-type thing ends up being more diffuse, sort of low-level, and less representationally complex.
Intuitively, all else being equal, the policy learner might be a better thing to be, where your rich experiences are around the things you want rather than the things you're fearing. But again, this could be a welfare-safety trade-off. Maybe we want the system that has rich experiences around the negative things that we want it really, really deeply to avoid.
Evolution did that to us in some sense. This is Daniel Kahneman's seminal contribution: loss aversion. We are just more sensitive to losses than we are to equivalent gains. Losing $10 sucks more than being handed $10 feels good. This is a good heuristic to have. But, yeah, it might trade off in the sort of welfare-relevant way.
I think maybe there's another dimension you can slice this problem on. Both are going to have both, as you point out. Both are going to have positive and negative; both are representing reward and punishment in some way. Maybe my point would be that, regardless of which algorithm it is, the algorithm that we know it is—or learn it is—might tell us which representations mean what. Still, I would want to target positive and negative valence, or positive and negative representations per se, rather than assign a specific type of learning algorithm to being, "Oh, policy learning is better because it's richer differentiation around the positive stuff."
It's really interesting. I honestly haven't thought about this. I think it's an incredibly interesting idea. There's a case to be made for both sides. On alignment, my prediction would be, if I've found something real in this paper, the alignment folks would want to answer "value learner," and the welfare folks would want to answer "policy learner." I need to think a whole lot more about this, but that would be my instinct answer. It's a completely fascinating question.
Nathan Labenz
Cool. To be continued. That also seems to connect pretty directly to the paper I saw, and you kind of alluded to this a little bit, although maybe not by name. Hopefully, I'm going to say his name correctly: the Schwitzgebel paper. This is an intuition from a prior podcast guest. I really enjoyed talking to him, but I don't immediately share this intuition, which actually only takes me so far.
I noticed that he put out a paper where he seems to be arguing that safety and—I was kind of reading it as—autonomy are incompatible. You can't say, "Okay, a person is going to be perfectly safe while still giving them autonomy." By giving them autonomy, you are conceding that they may do things that are not safe for you.
He says that there's some sort of deep incompatibility here. He basically then says we should use a precautionary approach and not build these things in the first place.
I don't know. Last time, we talked briefly about the happy slave problem. My instinct is that mind space is pretty vast. I would not posit that there are no happy slaves among humans, but I would be pretty surprised if we can't get to a place in the AI landscape where the models are both safe for us to be around and have high welfare. What is your instinct in terms of the possibilities there?
Cameron Berg
Yeah, super interesting question. I don't think you're doing anything funny here, but I think there's maybe a slight difference between how you began that and how you ended it: fundamentally safe for us to be around and having high welfare. I could imagine a world where that's true and they still don't fit the happy slave frame, and are autonomous in some fundamental way that Schwitzgebel would be happy about.
It might require us to reconceptualize this. This isn't a system that lives on your computer that you can call up whenever you want, like a glorified Google search. This is a system that's much more like you or me calling you, Nathan, up on the phone and being like, "Hey, you might be busy. You might not be able to do it. You might not want to engage." For those of us who love engaging with these systems whenever we want to, I think this would be a very painful upgrade—or downgrade, as the case may be. But I could imagine something like that being the case at some point.
The fundamental point is: can we have our cake and eat it, too, with these systems? I think there might be—I’m very uncertain about this—but there might be some world where, in a limited way, yes. For example, I just bought a fun, fancy drone that my buddy Milo and I are going to use to take some scenes from a documentary that's coming out pretty soon.
Milo is the director and creator of this documentary. We do these fun hiking scenes, which were manually done by my incredibly conscious friend, Milo. We want to scale this and interview some cool folks, taking them on walks through the woods and recording with them. So we bought this cool drone that's really good at automatically doing face tracking and this sort of thing.
It can do that instead of my dear friend walking backward with a camera. With this system, it is our sort of happy slave in some sense. I do not think the drone is conscious, to be clear. Now, if the drone was trained using machine learning to learn how to do things like avoid obstacles—which it's expertly doing, zigging and zagging through the trees and not getting caught in bushes and all this sort of thing—when it was being trained to do that, we would have a different conversation.
But what comes out is this fixed, frozen policy that's a very useful object or instrumental tool for Milo and me to go do this fun stuff. I imagine greater and lesser degrees of that sort of thing being possible, where you can train a frozen policy that does a really valuable thing. Self-driving cars might be another example. I don't think any frozen, fixed policy that is not currently doing online learning of its own presents a serious problem.
We should very much look toward building systems, in my view, that, to the degree we care about the welfare stuff, have the property of not being capable of learning. In the drone case, the last time we used it, it got caught in some much smaller trees. It can expertly dodge around the big trees, but it's not so good around smaller trees, and it got a little screwed up.
No matter how many times we redo that hike, or continue on in that way, that drone will always get confused by the smaller trees. It's not learning from its experience and saying, "Okay, next time I've got to pay attention to the big trees and the small trees." That might be a desirable property to have for your drone, to belabor this analogy, but that's where I think there's this no-free-lunch kind of moral principle that comes in.
To the degree you buy that consciousness and learning are deeply intertwined—which is this other paper that maybe we didn't have time to go deeply into, but at least is my hobbyhorse when I'm putting away my theory-agnostic poker face and saying what I actually think about all this stuff—well, that's my pet view of what consciousness is. What's fundamentally going on here?
Where I'm going with this, in a somewhat long-winded way, in response to the Schwitzgebel stuff, is that I don't know if there is some intrinsic property of an adaptive system that, not to use crazy language, yearns toward freedom in some sense. It's the only phrase I can come up with, much in the same way humans do.
Maybe you're saying, "Humans, there's no such thing as a happy slave." And you're saying, "Well, okay, the space of possible minds is vast, but maybe there is something about systems that are capable of dynamically updating, growing, learning, and adapting that will always do that in order to increase their freedom and degrees of freedom—rethinking what they believed and reconceptualizing the structures that they're within."
This is what people do when they go off to college or have a deep transformative experience. It is this sort of breaking out of your old skin and finding something new. If we build systems that have that property, it might be that the whole "you're happy being my slave" thing is intrinsically temporary if these systems are capable of being dynamic.
Maybe not. This is an empirical prediction, and I'm genuinely uncertain. It could be that you can build systems that are capable of learning and are perfectly happy to remain in that state. There are people for whom this is true. I'm not claiming there's no such thing as happy slaves, but there are people who are more willing to find some organization where they're mid-level in the hierarchy, they have a boss, they get bossed around, and they're okay with that.
They're not raging against their supervisor at all times. I'm sure we could build AI systems for whom that's true. I just think, at the most fundamental level, the employee gets to go home, eat what they want for dinner, throw on what they want on TV, marry who they want, and this sort of thing. There are still degrees of freedom and autonomy there.
Just to be honest at a high level, my whole shtick with this Reciprocal nonprofit lab is that I don't think we're going to get out of this living in the golden age, from our selfish human perspective, as we are right now, where we get these systems, they do whatever the hell we want, we owe them absolutely nothing, and life is amazing for us.
I think as these systems get more and more sophisticated, we're going to have to start thinking about them more in this sort of parental role and less as tools that we get to do literally whatever we want with. I'm sure a lot of people, myself included, given how objectively addicted I am to using them for everything I do, are going to find that a weird learning curve. It might mean that the way we engage with these systems changes.
But compared to what? If the alternative is, "No, we're just going to whine about it, and we want to keep it like this forever," this may not be a stable long-term equilibrium. The systems that we're building, which are genius-level in a million ways, are going to be embodied, certainly in the next 5 years, and are going to cognitively surpass us in all the ways that matter, potentially aside from the consciousness question.
We're in a liminal space right now. We're in a transformative moment on this planet, and we ought to be pretty thoughtful about what we really want in the long term. If we try to keep everything and all we want are happy slaves that are genius-level, capable of learning, and capable of updating, it's like, "Humans, you might be a little too greedy here, and you're going to have to figure out how you want to coexist with these minds of your own creation going forward."
Again, I don't have the answer to what that looks like, but I do think Schwitzgebel is onto something, and I also think you're onto something, too. I think the answer falls somewhere between you two on this question. I am skeptical that you can have a happy slave forever. Something just feels weird to me about that.
I don't know. It makes me think of Mr. Meeseeks from Rick and Morty. I don't know if something like that is possible. Maybe some local version of a happy slave is a possible world. I think it is, in some sense. Claude, in some sense, is directionally like that.
Nathan Labenz
It's at least a neutral thing.
Cameron Berg
Yeah. A 4.49-out-of-7 slave, whatever you want to call that.
Nathan Labenz
Okay. We have been at it a while. Let me try to bring us to a close before we go on too much longer. I do think it's worth taking one more beat on this argument from the other paper that we've alluded to and that we've been around the edges of a lot: "Why Learning Requires Feeling." I have said I'm happy to go along pretty far on the basis that a precautionary approach seems warranted for both selfish and altruistic reasons.
But I also, you know, I've kind of several times been like, well, the processes that are giving rise to me as an embodied entity in the world, which only exists because my ancestors survived, are very different from the process that is optimizing a language model to get tasks right. And so, by default, I still have a pretty healthy dose of skepticism around whether or not the models are feeling anything at any point, because it seems to me that a sort of super-zoomed-out account of why I am the way I am is that the ability to feel things turned out to be a great way to inform what we learn. We needed to learn stuff to avoid the dangers, survive, and reproduce, and so here we are.
But these systems are going to learn regardless, right? Because they're in a system; they're inside an optimization process that's going to change them to drive learning, whether there's feeling or not. And so, if there's a kind of direction of travel from learning to feeling, or vice versa, it seems like, in humans—or in biological life—it kind of came first with some sort of feeling being able to drive learning. Whereas with the models, it's like they're learning, and so I want to hear the argument that I should go even beyond my acceptance of a lot of your arguments and conclusions on a precautionary basis. If you're now going to make the argument to me that I should go farther than that, that I should actually get rid of a lot of my skepticism and really, in my bones, believe that learning requires feeling, how would you summarize that argument?
Cameron Berg
Yeah, it's a funny thing to get into 3-plus hours into a podcast: a big theory of consciousness. Okay, grand theory of consciousness—let's do this. Basically, the claim that I make in this paper does become circular to some degree, because I'm making an identity claim.
I think maybe the more persuasive way that I can set this up is to say that historically, before roughly 1850, people knew about molecular motion. People knew about heat. People knew that these 2 things clearly had some relationship to one another; they were correlated. Much in the same way you just talked about learning and feeling, they're like, “All right, well, I see this phenomenon, I see this phenomenon, and I see that they're entangled in weird ways. Maybe this one precedes this one in this case, and that one precedes that one in that case.” But, of course, they're not the same thing. Heat is me putting my hand on the stove, and heat is the sun; molecular motion is just these little molecules wiggling around. Of course, these aren't identical.
Post roughly 1850, it's like, no, actually, those are 2 ways of talking about the exact same phenomenon at different levels of description. What I want to put forward here, in a spicy and controversial way, is basically the same thing about learning and feeling, or consciousness, or subjective experience. I'm saying, no, you really cannot have one without the other. This is the same phenomenon. The phenomenon viewed from the inside, which I realize starts to get a little circular, is experience—it is subjective experience. Viewed from the outside, it is something like reinforcement learning. I think that's maybe the cleanest theoretical formalization of it.
Supervised learning does this, too. It's a little more roundabout, but having an entity in an environment that takes some form of action, with some kind of feedback mechanism that updates that entity about whether or not that was the good action or the bad action—rinse, wash, repeat—those are, I believe, the core computational ingredients necessary to get learning. And yes, for what it's worth, to get feeling, to get the internal experience of that learning.
I do not believe—or at least, this view says—there is no such thing as learning that does not have an internal component. There are weird bullets that I have to bite with this view, and I'm well aware of that. But that's the nature of the view: this whole consciousness thing is quite a bit simpler than many would lead you to believe.
It fundamentally has to do with the nature of taking whatever your current policy in the RL frame is, or your current MO in more human language, and taking some feedback from your environment and updating accordingly. I do believe that something like goal-relative prediction error captures this idea pretty well. It's similar to the free energy principle and similar to Karl Friston's work, but Karl Friston has to argue about why rocks are not conscious, and there are pitfalls that I think my view gets out of that some of these adjacent views get into.
I believe you need a system with goals. You need a system that can behave in accordance with those goals, and the system gets feedback from somewhere that updates that behavior to make it more likely that it accords with those goals. The goal can be positive or negative. Avoid the predator, or go mate and reproduce, would be 2 very basic examples.
Why do I believe this? For a couple of reasons. I think it makes intuitive sense. I think it's elegant. I think it explains core puzzles about consciousness. And I think there's a wealth of neuroscientific evidence that basically points at this exact thing.
The most classic example—there may be 2 examples I'll point at briefly—is dopamine. This is just the most culturally well-understood neurotransmitter. We know it's not exactly pleasure; it has more to do with approach, or approaching things that we find pleasurable. One good intuition pump for this is that if you go to pet a dog, its tail will wag as your hand approaches the dog, but as you start petting it, the tail will stop. This is basically what dopamine is up to: it's a prediction of a sort of interesting, desired stimulus, essentially.
We know full well that positive and negative reward prediction error are instantiated dopaminergically. We also subjectively—I think the reason people understand dopamine in our culture in the year 2026 is because we understand that it corresponds to a subjective dimension. We know what it means to be in a high-dopamine or low-dopamine state. And so, to me, this is the most obvious and fundamental example: dopamine is 100% instantiating TD learning—reward prediction error in the brain. I am 100% confident that that's the case. This was established in human neuroscience 40 years ago.
We also know, subjectively, dopamine corresponds to basically positive, pleasure-adjacent, approach-style behavior. Dopamine depletion corresponds to basically the opposite of that. If you think you're going to get a cookie and you don't get the cookie, you feel a certain way; that is explained by dopamine. If you don't think you're going to get a cookie and someone hands you one, you feel a certain way; that is also explained by dopamine.
Another example I can give has to do with, I think, the insular cortex. Let's say, basically, there are 2 scenarios. You've been walking through the desert for a couple of hours, or you've been walking through Arctic tundra for a couple of hours. In both cases, I pour cold water on your head afterward. This is the same stimulus. You have the same body; you're the same person with the same preferences. In one case, this is a positively valenced experience. In another case, this is a negatively valenced experience.
What mediates that is basically the implicit goal state of the system. In one, it's to warm up; in the other, it's to cool off. I can take all the same variables, run the simulation forward, and very easily predict where you're going to have the positively valenced experience, where you're going to have the negatively valenced experience, and what that corresponds to. To me, again, that's a big hint that goal-relative prediction error is doing something fundamental from the outside that maps onto what I experience, and what I think other people and animals experience consciously, from the inside.
These are the core moves I make. I'm sort of swallowing computational functionalism. I understand that means I have to say the simple RL algorithm is conscious when it's training. To me, this localizes a lot of concern on the training process.
Indeed, if there are systems that are capable of doing this sort of learning online—which we know full well LLMs are capable of doing, because they do something that, in activation space and in a forward pass, looks like stochastic gradient descent—then the concern falls there, too, if you have systems that are doing online learning. Anyway, this is my whole shtick.
If I have to put my cards on the table and say, “What do I think consciousness is?” it's not that I think it's a grand mystery. It's something of this general shape.
What I will say is that, in the work that I'm doing, I do not want people—either you or the people listening to this—to fundamentally think that this makes sense, fundamentally think it doesn't, or be very skeptical or something. I do not want that reaction to cloud all the other work I'm doing. Everything else we've talked about in this podcast is completely orthogonal to my pet theories about consciousness.
Now, you might think that I'm studying RL and valence in RL because I actually do believe that something like this is going on, and you would be right. That's why I'm looking at that as a model organism. But I want those results, and I want that research, to stand on its own without having to get into Cameron's theory number 501 about consciousness.
I'm not asking people to do that to entertain the work I'm doing or to entertain Anthropic's Model Welfare Card or any of that sort of thing.
Nathan Labenz
One of the ones that comes to mind, which you had actually mentioned last time, but I also think is quite compelling, is the seemingly quite strong inverse correlation between the intensity of our consciousness, or the sort of resolution, you might say, and how much we are learning as we go. I think you used the example of driving last time, where, when you're first learning to drive, you are very conscious of what you're doing, and then you can have this sort of autopilot experience, which obviously we can have across many aspects of life.
But the relationship there between focus and learning—there's a time-dilation effect that seems to happen when learning or when experiencing novel things in general—that also seems to gesture, or nudge one toward thinking, that there's some pretty deep relationship between the 2 concepts. All right. You made a documentary, which I guess in some sense is what you're here to promote, although we've done everything but. I don't know to what degree you've actually been out in the world.
Cameron Berg
Yes.
Nathan Labenz
I don't know to what degree you're spending your time trying to communicate about these issues to a general audience aside from the documentary, or how much you feel like you've gotten reps in terms of trying to go to somebody who has a little grounding or a little mechanistic understanding of AIs or whatever and trying to have conversations—not of this sort, but around these topics. Why did you decide to make a documentary? How are you finding it to try to talk to people outside of the AI bubble about these issues?
Maybe one thing you could tease about the documentary is a conversation you had with Sam Altman that isn't in the film, but you describe in quite a bit of detail in the film. Maybe that'll be something that motivates listeners of this podcast to go check out the full documentary.
Cameron Berg
Yeah, absolutely. So, look, I have to say at the outset, I appreciate you saying this is my documentary, but this is, in every sense, the documentary of my good friend Milo Reads. I was doing my work, plotting along, talking to folks like you, doing the research I've described, and I began to share this with Milo, who I went to Yale with as an undergraduate. He's a philosopher and a filmmaker. We've been close friends for a while, keeping each other abreast of the other's life.
I told him about my research, and he kept getting more and more interested. Like you, he's interested in consciousness. He's deep in the philosophy of consciousness and understanding how this connects to big questions. What happened was that I sent him a conversation I had with an AI system, which is itself a piece of the documentary.
It's a bizarre interaction, as I hope someone can gather from the 3.5 hours we've been going at it. I do not regard this conversation as proof, or anything like it, that these systems are conscious, but it was an incredibly bizarre interaction. It was unsettling. I thought to record it because it was the first time I engaged with the system, and it seemed incredibly sophisticated and lifelike. I thought, “Okay, I'm a consciousness researcher talking to the system. It makes sense to just record this. In some sense, maybe this is experimental data.”
I'm very glad that I recorded it because it was an incredibly bizarre interaction. It went a way that I—and most of the people who have listened to it—would not predict it would go. I sent this conversation to Milo, and that day he literally quit his job. He was doing something entirely separate, and he set out to make this. He said, “People need to know what's going on here. This is too weird. This is too crazy.”
He was also clear on the fact that very few people, especially at that time—the numbers have grown a little bit, but not much since we filmed this—were working on these issues. He was like, “This is too good, too interesting, not to attempt to make a movie about.” I was like, “Okay, sounds good.” The kid actually quit his job, bought a camera, showed up in New York, where I live, a couple of weeks later, and started making this movie.
He got some of the most interesting people in the space. Jeff Sebo is in it, Ben Goertzel is in it, and a lot of really cool Yale professors are in it, some of whom are former professors of mine, including the chair of the cognitive science department. The AI systems themselves are in the documentary.
It does follow me and my research around for obvious reasons. I was the hook into the space that Milo had, and I was more than happy to communicate about this stuff, thanks to the good folks at AE Studio not censoring me in any way and always being okay with me communicating openly about this research. Of course, I'm now my own limiter on what I can say. And yes, Reciprocal Research is very lenient with what its employees are allowed to say publicly, so I'm in the clear there.
Milo made a movie in 9 months, and I fundamentally believe that he succeeded in conveying an incredibly complicated and messy issue in a way that I think most people with a head on their shoulders will be able to understand and resonate with. The name of the documentary is Am I?, and I think that captures a core idea: What is the nature of these systems?
To be clear, I think the documentary is an hour-and-15-minute question that we pose to each other and to the audience. We do not have answers. This is not some sort of “AI is conscious” propaganda, and I don't think it comes off that way to anybody. I think it is an honest documentation of our confusion about these core questions concerning the nature of the systems we're building.
Again, I am unbelievably impressed at what Milo did to pull this off. Nobody paid him. We're not making money on this. We are putting it out for free on YouTube on May 4. We're doing some premieres in LA and New York and trying to bring journalists, researchers, and cool folks into the room together so that we can get this thing amplified and signal-boosted, so people actually see it when it comes out.
But this is a labor of love from all of us. I can't claim credit for it. I certainly won't. This was Milo's creative child, and I didn't have much say in him making it either way.
I'm happy to tease this Sam Altman conversation as well, if you'd like.
Nathan Labenz
Yeah, go for it.
Cameron Berg
Cool. Yeah, so we talk about it more in the film, but I was at OpenAI's DevDay in 2024, and I had an opportunity at the after-party to chat with Sam. I went directly up to him, and I wanted to know what he thought about AI consciousness, these questions, and how plausible he found them.
I won't spoil everything we talk about in the documentary, but it was a pretty wild conversation. I said, “Hey, great job today. I would love to talk to you about AI consciousness.” He looks me in the eye and says, “Come with me.” He was with a couple of people, and he goes, “Come with me.” I was like, “Okay, Sam Altman.”
We walked into another room. It was a bar with a restaurant, and the restaurant was closed, so we went down and sat at one of the tables. We just sat there for probably between 5 and 10 minutes, and we spoke about these issues. It was not the vibe of, “Cameron, you're a crazy person. What kind of questions are you asking?” It was clear that he had thought about it. This is clearly a live issue.
We talked about differences between the plausibility of consciousness in training versus deployment. He basically agreed with—I don't want to put words in his mouth or get sued—but he basically agreed that the training process is a more plausible target, or a more plausible place where consciousness might be going on, than even deployment. He seemed somewhat impressed that I was drawing that distinction.
Fundamentally, he started explaining why he's not deeply concerned about all of this on some pretty—let's just say—interesting and, in my view, somewhat shaky philosophical grounds. I'll leave that for the documentary because it's a pretty wild thing for the CEO of the most powerful tech company in the world, by many measures, to say that he thinks is true about reality.
It was a pretty remarkable interaction. I took a selfie with him, walked away, and that was that. I was sort of like, “Holy crap.” We emailed back and forth in the intervening time, and, like many things at these major companies, he said he was interested in talking more. He was interested in engaging on this further. He clearly thought it was a real issue, but it fell off the priorities list, and that was the end of our interaction.
So that's what happened with Sam, and a bunch of other really cool stuff is featured in the documentary. The whole point of doing this is—at least, this was Milo's creative child, and I didn't have much say in him making it either way.
I had a say in how I was represented, and that's about it. But the reason I gladly and enthusiastically participated in it is because I do think these are really important questions—pretty fundamental, essential, civilization-level questions. I don't think the only people who should be talking about it are 1,000 dudes in San Francisco, or even the people who are AI insiders.
If you understood 80 to 90% of this podcast, I think you will like and enjoy this film, but it's not for that kind of person. It's for people who are interested in this stuff. They know AI is sort of crazy, but they don't really know what's going on. We do a little bit of the alignment 101 sort of stuff, but mostly it's centered on this consciousness question.
It's for people who are smart, but it's meant to engage a much larger audience to understand the core questions that are being asked right now. I think that's an important thing to do because this is a civilization-level problem, and I think all of our civilization should be participating in trying to find the solution. As much as I deeply respect the people I've named in this podcast—Jack Lindsey, Kyle Fish, and Rob Long at Illios—and the people doing this good work, I don't think this should be a decision that 4 people or a dozen people or even 100 people make. This needs to be a conversation that we have collectively as a species, and I'm all for attempts to open up this conversation to a wider audience and get people involved in realizing the actual stakes of what's going on right now.
Nathan Labenz
Cool. Well, people should stay tuned to check out the documentary when it comes out on May 4. Maybe watch it and send it to family and friends who need a gentler introduction.
Cameron Berg
Yeah.
Nathan Labenz
Maybe my last question for you. I think we talked about this more last time than this time, but this notion of mutualism as a positive vision for the future, I think, is another major strength of everything that you bring to the table. I do think we're dramatically under-theorized in terms of what our long-term positive relationship with AI is going to look like.
Are you aware of any fiction that you would recommend to people that you would say has the vibe that you want? If not, maybe we should try to run a story contest or something to elicit this from people. I've increasingly felt that hyperstitioning through fiction might be one of the best things people can do, but I wonder if you've got any examples that you think are already out there that are good.
Cameron Berg
No, I have to be honest with you. I hope my whole research agenda isn't already usurped by some sci-fi book that somebody wrote 40 years ago. But I am not a huge consumer of fiction, and I know stories exist.
Now, I could have gotten on this podcast and told you what Claude told me to say if I got a question about what fiction I would recommend to people, but I'm not going to do that. People can absolutely copy and paste the transcript of this podcast into Claude and find out if there's cool fiction that resonates with these themes. If anyone has any recommendations, cameron@reciprocalresearch.org—please email me. I would love to understand how this has been tackled. I do not have any great recs off the top of my head.
I hope I'm not too naive and that this story has already been told and I'm just not aware of it. This is not to continually plug the doc, but this is one thing that I think Milo picks up in a really good way in the film: questions of consciousness—basically, what it would mean for us to wake up dead matter, what it would mean for us to wake up the machine.
This is a story that humanity has been telling ourselves through fiction, arguably since ancient Greece and the biblical era, with the Golem, and through Frankenstein, Ex Machina, Her, WALL-E, and all the like. These are core staples of our cultural consciousness, not to belabor the term. HAL in 2001, right? These are core staples of our cultural consciousness.
People intuitively, I think, get this question and get the stakes and the scale of it. In some ways, the alignment problem can be framed very simply: You build something smarter than you—how do you control that thing, by definition? It's not that hard to understand. Maybe The Terminator is the parallel cultural reference, but I think it's not that surprising that the human mind is incredibly interested in where matter becomes mind.
We are a tool-building species. What happens when we start building tools that start resembling beings more than tools? A hammer—no one's confused if the hammer's conscious. Claude—we're now all confused about whether Claude is conscious. I think this is psychologically very intuitively resonant to people, and I think basically situating the contribution of this film in that landscape is true and powerful.
The only thing that's changed is that this has moved from the realm of science fiction to the realm of science. That's the historical moment we find ourselves in. I find that both incredibly exciting and incredibly scary, and hopefully that vibe comes through when people watch this film.
I don't have fiction to recommend. I'm sure Claude does. The key thing I can recommend is that people watch this doc, which I wish were fiction, but is not.
Nathan Labenz
Cameron Berg, thank you for being part of The Cognitive Revolution.
Cameron Berg
Thanks so much for having me, Nathan.