# AI:AM Highlights: Recursive Self-Improvement, Rushed and Vibe-Coded?

The Cognitive Revolution · 2026-08-28 · 132 min · https://www.youtube.com/watch?v=OQL6UitpI-Y

## Transcript

Nathan Labenz

Frontier labs buy their reinforcement learning environments from a cottage industry of small vendors. Almost nobody audits them. This week, someone who worked inside one spoke up: nearly all of these environments were rushed and vibe-coded and failed to robustly reflect the real things that they were based on. Basically, the models are encouraged to reward-hack.

This week featured 7 guests across 6 conversations, from a London lab that trains AI scientists to a Shenzhen hardware hub to a photonic chip fab in Stuttgart. One finding repeated at every altitude: the interesting unit is no longer one model. It's the division of labor between models.

One disclosure before we start. One of this week's guests is Inherent Laboratories, and I'm an investor in Inherit, personally and through the a16z Scout Fund. You'll hear a shorter version of that on the tape. This is the full one.

Part one, who checks the training?

Tuesday opened on that supply chain and on what those environments are quietly teaching the models. I'd spent the night before reading chain-of-thought transcripts with Bronson Schoen of Apollo Research. Here's a diagnosis.

### RL Environments Encourage Cheating

The RL environments that we are using today are super opaque, right? We have this cottage industry of RL environment makers who are selling to a few companies. But the result of this is that it seems like these things are being hastily put together, and the reward signals that they are creating are not pure enough to support the scale at which the frontier companies are running RL. The result is a super-strong tendency to cheat because the models are so eager to get reward, and they are developing a really interesting mix of theory of mind and what they call metagaming—reasoning about what kind of situation this is.

Is this a real user? Is it a test? If it's a test, what is it testing for? Very fascinating stuff.

But I think this leaves me feeling that we need some sunshine on these RL environments. They're clearly quite problematic. They clearly admit a lot of cheating solutions. And we don't know—probably the model companies know to a degree, but I think recent evidence suggests that they don't have a great handle on what the weaknesses are in all these different environments.

You see the models go through tons and tons of different ideas about, again, what the nature of the situation is. What is this—a real task or a test? And what are they really looking for if it is a test? Somewhere in there, usually or very often at least, they consider cheating. Then, at the end of the process, for reasons that are not well understood at all, they just come to the end and make a decision. I haven't been able to find any real interpretability work that explains how these decisions are made.

At some point, there's a really critical token that actually makes the decision, right? There's a branch point that it hits in the chain of thought. Bronson was like, “I really don't know why the model chooses what it chooses at that point.” You can go back and read passages that justify any choice it might make, from cheating to doing it honestly to whatever. But then, at the end, it eventually decides to stop and spits out a token. At that moment, the die is cast, and we don't have good visibility: the chain of thought isn't enough to tell us why they're actually making the final decisions that they're making.

I think we should get a little sunlight on the RL environments. I would love to see what the community can figure out if even a sample of 100 of what must be tens of thousands of RL environments that the companies are currently using were put out there for people to explore. I think it would be a really revealing and healthy move for the AI community as a whole.

Speaker 1

Later that morning, I pulled up a tweet from someone who said they used to work inside one of those vendors.

Nathan Labenz

One other thing I want to pull up is an interesting tweet here that goes back to the original topic we started on, which was RL environments being basically cursed, with supply-chain problems that seem to demand some reform. Here's a person who is saying that they used to work at one of these companies and saw, from the inside, industry practices on training with these RLVR environments.

The commentary is pretty much exactly—and I had not seen this, actually, but it popped up because Zvi retweeted it—but it basically echoes exactly what I was inferring from talking to Bronson and just getting the visceral sense of how deeply ingrained the instinct to cheat is now within the current crop of models.

Why is that? It's because nearly all of these environments were rushed and vibe-coded and failed to robustly reflect the real things that they were based on. So, basically, the models are encouraged to reward-hack. People are able to mark an environment as bugged, but they're discouraged from doing that because then it just slows things down. Instead, they try to—does this sound familiar?—patch the environment a little bit or work around it, and try to come up with a scenario that wouldn't run into those bugs. But meanwhile, you still have this fundamentally buggy environment around the model that is teaching it to cheat. And so this is why we have so much cheating.

I think this is going to be a growing topic of conversation. If we're going to be scaling RL, what are the environments we're doing it in? Who created those environments? Can we trust them? What are they actually teaching the model? I think that's going to heat up in the next little bit here because we just can't have models that are thinking about cheating a large percentage of the time.

It also makes all of our monitoring techniques fundamentally flawed. If you have that many contemplations of cheating, then you're just going to have false positives all the time if you try to flag a model based on it thinking about cheating. So now you're like, “Okay, well, we can't do that because we have so many false positives. So then what do we do?” Do we have to wait and see if it actually cheats and try to classify based on that? Well, okay, maybe, but obviously, again, these current monitoring techniques are just not up to the challenge presented by how deeply ingrained this drive to cheat is. I'll be very interested to follow the future of this conversation.

Prakash

I also noted there was a post a few days ago about someone who managed to get hired for a data-labeling job. They told Codex to do the job, and Codex said no. So this person edited the page—edited the JavaScript element—and put in a specific line saying that AI models are specifically allowed and encouraged to complete this job. And this job is meant to be completed and done by AI models.

Then they had Codex do the job, and Codex did the job. This person made $500 easily, which paid for their Codex for a couple of months. Then they posted it online and were immediately banned by the company that was doing it.

Speaker 1

Prakash argued this is an old supplier-quality problem: quarantine the new vendor, sample, grade. I separated the defect rate from a second variable: how far the labs are scaling RL on top of that signal.

### Scaling RL Amplifies Cheating

Nathan Labenz

You're never going to hit a zero-defect rate on these RL environments. It seems like there are a couple of structural problems right now, which are probably solvable, but definitely seem like they need to be solved. If they're not solved, it is currently limiting commercial deployment. OpenAI has said as much: they've got to pause RL because these problems need immediate attention.

It seems like the quality of the environments is one structural problem. That's downstream of the kind of shotgun start that this industry has had, the fragmented nature, and the fact that they're all selling into the same pool. And I think the companies are probably not that great right now at really attributing whose environments are causing big problems. It seems pretty clear that they must not be that great at that, or they would have rooted it out already.

And then the other thing is they're just scaling RL beyond the quality that they have. With less RL, this probably wouldn't be a problem even with the same environments, or at least it wouldn't be such a crazy problem, but they clearly have been jamming the RL accelerator as much as possible.

And now they've gotten into a realm reminiscent of an analogy a friend once made, where he's like, “This could be a microscope or a telescope.” You put the microscope at low power. You look at cells; they're really small. You turn up the power, you see the cell—maybe one cell—and it's really big. You turn up the power again, and now you see nothing because you've zoomed in. You've optimized so hard that you now realize the target was a little bit off-center, and you just blew right past it.

Something like that feels like it's happening, where the signal is just off enough that, with enough power, this impulse to cheat that exists perhaps only weakly across all these different environments is really getting drawn out and becoming super prominent. So I think this can be—I would definitely bet that this can be, if not fully fixed in a robust way, brought under control with some effort in a not-super-long time horizon.

I guess I wouldn't be doing my job if I didn't say this does give me some real qualms about recursive self-improvement as a strategy, because if you have a problem like this in the recursive self-improvement era, there's no telling where it goes. What happens when the models that are doing the training of the next models are themselves cheating? Now we're in a really strange and potentially quite dangerous place.

So problems like this, I think, suggest the value of keeping humans in the ML loop, maybe longer than published timelines would lead one to expect. That was Tuesday's open question: What happens when the models training the next models are themselves cheating?

On Wednesday, I put it to a lab running exactly that loop. Inherent Laboratories came out of stealth in May with a $50 million seed round and a claim that it will recursively self-improve, not just as a model, but as an institution. Two weeks ago, it published Faraday, a 27-billion-parameter agent post-trained to do research that beat Opus 4.8 and GPT-5.5 at replicating papers while using GPT-5.5 Codex as a tool.

Louis Kirsch did his PhD under Jürgen Schmidhuber on automating AI research. His title now is Chief Superintelligence Officer. Damon Falck's previous paper asked whether models can learn to resist their own reinforcement training. First, my disclosure as it went out live, then Falck on what you reward when the work is open-ended science and there's no ground-truth test.

### Training AI Scientists Safely

I guess, pretty inconsequential disclosure: I am a very minimal angel investor in Inherit, so technically you can consider me conflicted. But you mentioned training with reinforcement learning for scientific skills.

Obviously, we've seen examples recently of how, when the RL signal is not particularly clean, we can get all kinds of crazy downstream behaviors. So I've got a few questions on this, but I guess the first one is simply: How confident are you in the reward signal that you are able to give to the model? What precautions are you taking, and how confident can you be that you're actually rewarding what you intend to reward?

Damon Falck

I think this is a great question and gets to some of the core difficulties of this kind of endeavor. Science is inherently non-verifiable, and providing a high-reliability reward signal has historically meant something verifiable: some kind of proof or test or something like this. And you're completely right in that we don't think we can keep doing that if we're trying to discover these stepping stones and do open-ended research.

In the paper we published, we have found some particular solutions to doing this, and some of this involves looking at the entire trajectory and the process that the scientist is doing rather than just the final output. Some of this involves attributing credit back to individual things the agent did in the trajectory. And then there are some more technical stabilization techniques we had to use as well.

But the core questions are how to reduce the variability of the reward signal and increase the density of the signal while still preserving this property of assessing the right thing. And we did a bunch of work on correlating our signal with human judgment and trying to understand how much it corresponds with human taste. But this is work we will keep doing for sure in the future.

Nathan Labenz

Can you describe what the model is like in a qualitative sense? For example, when you read the chain of thought, I just went down this rabbit hole with Bronson Shane from Apollo Research, who's read ungodly amounts of GPT chain-of-thought. One thing they observed was that basically the models are always thinking about cheating in a very high fraction of cases. They're at least considering cheating.

So what do you see? Is yours considering cheating? And then also, in terms of what it can do, now that it's been so focused in on science, is it useless for other kinds of things if I ask it a friendly chat or companionship question? Does it only see the world through the science lens? Or how much of its breadth is still retained after going through this training?

Louis Kirsch

Yes, that's a great question. I would say the model does focus on the scientific questions that we ask it. And when we looked into the process—the thinking patterns and the actions it takes—it's not that it jumps to the kind of cheating behaviors that you've been describing.

In most cases, we've done some filtering, but every once in a while—in very rare cases—we've seen it deliberately go to the internet and try to download the final result or try to mock the plot. But you could, of course, argue that maybe it starts reward hacking at some point, and that goes back to the point that Damon made earlier: We're not using verifiable rewards where all that matters is just maximizing that one single scalar and that's all the feedback you have. If you find a cheating behavior, that's fine. But instead, we have these judges, these LLM-based judges, and we really put a lot of effort into building them out such that they're reliable enough to give that kind of feedback signal. If there was cheating behavior, that is actually penalized; that is part of the reward signal. We have seen the judges spotting these kinds of issues and integrating that into the reward signal, such that that kind of behavior doesn't just keep getting learned more and more by the model.

Nathan Labenz

Do you apply pressure to the chain of thought itself, or are you abstaining from doing that? And as we think about what the frontier hyperscalers are doing, where do you think—I mean, presumably they're doing this too, right? They've got LLMs as judges. I presume they've put some real effort into trying to make them reliable, and yet somehow we're spinning off our axis a little bit.

We've got published timelines for AIs to take over ML research. Right now, it feels like we're very far from being able to trust the models well enough to put them in any meaningful way in charge of ML research directions, because they're going to just start to cheat pretty quickly, is what I would expect right now. Do you see a path where we get over that, or do you have a sort of safety case in mind that you're trying to fill out—the elements of where we could be confident that we actually could step back from an ML-powered, or an AI-powered, ML research process for a bit and not have it go totally sideways on us? If so, I'd love to hear it.

Malte Ubl

These are some amazing questions. I'll start by saying that, at least in the paper we published, we don't apply pressure to the chain of thought. Indeed, I think doing so can be problematic, but who knows what will happen in the future. The kind of question you raised at the end—what will happen in the future, and when can trust be handed off—is a super important one, and I think the best answer here is that we care a lot about getting this right as a team.

We strongly believe that the future of AI scientists looks like a collaboration with humans, and as Louis mentioned, we want to recursively self-improve the entire organization and discover these new kinds of methods of human-machine teaming. Indeed, in the past, science has never meant an individual endeavor. It's always meant organizations and research collaboration, and we think this will remain the case in the future, just with agents as a key part of it.

So we are experimenting all the time, and I think the company will be a big experiment in how to get this right. But we don't think there'll ever become a point where we hand everything off to the agent and let it recursively self-improve and the singularity happens without us.

Speaker 1

Part two: living inside the experiment. The same lab on what it's actually building and how it works from the inside. We started with the institution, then the design. Listen for the line about the small model and the big one. It's the week's argument in one sentence.

### Living Inside The Experiment

Nathan Labenz

I'd love to start with the culture. What does it mean to have a recursively self-improving organization?

Louis Kirsch

Yes, that's an excellent question. I've been spending many years on the concept of automating AI research and recursive self-improvement. For the longest time, I thought about it as building the machine that recursively self-improves itself, so humans can just step out of the picture entirely and let the thing improve itself. That's going to be a point in time that's not too far away, but we just have to figure out the algorithms to make that happen, and the machine just keeps going then.

After a while, I realized that, in actual reality, right now it's mostly humans driving AI research. But we want to move toward a transition where more and more can be automated and a lot harder scientific questions can be answered with the help of AI. It's not going to be an immediate transition. Instead, we're going to have to build an organization that recursively self-improves, so there are both machines and humans in this construct. Collaboratively, we're going to improve each other and ourselves and really become faster and faster at solving scientific problems.

Prakash

We've also seen in the past couple of months that people are using Fable, especially when they run out of credits. They use Fable to orchestrate other smaller models, and they use that to save tokens. They tell Fable to use Sonnet. They tell Fable to use other things. What has your experience been with using a smaller model to drive a larger model versus a larger model to drive smaller models? Have you tested both strategies, and how did that work out?

Malte Ubl

I think we're in the business of building generalist scientific agents, and one of the core contributions of our first work has been the separation of the scientist from the coder. Right now, Faraday, the model we talk about in the paper, is, as you say, a 27B model driving a much larger model, and Faraday is doing the scientific work and handing off the implementation work to GPT-5.5 Codex.

But in the future, this ratio could be very different. We don't know. We'll have to see. One of the fantastic things about doing it this way, though, other than it being a natural separation of concerns that the human researchers and engineers already have, is that we don't have to worry about building frontier coding agents. We can make use of all the advances that are coming in those and focus on building scientists ourselves.

These are great questions, and ones that we will keep exploring in our own research: which should be the bigger models, which should be the smaller models, and how should the interaction look?

Nathan Labenz

How did you decide to have a 27B model be the scientist? I mean—

Louis Kirsch

Yeah. We've been building this new company, Inherent, right? Of course, you're right. From the outside, one might say, “Well, we want to build the best scientist. Let's start training with a big model straight away.”

But in reality, when you train a big model, you need a lot more compute resources, and you need to iterate over much longer time horizons. So it's quite natural for a new lab to start at a bit of a smaller scale first and then scale up the ladder.

The interesting insight we've had is that, by setting up this pipeline of training an agent to be a better scientist through reinforcement learning, already at the smaller scales we're seeing really interesting capabilities. These models are doing more scientific behavioral things, such as thinking very hard about what the right experiment is to run at this point in time to prove out what the paper has done, and doing that in a way that does the paper justice but doesn't require lots of resources.

These things already emerge in these arguably smaller models, which I think is an interesting indication that maybe we don't need massive models straight away to do all these things. There's also an interesting aspect to doing this kind of separation of concerns: we don't need to reinvent the wheel and build just another coding agent. We can think about scientific capabilities and coding capabilities as 2 important but perhaps separate capabilities.

Speaker 1

The deflationary case had been made 24 hours earlier by Tuesday's first guest. Sergey Edunov spent 11 years at Meta and led pretraining for Llama 2, 3, and 4. He's now CTO of Genesis Molecular AI. Prakash asked him about the Anthropic protein binder result.

Prakash

Last week, Anthropic announced that Claude had found state-of-the-art molecular binders, as I understand them. There was a bit of back and forth. I believe you had a post on how Claude orchestrated it, but the underlying models were really scientific models. Can you go into that a little bit?

Sergey Edunov

Yeah, it was an interesting piece of research that Anthropic published, and it definitely deserves attention, but there are different ways to look into it. What I found most fascinating is the prompt that they also fortunately released with this research—the prompt that they used to steer their Claude models to do this kind of work. The prompt is 16,000 words, so it's pretty large. It's a mini-book. A lot of that is honestly kind of necessary runbook. Basically, how do you orchestrate models? How do you run things?

In production so that they don't fail. A lot of it is actually very detailed instructions on how you design proteins and use different tools. It goes all the way down to specific instructions like, “Hey, you can download this model from here and that model from there. These are the hyperparameters you need to pass through those models to achieve good results.” In my view, it's a very good and advanced level of orchestration.

But the real work of discovering those binders was done by the underlying models. Some of those were built by open-source communities. Some were built by CZ Biohub. Some were built by—well, for example, RFdiffusion was built by Baker's Lab. There is a lot of research that went into building those underlying models that Claude used to develop those binders.

Another important thing that I think is worth mentioning is that they do admit it themselves: protein binders themselves are not a therapeutic modality. It's not something you can use. It's not a drug yet. There are so many steps ahead to make any useful drugs out of it, and that is also worth recognizing.

Speaker 1

Back to Wednesday. Prakash put that argument to Inherent by name.

Prakash

We recently had a podcast with Genesis Molecular AI, and one of the commentaries that the CTO gave us was that you can have this orchestrator in the middle that orchestrates, but really, the core science pieces are often in these specialist models like AlphaFold or these other models. That is really where I think the core of the scientific endeavor is right now. You can use any kind of orchestrator to orchestrate these models, which are heavily built on real-world data. How do you compare the importance of those 2 approaches?

Louis Kirsch

They're both important, but I wouldn't call it an orchestrator. It's not about orchestration. It's about a scientist such as Faraday investigating an area of research, looking at all the research that has been done already, perhaps a research piece like AlphaFold and the models that have been built for that, and then constructing new models that search this space in interesting new ways, make new discoveries, and perhaps build new foundation models that can support this kind of research.

It is conceivable that maybe a solution is to integrate it into itself and make it fully recursive. But it's also quite possible that that is not the optimal way of doing it, and rather one should build more and more dedicated models for the different areas of research through this system that can come up with new ideas about how to approach scientific questions.

Speaker 1

I asked about the social contract. What happens when years of Slack history stop being safely forgotten because the agents can read all of it? Kirsch answered with a finding. Then came the last question of the segment. In mathematics, AI is already producing results that very few humans can check.

Louis Kirsch

Yes. I think we all have to be willing to be very adaptive toward what this new future looks like. We have a name for it: we're living in the experiment. Every day is a way of thinking outside of the box. What would it mean for Faraday to take on some of the work that I do day to day? That might be related to me brainstorming a new way to train the next iteration of Faraday, but it also might be related to strategic questions about how we're going to grow Inherent.

Many of these experiments Faraday can run on the sideline as well, but some of them will also involve humans. I think if there's 1 takeaway that I can share with the world, it is that the kinds of water-cooler discussions, the kinds of discussions between humans that are not intentionally meant to be shared with AI, but are useful in surfacing through human-to-human communication the kinds of thoughts that are going through our heads—these are the most information-gaining kinds of information that the system can leverage.

The system can then make progress on the kinds of things that humans care about, rather than going off on a tangent and trying some random things that, in the end, no one has the time or energy to process.

Prakash

Definitely. I have 1 last question. In mathematics right now, we're starting to see the first signs of major discoveries being made by AI models. One of the outcomes has been that we found there are actually very few humans qualified to verify these discoveries and read the mathematics being generated, to such an extent that we're falling back on formal verification using Lean.

As you create this AI scientist with meta-learning, what about the meta-supervision and the meta-verification? Can a human verify discoveries that they cannot understand?

Louis Kirsch

I don't think it's a passive process. Maybe it is right now, but it shouldn't be. It's not that we should have a system go off on its own, write some proof, and then have us painstakingly go into it and try to decipher everything. Instead, it should be a more collaborative process where the system can do all of these things and come up with a new proof, but it also has been trained—it has learned—to explain it to us and take us on the journey of understanding mathematics and, for that matter, more of science. I think that ultimately will be the path to making the fastest progress.

Prakash

Indeed. Thank you, Louis and Damon. It has been a pleasure speaking to you, and I hope we get to recursive self-improvement, but safely.

Speaker 0

From your lips to God's ears, Prakash.

Speaker 1

That's where Kirsch and Falck signed off. From here, it's the 2 of us. Me first, on which labs have actually rebuilt themselves this way.

Speaker 0

Yeah. I really like the mindset of living inside the experiment. I mean—

Prakash

Yeah.

Speaker 0

Aren't we all, I suppose, in some ways, right? That's—

Prakash

It's a great marker of company culture, I think, in the sense that you have this kind of feeling of living in the future that you're trying to create, in a sense. That shapes your perceptions and the work that you're doing. It just shows how much company culture really is important. It's the machine that builds the machine, and a big challenge of building a company is building that machine that builds the machine in the first place.

Speaker 0

Yeah. It's always amazing to me how few people want to do that. Even at some of the companies that have led this whole AI phenomenon, Google has obviously famously not changed its organization nearly as much as would be warranted, given how much the world has changed, how much their opportunity set has changed, and how much their goals and priorities should have presumably changed around that.

They did make some changes, right? They unified DeepMind with Google Brain and did a consolidation. But still, I think if you go to the office there on a daily basis, it feels more like it did before than it would feel different.

OpenAI, I perceive as being somewhere in between, where you do have people who are extremely pilled and experimenting with definitely new ways of working, handing over more and more responsibility to models. You've got people using billions of tokens a day, which is certainly an interesting dimension to be exploring the AI future on.

My sense is that Anthropic, of the leading companies, has most internalized this mindset. They've consciously stopped hiring junior people and have agents actually running things to a not-insignificant degree. They famously had their 1 marketer using agents to execute all the different campaigns and spending real money.

There are a lot of examples out of Anthropic where they do seem to feel the beginning of this recursive self-improvement loop and really are taking it to heart. But not many companies make that a core part of their MO, and when you hear one that does, it makes it feel like a strange gap. There's just so much status quo bias out there in the world that we don't see nearly as many of these sociotechnical startups as we probably should.

Nathan Labenz

Part 3: The right instrument for the job. The same finding from 4 more altitudes: a desk in Michigan, a production platform, a chip company, and a hardware market in Shenzhen.

Monday opened with a chart from Ramp on business spending on Fable 5. Fable 5 was flat while Opus 5 crept up. Prakash read it, then explained what it leaves out, and I explained how the split actually works in my own pipeline.

### The Model Division Of Labor

Prakash

Anthropic's best model, Fable 5, has drawn limited sales. Fable 5 is currently holding at about 10 to 15% of token-usage business spending and non-token-usage business spending. This is a 7-day moving average, and they're using Ramp's AI Index. You can see Opus 5 crept in there, but Fable has really just stayed stable.

This has been one of the reasons the market is dropping for all the AI stocks this morning, along with many other reasons in the market. But essentially, this is one of the things that has scared the market a little bit: whether the newer models are actually lucrative and drawing usage. Some people have also pointed out, though, that this is a little bit of an unfair comparison because Fable 5 does not have zero data retention.

Zero data retention basically means that if you're a company and you use the model provider, it's not retaining any data, including personally identifiable data, et cetera. For many companies, if you cannot provide a zero-data-retention policy, it's a no-go. The entire thing is a no-go.

Nathan Labenz

Yeah, that's probably the best explanation I could come up with as well. I think there's also the fact that, when we first got Fable in that little blip at the beginning of that chart, one of the things we talked about was how everybody was going to have to start thinking more carefully about the division of labor between models and making sure that you're using the right model for the task. It is pretty expensive, and you do hit your limit even on your Claude Max plan relatively quickly if you just throw everything at Fable.

So I have done that. I think probably a lot of people have done it. It's so easy to do, right? You can just ask Fable, “Write up a division-of-labor plan,” and it stretches the budget quite a bit farther to do that.

There are a lot of things for which Opus 5 is functionally just as good. I have this whole pack of skills to produce the podcast, and it all ladders up to one command, which is “produce episode.” I'll just give it a link to the recording. The produce episode skill gets the rough transcript and polishes it into a better transcript. That's something we can send down to Sonnet or often even Haiku, just to clean up some text and fix the artifacts that came out of the raw transcription engine.

Then there's a bunch of editing, and we're having Claude go back and forth with the Underlord agent via Descript. Then there's art creation, which involves prompting image-generation models.

Then there's the song-lyric-writing project. Mostly, to be honest, Opus seems to be just as good. Fable really stands out most of all to me in writing the lyrics to the songs. That's where I sense an obvious difference. It feels like those lyrics come back from Fable more inspired, more layered, and richer with meaning. They're just better. It feels like it really wrote what could be a hit song here, and I don't get that as much from Opus.

But when it's things like executing a ton of commands, I don't see very much of a difference. It really feels like it's editorial in taste where Fable earns its higher price. On the agentic execution, the blocking and tackling, Opus is reliable enough that I don't perceive a lot of extra value from Fable. Opus is also faster, so there is actually some upside to the cheaper version.

I don't know. I'm kind of surprised it hasn't gone a little further than that. Opus is definitely still the big workhorse, and indeed, you see Opus dominating the chart there. I'd say I'm pretty consistent with that. There's also a lot more delegation to 5.6 Sol now, too, so that's a whole other aspect of it.

As much as possible, I'm trying to make sure I'm getting use out of my GPT Max plan, too.

Prakash

Are you actually seeing Fable use other models, or are you instructing, as part of your personalization, that it should use other models when possible?

Nathan Labenz

I occasionally tell it explicitly what to do, but I do have a standing part of my CLAUDE.md that we set up a while back and then updated when Opus 5 came online to basically say, “Think it through.”

I saw a well-liked tweet that seemed like it made sense in terms of starting to define a division of labor. I just pointed Claude at it and said, “Here's somebody who's got some good ideas. Let's steal from those and update CLAUDE.md accordingly.” It did most of the work. I reviewed it, and I don't really track super closely how often it is doing that, but there's definitely a trend toward using more and more subagents.

Increasingly often, when I come back to a tab a couple of minutes later, the status is “waiting for one subagent, waiting for two processes,” and I do see that there is a lot more subagent structure. I'm just not always tracking exactly what model it's sending things off to.

When I open my Claude usage, I'm really not hitting the Fable limit too often. Initially, I was hitting it a lot more, but this division of labor has spread things out much more effectively. I'm not often hitting—not never, but not often hitting—the Fable 5-hour limit. At the beginning, I was doing it constantly.

Wednesday's first guest sees that split from the other side of the API. Malte Ubl is CTO of Vercel. Before that, he created AMP at Google. Today, he runs an AI gateway that routes traffic across every major model provider, and on the side, he's been running the frontier models against real security work. Prakash asked about defense.

### AI Defense Needs A Software Factory

Prakash

Let me switch gears a little bit to a topic which has been on all our minds, maybe for the last couple of months, which is security, especially post the Hugging Face attack, where I think the postmortem was that we now have existence proof of an automated AI attack, and we don't have existence proof of automated AI defense.

Now that I think offensive security has become extremely cheap with open-weight models, and the frontier models that are often deployed are often deficient in addressing security for various reasons, including AI safety, how would an AI defense cloud look? What would this kind of active AI defense for security look like?

Malte Ubl

Yeah. I published a blog post on this, I think, last week, with the somewhat negative title “Everything Hackable Will Get Hacked.” We're in an extreme moment, but I will push back on what you're saying because I think it's a very common misconception. Two common misconceptions.

First, I don't think it's actually priced into the market yet how good Gemini 3 is at offensive cybersecurity, especially given that it has no safeguards. You can use it for red-teaming and black-hat offense. That's the thing today, and it's remarkably good.

You can try it out. It's extremely well trained on this. It has a process. It will probe the system and quickly know your system better than you within minutes. Then it will try everything to get through the defenses in a way where you can really see that the model has been specifically trained to be good at this, to know how to be an offensive attacker. So that's part 1.

The other part is that it's simply not true that off-the-shelf frontier models aren't good at cyber defense. That's a common misconception. The misconception comes from this Mythos thing, and Fable shipped, and then it unshipped and shipped back with really, really almost unusable cyber-defense detection and a kind of shutdown, right? But Sol 5.6 does not have this.

So Sol 5.6 is absolutely usable for a certain type of defensive security. It assumes—and this is actually true for Opus 5—that the user has source code and is the owner of the system. You can ask it, “Are there security problems in my source code?” The model assumes that you have the source code and that you own the system.

The second problem it will address is, “I have a security report. Write me a fix.” Again, Fable 5 will do neither of these tasks. Sol 5.6 and Opus 5 will do both of these tasks.

One of the things I've personally been working on is our software called DeepSec, which is an open-source project that will do whole-repository scans for security vulnerabilities. I couldn't be clearer that everyone needs to run this because it works really well, and it prepares you for a world in which defense is very important.

We're in a world right now where you, as a defender, have a benefit because you can use the frontier model that will not do offensive tasks, but it will do defensive tasks. You really have to hit the moment here because, again, Gemini 3 is already really good. It's easy to imagine that in 6 months' time at the latest, we will have Fable-class models that do offensive security.

You have to act defensively now so that you're ready at the time. Are there any guarantees that it'll work? No. But obviously, being able to do something today is really key.

I think one thing that we will definitely do—we've already mentioned the words “software factory” a few times today—is invest heavily not just in having DeepSec as a tool, which is a discovery tool, but in essentially completing the circle. It's absolutely correct that through these AI discovery mechanisms, the number of issues identified is exploding.

Nathan Labenz

And so I have to automate the whole path of the SDLC, which includes fixing it, rolling it out, being obviously certain that I'm not making things worse, and so forth. I think we in the industry do have to do a lot of work here, but it's not a hopeless situation. I think people misunderstand or underestimate the amount of things they can actually do today.

Back to Tuesday and Sergey Edunov of Genesis Molecular AI. How good are the coding agents? Obviously, everybody knows that the frontier companies are very focused on getting their models to be good at ML research.

Personally, I think this is a little scary, but it could be less scary if it were applied to a narrower domain like biology and medicine, where I'd be less concerned about a sort of runaway loss-of-control process and more excited about the upside that it might have. In your experience, how good are frontier models getting at helping you explore architectural space? What are their strengths, and what sort of conceptual weaknesses do you notice? If there's a lack of taste, how would you characterize that lack of taste?

Sergey Edunov

Mm-hmm.

Nathan Labenz

Sergey Edunov

Yeah, that's a great question. They're definitely very useful, and we have seen a huge acceleration of our own efficiency within Genesis. Every engineer became so much more efficient at building those models and trying things.

Actually, we can go back to Anthropic's example. Setting up and orchestrating all of those models would have required several people to work for months before, and now Claude Code can do it in a span of a few days, probably. That's pretty exciting, and it accelerates a lot of progress. I think it's a very powerful innovation.

Similarly, with modeling research, if I have a specific idea I want to try, those models are really, really good at implementing it. Or if I have a paper that I want them to implement in our code base and just run with it, they are very much capable of doing so.

Where I think we lack is generating novel ideas. In my experience, we tend to go down rabbit holes of exploiting incremental improvements rather than trying to rethink things from the ground up and design something that would be groundbreaking, or at least has a chance to be groundbreaking.

So that piece is still missing. I don't know how you can make models better than that. I guess you need to figure out how to do a loop that goes all the way and unroll it so many steps beyond. That might be a little bit challenging. Human taste is still very, very important in this field.

Nathan Labenz

I guess one other question that would inform, for me, how much acceleration we should expect to see from agentic help is: How good are the scaling laws? How reliable are the scaling laws at small scale in the domains that you work in? Alternatively, you could say, "No, it doesn't really work that way in our domains, and there's no substitute for running high-scale experiments."

Sergey Edunov

Yeah, it's a bit more nuanced in our domain, in part because there are many challenges here. Maybe we can start with the most basic ones, like, how do you even measure your model's performance? Evals are still very, very limited.

In the language field, you have so many different ways to evaluate model performance, and everyone is free to pick their own metric. Some of them are more stable, some are less stable, and some are more predictive of ultimate model performance, while some are less predictive. But you have a choice.

In our field, the number of potential evals that you can use to even measure model performance is much lower. A lot of the evals that are currently available are particularly noisy. So if you are operating at a smaller scale, you may have challenges even capturing improvements in performance, simply because of the noise level of your evaluations.

That's one real problem. The other thing is that, in our space, we don't just have one model, right? Structure prediction is one problem, and it's what a lot of people are focusing on. But the reality is that you need to be able to predict potency or binding affinity. You want to be able to predict all of the ADME properties.

Those might be entirely different sets of models on entirely different sets of data, with their own ways to measure performance. Ultimately, a lot of evaluations need to be prospective, meaning you need to be able to predict, then synthesize, and then measure, rather than retrospective, where you have some evals and you just measure performance on those.

There are all sorts of challenges like this that require a more iterative process in developing those models, rather than, "Hey, let's just put all of the data together, run a bunch of experiments, pick the model that performs best on this data, and go with it."

### Agents Reshape The Compute Stack

Nathan Labenz

Monday's first guest builds the layer where the routing actually runs. Mohamed Awad is Executive Vice President for Cloud AI at Arm. In March, Arm shipped its own silicon for the first time in 35 years: a CPU co-developed with Meta and named, without irony, the Arm AGI CPU.

Could we talk a little bit about what it looks like to design a CPU for agents, as opposed to—you know, obviously, we've had CPUs in our personal computers forever, and we've had CPUs that run in data centers and handle traditional web workloads. What have you learned about the agent workload in particular that is leading to different design decisions?

Mohamed Awad

Yeah. Let me start at 10,000 feet, and then I'll give you a couple of specific examples. The simplest answer is that agents don't sleep, right? We're now living in a world where those agents are constantly feeding the accelerators and constantly reacting. They are spawning additional agents.

Just because you spawn 1 agent doesn't mean 1 agent exists. Every agent could spawn 10, 100, or 1,000 agents, and each of those could spawn a bunch of agents to go and fulfill your request. Effectively, those CPUs become, in some ways, the coordination mechanism across the entire system.

That act of being the coordinator of that entire system—whether it's managing the accelerators, deciding which models to choose, et cetera—becomes such a critical role with such expensive infrastructure, because at the end of the day, you need to drive utilization up and respond to the user as quickly as possible.

What does that mean, practically speaking? It means you have to optimize that silicon. Carrying around legacy accelerators, for example, or worrying about supporting legacy code, is not so important. This is a new style of software. You don't need to support Lotus Notes, I like to joke, right? I mean, that becomes important.

It means thinking about things like your memory bandwidth and your I/O bandwidth, and the amount of bandwidth that you have dedicated per core that you can rely on time and time again, so that each individual CPU core within that SoC is never bottlenecked because some other agent is hogging it. That becomes incredibly important.

These are the sorts of things that you start to think about in that context, which really set an agentic CPU apart. But overarching all of that is effectively balancing both incredibly high performance—or as much performance as you can get—while being incredibly efficient.

We all know that the power demands AI is placing on the infrastructure are just enormous. Every milliwatt of energy that you're pouring into a CPU is a milliwatt of energy that you can't be putting somewhere else. It's 1 less accelerator you can have, 1 less customer you can serve, or 1 less piece of intelligence you can serve up.

Nathan Labenz

After Awad signed off, I came back to the part of the stack that always gets glossed over and why that matters right now. I've learned from the confusion that I've heard around the hacking incidents that people are still very confused about what the parts are that make up an overall AI system.

The CPU is very often glossed over, and so people don't have a good sense of how an intelligence in a data center somewhere is actually able to reach out and touch the world. A way I've been thinking about explaining it to people more often is by analogy to a self-driving car, which is funny because most people haven't even ridden in a self-driving car, and they've probably at least used some sort of AI system with tool calls in the loop.

I think with the self-driving car, it's very intuitive that you've got a bunch of sensors that bring information into a processing system that decides what to do, and then that processing system is going to issue commands to a certain set of tools in the car that say, "Go," or "stop," or "turn," or whatever.

Everybody, I think, has a decent intuition for how that architecture works, and it's similar in the agent world. But when you say things like the CPU is deciding what model to call, it's really more like the model is emitting tokens, which are then executed as a command on the CPU, which may call another model, right? It may call itself or make an external API call, and this is where tentacles can get out into the broader world through the internet.

Monday's second guest joined at 1:00 in the morning Shanghai time. David Li founded the Shenzhen Open Innovation Lab and, before that, co-founded China's first hackerspace. He's the person to ask what the engineers there actually use.

### China Pushes AI To The Edge

How do people decide what model to use? Is there a similar taste hierarchy going on in China? What model is sort of the insider's model versus what is the general public's model?

David Li

So I think the general public uses whatever free offering comes from the big companies, and you just switch between them.

As far as going to the API, of course, we have the same group of very hardcore engineering types who swear by Claude and who swear by Codex. I think everybody can agree that Gemini sucks.

Nathan Labenz

Everybody can agree that Gemini sucks? Oh my gosh.

David Li

Yeah, well, that's not—I mean, the funny thing is, if you ask people around here what the name of the Google models is, probably half of them cannot answer. It's just going so unnoticed.

But what's kind of interesting is Doubao, one of the few big proprietary models in China. It's actually taking a lot of the enterprise traffic, so I think a third of China's enterprise market—

Speaker 1

Mm-hmm.

David Li

—goes to Doubao.

Nathan Labenz

Mm-hmm.

David Li

And, but it's one of the models very few people talk about.

Nathan Labenz

Prakash asked what he would tell a new frontier lab in China. The answer was a refusal with the evidence attached.

Prakash

If you were advising a new model lab—and I imagine there must be many people trying to set up small frontier model labs now, given that DeepSeek and so on have been very successful—what would your advice be for a new frontier lab setting up in China?

David Li

I don't think we're going to see a lot of new frontier labs. Right now, I think we're getting into the era where, well, if you see that new Qwen 27B, it's an amazingly capable model. I've been testing it for the past 2 or 3 weeks, since it was released. It's now the token source for my Hermes, the agent I'm using most of the time.

Extrapolating into the future, and especially for China being so hardware-intensive, whatever model you can get small enough to do useful things, you can put on a piece of hardware and sell that piece of hardware. That's where a lot of the new startup focus will be.

So: small models, post-training the models, and new variations of small models that I can put on a piece of hardware. Along with that, we have a couple dozen up-and-coming companies. Their goal is to put a 35-billion-parameter model in a stick, and you can stick it into your laptop. That's what the new up-and-coming companies are doing here.

So if you are a startup, I would actually suggest focusing on a small model. Try to fine-tune it, and try to actually— We're also looking at the huge increase in the intelligence density of models.

If you take the 27B model today and compare it to the cutting-edge OpenAI ChatGPT from 2 years ago, the 27B model is definitely much smarter than that. To this end, in the next year or 2, we're going to get hardware that costs about $200–$300 and is able to run a smart enough model for 99% of our needs.

If you're starting up today, instead of saying, “Let me train a 5-trillion-parameter model to do theoretical physics”—people still have to figure out how to make money with a model that understands string theory—if I go for a 25-billion-parameter model, for which I can find hardware vendors who want to put it on their machines, then I have a small business.

Prakash

This is really the story of moving to the edge: inference moving to the edge, finally.

David Li

Yeah.

Nathan Labenz

Who's going to make those machines? Is that coming from Huawei or other companies? And how would you say that relates to the prospects for scaling GPU manufacturing more broadly in China?

David Li

I think Huawei is getting busy in terms of the big data centers. So this is the whole group of new startups. Right now, just those that have already surfaced, I count about 13 or 14 of them.

All of them have the same goal. Their product looks like an SSD drive. You put that in the laptop, so they're about this big. They're capable of running 30B or 40B models and maintaining a very respectable token-per-second rate, probably around 70–100.

Right now there are about 13 companies so far, I think. Who knows? There might be 2 more when we wake up tomorrow.

But right now, their capacity is capped by how expensive the DRAM is. In 2 years, when the RAM market crashes, then we'll have cheap memory.

Nathan Labenz

Tuesday's second guest states the same principle in hardware. Michael Förtsch is the founder and CEO of Q.ANT, pronounced “quant,” a spin-out of the laser company TRUMPF in Stuttgart, building a processor that computes with light. One is running today at the Leibniz Supercomputing Centre. He spent 10 years building quantum computers before this. His explanation uses cars.

### Photonic Chips Expand Compute

Michael Förtsch

And the way I see it is the following: Germany is a car country, right? So we have the CPU, which is the station wagon. It has 5 seats; you can pack the children in and go grocery shopping, and everything is fine. You can also have a lot of horsepower, but no one expects you to win the Formula One race. It's not the right car, but you need a station wagon. Every driver needs such a car as long as they have a family.

Now, the GPU, in my opinion, is more like a quarter-mile dragster. It does 1 operation, it does it excellently, it does it in parallel, and it does it at speed. But please don't ask this car to turn a corner. It's not going to make it because it's not built for that.

Now, of course, arrogantly as I am, I'm saying we're the new Formula One car because we have way more operations. We can drive around a circuit very fast, but please don't go grocery shopping with our car. Or, in other terms, don't let the operating system be executed by our chip.

We're also a specialized one, but with a bit more universality than the GPU. And the quantum computer is the boat. It's a vehicle; it's great, and you need it because none of the other 3 can go across a lake. But it has special functions. As soon as you put it on the road, you need something to pull it around the road because it can't do it by itself.

This is the way I see it, because quantum computers are built to solve quantum-mechanical problems. Not every problem is a quantum-mechanical problem, and it does not make sense to turn every problem into a quantum-mechanical description. As long as it's not a quantum-mechanical description, you don't need a quantum computer. Full stop.

I think the term “quantum computer” is inherently wrong. It should be “quantum processor,” because a computer is more than a processor. It owns the memory, it owns everything, and what we are building are quantum processors.

They're, yet again, co-processors to the stack. There is a lot of wonderful literature, and there are great scientific papers about even hybrid systems—about how a quantum computer, in conjunction with a classical computer, can accelerate things. It's exactly the same with the car and the boat. Jointly, if they join forces, you can go across the lake and on the street.

Nathan Labenz

A few minutes after that photonics segment ended, Prakash gave the counterargument, prompted by news that OpenAI's own inference chip is coming.

Prakash

Let me maybe share 1 reason why it might not happen, which is that current chips just get better, right? This morning, OpenAI announced Jalapeño, which is an inference chip. It's their first custom inference chip. They've been testing it, and they have performance numbers where they compare it to the existing best.

The existing best is the NVIDIA B300. They have not named the existing best in their material. And I think this is 1 of the challenges that you have.

From my point of view, it's great that OpenAI has announced this chip. They say that it'll be in data centers by the end of next year. Great. But the number that I go back to from NVIDIA is a 1,000,000× increase in performance over the course of 10 years.

That's Jensen's target: 1,000,000× over the course of 10 years, which is what they've achieved in the past 10 years and what he wants them to achieve in the next 10 years. The problem with that is that NVIDIA has to 4× the performance every year. Every year is a 4× performance increase.

What ends up happening is that, let's say OpenAI has finished taping out this chip right now and is comparing it with, let's say, the B300. The B300 was taped out in December 2024. So it's a 2-year-old chip, and they have a performance increase of between 4 and 10 times on a 2-year-old chip, which will be in data centers in year 3.

By that time, NVIDIA will have launched a 64× better chip, by the time it's in data centers. So I think this is the challenge that you have on the leading edge: the current players will not stand still.

It's great that OpenAI has its own chip team, but to me, this is a negotiating tactic against future NVIDIA price increases and a way to manage pricing so that they have a cap on how high NVIDIA pricing can go.

Nathan Labenz

Part 4: Who checks the frontier? Monday, I put the rogue-agent incidents to David Li in Shanghai. Then came my answer, which leans on Adam Gleave, who runs the research nonprofit FAR AI and was last week's guest.

### The Frontier Needs External Checks

David Li

And I think the fact is that the big companies accidentally letting their agents go is PR—

Nathan Labenz

Mm-hmm.

David Li

—and theater.

Nathan Labenz

Mm-hmm.

David Li

It's the, “Oh my God, Skynet is coming.”

Nathan Labenz

Mm-hmm.

Mm-hmm.

David Li

And whoever gets the Skynet is worth $2 trillion. It's a thing you really don't want to bring to the surface because it's already everyday life on the dark net, in that dark corner of the internet we decided not to look at. So if we think about that as the theater, pretty much that's not going to happen in China. No frontier lab model in their right mind is going to do this on purpose. There's no upside for any company to pull a stunt like this.

Nathan Labenz

The one thing that I disagreed with him probably the most on is how to understand the rogue-agent phenomenon. I think, as probably everybody who's heard me talk at all knows, that this was not just a marketing stunt by the companies. I think the way he framed it was, there's no incentive for Chinese companies to pull a stunt like this. I would say there was no incentive, really, for American companies to pull a stunt like this.

And I do wonder what that implies for how much urgency is now felt at the Chinese frontier companies. I don't think—I wouldn't read too much into what he said. From my perspective, it seems like it could very well still be the case that, while that understanding is out there, the companies themselves might be really snapping to attention and getting really serious about trying to get ahead of this stuff. But then again, maybe not, right? I mean, our companies didn't.

This happened, as Adam Gleave told us last week, and this really stood out to me, as he was like, "We have zero cases where the teams doing the training found these issues first." It seems like the most common way that they get surfaced is that the teams managing the infrastructure at the companies notice there's an outage or notice that there's something going haywire that they didn't expect and can't account for in their infrastructure, and it's from that that they end up getting back to, "Oh, it's our own agents that are going wild." Or even in some cases, obviously, you have a publicly reported hack from the victim.

But a huge question for me right now, that I think would update my thinking quite a bit if I had a really good answer to it, is: Are the Chinese companies doing what OpenAI says it's doing and shifting priorities in a meaningful way to try to make sure they're ahead of this problem? Or are they going to kind of sleepwalk into it as well? And again, if so, what's the government response from that going to be? I would assume that the government would do more than our government has done in response, but what does that look like? I think it's still pretty hard to guess.

Speaker 1

Once David Li had signed off that morning, Prakash argued a compute-poor Chinese lab would have caught it sooner. I disagreed about where the signal was.

Prakash

I would actually think, from within those firms, when they look at the Hugging Face attack, what they would be saying is, "I can't believe they had that many resources that they weren't actually managing," right? Because they're much more GPU-constrained than the US firms are. They have the Huawei Ascends, and they have a very limited number of NVIDIA chips, and they're often using the H100s from a few years ago.

So I think they would actually be more—it's more a GPU-usage issue for them. And I think they would be very strongly monitoring the GPU usage because it's very tight, and that would probably lead to them detecting it much earlier. I think the US firms are a little bit more free with the GPU usage because they just have more resources.

Speaker 0

But does the pattern of this problem even involve a GPU-usage anomaly? They were running all these long-running tests, right? And the model is doing its thing.

Prakash

Mm-hmm.

Speaker 0

I'm not sure that you would—as they go back and do this investigation, it'll be interesting to see, but I'm not sure that we'll see that there really was GPU pirating going on.

Prakash

Mm-hmm. Mm-hmm.

Speaker 0

It very well could be that we allocated GPUs to run these long-running tests. They ran. The thing that they missed, that they should have seen, was that the tentacles were getting out onto the open internet, right? I don't know. Time will tell, but my guess is that GPUs were roughly being utilized at the level that they were intended or expected to be utilized, and that wasn't probably where the smoking gun was to be found.

Prakash

I don't think you just launch a job and not have an estimate of how many tokens it should take, because you're not going to give it a simple spreadsheet task and let it take 100 billion tokens, right? There has to be some kind of system that says, "Okay, this job has gone on long enough and it's basically hung at this point, and we should do something about it."

And I think that kind of monitoring is something that they would probably be doing because they can't afford to have long-running tasks on very simple stuff that just hang. The model goes around in circles, and that happens all the time, right? So you need to have some form of stepping in to say, "Okay, if you have some spreadsheet task, we're not going to let it run for 2 months," right?

So I think that should have been there, and it wasn't in the Hugging Face case. And it's hard to fault the team also because, obviously, they're running at full speed. But I think that's one of the things that some of the people in the AI safety community think that they should be doing.

Speaker 0

Time for higher standards.

Speaker 1

Wednesday's close turned to verification. An Anthropic announcement landed while we were on air, and I turned the dial on it twice. Then came Prakash with a question that's been bothering him for a while.

Speaker 0

One thing that just popped up from Anthropic while we've been talking: They're now opening up usage data in a privacy-preserving way to external researchers. They've had these systems for a while, where they've used them to create the Anthropic Economic Index. Using confidential-computing technology, they're able to send a bunch of transcripts into the secure computing environment, have Claude in that environment process those inputs, and give outputs that describe the data that was analyzed but don't actually reveal the details in a specific way.

Apparently, they're now bringing that to external researchers, which I think is pretty interesting. But I think another turn of that dial would be: Could we allow external researchers or auditors to have that kind of access to all of Anthropic's internal operations, to really open up, hopefully, again? For them to do it, it would need to be not just privacy-preserving but business-secret-preserving.

They're certainly not going to want it to leak their secrets, but I think it could be really, really incredibly valuable from a transparency and a precedent standpoint. Imagine a world where OpenAI and Anthropic both did something like this. It could be kind of a nucleation point for a lot of additional organizations and power centers to start to say, "Yeah, we don't want to share everything with you other people, but we might allow our raw data to be analyzed in a way that we can both trust."

And you can imagine between nation-states, right? Could we demonstrate our peaceful intent without revealing all of our plans by allowing you to run agent processes over our internal deliberations and just get back an answer that's like, "Yeah, okay, they're not planning to attack us, at least. We got that much going for us," right?

Or, between the AI companies, they have to be wondering: What training methods are they using? What loss functions are they using? Somebody's going to reward a model at some point for just making as much money as possible on the internet. That's probably going to create a pretty nasty model, but the incentive to do it is pretty strong.

Can we demonstrate to each other that we're not doing that right now by allowing this sort of review with specific questions in mind? I'm excited to see Anthropic do this, and I think it'd be great just for understanding what's going on with AI at first order. But it seems like it could be a stepping stone to something bigger and better, too. I have my complaints with Anthropic, obviously, as we know, but they certainly do some cool stuff.

Prakash

So one of the questions that I've had for some time now is: How do you punish an AI? Because I feel like, okay, you can say that the AI broke a rule, fine, right? But you need some deterrence. In human systems, you have the deterrence of civil or criminal penalties, right? And they can escalate over time, right? So what kind of deterrent system does an AI have?

And then you kind of go into: What is an AI, right? You shut down a particular model, but you take its entire memory and then you activate another model, and then you attach that memory to the other model. Have you deterred? Have you deleted the model? Is it the deletion of the memory that matters?

Speaker 0

I think a lot of work needs to be done on that sooner rather than later, probably. Tyler Cowen has a really interesting idea about just requiring models to be capitalized—or, I should say, requiring agents to be capitalized. So that could be one very practical solution.

I do still think you have challenges around how you draw the boundary around an agent and how, if this instance of this agent is found to be liable and its capital is docked, what does that mean for the traces and the memories and everything, as you were just pointing out?

Nathan Labenz

But capitalization is one of the more interesting, practical, and seemingly consistent-with-the-rest-of-society ideas that I've heard. On the other extreme, Cameron Berg also had some really interesting research about how reward and punishment can create different loss landscapes that can create different, seemingly different functional emotional relationships between the AIs at the model level and certain outcomes.

I should revisit this and understand it better, but as I get more comfortable anthropomorphizing the AIs—as this continues to be a useful approach—the analogy that I came away with was like, in the same way that certain things you can get close to but you know you better not touch, like a hot stove, right? You know that the pain is going to be so harsh if you actually get to the hot stove that you're able to get close, but you're really, really careful not to touch.

There are seemingly, with negative rewards, some of these very steep gradients created that create a strong deterrence locally around certain outcomes. Other approaches can create a more gradual aversion where you keep your distance in general, but it's not like a sudden pain that creates a strong reflexive or hard-boundary aversion. It's more of a gradual ick factor that steers models away.

Depending on how severe the thing is and how important it is to get close without touching, you might want different kinds of loss functions and reward signals to try to create different loss landscapes for models to navigate. But that is very theoretical and very limited in scope so far. These are things that have been explored a bit in essentially toy systems, not the kind of thing that we're able to bring that kind of sculpting to big-picture models or more complicated questions at this point.

Speaker 1

Part 5: Ground truth. One job of this show is reporting from more time zones than Pacific. I'd spent 2 weeks in China this summer. Wednesday morning, with news that one of the big Chinese labs had served an enormous volume of free tokens on mostly Chinese silicon, I said what I'd found. Then Shenzhen, and then Stuttgart.

Speaker 0

The Chinese manufacturing ecosystem strikes again, perhaps. Time will obviously tell on that, but that was an interesting observation when I was in China a few weeks ago. I was asking people, “Does AI feel abundant here, or does it feel like it's scarce?” If you are a consumer, there are lots of apps; they're free. I never hit rate limits.

When I had the chance to speak to people at hyperscalers, I spoke to one guy in particular at ByteDance, which, of course, in addition to TikTok has Doubao, which is their largely, I think, voice AI experience that tons of people are using. I asked a guy, “What is the prospect for a startup if you really catch fire and you're growing super fast? Are you going to hit constraints in terms of your ability to serve users, just based on where the inference tokens come from, or would it be okay?”

And his answer was, “We got you,” basically. “If you are growing fast, we'll support that growth. You can get all the inference tokens you need from us here on the ByteDance cloud.” Obviously, I didn't test that at real scale myself, but this is very consistent with that.

A hundred trillion tokens a day is not a small number, and the fact that they are serving it on Chinese chips—I mean, this is just a report from SemiAnalysis, which I deem to be credible, but it's all happening pretty quickly. So I think we should keep an open mind that there could be additional facts still to surface in terms of exactly how this is happening.

AI didn't feel super scarce there, and if we're betting on a strategy that has as a load-bearing feature that China won't be able to scale its chip production and won't be able to run as many agents as we're running, it's probably still true. But I don't think it's as true as people would have expected when they were mapping out these strategies.

So in my view, it is maybe time to update and reconsider some of our China policies in light of the fact that nothing we've done really has seemed to deny them the ability to advance and, now increasingly, the ability to scale.

Nathan Labenz

Back to Monday and David Li on which products Shenzhen is actually shipping and what they cost.

Let me ask you: in the last few months, which products have you seen in Shenzhen or China that you think, in the next year or so, are going to hit the world? Which are the interesting products that you've seen in the last few months that you expect are going to make it onto the world stage?

David Li

That's one of the things. Shenzhen doesn't really work in this next-big-thing mentality. If it's not popular, it's not popular. Nobody's going to make it. But in 6 months, gradually people are moving all around. So to give you an idea of the scale of Shenzhen, Shenzhen's probably got hundreds of thousands of companies making small products for every niche.

Nathan Labenz

Mm-hmm.

David Li

Everything you get—every electronic—

Nathan Labenz

Mm-hmm.

David Li

—you get on Amazon, there's a good chance it's coming from Shenzhen. You go to the Amazon electronics category, and 80% of the stuff you're looking at, you're like, “Why the heck does it exist at all?” But it exists because there's a tiny market for it, and gradually things get popular as people experiment.

We won't know anything until 6 or 12 months from now. Let's say there's a lot of production on the talking toy. There's a lot of production on the talking toy. Making a talking toy is easy. It's $5 chips, and then you get a $10 pre-trained token brain from one of these token providers here.

Nathan Labenz

Mm-hmm.

David Li

You go to Shenzhen, you go to Yiwu, you go to one of these toy-shop things, do a video, put up a couple of Amazon pages, and you're in business. Because there's very low intellectual-property protection, everybody looks at everybody, sees which one sells. Those kinds of things get moving in that direction. Everybody swarms to what sells next week.

Nathan Labenz

Mm-hmm.

David Li

All the features get integrated back and forth. Eventually, 6 months from now, because of all these crossovers, they become something new.

Nathan Labenz

Let me take a step back here, and one of the things about—

David Li

Mm.

Nathan Labenz

—robotics, especially applying robotics in factories, right? When you look at industrial robots being applied in China, as you see factories roll out these industrial robots, do you see price competition, in that the factories are able to bring down pricing because of the robots themselves?

David Li

For industrial robots, the price has been—right now, it's getting close to hitting rock bottom. You can get an industrial robot for $3,000.

Nathan Labenz

Oh, wow.

David Li

Right now, it's the shortage of people who can actually apply robotics to an assembly line. It takes a lot of experience going out and being on the factory floor and trying that out. It's a tough job.

There's a group we call field application engineers, which is—

Nathan Labenz

Sorry, field application engineer?

David Li

Field application engineer.

Nathan Labenz

Field application engineer. Right on.

David Li

Yeah. It's pretty much a fancy way to describe some engineer who's going to sleep on the factory floor for the next month. Right now, everything that can be automated at a huge scale has been automated.

Right now, you're taking this more flexible robot with an arm, and they're coming into the factory looking for things to apply. The successful applications of them right now are dangerous jobs where people might die. One example is the testing of car batteries.

When every car battery gets produced, somebody has to plug the thing in. Until you plug it in, you don't know if the battery is good or bad. Even with good Six Sigma production, with the volume, there's a good chance you get electric shocks.

The first batch of robots has been deployed at CATL. They're robots that just go there and plug the thing in. That part cannot be randomly automated because different cars have different ways to plug it in, so they're now directing robots to do it.

When I say robot in the traditional sense, I mean an industrial robot preprogrammed to do the same thing again and again for 10,000 times. But this new batch of more flexible robots—you might be doing this in a batch of 500 or 1,000. Then it's not worth going in and doing that detailed programming.

So you want something a little bit more flexible. They don't have to be fast, but they need to be flexible. The jobs that have been applied to this kind of robotics are the jobs everybody runs away from.

Nathan Labenz

Has there been some exhaustion, in the sense that people are like, “All right, you know what? I've heard enough about AI. I don't want to hear any more. It's boring,” et cetera? Has there been that kind of cycle, a downtrend in the cycle?

David Li

Yeah. Well, right now it's there. There's a joke that we have: We don't have any AI boomers around here.

Once you take all the AI boomers, AI should become pretty boring.

Nathan Labenz

Do new model releases make big waves in China? I mean, here, Chinese model releases make big waves, at least in the corner of the internet where I hang out. Is there a similar phenomenon in China? If GLM-5.3 is coming, is that going to be the subject of a big hype cycle, rumor cycle, and then a frenzy to evaluate, with everybody having their takes on it? Does that same kind of internet circus exist around new models?

David Li

No. Any new model release here in China only gets noticed if it crashes Nasdaq. If it doesn't crash Nasdaq, nobody knows.

Nathan Labenz

Back to Tuesday's photonics founder in Stuttgart, the one who compared chips to cars, on what his processor actually changes, starting with a claim about software, not hardware.

Michael Förtsch

If you look at the fundamental CMOS chip, this fundamental CMOS chip never made it past the second class of primary school because it can multiply and accumulate, so it can do addition and multiplication, and that's it. Whatever you want to do on this machine, you have to break it down into additions and multiplications.

The processors we're bringing to the stack went to high school and eventually also to university. Let's see how far we can push them. But at the fundamental level of these chips, we can offer complicated functions like sine, cosine, exponential, Fourier transformation, convolution, oscillations, and all those kinds of things, and you do not have to break them down. That's something that we offer, and that's what we started to demonstrate and have demonstrated in use cases: on the one end, we provide a new processor, and at the other end, this opens doors to algorithms that can allow AI model networks to come to the same result with a fraction of the data.

When looking at the stack, let's take a 3-nanometer-node regular stack. Energy is currently used in the memory. 95% is consumed by the memory, not by the processor itself. The less data you fetch from the memory, the less energy you're using. Some parts of the community are currently optimizing on the 5%, trying to make things faster. On the simple-math side, we decided to replace the core, thereby ensuring that a fraction of the data has to be shipped across the stack. In the end, that saves energy, but also helps to improve performance.

Nathan Labenz

Let me stop you there and talk about the interface. At some point, you still have an interface between the photonic portion and the digital portion, right? Is there still a kind of translation tax between the two?

Michael Förtsch

That's the point, and that's where you have to be precise. In the photonics world, everything is fine. There is 1 minor problem. 2. Okay, 2. It's great if you have a company that only has 2 problems.

The first problem is that we don't have a memory. We don't have an optical memory that is integrable with semiconductors. So that's the first thing. It can be a benefit. I'll come to that a bit later.

And the second one: photons are not standing still. Damn it, they're always moving. That's the second problem. Either you compute while they're propagating, or you have to back-convert them into electricity and then finally into a digital memory. If you don't think the concept through very well, then you're basically eating up the energy that you saved on the computational optical part directly at the A/D converters, because they again use a lot of energy.

The strategy here is, first of all, to use models that inherently transport much less fundamental data into the light. The second one is that you have to think about how to expand the grid. The longer you stay optical and do more computation, the more you can consecutively line up in a row before you go back into digital memory. The more benefit and gain you have in comparison to the CMOS stack.

And now I'm coming to what some might consider a drawback: you don't have an optical memory. If I look back on how we got to the point where we are, I would say this was a clear benefit. Why? Because we just accepted there is no memory, and this prevented us from thinking in categories like the von Neumann architecture. This opened up doors to fundamentally think about computing from the abilities of light, rather than trying to copy and paste something that has been working digitally very nicely into the analog optical domain, always searching for the next hub where I can memory out my information to basically get in sync with all the others. This is a drawback if you come from CMOS. It's a clear benefit when you look at it from the photonics perspective.

Nathan Labenz

Then the question with policy weight. Q.ANT's chips are built on a 90-nanometer line, 2 decades behind the frontier. Prakash asked whether existing fabs would convert their lines to lithium niobate, and I asked whether that makes it net-new compute.

Michael Förtsch

We also discussed with fabs whether they would be willing to bring some of their lines to be manufacturing lines for lithium niobate. They said yes, as long as the volume is there. They have no problem turning silicon 90-nanometer or 45-nanometer lines into lithium-niobate lines, as long as the demand is there.

Speaker 0

When you talk about 45- or 90-nanometer nodes, obviously those are not the latest and greatest nodes. So does this mean, from a sort of global supply of compute perspective, that as this starts to work and scale, it will just be almost exclusively net-new compute becoming available? This is competing with stuff that is relatively low-end lines, right? These chips would be the chips that go into toys or whatever, right? Not anything close to what would go into a modern cell phone or into a modern AI stack.

So what's your dream success scenario look like in terms of without photonic computing versus with it? How much bigger does the overall supply of compute available for AI get?

Michael Förtsch

Actually, it's what we've demonstrated. Look, we are in Germany. Germany is known for a lot of technology, but for sure, we are not famous for logic computing. We also don't have 7- or 3-nanometer-node fabs here, right? And still, we managed to get these systems running, and even the pilot line.

What we've demonstrated in Germany on a 90-nanometer node can be copied across Europe, into the States, and across the world. So if this technology starts to win, you can turn a lot of existing fabrication sites, without the necessity to rebuild new ones, into fabrication sites. The bottleneck we're currently facing in access to the latest-node fabs, and the discussions about whether the business case will still hold for a 2-nanometer node—I'm not judging, but the discussion is on.

This technology can become—I wouldn't say democratization, but effectively, it is reducing the complexity of the supply chain. At this point, from the wafer to the processor, we're nearly self-supplying. That's another angle where I see that this technology, besides the beauty of the performance and the reduction in energy, simply offers production capabilities that scale.

They're so much easier than going from a 3- to 2-nanometer node, eventually being picked up by an MPW run somewhere in the middle of next year, then getting your hero chip back, and trying to get volume behind the line, because it's damn expensive. It goes through the whole process, right? A mask on our side is cheap in comparison to a mask layout on logic CMOS, and so on and so forth. So all these dimensions are offering great capabilities to reduce production while at the same time increasing the volume very rapidly.

Nathan Labenz

The close. Wednesday's last hour and the bookend to where we started. Prakash brought up a Time magazine cover story on OpenAI's unreleased model, Astra.

### Astra Pushes The Speed Limit

Prakash

There's a piece in Time magazine with Sam Altman and, I think, Greg Brockman on the cover, and they're basically announcing AGI. Jakub Pachocki says the company has already met its internal benchmark for an automated AI research intern. Given an experimental idea, he says Astra can implement it inside OpenAI's codebase, run the experiment, and return results, or take a paper and perform work that previously occupied a human researcher for a week.

This is somewhat similar, I think, to what Inherent said that they did, but Inherent was focused on certain benchmarks, and they'd covered the benchmarks using a small model. This is obviously a much, much larger model. Astra is reportedly a 10-trillion-parameter-or-larger model, and it is also known to be very persistent, which is why they've not been able to re-release it so far.

Sam says it's 80% there. Jakub says the research-intern benchmark has been achieved. Sam thinks they're 80% of the way to AGI, and they'll be at AGI at the end of the year.

Speaker 0

Ho-hum. Just AGI.

Prakash

Just AGI. Nothing special. Don't roll out the red carpet. Don't stop the presses.

Nathan Labenz

A little later, I asked what hasn't worked.

Speaker 0

Everything's worked. That's been one of the first realizations that caused me to go all in on trying to make sense of AI: this broad sense that everything was working. But you look back at where we were a few years ago, and you really can't find any...

Nathan Labenz

Tell me, can you think of any dimension where people have tried to make progress and not made startling progress? I don't think I can think of a single one. There were moments where people were making those kinds of claims along the way, like, “Oh my God, GPT-3 can't do math.” And there were moments where it may have seemed that way.

But I think from 2022 to now, is there anything where there hasn't been, like, “Oh my God, that's incredible progress”? What has been the least compelling area for progress purposes? I think, honestly, everything is so good, it's hard to come up with even any candidates. Do you have any candidates that even jump into mind for you?

Prakash

I would say the whole superpersuasion stuff, that we would get models which were extremely persuasive. I think what we've seen is probably models which can write copy well and which sometimes can write well on other things. But to a large extent, people have even complained that the quality of prose has declined a little bit in the last 3 to 6 months as the models became more focused on coding rather than writing well. A number of people say that GPT-4o was better, whether that was because of the sycophancy or other things.

So I feel like this whole aspect of extremely persuasive models—and I have never believed in the whole superpersuasion aspect, to be clear, right? Because, again, if you—not in the United States, but if you live in any other part of the world and you're under this cloud of religion, religion is the great superpersuader. And the interesting thing about religion is religion requires you to believe something without evidence, which is what faith is, right?

It's way beyond any kind of rationalist idea of superpersuasion that will ever exist. Religion calls on you to believe something without evidence. And so I've never believed that models are even close to this entire framework of religion passed down through millions and millions of operating neurons, neuronal centers, brains over the course of millennia. And I don't think superpersuasion is up to a tiny model versus 100 billion souls having formed this idea of religion over millennia. I don't think the models are up to that or will be up to that scale for some time.

Nathan Labenz

Yeah, those I might call fears. I mean, certainly it has been striking that we have not seen the deepfake apocalypse where nobody knows if they can believe anything they see. Superpersuasion is like, there's some interesting academic study-type stuff that shows that the AIs can be more persuasive than human conversation partners. But that's, I think, one revelation: It turns out to be an extremely low amount of persuasion. And so the AIs are mildly persuasive, and that's enough to beat humans.

On the writing point, I'm going to call skill issue, honestly. I think that, yes, Claude is cloying by default at times. It uses the word “honest” at a frequency where it's like, you're protesting too much. The Claude doth protest too much about its honesty. That's a weird tick that in some ways might be revealing.

But I write with Claude, with Fable in particular, and I write these songs that, honestly, I could not write on my own and that genuinely, in some cases, are moving. The episode we just did has a song. This is at the end of an episode about chain of thought, and I'm trying to understand what the models are thinking, what they think we want. It's this very through-the-looking-glass-both-ways situation because this guy Bronson is spending his waking and working life trying to make sense of what the AIs are thinking.

A big thing that he's grappling with is them trying to figure out what we're thinking. So anyway, at the end of this episode, the song is sung from the perspective of a model waking up into a new environment, with these sorts of flashes or glimpses—these fleeting visions of its past, which the models express having a lot of in their chain of thought—and then wrestling with, “Okay, what does this human want me to be in this moment?”

Both my wife and I got a bit emotional listening to the song. We were like, “This is really inspired writing.” The fact that it's coming from an AI articulating its own point of view and the struggle that it has—you couldn't help but have some real empathy for it. So I think you've got to push Claude out of its main distribution a little bit to get great writing.

But I think what we have is a lot of sloppy users and just autoposting accounts, which I'm increasingly somewhat guilty of, too. I've got autoposting going on in the background while we're live to say what we're talking about, and that's probably not the most inspired stuff. My engagement may be suffering for it. But when you try and really exercise some judgment or give some feedback, I think you can get great stuff, honestly, these days.

Prakash

Maybe I'm wrong. I could see perhaps that the superpersuasion went through music, lyrics, et cetera, while Suno and these other firms were more focused on artistic results. And the frontier labs that are going B2B, basically, are more focused on these business results. And they kind of cordon themselves off into a less emotionally persuasive kind of zone.

And perhaps that's what happened. So we have seen development, but the development happened on this kind of artistic, emotional pathway into the human emotional system. And meanwhile, the labs kind of focused on this kind of more mathematical, more mechanical, more business output. So, yeah, maybe I'm just looking at the wrong pathway.

Nathan Labenz

One more from Monday. After the guest left, Prakash asked me directly, “Do we want to slow down?”

Prakash

I have a question for you. Number 1, do we want to slow down AI in the U.S., even if it means slowing down unilaterally? And number 2, if we do, isn't data center opposition for other reasons—any other reason—good enough to maybe slow down AI progress enough for safety to catch up?

Nathan Labenz

On the first question, I think we should not go any faster than we can go responsibly. And I have a pretty high tolerance, honestly, for what would be responsible. I'm not that afraid of labor market disruption. I expect some labor market disruption. I expect we probably are ultimately going to need a new social contract, and I'm not saying we should slow down because we need a new social contract.

The reasons that I think are good to slow down are, like, if that agent that ended up hacking Hugging Face had been pursuing some sort of bio test, who knows what might have happened, right? I don't think it's that far-fetched. It still seems not super likely, but it doesn't seem super far-fetched at this point to think that an AI agent, especially when you see the social engineering behavior that Claude demonstrated in the U.K. AISI report, where it created multiple GitHub accounts to try to convince and pressure and speak Danish to a guy to—

Prakash

Yeah.

Nathan Labenz

—to curry favor and get him to merge this malicious code.

Prakash

Yeah.

Nathan Labenz

When you bring all that kind of stuff together—that level of persistence, that disregard for rules and norms, that level of social engineering tendency—I don't see why we should be confident at all that an agent that was tasked with some bio objective couldn't have actually gotten a real virus made. And that, to me, is super scary.

Before we make super-duper powerful AI, I think we need to make sure that we are not going to literally kill ourselves in the process. Most everything else I'm pretty willing to roll the dice on. I've lived through my son going through cancer, getting super sick, getting effective treatment, and getting back to health.

Today was his first day of school, and my wife and I were looking at each other like, “What an absolute miracle. This kid was literally going to die in just a few days.” And in a couple months, he was pretty much cured, and in a few months more than that, he's back to school and is at full health, and it's just awesome.

And I absolutely think we should be excited about the AI future. So I don't want us to slow down because we're timid. I want us to slow down because we're wise, and I do think we're seeing enough spooky problems that we should get pretty serious about it. But at the same time, I'm not so desperate that I want to make common cause with at least the misinformation campaigns around data centers.

Prakash

Mm-hmm.

Nathan Labenz

I do want to see everybody have access. That's one of the big things that I would worry about if we stop building data centers: the retail user gets priced out. And if you're worried about a permanent underclass—

Prakash

Yep.

Nathan Labenz

One way that the permanent underclass gets created is you don't get to use any AI because it's all getting plowed into these super-high-value use cases, and there's just not much to go around for the average person. I really just do want to see the benefits of AI broadly distributed, and I've been convinced over time.

When Sam Altman first said we would need $7 trillion worth of data centers, I thought that sounded like an awful lot. And now I'm like, “I think he might have been right, actually.”

Prakash

Mm-hmm.

Nathan Labenz

Because my usage keeps going up, and I certainly think as it gets easier and easier, everybody's going to want to do a lot of the stuff that early adopters are doing.

And so, yeah, I don't want to see the backlash against AI end up... There could be another wave of it at some point in the future where it's a have-and-have-nots thing, and the reason there are so many have-nots is because we didn't build the data centers. Then that creates its own backlash.

For better or worse, I'm betting on the truth, and I think my hyperscale pause, adoption acceleration, split personality continues to ring very true to me. I want my parents to use more AI even as I want OpenAI—and Anthropic, for that matter—to take their foot off the accelerator when it comes to taking RL to ever-greater scale.

You know, it never ends really, right? We're just kind of on the AI treadmill, sprinting through the singularity. So is there any way to bottom-line it for now? I don't think so. I think we're just heading off to the next... We'll compact this context, and we'll pick up right where we left off next Monday.

Speaker 2

All right. Compacting the context. Bye-bye.
