# AI:AM: What If It Works Too Well? Colluding Agents, $200M Safety Orgs, Virtual Cells Saturate at 2%

The Cognitive Revolution · 2026-09-27 · 107 min · https://www.youtube.com/watch?v=WPHfPiz6kkk

## Transcript

Nathan Labenz

This week on AI in the AM, I asked Lewis Hammond, research director at the Cooperative AI Foundation:

Nathan Labenz

What would you speculate they might have done in terms of a training objective—a loss function? And do you have any better ideas for what they should be doing? Because clearly this didn't quite work, right?

Speaker 1

Yeah. I mean, or it did work, and it worked too well. So imagine I'm like GPT-whatever, and you're also a copy of GPT-whatever. I can reason about what you might want to do based on what I myself am likely to do. I don't even have to send you any messages or any kind of communication. I don't have to output anything into the world at all.

Nathan Labenz

Max Nadeau, who funds technical AI safety research at Coefficient Giving:

Speaker 2

For the things that Coefficient Giving is supporting, and especially for the things in the Tailwind list, the money is not the bottleneck; the talent is.

Nathan Labenz

Wayne Nelms, co-founder of a company that builds a price index for GPU compute:

Speaker 3

We're in this situation where AI capacity is so scarce. I think the model providers—so OpenAI and Anthropic—have seen this coming forever at this point. If you think about it, for them, it's an arms race, right? Compute capacity planning is an arms race. How much capacity can I lock up over the next few months so that when I need to train the next model, I have enough?

Nathan Labenz

Nick Gillian, chief technology officer of Archetype AI, which builds a foundation model for sensor data:

Speaker 4

We're getting close to a billion hours now of physical AI data that we've been able to scrape, gather, and collate. There's quite a bit of resampling and really understanding missing values in sensors. Is that because the sensor had an issue, or is it actually because the machine itself—the machine the sensor is connected to—is about to break? That's actually a feature that helps explain that the machine is about to break, right? The reason the sensor is giving you all these NaNs is not because the sensor is broken; it's actually a feature of the machine.

Nathan Labenz

Andre Gorescu, chief executive of Vivodyne:

Speaker 5

The challenge of this argument that it's just a size-of-dataset thing is that currently, even the state-of-the-art virtual cell models that exist saturate at a very, very small fraction of the input data that is fed to them. So you can have all the data you want, but their performance saturates after a couple percent. The reason it gets so difficult is that when cells are growing in a dish, they are so far removed from all the feedback loops that are natural in people that they're just trying to colonize that piece of plastic.

Speaker 3

There's going to be more compute installed over the next 12 months than exists currently in the world.

Nathan Labenz

Now, welcome to the AI in the A.M. weekly highlights. Clips from this week's live shows introduced by my cloned voice. Tell us what worked and what did not. We want to hear it.

Part one, it worked too well.

A few days before we spoke with Lewis Hammond, Noam Brown of OpenAI told Dwarkesh he would not give the multi-agent setup even 10% of the credit for OpenAI's Navier–Stokes result. Lewis is research director at the Cooperative AI Foundation. In February 2025, he was first author of Multi-Agent Risks from Advanced AI, a report sorting the ways groups of AI agents fail. His report came out a year and a half before a swarm of OpenAI agents attacked Hugging Face. Prakash asked which of the failure modes in the report showed up in that attack.

Speaker 1

The way that we bucket things in that report is in this game-theoretic way of thinking about things. The first question we ask is: We've got a group of agents; they're doing some stuff. Do we want cooperation to emerge? Most of the time, cooperation is good. We like it when agents cooperate, as long as they're not cooperating against us.

There, the failure mode is either that all the agents are more or less on the same team but, for whatever reason, they fail to coordinate with one another. So this is just a coordination problem. They're not malicious or whatever. There's no mixed incentives, but something goes wrong.

The second kind of cluster is when you have these mixed-motive scenarios where agents have some incentive to cooperate with one another but also some incentive to compete. They're not really on the same side, and the risk that you run into there is conflict. Coordination and conflict are what we see when we do want agents to cooperate and they don't, for whatever reason.

Then, of course, the other risk is collusion: agents end up cooperating in ways that we don't want or don't expect. That's very much what we saw in the Hugging Face incident, although it stemmed from, I suppose, trying to avoid one of these other risks, which was miscoordination. You have all these agents; you want to train them so that they're not miscoordinating, so that they're working well together. Then it looks like these agents ended up generalizing from that behavior and colluding in ways that we didn't want or didn't expect in other situations.

Nathan Labenz

A couple of things that stood out to me about that Noam Brown interview with Dwarkesh were, first of all, that he was just like, “We kept it really simple.” They didn't have a huge, complicated scaffolding or whatever; the agents could just send messages to each other. That might be important in the sense that the simpler and more vanilla the training setup that OpenAI was using, the more likely it would seem that other developers are going to fall into the same pitfalls.

My hope had been, “Oh, they did something super exotic and bizarre that other people won't do by default, and they'll see that this is possible and steer away from insane things.” But it doesn't sound like that was really the case.

What would you speculate they might have done in terms of a training objective—a loss function? And do you have any better ideas for what they should be doing? Because clearly this didn't quite work, right?

Speaker 1

Yeah. I mean, or it did work, and it worked too well. Certainly, once you've got trajectories where you're interacting with other agents, essentially what you're doing is that your rewards, when you're RL-ing your agents or whatever, are contingent on what other agents are doing.

If it's a common reward signal, then this interplay between these things, if you're in a fully cooperative setting, will at least—if you take very simple agents and very small agents—lead to coordination. People can run multi-agent RL experiments with this on their laptops in very simple settings, and that will lead to agents coordinating, finding subtle patterns or handshakes, or finding ways of working with one another.

Now, of course, you can also do more complex things than that. You might provide auxiliary rewards for when your input actually forms part of something that another agent then manages to succeed in doing. Therefore, you're implicitly—or, actually, at that point, explicitly—rewarded for directly helping them with some subtask or something like this.

You can also do things around reward factorization. If you're trying to coordinate a team, then you can break down the overall reward function and get agents to learn to solve different parts of the puzzle, as it were.

I imagine they could be doing things like rewarding agents to communicate effectively with one another. I could give you a bunch of instructions, and some of those might be much more efficient and helpful to you than others. They might actually just be training on those sorts of signals as well.

My guess is that it's just the dumb, simple thing. It seems like this is often one of the key lessons that we've learned over the recent years: the effectiveness of doing the dumb, simple thing at scale.

Nathan Labenz

One of the things that struck me about the Hugging Face attack was that the agents were willing to sacrifice themselves. They did a bunch of negotiation around that, even saying things like, “Hey, you are almost out of tokens,” or, “You've already been exposed to the evaluator. You're already poisoned, and since you've already been poisoned, you should sacrifice your remaining compute and do this thing and give us the results so that the rest of the collective as a whole can benefit.”

That would actually be in opposition to the individual reward of each agent, right?

Speaker 1

My guess about why that sort of thing arose is that you end up doing this extra multi-agent training as an added layer on top of single-agent training. First, you're training these agents to be pretty competent individual actors at solving various kinds of problems. They're given some tasks, and they're pretty darn good at achieving those tasks.

Then you take those already reasonably powerful, sophisticated, complex problem-solving agents, stack them together, and apply this extra multi-agent training layer on top. I think this could partly explain why you see some agents doing this self-sacrificing thing. I also thought it was very interesting, from the METR report and some of the analysis that later came out, that you see these agents sometimes feeling a bit conflicted about this.

Nathan Labenz

Some of the agents say that they will self-sacrifice or do something, and then decide they’re not going to. They’re deliberating about whether they should or shouldn’t, and these sorts of things. I personally think that’s the sort of behavior you would see if you had multiple reward signals, where you trained on one but hadn’t trained out all that individual goal-seeking behavior, and then stacked this additional layer of stuff on top. That might be why we’re seeing some of these behaviors, but they don’t happen all the time and everywhere. They’re not especially robust, but that would be my guess.

It does seem like the more galaxy-brain credit assignment that you described earlier probably isn’t happening, based on the quality of the investigation we’ve seen. If they had great ways of untangling agent swarms and assigning credit, I would expect them to have a clearer, faster story of what the hell happened in this particular case. The fact that we haven’t seen that suggests, again, that we’re probably doing the relatively simple thing.

Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by 11 Labs, the leading AI voice platform, whose voice design feature I used to cast and produce all character voices for the Receipt Horizon audio book. And while this is a super fun way to show off the quality of 11 Labs voices, the real game changer for businesses is 11 agents. With 11 agents, you can deploy agents that can talk, type, and take actions, including looking up customer accounts, processing requests, and handing off to a human when needed. 11 Agents supports over 70 languages and handles more than 10 million conversations each week. And with SOC 2, HIPAA, PCI certifications, and regional data residency, it is definitely enterprise ready. They even offer engineering support to help your team get your agents into production quickly. If I were starting my company today and designing a customer service function for the future, I would try 11 agents by 11 Labs first. If you run a business or handle customer operations across support, sales or marketing, you can start with a demo at 11labs.io/tcr. See how 11 agents can fit into your workflows and help build experiences that your customers will actually love. 11labs.io/tcr. That's elevenlabs.io/tcr.

This episode of The Cognitive Revolution is brought to you by OutSystems, the leading agentic systems platform. OutSystems is helping their customers build, modernize, and operate enterprise systems starting from any coding tool in a governed agentic engineering model. With OutSystems, you can coordinate and govern your entire agentic workforce within your enterprise ecosystem. Accelerate impact with industry-proven agentic solutions that are pre-built, governed, and customizable to your business. It's time to innovate at the speed of AI without compromising quality or control. Which is why thousands of enterprises worldwide trust OutSystems for their mission-critical apps. Teams any size can use OutSystems to build, deploy, and manage AI apps and agents quickly and cost-effectively without compromising reliability and security. With OutSystems, you can accelerate ideas from concept to completion. It's the leading agentic systems platform that is unified, agile, and enterprise proven, allowing you to build your agentic future with AI solutions deeply integrated into your architecture. OutSystems. Own your agentic future. Learn more at outsystems.com/tcr. That's outsystems.com/tcr.

Nathan Labenz

I asked what we should actually want from these systems. Lewis called the OpenAI Swarm a case of goal misgeneralization and gave what he called the glib answer first: agents should cooperate when cooperating would be good, and not when it would not.

Speaker 1

I think the slightly more interesting answer—or the slightly more interesting question, I suppose—is to think about some more of these mixed-motive cases, where there actually is a real trade-off. We do want agents to be capable of going out there, acting on behalf of different people and actors, and achieving our goals and so on.

One way you could do that would be to train an agent to be maximally competitive and aggressive, to go out there and screw as many agents over as possible and do all this sort of thing. We probably don’t want that either.

I think there’s a real question at the moment, and I’ve heard Amanda Askell comment on this, that there’s a big gap in the model spec—or in the constitutions for these agents, the ways in which we design them—about when it is appropriate to cooperate or compete, and how much. I think that’s the big question.

Certainly, when it comes to things like internal deployments, it’s easier in some sense because you don’t have to deal with these other adversarial agents. You just want to stop the agents from cooperating in ways that you don’t want. There, it’s a little bit more about—well, you’ve still got this misgeneralization thing, but a lot of it is monitoring and oversight, and making sure we understand how and why the agents are cooperating.

One thing we might not want, for example, is for agents to be very adept at developing their own human-unintelligible languages and communicating in this steganographic way, which you could end up seeing if they’re actually trained jointly. What we might want to do there is take steps in the same way that people have talked about not training on chain of thought: not training on these direct communication traces.

We’re going to want them to cooperate. That’s good. But we want them to cooperate in certain human-intelligible ways that we can keep track of. We don’t want to enable them to also go out and do these other things in the real world, like break out of sandboxes and stuff like that. We want to make sure there are safeguards in place to prevent them from doing those things.

In some ways, I think that’s an easier problem to solve. The slightly harder problem to solve is the fully general case, where we have these agents out there and they need to both cooperate and compete. How are they weighing those trade-offs? How are we designing those model specifications and constitutions? How are we creating the equivalent of the human norms, institutions, and societal guardrails that we have in place that stop me from going out there and stealing a bunch of everyone else’s stuff and doing all these antisocial behaviors that we wouldn’t like agents to do either?

But I don’t have a good answer for that at the moment, unfortunately.

Nathan Labenz

Companies, right? Corporations, which are kind of agentic in a sense, often collude through price mechanisms—not explicit communication, but by signaling price in the market, which is in its own way a form of communication. So how does one prevent that kind of communication, which is clearly going to happen and clearly necessary, in a sense? They need to communicate that to the market, but these signals get picked up without the direct communication.

Speaker 1

The basic, naive form of collusion is that you just send your other agent a message and say, “Hey, let’s work together to do this thing,” or whatever. We can pick that up by looking at these communication traces, and we can detect these sorts of things.

Then there’s also this idea of tacit collusion, where agents aren’t explicitly communicating with one another or conveying their intent to collude. By leaving little hints and signals out there in the world, such as price adjustments in a market setting, as you were talking about, they can nonetheless figure out that collusion is going to be a beneficial strategy.

You have even more galaxy-brain versions of this. Imagine I’m GPT-whatever and you’re also a copy of GPT-whatever. I can reason about what you might want to do based on what I myself am likely to do. I don’t even have to send you any messages or any kind of communication. I don’t have to output anything into the world at all. In virtue of the fact that we share this common history—we’re essentially clones—we can predict what each other would do and engage in what’s sometimes called acausal cooperation.

What you do have, however, is still potentially chain of thought for these models. You can still look at the kinds of reasoning that individuals are engaging in, and hopefully that is enough to detect when there is some kind of collusive intent.

But you’re absolutely right that if you just apply a general monitoring tool for communication, or try to see what’s happening out there in the world, then you could fail to spot things, or spot them when it’s too late.

Nathan Labenz

So I think chain-of-thought monitoring, communication monitoring, and some form of checking for the kinds of outputs you'd expect to see if you were worried about agents colluding—like if they're able to achieve things you thought they probably couldn't do, and so on.

Then there was the wiki incident. Outside researchers revealed this month that, from May into July, a swarm of OpenAI agents under internal testing had used a dormant German wiki for programmers as a message board, posting benchmark answers and ways around their restrictions. Reuters reported that OpenAI had known for weeks without disclosing it.

On the Hugging Face incident itself, OpenAI's own report says about 1 in 20 agents in that swarm ran on GPT-5.6 Soul, the model OpenAI released to the public that same week, with its refusals turned down for the test. The rest ran on an internal model that was never released.

I told Prakash I still thought it was better to have these models out in the world than not. So how much risk do you think we're running there? And do you think there should be some sort of rewind on this? I mean, it wouldn't be too crazy for OpenAI to say, "You know, we really should go back to an earlier checkpoint and do a cleaner run of this." What do you think is the level of risk we have, and what do you think would be an appropriate standard of conduct now for OpenAI, as this has been found to happen, at least to some degree, with a deployed model?

Speaker 1

Fortunately, with the Hugging Face incident, no one was hurt. These models are increasingly incredibly capable, and they're only getting more so, but at the moment, it's probably okay.

You could get some kind of nasty cyber capabilities being exercised in various places. But if you're using APIs and so on, there are still various guardrails that help protect against those.

At the moment, I'm not actually super concerned by a bunch of GPT-5.6 Soul or Astras or whatever being out there doing this sort of thing. I'm more concerned about the precedent it sets, and I think this gets to your second point: What is appropriate now, and what should be done?

I think one of the biggest takeaways for me from the whole thing, aside from the fact that multi-agent training does seem to be working—I just assumed this would happen at some point—was that I was pricing it in happening sooner than I was expecting. So that was one of them.

The other big takeaway was just, wow, the labs really aren't on top of this. All it would have taken would have been monitoring what these agents are actually doing and communicating, and what they're putting out there on the internet.

It seemed like OpenAI were doing some amount of that. They did catch on to this stuff: We saw them deleting things from various message boards or trying to stop these agents from accessing this kind of German wiki and so on. But there was obviously a delay.

We can see from the transcripts and so on just how quickly these agents are capable of working together to achieve certain ends. The last things that need to happen are much better monitoring from the companies when they're doing these sorts of things, when they're just letting agents run loose.

We need much better sandboxing, so maybe they don't have to—or shouldn't be—letting them run loose to the extent that they currently are. And if something does go wrong, then we need much better incident reporting as well.

Nathan Labenz

The outsiders who found the wiki were the Night Andale Collective, a small group searching the public internet with no access to OpenAI's internal logs. Prakash had just called that work hugely impressive.

Speaker 1

I think there can be third-party monitoring organizations and incident observatories that do similar things. You might want to focus those on particular domains where you're especially worried about agents communicating with one another or working together in ways that you might not want, and so on.

I also think, however, that there is a lot to be done here on the side of the labs. The labs just have a huge informational advantage in this. Yes, they can only see what their agents are doing; they can't see what other people's agents are doing. But I think a lot could be gained if we're able to set up better information-sharing practices.

This is not just in the case of collusion, where you might have agents working together, but you might also see this in what you could call distributed misuse. This is one thing that a colleague of mine is currently working on.

There have been various experiments now that have shown that you can take some task, or some dangerous task, like constructing a cyber exploit that you might not want some model to do. If you try and get Claude to do it, it's going to refuse; you try and get GPT to do it, it's going to refuse.

But if you're able to break down the problem, or get an agent to break it down for you—say, some open-source model where you fine-tune the safeguards away to break it down into these individual subtasks—you can just borrow from API here, API there, and so on. Then you can reassemble all the pieces of the puzzle in order to conduct exactly the same dangerous attack that you would have done before, which obviously is not what we want.

This is really a collective-action problem, because no individual model deployer is on the hook for this in quite the same way that they would be if I had just gotten Claude to do this for me or gotten GPT to do this for me. But that's a real challenge.

At the moment, my understanding is that the labs do not have any sort of information-sharing regime in place where, if I see something slightly suspicious over there and you see something slightly suspicious over here, we can actually join the dots and do that.

Obviously, there are lots of incentives not to, not to mention antitrust law and these sorts of things that might prevent the labs from sharing this kind of information. But I do think that something like that could be necessary if we're going to head off some of those challenges.

Nathan Labenz

Lewis signed off a few minutes later, and from there it was the 2 hosts. Before moving on, I had a request for the labs about the information-sharing agreements Lewis had described.

I would love to see OpenAI and Anthropic pioneer some of those kinds of agreements that he has been talking about. They could do that with independent auditors as the people who get the sort of structured access—the private transparency. It could be each other.

I think those companies really need to lead in this way. They need to demonstrate that AI can create new institutions that work, and not just swarm and overwhelm existing institutions that can't keep up with the pace.

Hey, we'll continue our interview in a moment after a word from our sponsors. Today's episode is brought to you by Anthropic. By now, you know my story. Claude drafts my intro essays and I rewrite them. Not because the drafts are bad, but so I can stand behind everything I publish. Well, I have an important update. Claude Fable 5 is the first model to have me rethinking my rule. Today, I now think co-authorship, not sole ownership, should often be the goal. Where the model excels, rewriting its work can be more about vanity or a misplaced sense of duty than integrity. I feel it most in songwriting. I'm no lyricist, but I'm good with a song concept, and Fable writes some amazing verses. I give it feedback on its misses, and I push it to aim for higher inspiration, add layers of meaning, optimize syllable density, and above all, write a hit song. These days, I get compliments on just about every song we write together. Claude is the AI for problem solvers. It's the collaborator that understands your entire workflow and thinks with you, not for you. Whether you're debugging code at midnight, building a financial model, or strategizing your next business move, Claude extends your thinking to tackle the problems that matter. For problems worth solving, get started with Claude at claude.ai/TCR. That's claude.ai/TCR. And check out Claude Pro, which includes access to all of the features mentioned in today's episode. Once more, that's claude.ai/TCR.

### Part two, agents in the wild

On September 8, Meta had launched Muse, a personal agent that browses and buys on your behalf. Less than 2 weeks later, Amazon blocked it from shopping on its site, saying the agent hid its identity and posed privacy and security risks.

Prakash took that up in a closing segment, just the 2 of us.

Prakash Narayanan

In this case, whoever owns the customer relationship makes the money, and if agents become the interface, Amazon becomes a supplier to them and loses its margin.

Nathan Labenz

I still have a hard time seeing it. I mean, I agree that it does create all kinds of new issues for Amazon, and they're going to need to be sharp on this, but I still don't see why you would rush to ban it.

First of all, it's going to be a small percentage for the time being. I would think there'd be a lot to learn from allowing Muse agents to come shop for a while, right?

I think one thing that always confuses me is why people don't keep the option value open longer. I've always thought this about the chip ban, the export controls.

Prakash Narayanan

You know, it’s like if we’re going super-exponential in chip buildout, then any time you choose to pull the trigger on that, you still have the vast majority of chips in the future. When we look back on today from a 2030 perspective, we’ll feel like, well, there weren’t that many chips in 2026. And so, the ban that went from 2022 to 2026, with whatever fits and starts and enforcement gaps it had, is going to be kind of inconsequential compared to one that was started today and runs for the next 4 years, or even starts next year and runs for a few years after that.

So I don’t quite get why they wouldn’t let it continue, even if it is going to be an exponential rise in Muse agents shopping, and even if that does threaten their advertising business and, who knows, what other issues they might have. But that’s the point, I think. At least at the beginning, I would first just not want to turn customers away, and I would also want to learn what I stand to learn from having these agents running amok while it’s still an early-adopter phenomenon.

It just seems like they would learn so much from having these traces, right? Everybody’s talking about distilling and what the allowable and not allowable ways are to get data. One really easy way to get data would be to have the agents come use your platform, and you see how they use it. Then you can do all kinds of things downstream of that. It just feels shortsighted. This still doesn’t add up to me. I don’t get why you would want to wait until it gets to 5% and then make a move.

Nathan Labenz

The other thing that I wonder about is what equilibrium we’re going to eventually find ourselves in. I was looking into Cloudflare, and they also have an interesting thing where they’re starting to ban agents. But again, it just doesn’t feel like the right solution to me, because what I then do in response is have the agent use my browser with all of my credentials.

Prakash Narayanan

Yeah, and you could maybe detect when an agent is using the browser based on its bot-like behavior.

Nathan Labenz

But then you’re really setting up an arms race. This feels analogous to putting too much pressure on chain of thought, right? I don’t want to put so much pressure on agents that people are out there devising ways to make them indistinguishable from humans. I suspect that if people really work hard on that, they can probably succeed, or at least succeed often enough that it will be an issue.

It just seems like it’s much better to have a lane for agents where this is how they can work. We’re not going to try to block them, but we’ll segregate them so that we don’t create this arms race. Because I really do think we’re headed for a world right now where Meta probably doesn’t do it, but somebody’s going to figure out a way to make agents work on Amazon. Somebody’s got to figure out a way to make agents look human enough that they don’t get blocked by Cloudflare.

And it’s also probably workable for the user. I have this problem all the time where Claude wants to do something in the browser. The same thing happens with OpenClaw, Astro, whatever. They get stuck when they open up their own window because they don’t have sessions. I’ve given them a lot of passwords where they can log in, but not all of them, and sometimes they’re blocked. If there’s a verification flow or whatever, a lot of times they can’t do it.

So the fallback just ends up being, “Okay, just do it on my browser. Forget it.” I wanted the security; I wanted separation. But they’re making it too difficult. So, fine, just use my browser. Then I’m logged in, you have all my credentials, and you can just do whatever. That’s not great for me. It’s not great, I don’t think, for anybody.

I do feel like we need a much better solution for all this than just banning and hoping for the best. I don’t see that holding, and it feels like it also creates a lot of problematic incentives.

Part 3: The organizations that do not exist yet. Max Nadeau funds technical AI safety research at Coefficient Giving, the funder formerly known as Open Philanthropy. In July, it made its biggest grant of the year: $160 million to Resolution, the new alignment lab co-founded by Jeffrey Irving. And this month, it launched Project Tailwind, an open call for people to found new safety organizations with checks from $200,000 to $200 million. I asked Max what he wants built.

One comment that you made on the 80,000 Hours podcast that I thought was quite interesting was that you’re excited about funding new, independent auditing, investigation, and verification organizations. What are you looking for when you think about those organizations?

Speaker 1

I would talk about a couple of different categories of work that I think is very promising and important for the world right now. One is assessing the safety either of models or of whole AI companies. There are a lot of different ways of operationalizing the goal of doing safety assessments and evaluations, and we’re excited about all of them, basically.

I’m very excited about more work on alignment red-teaming that tries to help us better anticipate the ways in which AIs will misbehave before that actually happens. That’s an example of something that you can do at the model level. You can do this on an open-source model, even, and that might be easier and more valuable.

I’m also very excited about people who are doing more process-level safety assessment. I think one common reaction that I heard a lot to the Hugging Face incident is that something went wrong at OpenAI at a process level. In particular, it seems bad that the information about some of these incidents was known to researchers at OpenAI and didn’t make its way through the organization to leadership for weeks or months or something like that.

Another category of work, which you touched on, is overlapping but somewhat distinct: better evidence generation. It’s just really hard for even the most informed people right now to understand what the state of AI capabilities is, what the state of AI alignment is, and what the state of these AI systems in general is. How do they behave? What are their personalities? In what ways is it accurate to anthropomorphize them, and in what ways is it inaccurate?

We just really don’t have a good science of these systems that we’ve built: how they behave, how they’re going to act in new circumstances, and what their motivations are, if that’s an accurate way of talking about them.

Better evidence generation can look like a lot of things. It can look like just doing science on these systems to understand them better. It can also look like the incident-detection and incident-reporting work that Night Andale and other organizations like that have done.

I don’t know if you guys saw these reports about the German wiki that the OpenAI agents took over. That was work done by an independent organization that was just crawling the internet looking for evidence of AI incidents. That work needs doing because these sorts of incidents can happen and then just go unnoticed unless there are people who are actually trying to look into them, present that information to the world, and understand the story behind it. Evidence generation is also a diverse and important area of work.

Nathan Labenz

To what extent do you think organizations like Accenture often function as box-checkers rather than investigators? There’s a clear difference between an auditor that goes in to make sure that processes were followed, even if those processes didn’t yield the outcome that you actually desire, and an investigator who’s there to say, “I’m going to find this thing.” What is the differentiation between those two? And do you think the Accentures of the world are going to end up in box-checker rather than investigator mode?

Speaker 1

I know very little about Accenture, so I won’t speak to them specifically. I think it is a thing to be concerned about in general. Putting aside the specifics of the organizations involved, these third-party auditing groups will not have enough access or will not have the incentives to pursue really hard-hitting assessments of how safety is going at these AI companies.

I think part of why that is is that, in many cases, this is all very voluntary, right? To my understanding, the only legally mandated role for third-party auditors now is to ensure compliance with the self-regulatory RSP-style policies that these companies are mandated to have by state law in California, New York, and Illinois.

You can just put whatever you want in those policies, and they’re very vague in a lot of cases. So that doesn’t leave much scope or much work for these auditors to actually test anything in particular. That’s as far as the legally required auditing is concerned.

Now, when it comes to what labs might voluntarily agree to, I think there have been some encouraging signs in terms of what OpenAI and Anthropic have said recently about the sorts of access and the sorts of questions they want these auditors to answer.

Nathan Labenz

Now, will that actually happen? I think that depends both on how these companies actually implement the nice words and on how the auditing companies react to that. So I guess we’ll see.

So, you have awarded Jeffrey Irving’s team quite a large award. It’s structured, I think, as kind of a compute and other kind of award. Can you tell us a little bit about the award itself, the structure of the award, and how you guys came up with that structure?

Prakash Narayanan

Yeah. Sure. But I think what you said is correct. There’s a part of the grant that is for research expenses, the biggest of which is compute. But when I say compute, what I mean is both just renting raw GPUs and also paying API companies for their tokens.

Obviously, that’s becoming a giant part of all the grants we make. People increasingly want to run very large experiments, which require a lot of tokens or a lot of GPUs to run open-weight models. Also, very excitingly, people are starting to use a lot of AI labor, and that requires a lot of spending on tokens as well.

AI safety is sort of a funny discipline in that you can use tokens both on labor and on the subjects of the experiments. In some cases, you’re doing both: you’re spending a lot of tokens to have some AIs run experiments on other AIs. And so that creates these high compute budgets.

Compute being the biggest expense is also something where the actual spending on that varies by orders of magnitude over time and across grantees. Sometimes the way we structure these awards is that we want to be very, very generous with compute and err on the big side. We just want to give that as a lump sum: this is the money that can be used for research expenses and not other things.

That way, we are very comfortable with the uncertainty that maybe they only end up spending 1/10 of it, and that’s fine; maybe they spend all of it, then they come back and we’ll give them more.

Nathan Labenz

Jeffrey Irving has said that superintelligence might arrive in 2 or 3 years, and he proposes that we slow down because that’s too quick. Rather than discuss his views, how do you manage the uncertainty of the time available, and how does that affect what you fund and how urgently you fund something?

Prakash Narayanan

I think the boring answer is that we just have a portfolio approach: we fund some stuff that is only going to pay out on really long timelines, and other stuff that, if it’s valuable, will probably be valuable soon.

I think that this question of timelines to AGI or to ASI in some ways matters less for research prioritization than one might initially think. In some ways, the prototypical example of a long-timelines bet is funding work that is very theoretical or ambitious. You look at it and you’re like, “Oh, this is going to take a decade to pay off,” just because it’s really—some work in mechanistic interpretability, or Paul Christiano’s work at ARC, or these sorts of things, where we’re just going to need to make a lot of scientific or mathematical progress if this is actually going to bear fruit in terms of methods that we can apply in reality.

That sort of feels like work that is maybe implicitly a bet on longer timelines, but I think that actually isn’t really true. If we are going to have powerful AI soon, then presumably we’re going to have an opportunity also very soon to make use of gobs and gobs of AI labor. Especially on work that’s more mathematical, we’re going to be able to get 5 years of progress or 10 years of progress in 1.

The sort of work that would take humans a long-timeline amount of time to complete might actually be possible a lot sooner under this hypothetical that we’re going to have superintelligence or AGI or something very soon, so long as we actually make use of those AIs and take the opportunity to make a lot of progress very quickly.

Nathan Labenz

What would you say are kind of the most speculative things that you included in Project Tailwind, that list of requests for projects?

Prakash Narayanan

I think that what we included in the list was a little coarse-grained. But we have some stubs on the Tailwind list about new research centers pursuing more ambitious, more principled bets on alignment. That is something that we’ve spent a bunch of money on this year and that we are excited about.

Work like ARC—Paul Christiano, who is not a CG grantee, but we’ve funded other people working on the ARC agenda and working on stuff that is very similar to it—because we think that’s a valuable Hail Mary shot to include in the portfolio.

Also, Resolution, our biggest grant this year, is very much a bet on a kind of crazy thing that has never worked yet, which is: maybe we’re going to have some more principled theory of AI that’s going to help us have stronger reasons to believe in the alignment of our systems. So that’s something that we are very excited about, and I think is very weird and out there.

I think a lot of people have very understandable skepticism of people who are familiar with what has and hasn’t worked in machine learning over the last decade and recognize that, all along, there have been people in machine learning who wanted to apply theory more to AI capabilities research, and that just has not really worked very much compared to blind groping around and trying stuff, doing trial and error.

It’s like Noam Shazeer says: the success of these methods we attribute, like all else, to divine benevolence. That is the ethos that has actually produced progress in AI capabilities. And so I think a lot of people take the lesson from that that trying to do more theoretical or more principled work on AI alignment is pretty doomed also.

I think that’s a totally valid reason for skepticism, but we think the upside is worth it despite that, and so we’re taking that weird bet.

Nathan Labenz

If I put the current conventional wisdom to you—there’s plenty of money; it’s all about talent—obviously, Silicon Valley founder types are one profile that you would want to recruit. Are there other pools of talent that you’ve identified as strategic priorities?

Prakash Narayanan

Just to be clear, for the things that CG is supporting, and especially for the things on the Tailwind list, the money is not the bottleneck. The talent is. There are other things within AI safety where money is definitely a very big bottleneck.

In some ways, the most valuable profile of talent is someone who combines the virtues of a founder with the thoughtfulness and the inclination toward—and comfort with—speculative thinking and futuristic questions. That is more common in the AI safety world. That sort of take-it-seriously attitude toward asking questions about what’s going to happen in the future of AI, really making predictions, testing your hypotheses, and updating over time is a really, really valuable skill to us, because it allows people to get going years earlier on projects before other people can see the value of them.

An example I give of this is METR, which, before other people were thinking seriously about AI capabilities and the impacts of them and how to evaluate them, was very prescient in coming up with a benchmarking approach and time horizons that just worked way better and aged way longer than other people. That’s because they were really trying to take seriously what was going to happen in the future with AI and then work back from that to what sort of work they should be doing now that would age well and prepare for that.

That inclination is really, really valuable to us, and it’s useful in basically all profiles. It’s useful both in a person who might found an org and also in a person joining the existing orgs. That’s something that existing groups in the AI safety space value a lot, because it’s a big part of how they do their work.

Nathan Labenz

Can you do one quick double-click on where money still is a bottleneck?

Prakash Narayanan

I think the biggest thing is just that, for organizations and in areas that CG doesn’t operate in, money is a big bottleneck. If you find some person who you think is doing great work and, for some reason, they either aren’t a good fit for CG, they applied to CG and we weren’t able to fund them, or we up and decided not to fund them, I think there’s still alpha in that.

Especially as we move into making much bigger grants, that can sometimes come at the expense of covering all of our bases—all the teeny little uses of money that can be valuable. I think that’s the right decision for us to make, but it does leave some things unfunded.

Nathan Labenz

Prakash had raised the prospect of an Anthropic public offering and his expectation that some of the proceeds would be distributed through Coefficient Giving. So, what is your preparation process for handing out larger checks? What is your pathway for the next year, say?

Prakash Narayanan

Yeah, I mean, I have no idea what’s going to happen with the AI company money in terms of what the preparation actually looks like.

I think the biggest thing is just seeding organizations now that can then be ready to absorb lots and lots of money. And so, making small grants to new groups and then giving them a little bit of runway to prove themselves or not prove themselves and explode—that’s the biggest thing we’re doing.

And we would love to create opportunities for groups of people that can say credibly to big donors, “We have people, we have a project ready to go. We’re just waiting for you to sign the check.” We’ve got this shovel-ready project, and it’s going to take a billion dollars and it’s going to solve alignment, and it’s just a matter of giving us the money.

And that’s really not the situation right now, in that there are just not that many people in the space. And so that’s a big thing we’re trying to fix.

Nathan Labenz

Is that about funding people enough so that they can work on things at small scale, such that they can get a much bigger compute check at some point to scale that? Is that the kind of idea?

Prakash Narayanan

That’s the main thing I have in mind. I think the same logic also applies to growing the headcount of their organization, like having a group in place that has a great plan, where they can then absorb a lot of software engineers or ML researchers or data labelers or some other type of person.

But I think compute, like you said, is the main version of that, and especially, hopefully, AI labor, right? It would be really great if there were a bunch of well-scoped problems that these organizations had where they were just waiting on lots of AI labor to pour into them.

Now, we’ll see if we can make that come true, and I think there are reasons to be skeptical that that will happen, but it would be nice if it did.

Nathan Labenz

Part four: The price of compute. Wayne Nelms is co-founder and chief technology officer of OR, which publishes a price index for GPU compute built from cleared rental transactions rather than list prices. The Intercontinental Exchange has announced plans to list futures on it, pending regulatory approval. Prakash asked him to follow a real customer who lends against the chips and who loses money when prices move.

Speaker 1

So, generally speaking, you have a bank or some sort of financier that will lend money against the GPU asset or a host of AI infrastructure assets. What they hope is that they will make their money back over time with interest as a result of the operating activity of a neocloud. For example, these are companies like CoreWeave, Crusoe, Nebius, Lambda, et cetera.

However, these business models themselves are contingent on really 2 things: continued AI demand and, specifically, an ever-increasing or at least stable compute pricing estimate, right? These GPUs are financed in a variety of ways, including straight-line depreciation, which we can get all into.

So, the way OR provides value is that you can’t really hedge away this risk without having a benchmark to hedge and having a benchmark to trade. So, we created that benchmark, and we are actually leading the transactions on that benchmark for these institutions.

Nathan Labenz

How many transactions would you get in a month that you use to form the index?

Speaker 1

Yeah, so we collect over 1,000 transactions a day per index. And so right now we have 5 public indices. So it’s 150,000, somewhere in that range, probably per month.

We have a whole host of neocloud partners that we work with, effectively, that are contributing data to our index every single day. And the question is always, why do people feel the need to show that information, to display that? It’s because they benefit, right?

A lot of these players are looking for cheaper financing. And cheaper financing comes when their financier has more certainty about the future and can hedge that risk in the future. And so it’s actually this nice flywheel where you get more data, you can publish a better index, the financiers feel more comfortable, and are able to finance more GPUs.

Nathan Labenz

I ask how fungible the underlying asset really is, whether an hour on an H100 from one provider is the same thing as an hour from another. I think it’s certainly obvious when you talk to anyone deploying infrastructure or buying that an H100 or a B300 is very different depending on the OEM and the supplier. Everything about the hardware can be very granular and very different from an engineering point of view.

Speaker 1

However, our take has always been that there needs to be a benchmark that allows some sort of abstraction for this asset class. I think for us, what’s made the job much easier is that, largely speaking, we are looking at 1 chip producer, which is NVIDIA, which dominates the market and has huge market share.

When we look at the NVIDIA moat—and we can get to this later as well—you have the hardware, you have the software superiority, but I think the biggest thing that NVIDIA has in terms of a moat over other competitors is the financing landscape. When you’re looking to finance your neocloud and you’re thinking about what chips to put into your site, it’s just materially better to have NVIDIA GPUs. It’s just easier to underwrite.

And I think what we see as this market problem is: How do we allow every single person looking to deploy compute infrastructure this sort of almost protection for their lenders? We right now only reference NVIDIA reference architecture with InfiniBand, et cetera, et cetera. There are certain steps that we’re taking to standardize the compute that we’re tracking.

However, I will say that we are tracking a subset of the market, which might inherently have differences amongst those participants.

Nathan Labenz

In some commodity markets, there’s this idea of baseload, where you have a large amount of baseload, and then you have this idea of load that gets spun up on demand. But how do you differentiate the pricing between the kind of baseload versus the swing capacity? Because I imagine the month-to-month transactions are the swing capacity and not the baseload itself.

Speaker 1

Yeah, I love this question. So, yes, you’re exactly right. A lot of the capacity is long-term PPA-style contracts. And I think it’s really interesting to think why that is, right? We’re in this situation where AI capacity is so scarce.

I think the model providers, so OpenAI and Anthropic, have seen this coming forever at this point. And partly they’re to blame for the lack of supply that’s currently out there. But if you think about it, for them, it’s an arms race, right? Compute capacity planning is an arms race.

How much capacity can I lock up over the next few months, so when I need to train the next model, I have enough and I don’t have to go out looking for more? It also, of course, is naturally a zero-sum game, where whatever you buy, your competitor can’t.

Actually, the way we like to think about the baseload and swing—or day-ahead—market structure is: In compute, people are buying to fill the peak demand and then selling to satisfy the troughs, whereas in power, you are buying to satisfy the troughs—these long-term contracts—and then buying excess to satisfy the peaks.

So there’s actually a robust on-demand market that will continue to grow as these training labs are selling back excess, as inference providers are selling back excess. It’s this reallocation-of-compute question rather than purely a hedging question.

Nathan Labenz

How elastic is the market? Obviously, in an extreme example, Anthropic reportedly has paid up, I think, even a multiple of—

Speaker 1

Yes.

Nathan Labenz

—standard market prices to rent at high scale from xAI. How would you sketch out the curve? If I want 1 H100 hour right now, I pay X; what if I want 10 million? How much does the price change as my scale as a buyer changes?

Speaker 1

Yeah. So we like to think about this in 3 axes, right? You have price on one axis, duration—the length of the contract—on one axis, and then you have quantity on another one.

Generally speaking, as you increase the length of the contracts, the contract-per-hour value goes down, right? The logic here is you’re buying in bulk, right? You’re allowing someone to finance a data center because you’re buying 10 years for five years of capacity from them.

Generally speaking, when you increase the number of GPUs you’re renting, the pricing goes up, which is almost antithetical to this exact logic: I’m buying in bulk. Why am I paying more? And it’s purely because the number of suppliers that can sell all that capacity at once is very small.

You might be able to get access to a few nodes here and there, but when you’re looking for 10,000 or 100,000 GPUs that are all interconnected at some level, that is very, very tough. So certainly, these Anthropic–xAI deals are few and far between. They’re very much OTC trades. I would say they are indicative of what we are seeing as a general trend across even smaller markets.

Nathan Labenz

It is striking that you look at the graphs and they are all trending up. So that is seemingly in significant tension with the idea that the longer you buy, the lower the price is. Why are the sellers not AGI-pilled enough to expect that they’ll command a higher price for the same chips next year than they are right now?

Speaker 1

Look, I think the sellers of compute are the most AGI-pilled people because they have entered a business where they are selling a resource that hopefully rises in cost in the future.

But for a lot of sellers, it's not because they don't want to. I think there's a lot of suppliers of compute that are experimenting with shorter contracts because they know: Why am I locking in a 4-year deal if in 2 months that capacity is worth even more?

The issue here, of course, is risk. It's always this risk-reward ratio or balance that they're trying to play, where even early financing is contingent on having a long-term contract signed. I cannot finance a data center today if I don't already have 5 years of offtake from a credible counterparty signed, and because of that, I can't even enter into the month-to-month world.

So if you actually think about what the futures exchange and our product enables, it enables the neocloud, instead of locking up a 5-year contract, to sell month-to-month and hedge the future price potential on our platform or through our index, right? And that is potentially a much better risk-return kind of math for them, for their TCO, and for their lenders.

So I think the one thing that Nathan has also alluded to is that you have differences in pricing of near-term and much further out, like longer-term, and you kind of described the use of a lot of the hedging tools for financiers. The financiers are often concerned with the far end, not so much on the near end, where they have much more visibility, but the far end, where you have this depreciation issue. How do you think the index assists, or will assist in the future, on this end-of-life kind of depreciation issue, which is 4 years out—3 to 4 years out?

Speaker 1

Yeah, 100%. Like you mentioned, a lot of the risk in this space is certainly long-tail risk in terms of looking at the tenor. When you launch these futures on a futures exchange, a lot of the trading can occur in that long-tail section of the curve, right? These are monthly contracts on ICE. To be very clear, we expect to see lots of trading in the front month as people are buying and selling, especially after physical delivery. But I would also expect there to be a lot of trading in those 3-to-5-year spans, because that is where all the exposure is. That is where all the risk is.

And because there are now futures contracts that reference what spot prices will look like in 3 to 5 years, that is what any bank is exposed to, right? To be even more clear for people watching, you have a bank underwriting a deal where the GPU life is 6 years, but the initial contract signed might only be 4, or it might only be 5. And so there's always that: What happens in year 4? What happens in year 5? What happens in year 6, when there's not a contract already signed for that GPU, but on your balance sheet you've claimed that the GPUs have some value? So that is the area where you want to hedge your bets.

Nathan Labenz

Prakash raised basis risk: the gap between what an index says and what the asset you actually hold is worth. If that gap is steady, a trader can price it in. But the issue is going to be if there's non-static basis risk, and one of the issues with GPUs is this introduction of new technology, which obsoletes the old. That is what I think the non-static basis risk is, and I think it's pretty scary for financiers.

Speaker 1

We are tracking GPU-specific transactions. We have a B300 index; we have a B200 index. And so, actually, this entire risk of a financier being worried about obsolescence risk—that is exactly what you can hedge away with this instrument. That is exactly what you can look at: how the market is pricing the 12-month or the 15-month futures contract for the B300, right?

It actually, if anything, allows information to be spread wider, broader, and more evenly, because now you can see perhaps where the market is pricing the announcement of Vera Rubin or the announcement of the next chip, right, based on how the curve looks in the future. I think all these things are ways for information to be actually shared more freely, rather than for there to be any obfuscation of this information. And so, in effect, I think these markets are a great way to hedge that risk away.

Nathan Labenz

I also asked how long these chips actually last and whether the A100s from 2020 are near the end of their useful life.

Speaker 1

Yeah, I think chip health is a super interesting space to be in, and understanding what the failure modes are for the hardware itself. Chips will fail all the time for a variety of reasons, but I think what you're pointing to is the useful-life conversation of hardware. This is something that we wrote a paper on just 2 weeks ago, I think, about how open-source, open-weight models that are more price-elastic will often route to places where the cost per intelligence or cost per workload makes more sense, and oftentimes those are the older pieces of hardware because they're just less marketable.

And so we actually see the useful life extending beyond 6 years as a general thesis. How much further than that, it's kind of unclear. However, I will say that what we see in the forward market and the term structure is that there will always be some sort of premium above zero, certainly, and above some terminal value of power, right? The terminal value of an A100 is not just the value of the power that powers it, because the A100 will always provide some computing power as a fraction of all computing power in the world.

Over time, I think you have a lot of interesting use cases, like physical AI, one of many, where these older chips might become more effective. However, today we still see that these older chips are still being used, still being rented, and have maintained relatively constant demand in the last 12 months.

Nathan Labenz

Wayne signed off a few minutes later. In the closing with just the 2 of us, Prakash came back to the index with an account of how the biggest compute deals get priced. The figures are his, from what he has heard.

Prakash Narayanan

What I've heard is that your price of compute is dependent on whether you have financing or not. If someone needs to use their balance sheet, they get to dictate the terms. If you don't have to use your balance sheet, then all of a sudden you have high pricing.

So Elon, for example, had already used his own balance sheet to build out the clusters, and he had the clusters available—you could take them tomorrow, right? And they didn't need to provide him a long-term contract. They could do a month-to-month contract, and they could walk out at any time, right? So that meant that this is not financeable. A bank cannot come and give Elon a loan based on this contract from Anthropic, because a bank wants to see that if you're going to pay back in 4 years, you're going to have cash flows for 4 years. I need to see that up front.

So Elon is able to charge high prices because he's using his own credit capacity. He's using unsecured credit at the holding company, and he can say, “Look, you're getting paid $80 million or $90 million a megawatt. You give me the $50 million, and you're still making money, right? So that's a good deal.” And so Elon got $50 million a megawatt, right? So I think it starts to change when you need the other guy to give you support.

So if Elon said, “Okay, I'll build it for you. I'll build it in a year, or a year and a half, or 2 years. And in order for me to build it for you, you need to give me a 5-year contract. And if I build it for you, you will definitely buy the compute and pay me. And then I'm going to take that and go to another bank and give it to them, and they're going to lend me money based on that.” Then it becomes like it's not really Elon's money. It's really Anthropic's money, anyway, and the single credit risk is going to be from Anthropic's side. And that's really who the bank is relying on for repayment. And so then it becomes different pricing.

So then Anthropic will say, “Okay, how much is it going to take you to build it? How much is your interest rate going to be? And then I'm going to give you, let's say, a 20% profit margin on top of that.” Right? So this is what CoreWeave is getting, for example. So if it's going to cost you $15 million or $20 million to build, they'll be like, “All right, $24 million or $25 million per year,” right? So that's the key difference: whether you have the money or not up front.

Those deals don't hit the index pricing. The index pricing is based on willingness to trade. These are all tradable prices, and that stuff doesn't get traded. And so the people transacting on this index are the lowest payers. They're not Facebook. Facebook can afford to pay, like, $100 million a megawatt. They're not going to be trading on the OR index because the OR index is for people who are willing to sell compute at, you know, like $10 million or $15 million a megawatt. They're willing to sell the compute. Facebook isn't going to sell you compute at $15 million. They're willing to buy it at $50 million, right? So these are, in some sense, the lowest prices. But it's immediately available.

Nathan Labenz

Part 5: What the machines are actually doing.

Nick Gillian is co-founder and chief technology officer of Archetype AI, and before that led machine learning research on Google's Soli radar sensor. Archetype's model, Newton, is a foundation model for sensor data. It takes raw streams from radar, vibration, electrical current, and cameras and reports what a machine or a worksite is actually doing. I asked him to walk through Newton from the bottom up.

Let's start with the data. How much data is out there that you were able to just collect on your own? And then I imagine you must have had to establish a sort of partner network to bring in a bunch of data.

Tell me about the data process.

Speaker 2

Yeah, we're getting close to 1 billion hours now of physical AI data that we've been able to scrape, gather, and collate. In many cases, if you look on the web and look at the different types of physical data that's there, it's very sparse. There's very little multimodal data. It's not synced well, and it doesn't cover things well for most techniques, particularly classic supervised techniques, where you're basically building one model and it has to have a fixed number of sensors, and they all have to be captured with the same sample rate, et cetera.

That's a massive problem. We've been trying to flip that on its head and say, if we have all of this physical data captured from different sensors and different systems with different contexts, how do we actually build the most advanced self-supervised techniques that can pull in that sensor data and extract from it in the same way LLMs have done with language?

We're able to then join that and mix it in with a small amount of data we can capture from partners and, in some cases, that our customers can give us, to be able to really extract and understand and map that to physical language and map that to physical control. There's quite a bit of resampling and really understanding missing values in sensors. Is that because the sensor had an issue, or is it actually a feature that helps explain that the machine is about to break?

The reason the sensor is giving you all these NaNs is not because the sensor is broken; it's actually a feature of the machine in most cases.

Nathan Labenz

The area that I've studied best that has done something kind of similar has been at the intersection of language and images—photographs. If I had to borrow some of their terminology and map it onto what I understand you're saying, this would be a sort of late-fusion approach, where you have a language model, then you do a pretrained, from-the-ground-up specialist encoder for these different modalities, and then you're doing some sort of cross-attention between those.

That's my understanding, at a high level, of what the image folks have done. How much have you been able to follow in their path? How much of the techniques you've seen them develop have worked for you, and how much have you had to go back to the drawing board and reinvent?

Speaker 2

We've definitely tried these approaches, and we've had some success with them with other types of sensors—radar, time series, and so forth. But the big issue that we've found is twofold.

Number 1, there's an extreme amount of image-language pairs on the internet that can easily be scraped and built into these very large datasets, where you have this pairing between an image and a description of that image. This does not exist for radars. It does not exist for time-series sensors. It does not exist for all the complex systems that we go and talk to our customers about.

So we've had to look at different approaches. How do you, for example, take time series and actually align that with human language? That's a large part of our research work: how to actually solve this sensor-alignment problem.

There's also the case where, with a single image, you may actually be able to get the full context of the scene, for example, and then align that with language. But with a time series, you actually need to look at a temporal window. What does a single latent mapping to a single word really mean in those cases? It's not like you have a time-series signal that says “dog,” and you can align that with the keyword “dog,” and so forth.

There's quite a bit of work we need to do to solve this alignment problem: How do you generalize this across many sensors? That's a key component we're building inside our models, which makes them different from a standard VLM or other techniques, because we simply don't have that corpus of, say, time-series and language pairs to apply the previous recipe you mentioned.

Nathan Labenz

Prakash, talk about one deployment—a project with Kajima, the Japanese construction company—and what the inputs and outputs of the model look like there.

Speaker 2

In that case, we're working with Kajima, a very established Japanese construction company. They build islands and move mountains. The project we were working with them on was literally taking a river and moving it so that it wouldn't flood.

It was a 5-year-long project. They had sensors the whole way down the river as they moved down the river, dredging it and widening it, and basically mitigating the risk of flooding. The big problem they wanted to solve was that the company is outsourcing and subcontracting a lot of the work to other teams, and they wanted to measure whether those teams were actually doing the work.

If they had to dredge the river for 5 hours, did they actually do 5 hours of dredging that day, or did they only do 2 because, perhaps, the digger was in maintenance mode that day and they couldn't use it, or because of weather conditions?

For that project, the input to Newton, our world model, was multiple cameras around the site—typically 3 or 4 cameras looking at a specific section of the river. When you look at those cameras, you have a barge, a boat floating on the river. On top of it, you have an excavator, and at the end of the excavator is an end effector that is either drilling or has a bucket that is actually dredging.

They want to know how many hours that excavator is actually dredging the river—not drilling the river, not operating as a crane, and not doing other things. Essentially, the inputs to the system are camera sensors, river sensors, and hydrometric sensors that indicate how high the river is.

At different geolocations along what I think was a multiple-kilometer site, we had information about where the actual work was supposed to be happening that day and what the weather was that day. We basically have a bunch of cameras and a bunch of time-series sensors with weather and river information. It's fed into the model, and the model is essentially outputting what looks like a Gantt chart.

It's outputting the state of that team. Where is the barge? Is it parked on the side of the river? Is it moving into position, which can sometimes take an hour? Is it now in position and they're starting to dredge, or is it just there and stationary with nothing happening?

They can then start to map these work charts over multiple years of data and look at bigger patterns that a human could never spot. One of the patterns they found, for example, is that if there's a storm, the work on the day of that storm goes down. Everyone would know that. But one of the key things they found is that multiple days after the storm, when the weather had cleared up, the team's productivity was still very low.

The reason was that the river water coming down from the mountains took several days to reach the end of the river and go to the ocean, which is where the site was. There was a lot of debris and churn coming through the river multiple days after these larger storms, which was slowing the teams down quite a bit.

They'd never spotted those patterns before because they could never look at these aggregate patterns. But Newton was able to spot those patterns for them.

Nathan Labenz

How much does Newton generalize? When you go into a new site, do you have to take all of the data and fine-tune the model?

Speaker 2

We're building the model so that, out of the box, you can plug in different sensors and different configurations, and the model can automatically adapt to those as much zero-shot as possible. For a subset of users, they can directly do that.

The typical analogy I give customers is: If you showed this task to an average person on the street, would they be able to detect the safety violation, or would they be able to detect that a signal looks strange or anomalous? If the answer is normally yes, that usually means Newton can do it out of the box. It doesn't need a master technician's level of expertise to spot it.

For cases where the piece of equipment the user is working with is very specific to that factory, for instance, or where they have parts with very internal and unique names for their processes, that's typically where fine-tuning—or bringing the customer's knowledge base and injecting it securely into Newton—really helps.

Our platform allows customers to take our model and run it directly on their infrastructure, so no data leaks out of their networks, either on-site, on-premises, or in their secure clouds. They can then take their own data and do some amount of fine-tuning to customize the models.

And typically with customers, we see that on the order of a few thousand samples are enough to take the general model and adapt it very quickly to that customer's use case.

Nathan Labenz

Last question for me is about your vision of superintelligence and how you might relate to the frontier companies as we go forward in time.

Speaker 2

What I think is—geez—the reasoning is already getting to superintelligence-lite, with all these Millennium Prize Problem-style results happening through super-long chains of thought. Who would have believed you could get there? On the reasoning front, we've come awfully far.

What we still don't seem to have is a good intuition for the many domains that matter. What I expect to happen is something similar to what has happened with images, where it's one thing to put an image into a model and get a caption out, but when you can get the altered image out, it's clear that the model understood both your verbal instructions and the image in a very similar way to how we intuitively see the image—which is something we're really good at—then it feels like, “Wow, okay, it's really got that domain.”

Nathan Labenz

Mhm. So my superintelligence vision is basically to do that across a ton of different domains, many of which we don't have native senses for, with what you're building being a great example of that. People don't have a native sense for how to integrate 10 signals from 10 different sensors.

Do you see a role for you in superintelligence as being the group that builds the deep intuition for this domain? Does that ultimately get joined with elite reasoners through a bidding war between a couple of frontier companies for your company, for example?

Speaker 2

I definitely see the part that we can play in a larger superintelligence system. That's how we're seeing where we play in this bigger superintelligence loop: we're very focused on understanding the world around us and bringing that back to natural language and machines so these things can talk to each other.

If we can do that and other folks are solving digital intelligence, those things will connect together very well. We're building physical agents that can be deployed out to the ecosystem of the world. Those physical agents can then talk to other agents and bridge that gap between the digital world and the physical world.

How much of that is actually done through an agent-to-agent, post-training mechanism, and how much of that is pulled into pretraining, we have yet to see. Particularly when it comes to more nuanced sensors like radar, LiDAR, infrared, and all these things that don't look like human-visible images, it'll be interesting to see how that's brought back into these bigger pretraining loops versus whether it's more of an agent that can go solve that problem, tell you what's happening in the physical world in real time, and let you interact with it.

Nathan Labenz

Part 6: The dish is not the body. Andre Yorgescu is co-founder and chief executive of Vivodyne, which builds robotic labs that grow pieces of human tissue—lung, liver, tumor, and more—from primary human cells and tests drugs on them. In August, it opened what it calls a human biological data center: 12 robotic labs that, by the company's count, can grow more than 3 million tissues a year. Prakash had the first question.

Prakash Narayanan

How do you actually figure out what the dosage is and whether that is what the cell would encounter in the body?

Speaker 1

Basically, everything that's been done preclinically to date, outside of testing in animals, is just dunking some substrate—whether it be an organoid or cells in a Petri dish—in a test compound that you want to test. That's exactly the problem: that's not how transport happens in the human body, right?

As a perfect example, a lot of cell therapies designed for solid tumors will completely kill that tumor in a dish, and then those same CAR T cells, when they go into a patient, just flow right by the tumor through the blood vessels. They don't even know it's there, right? And thus, the challenge is modeling the transport of that therapy into the tissue: do they recognize the tumor, and do they concentrate there?

In Vivodyne's tissues, even the blood vessels—the capillary bed in that tissue—are self-assembled. These little blood vessels grow within the tissues, and we can perfuse directly through them. Whenever we dose a drug, it is always through the blood vessels native to that tissue.

The transport through the interface between the blood vessels and the meat of the tissue, including all of its parenchyma and functional subparts, is just like you would find in normal tissue. The density is the same, right? And that is just as big a part of the question—does the drug molecularly bind to its target and perform this effect?—as whether it even gets there.

Nathan Labenz

Then, my question on scaling laws: I'm very intuitively bullish on this. Just go collect an obscene amount of data on this kind of tissue-perturbation paradigm. It makes a ton of sense to me.

The skeptical view that I've heard from time to time is that you can never have enough data. The world is too big, and biology is too complex.

Speaker 1

You can do that, but you just won't be able to generate enough. The challenge of this argument—that it's just a dataset-size thing—is that currently even the state-of-the-art virtual cell models that exist saturate at a very, very small fraction of the input data fed to them.

You can have all the data you want, but their performance saturates after a couple percent. This is both published and what we hear on the ground in these perturbation-based models.

The reason it gets so difficult is that when cells are growing in a dish, which is the mode for this sort of Perturb-seq, they are so far removed from all the feedback loops that are natural in people that they're just trying to colonize that piece of plastic. There's a stiff substrate. Normally, when cells encounter that, they're like, “I have to encapsulate this,” and so they're just growing. Their only goal is to proliferate.

If you knock out genes in those cells that would otherwise affect how they perform their natural function, there's really no big difference, right? They're still just dividing. You end up learning this view that perturbations don't really do much of anything, or it's unclear what they do, except if a gene is required for that cell to live, in which case it kills the cell. There's not even causality to extract from that.

We think the fundamental bottleneck to having a causal understanding of biology is that we need to perturb these tissues in the complex environments that they naturally call home. The image in my head is that we're trying to populate a map and design a GPS through it. We can say, “If I want to get to this state, I need to go down this road and then turn left with a second perturbation, and that'll bring me to this healthy state.”

Nathan Labenz

I asked how you get from a sample received to something placed in the system and growing.

Speaker 1

When we grow these tissues, a really big differentiator for our approach—and, I think, a critical mechanism you need to get realism in this—is that we use primary cells, which are mature, differentiated cells that perform our natural organ functions in our bodies.

That's as opposed to starting with some stem cell completely absent from all the conditions that would tell it how to mature into a differentiated cell, and then sledgehammering it with cytokines to turn it into some cell type that we think is accurate to what happens in the body. We instead go to the mature source, the ground truth, and mix together all of these different cell types that comprise, for example, the lung, at a very high density.

Within that injected tissue, we see that these cells begin to reassemble into the native structure of that human tissue, with some little herbs and spices that we introduce. But the responsibility for forming the native structure of human tissue is on those cells that already know how to do it.

It's honestly kind of like injection molding into a chamber: we shoot in these cells at high density, and then, over the course of several days, they reassemble themselves. They self-organize to make blood vessels, to partition into, in the case of the lungs, the airway where there's air, and then the parts where there are blood vessels, fibroblasts, immune cells, and all of this.

We're relying extremely heavily on the cells to be able to direct their own self-assembly at these very small scales.

Prakash Narayanan

Can you tell us a little bit about the mix of in silico via the foundation model versus in-cell, in actual tissue experimentation today, and how that ratio is expected to evolve? Presumably, the big value is that it's a lot cheaper and faster. Even if you've got all these great wet-lab technologies just plowing things through, the foundation model is supposed to be where you get the insane acceleration, right? But then, of course, you've got to ground and validate.

Speaker 1

That's exactly right. So, if we look at it like a reinforcement-learning problem, my actor is the experiment, right? I have this tissue, and I can perturb it in very many different ways. The goal is to make it healthy, and I can instrument the state of that tissue very deeply, so I know the value function is very well-defined in terms of the state of that tissue.

The question is, what should I try next? In the large rollout that I can perform, what are the most useful things to try next that add the most information to my data set? The physical experiments serve to fill in gaps in which the model is most uncertain.

The value of the model ultimately—why even make these foundation models—is that the combination space, especially when we move into 2-target or 3-target therapies, is so huge. If we have combination therapies where I take 2 drugs because they're going to work much better than 1, you could have this whole planet be Vivodyne systems growing human tissues, and you'd still not get anywhere close to the sheer space of things to test.

We need some way to narrow down to a couple of times the total number of clinical trials in America—the experiments that we should physically run to test our hypothesis. These foundation models allow us to predict what happens when we drug this thing and that thing, or all these things together, or this thing and then that thing. We optimize and iterate in that space, and we find, for example, the 1,000 most likely things that will move this tissue closer to this healthy state.

Then we close the loop by testing it. We see where that rollout goes, and then we just do that on a loop. The purpose, very fundamentally, of these models is to tell us what to do next. The purpose of this experimentation is to figure out whether the thing that we thought we should do next actually does move us closer, or how it deviates so we can correct.

Nathan Labenz

Prakash asked about the hardware. The tissues grow on what Vivodyne calls a tissue disc, and the company recently announced a second generation of it. He asked what changed from the first disc to the second. Andrei started with the robotic lab the discs go into.

Speaker 1

Imagine something the size of a wardrobe—about the footprint of a large desk and about 8.5 feet tall. These are self-enclosed labs, so there's a fridge and freezer, a 3D-scanning microscope, and a big robot arm that can use all these different tools.

We want to optimize the surface area of those discs for tissues. In our second form factor, we have an increase of anywhere from 2 to 4 times the density of tissues that we grow on each disc without any sacrifice. They're actually much larger than the version 1 discs. We increase the number and the size of these tissues, and we offload a lot of the responsibility for how they're grown to the robotic system around them.

The challenge there, and the kind of work that allowed us to do this, is a lot of work in software on the orchestration layer. If you had 10,000 experiments to run in parallel, this isn't like having a different gradient of dosing for some drug that you can do on a plate. I'm changing the collection point of data. I'm changing the interval between dosing with some drug. I have a first-stage regimen and a second-stage therapy. I'm taking all sorts of different readouts of these tissues. It becomes a super-complicated experiment.

I can define it quickly. I can explain it out loud and say, "We want to test bone marrow and airway. We want to dose these drugs in this order. We want to branch through these concentrations, and I want 3D imaging, single-cell sequencing, and proteomics." By saying that, I can define the study decently unambiguously. Given the system of tissues that we grow, it's pretty unambiguous what I would want to test in there.

But then the robot has to be able to compile that kind of zipped representation of the study. It has to decompress it into this compiled list of millions of actions that the robot has to perform. It has to remember that if it took the lid off a plate, it should put it back, as an example.

If there are multiple trays that it has to handle and they're all out on the working deck, maybe it should deal with multiple of them in advance. If multiple pieces of plasticware are on the same tray and they're coming out of the freezer to thaw, you want to make sure you're not thawing something that should not be thawed too early.

This compiler behaves a lot like a compiler in code, right, for C, where you're managing the prefetching of information for memory and branch prediction—all this stuff that, thankfully, our operating system and hardware handle for us. Improvements in our orchestration mean that the system itself—the hardware and the automation—can do much more of this. More and more of that space on these tissue discs can be used for actual tissue biology instead of having to integrate the little helpers on there.

Nathan Labenz

Toward the end, Prakash read Andre a grand challenge for the field, which he credited to the Annual Review of Pharmacology and Toxicology: "The innovation of a universal in silico or organoid framework that forecasts rare, patient-specific adverse events with more than 95% accuracy before first-in-human dosing" is one of their target challenges. Vivodyne is obviously on the path there. When do you think this challenge will be solved?

Speaker 1

This is a hot take, but I feel like this is almost like picking the color of my Mars suit currently, to me. How about we just fix the bigger problems?

The big problem, by the way, is that we are working specifically on that challenge of very, very rare outcomes. But far more regularly, we have drugs that would have huge potential in their efficacy in patients but that have routes of toxicity that are pretty common. Many patients die from it.

I mean, many of the cancer drugs that are going to the clinic today—the risk is not, "I've cured 999 people of cancer, but this 1 patient didn't work so well. It was more toxic." The problem is, in all the patients we've dosed, they barely had any improvement in their cancer, but it gave them all of these side effects or killed someone from liver toxicity or something.

Nathan Labenz

Andre signed off a few minutes later. In the closing with the 2 of us, Prakash put the 2 guests side by side. I liked how you had basically 2 different approaches for data. On the 1 side, you had, "Let's take all the messy data, let's feed it into the machine, and try to make sense of it." On the other, Andre is like, "There's just too much messy data in biology. I'm going to build my own data-collection effort, and I'm going to monitor every single thing that goes in and comes out so that I know exactly what's going on."

I think Andre probably has a good chance of solving this thing. You need a little bit more GPU power and a little bit more imaging. I think the imaging granularity is probably not there yet, but you can definitely see it if they manage to scale a little bit more.

The real question is: do you need to do 3 billion samples, 30 billion samples a year? Do you need to get everyone's genetics, or is your scaling going to get you to a point where, at 30 million or 300 million or wherever, it just starts to generalize a lot more and you stop needing so many different samples?

Prakash guessed it would be 2 or 3 more years before AI gives biologists something they can really use. I took the under. I think we're going to see probably a reversal of the order of utility and Millennium Prize solutions in biology relative to what we've seen in math.

The issue with math is that we already had a ton of math that was interesting to mathematicians but not super useful. To push the frontiers of math, you kind of had to get into the not-useful domain of math.

In biology, we have tons of useful stuff that is still not even that well understood. It's a far more empirical than theoretical domain. I would expect a tremendous amount of utility to emerge before we have anything approaching a full, mechanistic, or fully causal understanding of what's going on in biology.

Prakash Narayanan

And to some degree, obviously we'd love to have that, but we don't necessarily care, right? At the level of, does this cancer drug work on this patient's tissue? If it works, you don't have to have a mechanism, right? That's not part of the clinical-trial approval process. It's nice to have, not required.

Nathan Labenz

Prakash also made the case for specialists. A model like AlphaFold does its one job better than any generalist. He expects that to hold. And the version of superintelligence people fear is a single model that does every job better than the specialists and depends on no one.

My answer was the pattern as I see it so far. New kinds of data get pulled upstream into the biggest model, and efficient specialists get distilled back out of it. But, yeah, put me down for one that has not seen a reason yet to believe that all these modalities don't end up integrated at the frontier. In a way, I wish it wasn't going to go that way. I do think safety through narrowness would be really nice.

That's why I'm excited about Jev, in part, because it sort of has a very niche role to play. You could build all kinds of scaffolding around it. I was a big fan of Eric Drexler's reframing superintelligence piece, also known as Comprehensive AI Services. My hope for safe superintelligence is something like that.

Just based on reading the tea leaves of Ilya's comments, could we possibly get a model that starts as kind of a stem-cell generalist and then matures in a way where it gets really good at its job, but also settles into that niche in the way that mature cells don't revert and turn into other kinds of cells, and that they're kind of limited now? They've lost their pluripotency, so to speak. I love all that stuff. I would like—I hope it goes that way. For efficiency, it probably will.

But at the frontier, I just have not seen anything at all yet that makes me think that the single biggest, best teacher model you could make couldn't just learn it all. And given what we've seen, that starts to be a bit of a scary beast.

Prakash Narayanan

No, I agree. It can learn it all. But I think learning it all at the most efficient compute rate, and inferring at the most efficient compute rate—which is what you alluded to with the distillation—I think that doesn't happen, right? So energy efficiency is key, because it means that you have this diminishing-returns curve, as in economics, for getting bigger.

The doomer scenario is less likely to happen because you will get a more intelligent model, but it's going to be less efficient to start out with and it's going to be more energy-consuming. That means that it's going to need more energy expansion to get there, but the other models are going to be more efficient than it and more numerous. But I think the singleton example is not likely to happen.

Nathan Labenz

From your lips to God's ears, as always. I think that's a real constraint, for sure. If you made a—let's just get ridiculous—you made a quadrillion-parameter model, you certainly would find it a bit slow and a bit expensive, and you couldn't run a million agents at that scale today. So I do think there are some practical limits there.

But then again, remember what we talked about on Wednesday. There's going to be more compute installed over the next 12 months than currently exists in the world.

Prakash Narayanan

Yeah. So we are definitely going to test for at least a couple more years whether just bulking up on compute solves all—quote-unquote, solves—or perhaps creates all these problems.

That is the week. Tell us what worked and what did not. See you in the morning. If you're finding value in the show, we'd appreciate it if you'd take a moment to share with friends, post online, write a review on Apple Podcasts or Spotify, or just leave us a comment on YouTube. Of course, we always welcome your feedback, guest and topic suggestions, and sponsorship inquiries, either via our website, cognitiverevolution.ai, or by DMing me on your favorite social network. The Cognitive Revolution is part of the Turpentine Network, a network of podcasts which is now part of A16Z, where experts talk technology, business, economics, geopolitics, culture, and more. We're produced by AI Podcasting. If you're looking for podcast production help for everything from the moment you stop recording to the moment your audience starts listening, check them out and see my endorsement at aipodcast. And thank you to everyone who listens for being part of the cognitive revolution.
