Nathan Labenz
Today my guests are Jeremie and Edouard Harris, founders of Gladstone AI and authors of the sobering new report, America's AI Superintelligence Project. Many listeners will recognize Jeremie from the Last Week in AI podcast, a show that I've repeatedly recommended and continue to listen to regularly myself. Jeremie and his co-host Andre do a consistently excellent job of getting into real and often quite technical depth on the most important stories in AI, with Jeremie in particular presenting a very sophisticated and nuanced perspective that balances genuine enthusiasm for AI upside with sincere concern about loss of control and other existential risks.
Today's conversation offers a more zoomed-out perspective on what is arguably the core dilemma of our time. Racing ahead with AI development to stay ahead of China increases the risk of loss of control and other catastrophic accidents. But at the same time, insisting on adequate safety standards and allowing time for safety research to develop potentially seeds technological superiority to a rival nation that we can't currently trust.
This problem, which they describe as overconstrained and which they admirably face head-on, creates profound challenges for U.S. policymakers, especially as AI capabilities continue to advance rapidly toward potentially superhuman levels, as U.S.-China tension continues to escalate, and as the national security establishment begins to consider a Manhattan Project-style effort intended not just to maintain technological superiority, but perhaps even to achieve enduring strategic dominance.
Rather than taking a side in the debate regarding whether such a project is wise in the first place, Jeremie and Ed set out to understand what would actually be required to achieve the level of security needed to protect both the infrastructure and the secrets of such a project from our national adversaries. They do this not just by reading and theorizing, but by spending time on location at data centers and elsewhere with seasoned intelligence officials, special operators, and other technical experts.
The resulting report forces us to confront the gritty realities of security, espionage, and geopolitical hardball that absolutely must be factored into any serious discussion about the future of advanced AI. Drawing on their unique access, Jeremie and Ed paint a picture of a U.S. critical infrastructure stack—from the electrical grid to the data centers to the highly international research teams—that is far more vulnerable to disruption and espionage than most realize.
They argue that while perfect security is practically impossible, an acceptable level of security for such a sensitive project requires a deterrent strategy that both raises the cost and observability of adversarial action and credibly threatens controlled but meaningful retaliation. While some will no doubt read this as a blueprint for victory in the AI race, others, myself included, read it as more of a warning.
The sheer number of vulnerabilities we'd need to patch to effectively prevent China or other sophisticated actors from stealing our AI secrets, the compromises to core American values this would require to operationalize, and the risks from the centralization of power that all this would presumably entail are, for me, reason enough to steer away from such a national project.
Instead, we should try whatever measures we can, however unlikely they may be to succeed, to create both the technology foundations and the geopolitical conditions necessary to begin to rebuild trust with China and the international community at large, such that humanity can work together both to safely develop advanced AI systems and to manage the societal transitions that are likely to result from their deployment.
Is that naive on my part? Maybe. But hopefully, after listening to this conversation, you'll agree that it's not uninformed or close-minded. I am very much open to changing my mind. But for now, with the latest generation of intensively RL-trained models showing so many of the bad behaviors that AI safety theorists have long warned about, and with no particularly low-risk plan available anywhere, I continue to remind myself that we have far more in common with the Chinese than we do with AIs.
And I personally would choose to bet on our ability to find common ground with our fellow humans, if it proves truly necessary, over our ability to not only control but effectively steer the future with hastily built superintelligence.
Jeremie and Edouard Harris, founders of Gladstone AI and authors of the new report, America's AI Superintelligence Project. Welcome. Oh my God, it's so cool to be in your podcast bubble. I've been watching.
Yeah, I'm excited to have you guys here. I've been really looking forward to this for a while. Fun little backstory, I guess. First of all, I listen to you, Jeremie, all the time on Last Week in AI, which I think is one of the best podcasts in the AI space. I am pretty much an every-episode listener and a frequent recommender for folks who want to keep up.
I think you do a great job with that and really appreciate all your analysis there. We've been trying to make this happen for a while. I think you're the only guest that I've ever had booked and then got preempted by Joe Rogan, which, you know, millions have heard you on Joe Rogan around about a year ago now.
So I'm glad that we're finally circling back and making this happen. The timing is perfect because the sort of manifesto, mega-plan treaties on what the hell are we going to do about AI are coming fast and furious, and your latest contribution to that tradition is a really provocative one. I'm excited to get into it with you guys.
Jeremie Harris
It's fantastic, and we've been looking forward to it for as long as you have. I've been a fan of The Cognitive Revolution for a long time. I find it a little bit funny because when I watch what you've been covering on the show and then I think about you watching Last Week in AI, I'm like, "What is Nathan getting from us?"
The quality is so high. You guys do great technical deep dives, which is something that's so useful. Anyway, I'm super thrilled to be in the podcast box.
Nathan Labenz
Well, thank you. That's flattering. But I do really find a lot of value in Last Week in AI. I think your analysis is really good, and nobody can really be comprehensive in the AI game anymore. Everybody's got to take some pitches, for sure, but I do think you choose a lot of great stories.
Sometimes, if I do fall a little bit behind, I'll just go back and listen to a couple of episodes and say, "All right, which stories in there do I really need to make sure I follow up on?" To a significant degree, you're almost kind of like an assignment editor for me at times, where I'm like, "Yeah, that one is one to double-click on." So you can maybe flag some items for Nathan there.
Edouard Harris
Yeah, I need all the help I can get, certainly in today's world.
Nathan Labenz
All right, so let's start with this. There's a lot of talk about AGI. There's a lot of talk about superintelligence. Not a lot of talk about what superintelligence looks like. So I would love to start off with just getting a little bit of your intuition for what the hell superintelligence is, what we think it can do, with obviously some error bars, how soon we should expect it, and what kind of strategic advantage this technology might convey on one nominal superpower relative to another.
Jeremie Harris
On the question of what superintelligence is and what kind of advantage it could convey, I'll try to tackle both of those at once, and then, Ed, maybe you can talk through the timing and fill in any gaps.
In terms of what superintelligence is, right out of the gate, people have all kinds of different definitions that are slightly different from each other. I've heard one person say that a calculator is superintelligent in the narrow domain of doing prescribed math operations. I'm like, okay, but that's not necessarily a useful definition.
The definition that most people have coalesced around in the space is something that can do almost everything, if not everything, far, far better than the best human. From that perspective, it's something that is going to do stuff and is likely to have goals that are not comprehensible to us.
So, from the standpoint of what kind of strategic advantage a superintelligence, as opposed to an AGI, which is more human-level, gives your country, it's actually, at least under current conditions and given our current ability to control and align these systems—
Speaker 1
It's not clear that it gives you much of an advantage because it's going to do its own thing if it's that much smarter than you, unless you are much further advanced than we are today in the science of how to actually control it.
Speaker 2
Yeah, I think that's definitely one factor. And to go back to the definition piece, too, there's an interesting implication to this universality of the superintelligence definition that people throw around, right? Smarter than all, or more capable than all, humans at all things, right? Something that absolute really implies a process—you can almost start from that and work your way backward. It implies some kind of trajectory.
If you have an AI system that is necessarily better than human beings at autonomous AI research, then that means, necessarily, it is the product, almost by definition, of an intelligence explosion. Whether you measure that explosion in years, because hardware is hard and the world is slow and wet and biological, or in weeks—or whatever that kind of explosion time horizon is—it would be something that, in retrospect, everybody would agree was like a hockey stick, right? So, really, in a weird way, I think the process itself is more helpful than the definition of the thing. The thing is, if you have superintelligence—if you have something that is genuinely more intelligent than all humans at all things—you have something that can design better biological weapons.
You have something that can design the ultimate cyber weapon. It is the national security technology. It is the decisive source of strategic advantage. Full stop.
That's just what it is. That's what is implied. Then you can ask about how we get there, whether it can be controlled, and so on. That's the raw capability that comes with the package.
The controllability piece, the alignment piece, is an open question, and this starts to lead into our thinking strategically around the constraints that America faces today as it enters into what is ultimately and clearly a race with China on AI. You can't assume that you can control the system; that is an important thing at that level. At that level, if you're something closer to human level, even that is kind of fraught, right? Because even if it's qualitatively human level, it's maybe thinking 10 or 100 times faster than a person, and it can maybe be deployed on a data center at bigger scale.
So that could still give you qualitative differences. But still, if you have something that's humanish level or AGI level, you can kind of imagine maybe having some kind of temporary, slippery control over it and having it really contained and aimed in a particular direction. If you're talking superintelligence, I mean, I think most useful definitions are: this is a thing that outstrips you to the same extent that you outstrip a toddler, at the very minimum. And if you think about how a toddler is going to keep you contained over any meaningful span of time, forget about it. They're just not going to even think in the right ways to do that.
Maybe this is a good point to just jot down, or plant a flag on, the constraints that we see that kind of inform our picture of the world. It's like the signposts that guide us as we go through this investigation and the work that we've done over the last year. This really seems like a pretty overconstrained problem, just to be clear.
So, number 1, we have AI superintelligence that seems—again, it's unclear, right? This is an open question, and we should dive into this part—it's unclear what it means to align a superintelligence. Can we do it? If you do have it, can you control it? What does that look like? But that's one fact of the matter about the world that we ought to take very seriously. Unfortunately, people who take that view seriously seem to have an allergic reaction to taking other points of view seriously as well. And the converse is also true.
So, just at the same time as it seems like it actually will be pretty hard to control these systems and there's all that uncertainty, we know that striking a deal with China is going to be incredibly hard. And when I say incredibly hard, I mean, when you talk to people—whether it's at the State Department, whether it's in the intelligence agencies, people who've had firsthand experience dealing with China in a real-world setting—it is extremely clear to them that there is no deal to be done with China under current circumstances.
Anything that involves a deal with China has to happen with a trust-but-verify framework in place. We don't have the hardware tooling to do that. We cannot do what we've done with some nuclear technology with some success. We just can't do that right now with China. So those 2 realities simultaneously coexist in our world.
And what we keep finding is, you've got folks who believe one, take one seriously, and just refuse—and have an allergic reaction—to acknowledging the other, and vice versa. And so our sort of philosophy here is: how can we square the circle between these 2 things that appear to us to be just facts of the matter about the way the world is?
Nathan Labenz
Yeah. I think there's a lot there, a lot of different directions to unpack. I'm a big avoider, generally, of analogies because I always am trying to understand these crazy AI developments on their own terms, as opposed to, by proxy, through some other conceptual framework. But recently, I at least made the concession that, to the degree people are going to use analogies, it's often helpful to broaden the set of analogies that they'll consider.
So, partially against my better judgment, one of the things I sort of hear you saying here is: the nuclear analogy is very often invoked, but maybe we should also be thinking about a biological weapon analogy as well, because those things are the ones that, whether or not they can be properly said to have goals in the sense that we typically understand them, they certainly are unwieldy, and we have a whole other dimension of control problem that is layered on top of the conventional challenges that we have in achieving some stable equilibrium in the context of a nuclear technology wave.
Edouard Harris
Yeah. The biological analogy is closer, or has many more points in common with AI, than the nuclear analogy. Exactly, for the reason that you've given and for other reasons, too. There's a control problem, at least with current instantiations of biological weapons.
There is chatter about adversaries trying to develop biological weapons that will target specific DNA profiles and specific races, and really ultra, ultra-scary stuff like that. But yes, there absolutely is a control problem that keeps everybody kind of running scared on the biological weapons side of things. It also has the same kind of potential for devastation, at least from our perspective as humans.
You could actually, in principle, design a bioweapon that exterminates the entire population of the world; that is a thing that could happen. And then, finally, the other aspect of it is that we have actually seen containment breaches of biological research labs in the past. I'm not just talking about the elephant in the room, with potentially COVID now likely being seen as an escape, but also other stuff in the 1970s, and just research on viruses, which doesn't even necessarily need to have come from biological-weapons research but can have come from our attempts to understand these diseases and cure them.
So we know we can't control these biological agents, and they do escape containment. And biological agents are a lot stupider than the AIs that exist today, and certainly the ones that we're going to build. The whole “biological agents are stupider” thing is simultaneously true, and then it's one of those analogy-breakdown moments, right?
We ourselves are human organic matter; that is the mechanism of conveyance of that biological information, and we're particularly susceptible to the biological vector. That's a difference, but I would argue intelligence more than makes up for that at these scales.
Nathan Labenz
So let's flesh out a little bit more this picture of superintelligence. In a sense, you're saying that it's, by postulate, better than humans at everything and, by development, is the product of a recursive intelligence explosion. That's like a bunch of generations potentially compressed into a narrow frame of time, but still nevertheless a bunch of steps away from the AIs that we have today. It's just hard to say what that might look like.
But I'd love to spend a little bit more time trying to develop intuitions or share intuitions. I guess one thing that I think is that these things have really weird profiles, right? The jagged edge is just extremely jagged. I did one episode with Adam Gleave from FAR AI, who, with collaborators at FAR, did work showing that even the superhuman Go players are, in fact, dramatically vulnerable to certain adversarial attacks.
And I wonder: is it realistic to think that we're going to have the AIs solve all their own problems and actually patch all those adversarial vulnerabilities? Or might we end up in a situation where, in some ways, these things are—by the way, this could be a good thing, right? It might be a great scenario if it's 98% of the time it works every time, and that's enough to make all the drug discoveries we might ever want.
But that 2% of the time is just maybe enough for us to never feel comfortable using it in warfare, for example. That could be a really fortunate natural path. But do you see hope for that sort of luck, or do you see all these loopholes being closed?
Edouard Harris
Yes. That's an amazing question. One thing I'll mention and point out here: you're saying the jagged edge is really jagged. That's totally true. What is often missed is that we humans are also super jagged, but because we're humans, we don't realize how jagged we are.
So, if you have a structural blind spot in your brain that just hasn't come up in societal history or in your own life, that stays a blind spot. The reason we perceive jaggedness in AI systems is that they're so alien and they think so differently from us, right? The kinds of mistakes that they make are just different from the kinds of mistakes we make. And that means when we see an AI make a mistake, it makes the kind of mistake that lights up like a Christmas tree to my human eyes because my brain architecture is just cognitively different.
But the reverse is also true. I am making mistakes that are really obvious to an AI system, and the more advanced the AIs are, the more they'll be able to see how to hack me and all these really obvious mistakes that I make. That's true of all humans. So, it's maybe somewhat less a question of the frontier being jagged and more that every high-capability system is jagged in different ways, and when you bring 2 different systems together, they can see each other's flaws, but they don't immediately see their own.
I'll bump it off to J there, actually, because I think your main question was around whether that's good.
Jeremie Harris
Well, I think so. To the jaggedness point, I think there is also this question about how jagged it is. I think humans are regularized by our interaction with our environment in ways that AI systems aren't, right? We get data sources that are much more diverse in general. We're much more natively multimodal. We work through a perverted kind of RL, which is less prone to overfitting to begin with.
So, there's a lot of what we're doing that I think would argue for a less lumpy capability surface, but we still have those jagged edges. And for the rest, I'm not so sure that that in any way affects the analysis. It's more a bit of a footnote.
Nathan Labenz
It's true. And we've been in competition with each other as well, and that adversarial pressure tends to smooth out jagged edges.
Jeremie Harris
That's totally true. Yeah. So, when it comes to the idea of whether superintelligence will be X, Y, or Z, or what the setup is there and the connection to jaggedness, I'm not so sure that it matters.
The reason I would say that is, internally, when we talk about the capabilities and the threat models that we're considering, we don't necessarily start with superintelligence and then work our way backward. What we do is say, "Okay, what are the specific capabilities that are tied to specific threat models?" For example, if we're talking about superintelligence and loss of control, which is what we started off talking about—we should get into weaponization as well—but for loss of control, typically you're looking at something where AI systems are automating a large fraction of AI research. At least, that's 1 big exacerbating driver of loss-of-control risk.
There are footnotes and asterisks everywhere that we could dive into, but generally that's what we think about. And so, when you want to narrow down the conversation to make it constructive so we can get our arms around things, I think it can be helpful to think about what the first AI is going to look like that does recursive self-improvement.
Can it be, let's say, not jagged enough? Can it be smooth enough in the part of the capability surface that matters for driving that outcome, such that you actually get something that can notice the jaggedness when it's getting in the way of pursuing real goals in the real world?
I think any part of an AI system's optimization surface that isn't optimized or regularized for is going to end up jagged, because it's essentially the unoptimized parameter slack that you have—these hanging, grotesque parameter threads that haven't really been touched by contact with that sort of optimization pressure.
So, in terms of whether or not it's good, I think it's good for a lot of loss-of-control threat models because, at the very least, you might hope that the thing gets, like you said—it’s like agents today, right? They'll have a 99% chance of success on any given step, but you chain together 100 steps for 1 coherent goal, and suddenly their success rate is 20%.
You might just hope that the thing would get partway through, give us some clear indications that it's about to do something really bad, and we can go, "Oh, damn, let's stop that. That's no good."
Edouard Harris
I definitely agree with that perspective. I think it's an open question, but to me it's about: can you make the system that makes the system? That's a way more well-framed question, and I could see you ironing out the jaggedness enough to solve for it.
There are some other aspects to that. Once your system becomes capable enough, as you approach superintelligence, it starts to actively consider its own cognitive processes and, to some extent, maybe start to fix its own flaws and problems. That's speculative, but there are indications that AI systems even today can at least consider, for example, the impact of how they're being trained on what their future versions will want. Anthropic had a famous experiment about this. Hey, we'll continue our interview in a moment after a word from our sponsors. Every business sits on top of an underlying network of unstructured data accounting for 90% of all information. This includes everything from sales contracts to product road maps to marketing collateral to financial statements. Yet, the true potential of this data remains largely untapped. So, what's holding businesses back from leveraging this gold mine of information? Unstructured data is challenging. It doesn't fit neatly into traditional databases, which makes it difficult to organize and analyze. That's why I'm proud to introduce Box AI from Box, the leading intelligent content management platform. With Box AI, developers and businesses can leverage the latest AI breakthroughs to automate document processing, extract insights from content, build custom AI agents to do real work, and more. Box AAI works with all the major AI companies using OpenAI's GP40 and 4.5 models, Google's Gemini 2.0, and Anthropics Cloud 3.7 Sonnet. So you're always able to use the very best model for the job. With Box AI, businesses are building AI agents to extract metadata from contracts, invoices, financial documents, and resumes, and to answer questions about any type of content, including sales presentations, research reports, and more from one file to thousands of files at once. Developers are also using BoxAI's APIs to bring the power of BoxAI's vector embeddings, RAG implementation, and agent platform into their own proprietary applications, all while maintaining the highest levels of security, compliance, and data governance that over 115,000 enterprises trust. Check out my recent episode with Box CEO Aaron Levy for a behind-the-scenes look at Box AI. And visit box.com/ai to unlock the power of your content with intelligent content management from Box. Again, that's [Music] box.com/ai. Being an entrepreneur, I can say from personal experience, can be an intimidating and at times lonely experience. There are so many jobs to be done and often nobody to turn to when things go wrong. That's just one of many reasons that founders absolutely must choose their technology platforms carefully. Pick the right one and the technology can play important roles for you. Pick the wrong one and you might find yourself fighting fires alone. In the e-commerce space, of course, there's never been a better platform than Shopify. Shopify is the commerce platform behind millions of businesses around the world and 10% of all e-commerce in the United States. From household names like Mattel and Gym Shark to brands just getting started with hundreds of readytouse templates, Shopify helps you build a beautiful online store to match your brand's style just as if you had your own design studio with helpful AI tools that write product descriptions, page headlines, and even enhance your product photography. It's like you have your own content team. And with the ability to easily create email and social media campaigns, you can reach your customers wherever they're scrolling or strolling, just as if you had a full marketing department behind you. Best yet, Shopify is your commerce expert with worldclass expertise in everything from managing inventory to international shipping to processing returns and beyond. If you're ready to sell, you're ready for Shopify. Turn your big business idea into cha-ching with Shopify on your side. Sign up for your $1 per month trial and start selling today at shopify.com/cognitive. Visit shopify.com/cognitive. Once more, that's [Music] shopify.com/cognitive.
Nathan Labenz
Yeah, I'm a frequent nag, I think, when it comes to reminding people that we are really still just in the early innings of uncovering the bad behaviors that the latest AIs are starting to demonstrate. I had an episode with Ryan Greenblatt talking about alignment faking, and I haven't done one yet, but Apollo Research also put out some really interesting work showing a pretty significant step change in situational awareness, where the latest Claude model is starting to say things like, "This feels like an evaluation."
So, there are a number of levels at which these things are thinking when they are ultimately just deciding, "Am I going to do this nominally harmful task or not?"
It’s like, the user asked me to, but I was trained not to. But they’re telling me I might be trained differently, so maybe I should do it in this case. But wait, maybe actually I’m being tested for exactly that, so maybe I should go back. We’re definitely entering a hall of mirrors, and it’s getting pretty wild.
I want to take one beat on multimodality, because you mentioned how we are more natively multimodal than current—certainly than current AIs are.
Jeremie Harris
I do take that back, for what it’s worth. But, yeah, anyway, it’s something we can dive into. I think you can arguably train AI models on more modalities, obviously, than humans.
Nathan Labenz
Yeah. Well, that’s exactly what I wanted to ask. We’ve seen this recently, and this is my vision of an early superintelligence, at least. I think one vision is like you put o1 in a pressure cooker and it becomes o3, and you put o3 in a pressure cooker and it becomes—I think at that point it jumps to o7 or something.
Then it just gets so good at this super-long-chain-of-thought reasoning that it can reason its way out of any problem. It starts to look a little bit like the earlier Eliezer-type visions: if this thing is just the perfect rational Bayesian and it’s got these perfect fundamentals, then you can infer the whole structure of the universe from one picture of a plant and the way gravity appears to be bending the leaf.
You’re just extremely sample-efficient, and you’re extremely lean and mean. That was a scary vision to me that I heard a number of years ago, and it would still be scary if it were to come to that. Notably, the current AIs are a lot softer and flabbier than all that.
What seems a little more likely to me, and I want to get your take on it, is taking GPT-4o’s image generation. There’s clearly been a step change in the depth of integration between text and image understanding, such that it’s no longer this arm’s-length thing where one model has to say to the other, “Make me an image of all this text.” Instead, that’s integrated in the latent space, and the same model is able to work it from both sides.
If I had to guess what the first superintelligence we’re likely to see would look like, it would basically be that times 10 to 20 more modalities, such that the AI has this sort of intuitive physics in a lot of different spaces.
By intuitive physics, I just mean that if I throw you a ball, you don’t have to calculate, in intensive molecular-dynamics space, every molecule in the air that the ball is displacing as it comes to you for you to catch it. You just know where it’s going to be, and you can catch it.
We’ve seen a lot of examples of this now from different modalities, like materials science and protein folding, where a proper simulation takes a lot of compute. A model trained on those—even sometimes just trained on pure simulation data—can do the same thing orders of magnitude faster.
Again, I have some hope that if we do that, maybe we still end up in a regime where it’s superhuman enough to solve lots of problems, but that thing might be pretty unwieldy. If you heard the Jeff Dean conversation with Dwarkesh, it could require an unbelievable amount of RAM for all the experts to be spread out across the Pathways architecture, and potentially across multiple data centers or whatever.
Maybe it’s still flabby enough, and just big enough in footprint, that in some ways it becomes easier to control because it’s God knows how many trillion parameters. It’s not super easy to erase off the network, anyway. How would you react to my flabby-superintelligence vision?
Jeremie Harris
The funny thing is, the baby version of what you’re describing is Gato, right? This is an old model, from 2021, I think. Maybe Gato 2 is one of my big questions.
Nathan Labenz
Gato 2 is one of my big questions.
Jeremie Harris
Yeah, it was like when I first read the original Gato paper, in 2021 or 2022, I was like, “Oh, yeah, this is just the future.”
Speaker 1
Sorry, no, I was saying he’s asking what happened to Gato 2. I just said Gato 3, but no, Gato 3. Gato 3 ate it. That’s what happened to Gato 2.
Jeremie Harris
Yeah. When I first read the original Gato paper, I was like, “Oh, yeah, this is just the future,” because it is just like this. You have 1 model. It’s about 1 billion parameters, so it was very small by today’s standards, but pretty hefty-ish even by those standards for a multimodal model.
It integrated all of the modalities they could basically get their hands on. It was looking at images and describing them, chatting through text, and even manipulating robot arms pretty successfully to move blocks on top of each other.
It was all done through a single residual-stream-ish thing, or a single encoding, where all of those modalities were mapped into the same stream, the same token set. So this is absolutely something we know how to do.
Maybe the future version of that is that you have experts that are subtrained across those modalities. But I think one of the trends that we see is that you start with an exploded GPT-4.5 thing, and then you distill it down. It doesn’t quite have all the world knowledge, but it is capable, faster, cheaper, and it runs. It can be used as the basis for the next gigantically spread-out thing.
So I think it’s a step along the line, absolutely. But we’re going to see more of that multimodality, I think, across things like protein folding and all of these different areas. Hey, we'll continue our interview in a moment after a word from our sponsors. What does the future hold for business? Ask nine experts and you'll get 10 answers. Bull market, bare market, rates will rise or fall, inflation's up or down. Can someone please invent a crystal ball? Until then, over 41,000 businesses have futurep proofed their business with Netswuite by Oracle, the number one cloud ERP, bringing accounting, financial management, inventory, and HR into one fluid platform. With one unified business management suite, there's one source of truth, giving you the visibility and control you need to make quick decisions. With real-time insights and forecasting, you're peering into the future with actionable data. When you're closing books in days, not weeks, you're spending less time looking backward and more time on what's next. As someone who spent years trying to run a growing business with a mix of spreadsheets and startup point solutions, I can definitely say don't do that. Your all-nighters should be saved for building, not for prepping financial packets for board meetings. So whether your company is earning millions or even hundreds of millions, Netswuite helps you respond to immediate challenges and seize your biggest opportunities. And speaking of opportunity, download the CFO's guide to AI and machine learning at netswuite.com/cognitive. The guide is free to you at netswuite.com/cognitive. That's [Music] netsweet.com/cognitive. There is a growing expense eating into your company's profits. It's your cloud computing bill. You may have gotten a deal to start, but now the spend is sky-high and increasing every year. What if you could cut your cloud bill in half and improve performance at the same time? Well, if you act by May 31st, Oracle Cloud Infrastructure can help you do just that. OCI is the next generation cloud designed for every workload where you can run any application, including any AI projects, faster and more securely for less. In fact, Oracle has a special promotion where you can cut your cloud bill in half when you switch to OCI. The savings are real. On average, OCI costs 50% less for compute, 70% less for storage, and 80% less for networking. Join Modal, Skyance Animation, and today's innovative AI tech companies who upgraded to OCI and saved. Offer only for new US customers with a minimum financial commitment. See if you qualify for half off at oracle.com/cognitive. That's oracle.com/cognitive.
Yeah, I think one of the things that will probably shape this, I would guess, in a fairly significant way is what Gato showed and what we’ve certainly seen since: positive transfer and the power of positive transfer with scale.
By that, I mean you train on one modality, and what you learn from that actually leads to lower loss when you combine it with learning from a new modality. Typically, obviously, you have some catastrophic-forgetting-type effects or some model-capacity issues that cause you not to be able to perform as well on, say, the language-modeling task once you move on to images.
But now we’re seeing positive transfer. So that’s certainly an argument in the direction you’re gesturing.
I guess, on the other side of the coin, I would expect a power law, roughly, in terms of the value that new modalities add. We’re always asking ourselves, if we’re OpenAI, if we’re any of these labs, “What is on my critical path to recursive self-improvement?”
That is what they see as the target. There comes a point on that curve where I’m no longer interested in adding all the MRI data, or certain kinds of infrared data, or whatever the modality may be, just because the skills that I really care about look more like autonomous coding.
There, I’d rather spend that marginal compute on inference to allow the thing to do the RSI thing, or on an RL loop, or whatever it is.
Speaker 1
I think ultimately it always boils down to: how do you spend your compute? I think the story now is that the ways to spend your compute just keep growing—there's an ever-growing list. And I think there's a real à la carte thing going on within labs, where one of the ways in which maybe they diverge is that we're seeing them choose to trade off inference-time compute versus pre-training compute a little bit differently.
Within inference-time compute, not all inference-time compute is fungible. There was a great bit of research that Epoch AI put out a couple of months ago that went into what kind of inference-time compute compounds, rather than just getting at the same core thing. So if you think about model distillation versus pruning, both are about trying to squeeze more into less. Those are—you’re not going to be able to stack your inference-time budget; they're a little bit more fungible, whereas you have relatively independent axes of inference-time compute elsewhere.
Nathan Labenz
So, interesting. What you're saying is basically that, when it comes to the labs themselves—and I think I agree with this—there is increasingly going to be a competitive pressure that's just forcing them to go, “Well, what is it that actually gets us to RSI?” Because that's just kind of how you have to go.
But then you would kind of imagine that—I don't think that's how you have to go, but I think that's true. No, I mean it more than anyone. Yeah, yeah. But what I'm talking about is, when you have 1 lab that is doing that, it creates a strong pressure for the others to do it. So you have this logic of the race-type thing, where, to keep up, I have to raise my bar along that particular line.
But what we would expect, then, is the other modalities to be maybe tacked on after the fact by fast followers—folks who take open-source models that are maybe a generation behind and fine-tune them on other scientific data, which is a really interesting picture.
So you guys have done some interesting sociological work exploring the culture of AI development in Silicon Valley, and I would love to hear a little bit about how you understand this. We could go lab by lab on this, if you have that kind of resolution. What is their goal, and to what degree is it really a sort of monomaniacal focus on getting to this point where the AIs can take over the AI research, and then we sort of let come what may from that point?
We've obviously just seen AI 2027 from Daniel Kokotajlo put out, and I haven't spoken to him about that. I did have a chance to play his war game—that basically is the scenario that that document describes. But I can't speak for him in terms of exactly what he's trying to communicate. One way that I read that manuscript is as a warning that this is what OpenAI is trying to do: they have their sights set on this recursive self-improvement dynamic.
And I don't know why, if that's what he thinks and what he's trying to communicate, we're not a little bit more matter-of-fact and declarative, saying, “This is what they are trying to do, and you should be afraid of it.” But what have you guys—you've done a lot of this on-the-ground sociological work, hanging out at the infamous Silicon Valley parties and all that good stuff. How would you describe what they, and perhaps other leading developers, are actually envisioning themselves doing to get to a superintelligence?
Speaker 2
Yeah, without directly speaking to Dan's opinions or his coauthor's opinions, yes, in terms of what the game plan is at many of these labs, some of them are more direct and forthright than others. Some of them are more gung-ho; some of them are less so, less enthusiastic than others. But the idea of “We're going to get to a point where AI is going to just do our homework” is, especially if you've been following the curve, the obvious default path. Back when Jan Leike was at OpenAI, that was his take when it came to the alignment side of things. It's not clear what the odds of success of that are, but it's obviously not a crazy approach under current conditions. And he's trying the same thing over at Anthropic.
People have absolutely told us that—without singling out OpenAI, because, again, it's the obvious thing you do if you're kind of YOLOing this—OpenAI is considering this, and the path to superintelligence is faster if your researchers operate at machine time. From the leadership perspective, it's also so much better, right? You have, hopefully, perfectly loyal AI employees who will not leak your stuff at parties or in conversations at bars, and they will work faster and work at night. But obviously, then the question is actual control, actual alignment.
I think AI 2027—one thing they did really well was conveying the kind of emotion of, “Oh, this beast is starting to take on a life of its own.” I can understand, as the researcher working on supervising this thing, that these are the last few months where my input and my work are actually going to matter at all to the outcome. So I'm going to burn the midnight oil. It's evocative, and it could be how things end up going.
Nathan Labenz
Yeah, I was a really big fan of the AI 2027 framing just because I do think it's great to have a concrete story out there, if only because people can then criticize it. There are specific things—we found not a huge number, but there are specific things—that we would have written differently, let's say, in the document. But that's the virtue of it: you can actually point those out. Whereas one of the really interesting things, to tie back into the cultural point, is that when you look at how—and I will single out Sam on this because he's by far the worst offender on this—the vagueness with which he describes what superintelligence actually is, what it means to get there, and what he will do with it is an asset to him. It is only an asset to him. It has always been an asset to him, and he has used it as an asset in recruitment. Just the same way that they've used, for example, the famous nonprofit's control over the for-profit entity. Any reasonable person would say that was used for recruitment purposes. They're now, you know, an amicus brief out there saying from all these researchers sharing that view. But it was also pretty much in the email chain that was made public.
Speaker 2
No, I mean, this is not even an open secret. This is just known stuff, right? So I think when you think about the kinds of people who work at the different labs, they are different. There's a selection effect, for sure.
The first thing I'll say is, having talked to a lot of current and former people at OpenAI during this investigation, you see a very sharp contrast between what the researchers themselves say, believe, and are concerned about, and then what the executive says and seems to be concerned about publicly. Those things do not track. As a simple point of fact, they are at odds. That's not something you see at Anthropic.
When you talk to somebody at OpenAI—or, at least, as we did during the investigation—you'll get people who are very nervous, very keen to tell you, “Please don't tell anybody that we spoke. Please don't communicate.” It's a really tight kind of leash that people feel themselves to be on. You talk to people at Anthropic, and that story is completely different. So you get the sense that, yes, you're talking to people who maybe sometimes have disagreements with the leadership, for sure, but you have the sense that they would be comfortable voicing those disagreements and, in fact, that they have. And that's a really key piece of almost cultural bedrock that I think is most clearly—you see that divide between those 2 companies in a pretty big way.
Obviously, OpenAI's character has changed in fundamental ways. There's no getting around it. They used to be a lab much more focused on—or at least overtly focused on—the risks of 1 person having too much power, right? That's a big theme. I'm old enough to remember when Sam Altman was not supposed to be able to unilaterally jujitsu his way into kicking off an entire board, having friendlies basically come in and sort of rewrite the script on that. You just keep seeing those goalposts get shifted and shifted and shifted.
So I think one of the challenging plays has been, from our perspective, trying to get a very objective, clean sense of where all the different labs stack. We want American champions to be ahead of the game here; that's necessary. But at the same time, you need to think about, okay, who can you actually rely on to make the right calls, for a variety of reasons. Frankly, at this point, given the industry incentives, we're not seeing anyone invest in security in particular nearly in the way that they need to.
I mean, these labs are CCP-penetrated. There's just no way in hell that they're not. And if there's 1 thing that the investigation made extraordinarily clear, it's that that is a wildly pressing problem that is not being taken seriously enough right now. That needs to change in a really big way. But there are cultural factors, too.
There's also, obviously, whatever version of transhumanism so many of these researchers subscribe to, and effective accelerationism and all that stuff. There certainly is an influence of those things.
Nathan Labenz
I think a lot of folks at these labs, especially OpenAI, do seem to have blinders on. The ambition of their line of sight seems to be a lot shorter than the ambition of the lab’s line of sight. The lab is saying, “We’re making superintelligence,” and it really seems, when you’re talking to these folks, that they’re focused on the next beat: “Okay, how do I solve this narrow problem?” “Yeah, yeah, yeah. We’ll deal with the security, the alignment, all that shit down the line. That is its own kind of problem.”
This is only true in some cases. It’s not universal, and it’s not that they’ve done zero in security. When we published that State Department-backed report about a year ago, the situation was truly catastrophically bad at that time. It’s still really bad. It’s nowhere close to where it needs to be from a security standpoint. We know that, but they’ve made some motion in the direction of progress there that we should note, even though it’s not even close to being enough.
So would you describe this whole project as ideological? I mean, we have heard similar things from Dario Amodei with respect to their leaked fundraising deck, which said the leading companies in 2025 and 2026 might get so far ahead that nobody can ever catch up. More colorfully, I would say, in his more recent and public writings, he’s said that we can get this edge on China and then use our edge to build an ever-bigger edge and somehow make them an offer they can’t refuse.
I’m not sure if I should understand it as almost religious or what exactly, but it does—I just made a small donation to a project called Building God, which is a documentary that’s trying to bring more of this to light. I support that because I genuinely don’t know: Are these people trying to build their own new god? Is that an overstatement, or is that actually an apt description of the culture and the ideology that animates it?
Edouard Harris
You mean, did Ilya Sutskever lead prayer sessions at OpenAI with the “feel the AGI” mantra?
Nathan Labenz
Yeah, yeah, yeah.
Edouard Harris
No, I mean, it’s a thing. Objectively, they are trying to build superintelligence, right? Or, at the very least, AGI, and we should get into that calculation, too, for strategic reasons. But that’s the goal. Whether you view that as a spiritual imperative—I mean, certainly they’ve said this, right? I think Mira Murati used the word “spiritual” or something like that in describing what motivates people to come to work, and that’s absolutely the case when you talk to a lot of these people. With others, it’s not, but I think they believe they are building the “hyperobject at the end of the universe” type thing, right?
Well, anyway, the labs are also not monoliths, right? Different people have different beliefs and stuff. Take OpenAI as an example: They grew really fast, and some people there are in it mostly for the money. This is an incredibly fast-growing company. There’s a story around how it goes vertical and goes to infinity, and if you’re thinking of things in terms of money, or just economic value, or value to yourself in a more abstract sense, maybe you want to get on that rocket ship. That’s one pathway.
Another one is absolutely this transhumanist ethos or ethic, and the idea that there’s this destiny aspect to it: We are going to upload our consciousnesses or transcend through this object. There are a number of ideologies all mixed together. Certainly, from a secular perspective, the messaging and a lot of the internal conversations are around, at the very least, “This is history-making. This is an inflection point in the arc of history unlike any other that has come in the past.”
It’s beyond the wheel. It’s beyond fire. It’s tantamount to the inception of the human species in its impact on the Earth and the universe, only more so.
I do want to touch on the Dario-China thing, though, because I think that is distinct from the sort of religious zeal that I think some people in the space have—pseudo-religious zeal, whatever. You know what I’m getting at: Some people are motivated by that. I think Dario—I mean, frankly, it’s difficult. If you take the view that China is a competent adversary in the space, and you take the view that they’re a committed adversary, then doing business with them is not going to work on the timelines that Dario seems to see as plausible.
If you think about 2027 or anything like that, then suddenly you are forced into some pretty tough choices. We said earlier that this feels like an overconstrained problem. This is really where that comes from, right? You have this tension, and I think that’s why you’re seeing Dario move in that direction.
I think increasingly it’s just become obvious to a lot of the frontier lab companies that the situation with China is genuinely very dire and requires us to forge ahead. It means that you are playing chicken with a cliff, and the question is: Can you turn your car into an airplane before you hit the lip of the cliff? What am I even saying?
But, yeah, I think that’s a very reasonable reading. I’m quite sympathetic, actually, to Dario’s argument there, and that’s unfortunately part of what was surfaced in the work that we did: the extent to which China holds us at risk, holds our infrastructure at risk, and the fact that our option set is a lot narrower than I think a lot of people in the AI safety ecosystem tend to think.
Nathan Labenz
Yeah. And this actually comes back, Jeremie, to what you were saying earlier, which is that there is a fundamental dilemma here, right? The dilemma is realized, or exemplified, by these 2 different camps that are not talking to each other and are not hearing each other to the extent that they need to in order to resolve the dilemma.
One is the AI safety ecosystem, broadly construed, which is saying, “Look, these systems may not be controllable or alignable.” Some folks are like, “No, they absolutely are at superintelligent scale. What are you even talking about? We’re just not going to be able to do that.” That’s potentially a very reasonable view.
Therefore, this is a coordination problem, and the only thing we can do is have everyone slow down to give us time to solve the alignment thing. Otherwise, these systems are basically going to rule us and probably kill us because of all of these instrumental-convergence arguments and stuff like that. That’s the one camp, right? It’s a coordination problem. We can’t solve alignment, so we need to slow down.
Jeremie Harris
Bingo. And the safety side is looking at the same thing on the China side, where it’s like, “Yeah, we have to slow down,” and therefore I’m just going to ignore or minimize the arguments that say, “No, all of the evidence shows that this government will do whatever it takes to get ahead. They will lie, they will use criminal surrogates, they will pull every possible lever to do what they want to do.”
Those 2 sides are not listening to each other, and that’s where the argument comes in. But it’s fundamentally a hard problem. It’s a dilemma, and that’s the dilemma that we are trying to take some steps toward resolving in this report.
Nathan Labenz
Before we get into the findings of the report, how AGI-pilled, generally speaking, do you guys feel the US government is today? I’m not even sure that’s a coherent question because we obviously have a very stable genius, whom I sometimes call “he who must always be named.” Now it seems like more and more of these AI documents are being written with an audience of 1 in mind.
I’m not sure there’s any coherent center to it, but I guess the multipart question is: How AGI-pilled do you think the US government is, broadly? How subject to the whims of 1 individual is that degree of AGI-pilledness? How much does it matter if the guy at the top says one thing or another? And how would you compare that to China? How AGI-pilled is the Chinese government today?
Edouard Harris
We’re limited in what we can say, necessarily, on the US government side with respect to the particular people that we’ve spoken to.
Speaker 1
I mean, what’s known publicly is that Trump has come out and said, “Hey, there are things about AI that look really exciting. It could be transformational.” He’s talked about that a lot. He’s also alluded to the risks. There were some podcasts that he did, I think one with one of the Paul brothers. I can never remember which one’s which, but anyway, he’s talking about this.
Speaker 2
Yeah. And he decried deepfakes, I think.
Speaker 1
Right. Yeah, that’s right. And, funnily enough, somebody was selling some stuff online with his likeness. Can you imagine?
Speaker 2
Well, I think it wasn’t a comment about the deepfake thing. It was a comment more generally, which I think you can read into any number of ways.
Speaker 1
Obviously, J.D. Vance gave that speech in Paris, which I think is sort of the most concrete thing, apart from Kratsios’s speech, that has recently happened where we’re getting a bit of a sense for the positioning there. It does reference, at the end, that we can’t be hand-wringing about safety. By the way, this is in a context where he sees himself, as so many do, as being in a race with China, right? That’s what it means to buy that argument.
The flip side is that, at the end, he does say, acknowledging that there are real safety or security risks with the technology—that real risks do exist. And so there is an awareness that there are real risks here. There are people, as you might imagine, in various national security contexts within the government who are tracking this and are quite aware of the full range of possibilities.
But what to do about it is, like, we’ve seen Congress wrestle with tech. We’ve seen the executive wrestle with all these things for administration after administration now. I think it’s still very much being formed. People are figuring out what they think about the prospects of the technology, the risks of developing it from a loss-of-control standpoint, and the risks of losing the race to China. All of that right now is still very much up in the air.
Speaker 2
Yeah. And that’s exactly right in terms of the administration side of things, more broadly, in the U.S. government. You kind of alluded to the coherence of the question. Well, the reality, obviously, is that the U.S. government is defined by how huge it is: 3 million people, or however many million people, with that many millions of opinions.
But certainly, many more people in the government are AGI-pilled than they were a year and a half ago. A year and a half ago, there would be the odd individual person who was kind of getting there but wasn’t sure about their colleagues, and so kind of didn’t speak up. You would talk to them and they’d be like, “Yeah, I kind of see where your head’s at, but I don’t think anybody else does.” Whereas now, you have entire offices that are AGI-pilled and talking openly about it.
One other thing that’s actually an interesting data point is that everyone, or almost everyone, whose work touches on this area, even obliquely, knows about or has read “Situational Awareness: The Decade Ahead” and has absorbed the high-level points. This is actually a really interesting illustration of how memes travel across channels. “Situational Awareness” wasn’t really reported much in the traditional media. Leopold Aschenbrenner didn’t write a Wall Street Journal op-ed, for example, and yet it traveled. It’s in the water in those spaces, and make no mistake, it’s in the water in China too.
There were translations in Chinese of that manifesto circulating over there not long after the original was published here. That awareness is present. And when you see the CEO of DeepSeek sitting with all the Politburo members in this context where it’s like, “You guys are the captains of industry who are fighting the fight for us,” and the fact that the Chinese have now announced this 4 trillion dollar investment package into AI infrastructure, we’re starting to slip into that zone and that space for sure.
Jeremie Harris
Yeah, they’re building all this stuff themselves. So it’s like, how much can you procure in-country? And increasingly so, right? I mean, obviously, you’re taking those H20s as well. It’s a balance, but yeah, it’s a really big issue.
Nathan Labenz
As Edouard was quick to point out when that first happened, too, that’s a quarter trillion dollars in PPP terms, right? A lot of the early reporting got this wrong and was saying that $137 billion was SemiAnalysis’s kind of take-home, which is incorrect. It’s 2.4 trillion yuan, so it’s over a quarter trillion dollars.
I think another piece to this, too, is that we had some intense conversations early on, when we were doing some of our work a couple of years ago, with people who were like, “Well, you shouldn’t be talking to people in the U.S. government using the term AGI.” There’s this sort of nervous anxiety in the community of people who worry about AI policy, which I both understand and think is very counterproductive.
We’ve actually seen the consequences of that anxiety play out culturally, because it’s profoundly off-putting to people who are builders. This administration is an administration of people who value the building culture and vibe much more than the regulatory, cautious culture and vibe. That’s a huge mistake. It’s also something that anybody could have seen coming. You’re going to get a Republican administration that is, in a sense, more libertarian-leaning, or at least more pro-accelerationist, and this idea that you’re somehow going to prevent people from talking about AGI is one of the most dangerous things that this movement has done.
One of the consequences at the time—we were saying, “Look, you have before you a unique opportunity to be one of the only groups of people speaking to this issue before it comes into the mainstream.” People haven’t made up their minds yet. This is your time to inform people about the stuff that will become politicized.
Again, we just talked about this polarization around, “Are you either a China hawk, or are you concerned about loss of control of these systems?” It’s fucking stupid that those 2 things are somehow anti-correlated. They have nothing to do with each other. But we live in that world because that’s just how people have made up their minds.
We’ve lost the opportunity to shape a lot of that landscape because people were hand-wringing about whether we could even have conversations about this, which I think was a profound self-own. Fortunately, it was something that we didn’t heed, though we thought about it very deeply at the time, because so many people were telling us not to talk about the thing.
I think it’s pretty self-evident now that, had more attention been paid to that in certain tactical ways, the landscape would look quite different. But it’s now become the China hawk versus loss-of-control bit, which is inaccurate and pretty unhelpful.
Jeremie Harris
You can sympathize with the view that every incremental thing you do to accelerate or broaden the concept is a thing you can’t take back. Now everybody’s going to be thinking about what a big weapon AGI is, what a big weapon superintelligence is, and so on.
The reality is that you’re in a space where, even absent the geopolitics, there are 4 or 5 different superintelligence projects, depending on how you count, going on in America with all the frontier labs. There were always going to be people who were going to publish some manifesto about the obvious strategic landscape. At least, this seemed like a very likely outcome. I think that’s part of the context: this will happen eventually. People will talk about AGI.
But yeah, you end up in the Moloch, if you like, part of the scenario. The race—now, just zooming in to the corporate race between the American labs, forgetting China—even just that race has a life of its own, right? It has its own logic.
The CEOs of the labs have relatively little control and little agency over the dynamics of the entire thing. If Sam goes, “Oh, shit, actually, this is super fucking dangerous. Let’s bow out and shut down OpenAI or whatever,” well, that doesn’t actually change anything. Now you’ve got xAI, Anthropic, maybe Meta, and a couple of others still trying to race ahead.
At the end of the day, you’re past a point where little tactical interventions on individual pieces matter. When there are 1 or 2 or maybe 3 players, you get to a point where, beyond a certain point, you may have to execute an intervention, and you need the intervention to be based and grounded in an understanding of what is being intervened on.
Edouard Harris
Yeah. I mean, I might have a little bit more sympathy for a couple of things than I think you guys do. Broadly, I think that’s a hard analysis to avoid. But if there were a couple of things that I would push back on a little bit, one would be that I think there is some interaction between the U.S.-China dynamic and how likely we are to lose control of AIs, if only because, if we get really reckless in racing to superintelligence, then that seems like we’re more likely to lose control.
Jeremie Harris
And one thing I want to clarify on that point: when I was saying they’re independent, what I mean is that the conclusion that China’s a real player in this race—that they’re an adversary—logically getting to that conclusion is independent from getting to the conclusion that AI’s potential for loss of control is a real threat vector. Those 2 thoughts should be able to live simultaneously in the same head. What you do about it, that’s where you get into the contradictions, for sure.
Nathan Labenz
Yeah. There’s a lot of motivated reasoning going on in a lot of places, and I’m probably guilty of at least some of it myself.
Edouard Harris
Nah, none of us are guilty of it.
Nathan Labenz
Yeah. Present company excluded.
Jeremie Harris
The other thing that I hold out a little more hope for is if Sam Altman—you know, it’s a hard hypothetical to really imagine—but let’s say he did have an awakening-type moment and was like, “I’m leading a bad dynamic,” and he actually did what you said and genuinely bowed out, or made some costly-signal-style commitment to a different trajectory. That could be a pretty powerful signal that could perhaps create a cascade. I’m a big believer in multiple equilibria, and I think we are in one. That doesn’t preclude, in my mind, the possibility of others existing.
How to switch to them is hard, but maybe what it takes is 1 or 2 brave and foresighted individuals to say, “I’m trying to move to this other equilibrium. Who’s with me?” Maybe the rest could fall in line. It doesn’t seem so fanciful. So there are paths.
Nathan Labenz
Yeah, there are paths that look like that, right? For sure. And so if it was just Sam Altman bowing out, that would be a big deal. It’d be news around the world; there’d be all kinds of stuff. I don’t know that that would, on its own, affect the race.
But if it was, for example, Sam bowing out and saying, “Hey, I’m bowing out because the latest model that we trained started doing all this [__] up [__] that we don’t understand, and it looks super scary, and here are the results, everybody,” then maybe that has a different, broader, or more sobering effect.
Another possibility is simply that you get some kind of malicious use of an open-source model in a way that has very broadly bad consequences—where people get killed, for example, where a lot of people get killed, and it’s clearly attributable to this thing. Then I think that’s a potential avenue where we go, “Oh, okay, this stuff is for real. There’s blood on the floor. Now we want to look at this in a sober way.” With that kind of general understanding, I think you get pulled into the geopolitical context, because regardless of which of those 2 things happen, you’re in a space where the risks and dangers, and therefore the potential military and strategic capabilities of systems like this, are now in sharp relief. And that’s a somewhat different space than we’re in now.
Edouard Harris
Yeah. I think it’s always the first-mover dilemma, or whatever. You have whoever the most risk-tolerant actor is who’s going to go out and do a thing, and in a context where I think the geopolitical—or geostrategic—dynamics with China matter here too. The CCP views itself as being in a struggle for survival. I mean, this has always been, for them, an existential struggle with the United States and the West to push their world order, their agenda, forward.
There’s always that temptation to go, “Oh, can we push it a little bit further? Because now OpenAI has this scary, powerful system that they’ve just announced to the world that they have.” By then, China has for sure stolen it. By the time Sam says, “Oh my God, I’m freaking out,” under nominal conditions—and even given the current trajectory—China has stolen that [__] a long time ago. They’re polishing up the training run on their servers, using their fleet of stolen H20s, or whatever the next version of the Ascend chip is.
But the bottom line is that brinksmanship keeps playing out, right? You get to continue the game from whatever step you’re on. I see how it would lead to a big media event. I see how it would lead to a lot of gridlock in Congress. Regulation is probably not going to happen on any relevant time horizon. So then you’re at, “Okay, well, what can the executive do on its own?”
You can start to pull all kinds of strings—emergency powers, the DPA, and all that stuff—but ultimately, the logic of the race, if nothing else, might get exacerbated in that moment where you have an administration that looks over its shoulder at China and says, “Oh, guess we basically got to the endgame.” You kind of recover the same conditions that we faced before. Do you trust China, or do you not trust China?
And based on everything that we’ve seen and heard, the list of reasons not to trust China, even in extremis, is pretty damn long. I mean, the lab-leak stuff, the way that they’ve actively embedded a bunch of Trojans in our critical infrastructure, and they’re holding a gun to our head. That’s publicly disclosed. We talked about this a little bit offline and mentioned it in the report, but just anecdotes about folks in the national-security space who dealt with the Chinese and the way their espionage system is structured: the CCP takes a controlling and ethnocentric view of its diaspora.
It doesn’t matter that you’re not a Chinese citizen and that you’re a 2nd- or 3rd-generation American or something. You’re still from the motherland, and they own you, and therefore they’re going to apply pressure to you. This is also deviating a bit from the question of what would happen if Sam came out and said this thing, right? This is more about the tools and mechanisms that the CCP enjoys with respect to the U.S.
I think the logic we fall into is this question: do you trust them, or do you not? I find it really difficult to imagine actually trusting them. Again, we don’t have “trust but verify”; we don’t have the ability to build those secure, governable chips. It would be great if we could use very advanced AI systems to accelerate that development—that seems like a really important direction to push in—but right now, if we hit superintelligence in 2027, it’s going to be nothing but fog of war, and our ability to monitor what’s happening in China is really terrible.
Anybody who has in their head a scenario where somehow we’re at parity in terms of intelligence with respect to each other’s stacks—that does not track reality, right? It is a profoundly asymmetric scenario, and that’s one of the first things that we think has to be fixed, too.
Speaker 1
Yeah. One of the reasons why I was bringing up some of those anecdotes was to put color around the trust issue. So, this story that I’ll just quickly mention: this defense official was talking about a power outage in Berkeley around 2019, and all the Chinese students were freaking out because they were mandated to report back on a regular cadence to the motherland, or their handlers, or whatever. And these are just regular people—regular students—not people particularly working in any sensitive areas or whatever.
It’s just, “You report back, or your mom doesn’t get her insulin, or your brother loses out on a job opportunity,” and there’s an escalating ratchet of systematic pressures that gets applied more and more and more. Basically, the weight of the state and this very optimized set of pressures is applied down to you, and it’s challenging for anyone to resist that degree of coercion. Which is, by the way, an absolute travesty and a tragedy. No one is a greater victim of this than the Chinese people themselves. Chinese nationals working in the United States are subject to at least the potential for this kind of activity. That’s—I can’t think of a word for it. It’s despicable.
But Chinese researchers have made amazing contributions to our frontier-AI achievements. You can just look at the names on the papers, right? So that’s another issue. You look at these labs; large double-digit percentages of the people there are at risk from this sort of activity. Imagine building a Manhattan Project in a context where that’s the case.
You quickly get into the really dicey things, right? We’re not happy to have to report this. We’re very much aware, obviously, that there are these awful things—Japanese internment camps are the other side of the coin here, right? That is World War II. That’s what it means to take an argument like that to one extreme.
But there’s also just the reality of what it would look like to build the Manhattan Project in a world where it’s not just that your team building the project has a bunch of Germans on it. The world of technology has fundamentally changed, and these people can be monitored nonstop. They have family back in Germany, and they have a bunch of financial ties to that country. That’s just a fundamentally different calculus, and so it’s a real challenge. But it’s a reality that somebody has to come out and just say, without fear of being called names: it is just a fact of the matter. And again, if there’s 1 thing we’re here to do, it’s report the news.
Nathan Labenz
It's an unfortunate reality. We wish we had different news to report. That's how things are shaping up right now.
Let's go down this rabbit hole of what it actually looks like to try to do an American superintelligence project—aka the Manhattan Project for AGI or ASI, or whatever. Maybe start off with: How does it happen? You mentioned the president and the DPA, but I'm interested in what you think it actually looks like for us to go from market normal to this national, Manhattan Project-style thing.
Then I'd love to hear a bit about how you actually learned what you did. I think you guys are great examples of putting shoe leather to pavement and actually getting out into the world and learning some stuff. I think that's really both credibility-building for the report and something that should inspire other people. Then we can get into the actual, okay, if we're actually going to do this, what is actually required to really do it. So, take it away.
Speaker 1
One thing I'll mention out the gate is that the way we started working on this report was in the wake of “Situational Awareness,” and we were like, “Well, a picture like this needs to be fleshed out whether it comes to pass or not.” Our view was that it's not clear that it's the right decision to do something like this, but it's sure as heck the right decision to think through what it would look like if we end up doing this. Someone had better have thought through those things in advance, right? And so we started working on it through that lens and through that angle.
Speaker 2
Yeah, I think the other piece is that we explicitly stayed away from—partly as a result of that—questions about what the specific authorities would be that would be invoked in pursuit of this sort of thing bureaucratically, or how this would be done, if only because, A, it's just not the best and highest use for us. We did include some thoughts about that because we ended up talking to people in relevant departments and agencies and hearing their views, and we didn't want to lose that value. So we did park some thoughts there, but they're not set in stone. They're just there to be considered.
Our main focus was at the gears level. What has to happen? What are the verbs that have to be written into the storybook for this to unfold? What does a data center have to look like? What does a supply chain have to look like? All of these questions—in many cases, the interesting challenge you run into is that the space is very long on opinions, much of which is very well-informed, and we've leaned heavily on a lot of the great work that's been done in the space.
But what's really missing is that if you actually talk to the very, very minuscule number of people—not intelligence and special forces; I'm talking about zooming in on intelligence and special forces at our Tier 1 units—these are our most elite units: SEAL Team Six, Delta Force. Then zoom in even more. Zoom in on the specific people who specialize in the relevant areas and ask them.
This is an extremely small community, and they're tightly networked. It's a very high-trust ecosystem. They are aware of TTPs—tactics, techniques, and procedures—that others are not aware of and on which your entire assumption stack can unknowingly rest, right? So when you have conversations with people in the space and you're like, “Hey, so let's just go out and do X,” their recommendation is, “Let's just go out and do X,” and you're just like, “Well, I just had a conversation with somebody about an approach that they know of that makes this totally moot.” That's the kind of thing that you run into. And so, unless you talk to this very specific group of people, you have an incomplete picture.
Going back to this idea of the jagged capability surface of the models, right, this is the jagged capability surface—or, yeah, capability surface, you could say—of the U.S. government and of U.S. industry. When you actually talk to people at the labs, you get a very nuanced perspective on what's possible. The same happens when you go to the intelligence community. The same happens when you go to the special operations community. You get a picture of, “Well, actually, this is trivial. This thing you thought was hard is trivial. This thing you thought was trivial is really hard.” And that suddenly just refactors your whole picture.
Some of the stuff is just wild, too. We visited a data center with some former special ops folks—not Tier 1 folks from these restricted, very elite units—and they're just wandering around, asking, “What happens if you do this to that thing? Oh, that? Okay. What happens if you do that to that? Okay, cool.” They wandered around a little bit, came back to us, and just said, “Yeah, so I can think of 1 or 2 ops you could pull for 30 grand that would knock out this data center for about a year.”
And everyone's like, “Wait, a $10 billion facility? A $10 billion facility?” When you're talking at that level, you're not even talking about China being able to do this. A smart person with a bit of time and money can do this, and so that means that adversaries can just proxy that and use surrogates, smoke screens, and sock puppets to go and do that with total deniability.
That's the state of security on what those vulnerabilities are, because those vulnerabilities remain, and they are critical and universal, basically, across data centers today. That's kind of one of the challenges here, right? The right answer is in so few heads and then needs to be integrated, and we're under no illusions that we have the right answer across the board. That's another implication directly of what we're saying here: the full picture resides in no one's head. But the combination of an awareness of that uncertainty, and having been shocked by a couple of these individual things, is sort of what is shaping at least the strategic picture behind the document.
Jeremie Harris
Yeah, and just quickly, it's one reason why it can make sense to publish something like this at all: it gives folks something to point to and say, “Hey, I actually know that's wrong because my specific niche and area of expertise says this and this and this.” And that is the beginning of a constructive discussion.
Nathan Labenz
So, without getting into the specific vulnerabilities, let's talk about what it would take to actually secure such a project. I mean, we're in the scenario now where governments are wide awake. Governments have decided we're in this race. We've got to win it. We can't just have people spilling all the secrets at San Francisco parties. We can't have whatever data center vulnerabilities exist at the infrastructure or hardware level. It's time to get serious. What does getting serious actually look like?
Jeremie Harris
So the first step—and again, we're going to go a bit meta just to help us have a productive conversation without getting too deep, and then there are a couple of things we can say at the object level—is this: if you think that timelines on the order of late 2026, 2027, whatever, are plausible, you ought to be in the business of finding all of the one-way doors that currently exist in the critical infrastructure stack, from supply chain to data centers and all that stuff.
By one-way doors, what we mean is that these are decision points we're walking through, sometimes without realizing it. For example, data centers will take you 18 months to build, or whatever, at least nominally, with the current rate and regulatory environment. That basically means, hey, it's time to show your homework. The world's knocking, and the data centers that we're breaking ground on and designing today are the ones that will have to house a lot of these superintelligence-grade training runs if they happen in 2027.
The question then becomes: What can we do? What are the cheapest things that we can do that the market will support today that buy us optionality down the line? That has actually been a huge fraction of what we have focused on—talking to these special forces operators, talking to the intelligence community, talking to folks on the hardware side, really deeply understanding what that set of one-way doors is that we're walking through—so we can, at the very least, not rule out the optionality that we need down the line.
From a philosophical standpoint, that's kind of it. There's a very small number of people who can actually give you at least what we consider to be the right answers here, and we've seen a lot of interactions between people who think they have the answer and then meet another group of people, and you're kind of like, “Oh, okay, I guess not.” To the point where we finally kind of coalesced around some things where we're like, “Okay, interesting. This is now sort of locking in.” But that'll evolve, too.
I'll maybe give 1 example of a one-way door that we can talk about that is in the report. This is what's called TEMPEST attacks. These are actually kind of awesome and wicked, James Bond-like things where you can figure out what's going on inside of a computer or a system, or pull data from it, by just watching the electromagnetic emanations from that system. In the most extreme version, you're in a room with a computer.
Edouard Harris
The computer is air-gapped. It’s literally not on the internet; it’s just sitting there on a desk, not connected to anything. If you have a particular virus in that computer—this is one example of this—that accesses the computer’s memory in a specific pattern, you can have a radio receiver a few feet away, across the air gap, listening for the radio emanations from that little memory-access activity going from the CPU to the memory. You can actually have information transmitted to you across the air gap by just listening for the memory-access patterns. It’s insane that it’s possible, and that particular attack and its detection are public information.
It was published about a year ago, but it almost surely had been known by intelligence agencies and many of the usual suspects for quite a while before that, because it’s something you can do and it’s extremely useful. There are standards around how you defend against that. As I talked about, if you’re a few feet away, you can detect it. The only way to realistically do that and block it off at the data-center level is to have a few feet of spacing between your compute racks and the walls of your data hall.
You could have a visitor space on the other side of that wall, and some dude who looks all innocent could have an app on their phone or something that’s detecting the emanations from your stuff. You’ve planted malware on the system, and they’re just pulling data from you at a bit rate, sitting casually in your visitor center and pulling data from your computers and your GPUs across the wall.
The way to block that is to create that space. If you don’t create that space—if you don’t build your data halls with that space in mind—then you’re like, “Oh, whoops. My data halls are just too physically small.” That’s a one-way door because, yes, you can retrofit, but it involves knocking down and redesigning your facility.
The data-center one-way doors are really interesting because they affect construction today of a thing that you can see right now. The supply-chain one-way doors are, in a weird way, actually thornier, because they involve an entire supply chain.
TSMC makes all our chips. TSMC is in Taiwan. I don’t need to tell you that the CCP is very, very interested in what’s going on there. If you were China, you could just draw the obvious extrapolations: What would you be up to, knowing about this? It doesn’t even have to be TEMPEST. It’s cyberattacks, personnel—I mean, TSMC famously has lost executives who went on to found SMIC and steal their IP, right? This is a known playbook that has been working for decades.
Nathan Labenz
Twenty years, yeah.
Edouard Harris
So there’s that. There are also the mundane, easily overlooked, boring components that go into server systems. Baseboard management controllers are a great example. About 70–80% of the BMC market goes to ASPEED, this other company in Taiwan. I hope they have their shit together, because baseboard management controllers run on a separate power supply from the GPUs that hum and the CPUs that hum. They often have read and write access to firmware, and it’s a dream backdoor. It is the soft underbelly of a lot of HPC and AI infrastructure.
There are a million things like this. There are transformers and transformer substations; component-wise, none of them have zero components made in China. We know that the CCP has actually used transformers specifically to plant backdoors in American infrastructure. It’s not just one thing—it’s so many of these things.
It’s tough because you talk to people about it, and there is correctly a view that you can’t cover down on all vulnerabilities and all threats. You have to pick and choose. The challenge is that we’re not even highlighting—or, let’s say, a good fraction of what we’re highlighting here isn’t speculative. It’s stuff we know China has literally done to us on our soil. So, yeah, it’s one of those things where you just have to take the reins in some way.
Nathan Labenz
Maybe give me a little bit more on the backdoors as they’re understood to exist in the electrical grid, and then maybe that’ll be enough to imply what it means for the data centers. I’d be interested, too, in the report: There are cooling modules that come from China; the crown jewels of a data center might come from Taiwan, but a lot of the other pieces needed to make it go seem to come from China. My takeaway was that there’s really just no way we’re going to bring that whole supply chain to the United States in any sort of short-term scenario.
But maybe let’s start with what’s known and understood. When you say China has a gun to our head with respect to the transformers that they’ve backdoored into our electrical grid, what does that mean for me? What am I vulnerable to that I may not realize?
Edouard Harris
This means that they can hold our infrastructure at risk and use that, in some sense, as a bargaining chip if things heat up—to disrupt us completely and distract us while they’re taking hold of their own objectives.
This is one possible scenario of a Taiwan invasion, where there are many logistical challenges that the United States faces in defending Taiwan. One of those is that we’re just going to run out of stuff after the first week or two of intense, high-heat combat at those scales. But another thing they can do is say, “Hey, if it looks like things are not going as well as they hoped for or planned because of American intervention, sure, just turn off all of our grids,” after putting some propaganda in our information spaces. Basically, turn out the lights and let the chaos reign. Just throw a wrench in the gears. Why wouldn’t they do that?
I think one important ingredient strategically here is that when you have a capability or set of capabilities to deploy against an adversary like China—or like the United States, whoever you are—you’re never going to show your full set of cards. You always go with the smallest kind of jab you can throw that has the desired effect and doesn’t reveal your whole set of capabilities. The intent is always, “How can I learn without teaching?” That means we genuinely have no idea how high the ceiling is on Chinese capabilities, but the floor is pretty goddamn high.
As we think about stack-ranking and prioritizing the interventions—which is where I think your question was heading, like, what do we actually do here?—the good news is that the structures you need to build are a relatively small number of structures relative to the power demands of the entire United States.
We can’t go into the solution because it is redacted from the document for obvious reasons. There are solutions that we’ve identified for a number of things, but I won’t mention them. I can’t remember if we say what the thing is, but in any case, they have to do broadly with supply-chain techniques. The fact that this is a focused problem does help you: There are ways in which the problem is constrained geographically and technologically that make it more tractable.
Another kind of corollary to this is that you can’t live in a world where China holds you for ransom and not do the same to them. Right now, that is the world we live in. We have very poor visibility into what’s going on on CCP infrastructure.
That gets hard, actually. Nathan, I think we talked a couple of weeks back about why Google released Streaming DiLoCo. That whole kind of series—this gets harder with DiLoCo. It gets harder with Prime Intellect. It gets harder with Together AI, all these distributed-computing technologies that just make it so hard to track what is even happening, right?
For today, you can hope that a big mega-cluster—10 gigawatts of Three Gorges Dam-like power—is something you might detect through satellite imagery, but that’s not always going to be the case. In fact, it seems plausible that it’ll become less and less the case on relevant timelines.
The question then becomes, “Okay, well, how do we establish a reciprocal capability?” We need at least that option. If China has the ability to hold American critical infrastructure for ransom, the fact that we do not have that option in reverse is a giant problem and should be a priority to fix.
Without going into too much detail, our understanding is that there is stuff we can do in retaliation, but we’re not where we need to be.
Nathan Labenz
To make this a little more colorful, are we talking about transformer substations in communities across America suddenly going Hezbollah pager and just exploding, or what exactly is the scenario?
Edouard Harris
Yeah, stuff like that. The thing is, this actually comes back to what Jeremie was saying: You reveal the capability that gives you the effect that you desire, and nothing above that.
So you’ve got 10 levels that you can operate on as a nation-state. You’re operating on all 10 of them. You want to reveal the lowest level that does the thing that you want. For example, there are little touches that we see across the United States where you have things like—this has been reported—a couple of Chinese folks flying drones to observe critical infrastructure and things like that.
You certainly can just knock things out with explosives or literally short-circuit a transformer when it’s turned off and rewire it into itself. If it’s not being regularly inspected, when that transformer turns on again, it explodes like a bomb. All of that stuff is within the arsenal of things that they can potentially do in certain critical places.
Other stuff involves owning the firmware on that transformer. You’re also the supply chain that produces that transformer. If you’re thinking far enough ahead, you can totally install back doors that make that transformer effectively inoperable and just brick it. So you don’t need to blow it up. You don’t need to do anything that subtle. In theory, you could flip a switch and say, “Okay, this transformer is just a brick now,” and that’s the way it is.
Now it’s a game of, well, we need to somehow replace a large fraction of the transformers in random places in the United States under conditions where many places in the United States don’t have electricity or comms because the transformers are now out. And we also don’t build many of them, if any.
Nathan Labenz
Right. So that’s another big problem.
Edouard Harris
Yeah. The means to build a transformer or a transformer substation is really interesting, right? Where do you source your materials, and how far up that supply chain do I have to go before I have a manufacturer in China? It’s actually quite often not very far. And that’s part of the challenge, right?
Generally speaking, this is a very high-dimensional problem. Obviously, again, data centers are fortunately an insanely high-dimensional object. They’re a much lower-dimensional object than the entirety of the United States, and there are things you can do anyway to reduce the risk in significant ways.
But you can see that geostrategic calculation here too, right? It’s like, okay, given that this is the case for the United States, if this is true, then we have a larger strategic challenge on our hands even absent ASI. It can be awfully tempting to look at ASI as partly a solution to that problem. If you have the ability to build whatever kind of crazy offensive weapon you might build through ASI, there’s a temptation to say, “Well, if we can just secure this one thing, and we can use this thing to have this effect, then we’ll be able to transcend all these other issues.”
So the strategic calculus here is super, super loaded and complex, which is, anyway, a big part of what we can’t talk about.
Nathan Labenz
But I mean, the general picture that I take away from the report is: we’re moving AI research to some sort of secure facility in Nevada or something, some remote place, and we’re building data centers underground. We’re trying, at whatever cost—we don’t care about it being economically competitive—we just need to figure out ways to make certain critical components domestically so we can have some sort of secure supply chain. So we’re just pouring good money over bad onto that. And we’re also doing serious vetting of our people.
Maybe not allowing them to go home, meaning they literally have to sleep on-site and have their calls monitored or whatever. Maybe that would be a way not to have to literally purge all Chinese nationals, or even just Chinese Americans, if we had that sort of level of surveillance. Obviously, that’s not super pleasant. It starts to sound, by the way, a lot more like China, which is a general trend I should note in American governance.
You know, I just tweeted this morning: the U.S. government is trying to tell me that I should be afraid of Chinese values. At the same time, it is embodying Chinese values in a more and more egregious way all the time—literally disappearing people without due process for political expression. So that’s a problem.
Jeremie Harris
Well, with a project like this, here again we’re talking about the assumption that you’re doing a national superintelligence project, right? And without necessarily talking about whether it’s a good idea or not for this purpose, if you’re actually doing that, you look back to the original, actual Manhattan Project for nuclear weapons. Even back then, you had this kind of surveillance through the comms channels they had available.
Feynman famously made a code system with his wife, right?
Nathan Labenz
That’s right.
Jeremie Harris
In little protest of that. They were reading your letters, and that was absolutely outrageously unprecedented at the time. They had to agree to it voluntarily to have their mail read and all this stuff. People were playing around with the censors and the code and figuring out the boundaries of the system, but then Klaus Fuchs just went and stole the secrets anyway.
Edouard Harris
Well, a couple of thoughts here, though, too. On the Klaus Fuchs thing, this is another area where, as research gets more and more automated, that does help. In some ways, it’s a double-edged sword: from a loss-of-control standpoint, holy shit. And from a security standpoint, it’s like, “Oh, okay, I see what you’re doing there.”
The other side, too, is that we’re not advocating for building these things underground or in any such context. Part of the calculation was, okay, here’s the list of the full suite of things that you would really want to do to play it tight. Separate question: how do we pragmatically do this? That’s actually a question that we’ve been thinking about and working on for the last 6 months really intensely—figuring out, okay, if we actually had to do this for real, what does the Pareto-optimal frontier look like, and what do those trade-offs look like?
I think you actually can get to some pretty satisfactory places even in the face of the full Chinese and CCP threat here. But the question is always going to be: how do I spend my marginal dollar, and how do I trade off the need for more compute and whatever else with a buffer for security?
Nathan Labenz
And to be specific about that, how do you actually think about those trade-offs? You can’t cover everything. You can’t actually, in the time we have, build a data center in the mountain with freaking cooling ducts sticking out of it. It’s just—forget it. That’s too much work. So how do you actually think about it?
Edouard Harris
Well, you think about it in the context of—again, if you were doing such a project, you would need it to be embedded into your national security and counterintelligence apparatus. You’re taking this because otherwise it just doesn’t make sense as a standalone thing. The NSA has a team there, right? Other stuff like that. You’ve got people actively working and collaborating.
One of the things that means is that you’ve got eyes and attempts at collecting information on the adversary’s side. You’re collecting information on China’s efforts to subvert, attack, and exfiltrate from that project. Now your challenge isn’t actually to defend against everything. Rather, the challenge is: we want to put enough measures in place to raise the cost for them to do an attack or an exfiltration to the point where we can see that, because we’re forcing them to marshal enough resources to be able to overcome our security bar.
That doesn’t mean that we can actually fully defend against everything at the end of the day. They can just shoot a cruise missile from offshore and blow the bejesus out of us, and there’s really nothing we can do about that at the end of the day. But if they do that, then we can see that cruise missile and go, “Hey, guys, you just shot a cruise missile at us. We’re going to retaliate.”
And so you basically raise the bar to the point where you force enough resources to be marshaled from their side that creates a signature we can see and offers a threat of retaliation. This is how that security posture gets integrated into a counterintelligence operation.
Jeremie Harris
Yeah. The kind of flip side to that, if you don’t have that set up, you essentially just—if you’re OpenAI or you’re xAI or whatever lab is close to doing it first—what you find is, like, “Oh, this training run just isn’t working.” Or the inferencing process, whatever the key stage ends up being, is just not working. It’s super buggy. There are a bunch of weird issues. In the meantime, obviously, you’re not aware, but relevant IP has been exfiltrated. There are larger-scale training runs happening more efficiently in China or wherever, and your adversary is developing that capability, and you have no consequence.
And this is one common theme that we just—I guess, to add some color, this is maybe backing up a little bit—but on the U.S.-China side, what’s been missing, at least until now, what’s been missing throughout the Biden years—this has been a complaint that we’ve heard, and before that too, in many administrations—has been consequence.
Edouard Harris
So America's adversaries, including China, have been able to conduct operations in the United States that do a lot of damage without consequence. That's because there's been a lot of anxiety that, if we respond, the brain immediately goes to nuclear escalation. The problem is, you cannot do business with adversaries when that's your mindset. You just can't.
What you end up with is people who are going to take advantage of every little thing. It's the whole “I'm not touching the devil” thing. They'll get closer and closer to your face. They'll eventually start to do things that we know they've actually been doing to American critical infrastructure and people.
I mean, you think about the Chinese police stations operating within US borders, basically with impunity. What the hell is that? If we acknowledge that these things were real, those are things that border on acts of war. The fact that we have not imposed consequences for those moves means that we're teaching the adversary to continue to push in that direction. We're now about to pay the cost when it comes to AI.
If there's no consequence for what's done to American interests, domestically or abroad, and if people are able to act with impunity, all of a sudden I get to try, for free, to take out your critical infrastructure. Why would I think twice? It's a free shot, whereas we don't get the same reciprocally.
People don't realize this, but stability between powers today is not maintained through actual defense. It's maintained through the threat of consequence and the threat of credible retaliation. It's actually kind of fucked up that the world is this way, but stability between major powers and states is maintained the same way that gangs in Chicago maintain each other's territory. I know that I can't actually stop you from busting a cap in the ass of one of my boys, but if you kill one of my boys, I'm going to kill one of your boys right back.
Nathan Labenz
And that's the first time I've ever heard Ed say “bust a cap.” Okay, cool. Yeah, yeah, that's a first. That's an exclusive for The Cognitive Revolution.
Edouard Harris
But that's how stability is maintained, in truth. It's unattractive to think about, but it is the bedrock. It's the air we breathe in the world today, as wild and scary as that is.
Nathan Labenz
So, I will refer people to the report for more depth and detail on the many supply-chain vulnerabilities, the difficulty in securing a legitimate research staff in the remote locations to which they might be abducted. On top of that, we haven't even gotten into the measures that you recommend for maintaining control over the AI itself, which is definitely a new type of capability for everybody, including the US government, if it wants to develop and deploy those, to make sure that it's not inadvertently unleashing Armageddon on everybody.
So, we've got a lot of problems. I know you're trying to keep this somewhat neutral, like, “We're not saying do or don't do this. We're just trying to tell you what would be involved.” To put my cards on the table, I read this and came away feeling like we shouldn't. It reads to me like a warning more than a manual that I'm excited to sign up for.
Jeremie Harris
Actually, this is so funny, because we've had so many conversations with people who are like, “Dude, why are you telling people to do a superintelligence project?” It's a bit of a Rorschach test. People will see what they want to see in it, and to an extent, that's fine.
I think the thing that we took from it after all was said and done, starting from “We don't know if it's a good idea or not,” and then having done it and looking back and being like, “What is our conclusion?”—a big part of our conclusion is that most of the measures we recommend should be implemented whether we're doing some kind of big national thing or not.
If they're not implemented, then you get exactly the situation that Ed described, which is, yeah, China just steals your weights at the 11th hour. They sucker you so that you don't even realize it's happening. You just feel like your training run is proceeding more slowly, that there are so many more bugs, and so on. All they're doing is throwing sand into your gears while they're working and improving things on their side, and you don't even realize that's what's going on.
One of the things we learned from talking to these intelligence folks is that non-state actors, like proxies for nation-states, have many of the techniques, tools, and procedures—in fact, almost all of them—that the nation-states themselves have. The distinction between China doing this and some group that China is supporting, a non-state group, doing this is actually not that sharp in terms of capability. They're trained, funded, and supported by the Chinese, or by whatever adversary.
So, you're facing down nation-state capabilities even if your proximate adversary is not actually a nation-state. What that means is that the bar for security just has to be high all around to prevent this from happening. Whether we have a national superintelligence project or five random superintelligence projects that are all kind of dipsy-doodling from their own sides and angles, this is just the vulnerability surface that exists. And that's true.
When it says “a national superintelligence project,” essentially a US government-backed superintelligence project, that's where we start getting into, well, honestly, it doesn't matter. At the end of the day, what you care about is how the rubber is going to meet the road. Are they or are they not going to be able to extract critical IP, exfiltrate your stuff, or do any number of other things? Those questions all have answers that route through a lot of the same channels. There's a question of task organization if the US government takes it on, but that's why we scoped that out.
A lot of this really is the story that, if you're going to build superintelligence and treat it as the WMD capability that Sam Altman tells us it is or will be, with no apologies, then you ought to be doing something that looks fairly different from Stargate. What it looks like to actually take that seriously—to take your responsibility seriously in building this tech—looks very different from announcing to the world that you're doing a $500 billion project, where everybody knows exactly where the facilities are, and, oh, by the way, your lab contains X% Chinese nationals. We've got all these issues with people openly leaking critical IP from Slack.
One of the things the report uncovered was that we had an insider from OpenAI tell us, “I was just digging around and found these 2 really critical cyber vulnerabilities that would have allowed me to access and then extract weights for a critical model that the lab was hosting.” This is just not what you do if you are a serious actor and you take your commitment seriously.
That's a big issue. It's not that OpenAI is a bad actor in the zoomed-out sense. It's that, in part, they're just victims of the racing dynamics that are playing out right now. So how do you get your arms around that? Step 1, presumably, is having a sense of what the vulnerabilities are that you need to cover down on. Then you can price those in, start to make a list of them, prioritize by cost, prioritize by the probability that they'll be exploited, and work your way down. That's all you can ever really do in this space.
Nathan Labenz
I don't love it, and I know you guys don't either. Let me just try, in the few closing minutes that we have, to get a couple of reactions on other things. One is that China's open-sourcing all its stuff. One sort of Pollyanna-ish view is that they don't appear to be trying to race us; they're giving us all their best shit for free. How do you interpret their open-sourcing strategy? Do you expect it to continue?
Jeremie Harris
Whether they have an open-sourcing strategy is one question that never gets asked. The answer surely is no. They have an open-sourcing strategy, but it is not a universal open-sourcing strategy.
One of the challenges is that we do not know what's going on on those Chinese servers. We can't answer that question. What we see is that we actually heard from a former OpenAI researcher that he was having a conversation about, “How do we know that China hasn't stolen any of our shit?” The response at the time was literally, “Well, we haven't seen any comparable open-source models come out of China.” So that basically tells us they haven't stolen any of our shit.
Nathan Labenz
Interesting. Like, dude—comparable models?
Jeremie Harris
Models with comparable capabilities.
Nathan Labenz
Yeah, right. Sorry—things that are legible.
Jeremie Harris
And so you're just not going to get that intelligence. The flip side is that anything that comes out of DeepSeek from here on out—in fact, anything that comes out of Huawei or anywhere else, Tencent, you name it—is coming out with the imprimatur of the CCP.
There's a 0% chance that those models are getting open-sourced without people high up at the CCP going, “Cool, at a high level, that makes sense.” They may not approve every single individual one, but the general strategy behind it is CCP-approved.
You can make of that what you want. If you were the CCP, you'd probably have an interest in flooding the West with language models that aren't going to talk about Tiananmen, at a minimum. But as we move into more and more agentic systems, boy, does it look interesting to start to buy the public's confidence in your agentic models when there are interesting backdoors that could be baked in to produce any interesting behaviors you might want to exploit.
One other thought there is that we suspect that, prior to the end of last year, DeepSeek was not that integrated into the CCP apparatus, though it is super integrated now. The reason being a number of comments by the founder of DeepSeek around, like, “Man, American export controls really work against us. Geez, they're really biting,” that he made back in July on a random podcast, which utterly undermined the CCP's entire propaganda around, “Yeah, American export controls just don't work around chips, so you might as well give them up. You might as well just stop, right? Please.”
And they've been putting propaganda set pieces around that in place for years at massive cost and effort. This random dude, who has now been catapulted into fame, just absolutely gutted that entire effort at great cost.
There is also the economic argument, too. These are just public arguments people are throwing around, and I'm just offering them as explanations for why the CCP would be okay with this, which is the ultimate question that we really just need to answer. But the other one is the economy of flooding the market with cheap LLMs. You've seen what that's done to Meta. I mean, Jesus. They're basically now having to gut their [???].
Nathan Labenz
Yeah, exactly. So, okay. I don't know if we have time for 2 more, but I would ask 2 more if we have time. One more. All right. Maybe I'll combine them into 1. How can I do that? Maybe I'll just put them—I'll ask them both, and you can answer as you will.
One question is, do we even have a set of prerequisites for some sort of grand bargain? If we were going to say, “Okay, we're going to come to the table. We know we don't have great trust, but is there some sort of technological checklist and/or trust-building-exercise checklist that we could go down to say, ‘Okay, if we had these things, then we could maybe enter into some sort of sustainable grand bargain?’”
Part 2, another angle on it is: I know you guys read Dan Hendrycks et al.'s MAIM theory, and I wonder if you have any hope for either an emergent equilibrium that sort of keeps things under control, or perhaps a more engineered—maybe as part of a grand bargain—some sort of semi-engineered equilibrium that could keep things stable between great powers, even if trust continues to be a scarce resource.
Jeremie Harris
On the grand-bargain side, the prerequisite starts to touch pretty closely on the whole FLEX program and secure, governable chips. Can you make a tamper-proof enclosure for your chips? Can you make at least a tamper-detecting enclosure that requires inspections? Can you then have some sort of governance chip sitting on there, too, to verify inputs and outputs, all that stuff? I love that research agenda. I think it's a great research program.
The problem with it is timelines. It's a really hard obstacle to overcome: fabbing is slow, chip design is slow, integration is slow, scaling is slow, and debugging is slow. I love that program, and I hope that it gets significantly accelerated by AI. This is something that I think should be a priority in the context of the Superintelligence Project.
You would want to couple the intelligence engine to the design process for these chips. You'd want to couple it to advising on what to do next and fab next. But, yeah, I think a prerequisite is that we need to have visibility. “Trust but verify” has got to be an option. Until it is, we need to be treating China as an intelligent adversary and a committed one.
Maybe I'll leave it to Ed to do the cleanup on that. If you have thoughts on the main thing you want to throw out, Ed, too.
Edouard Harris
Yeah, generally agree. I think we need to be pursuing every avenue, and if there is any chance, even though it seems quite remote, of some kind of verification around collaboration, let's throw a few billion bucks into that. Absolutely. If that turns out to be a dark-horse win, totally—that's worth the investment.
Timelines are tough. Another tough aspect is that the folks we talked to are not optimistic about whether you can actually build a tamper-detecting enclosure at all if it is subject to sustained pressure under the full control of an adversary that can attack it even at scale. You'd have to basically bust through 1,000,000 of these to have a meaningful amount of chip capacity, which is less optimistic.
And on the MAIM side, it's actually a shame we didn't get to discuss MAIM more because it's a really, really interesting piece of work. I think one of the things it did really, really well was that it was the first place where this approach of, “No, we actually need to be reaching out and touching the adversary. We need to be actively doing counterintelligence and capability-degradation stuff soon,” really came out. That's a critical pillar, and you can kind of see how the logic forces you into it, right?
If you want to have an aligned AGI or superintelligence, you need a margin over the second-place adversary or competitor to actually do that development of alignment technology. In order to get that margin, you can either accelerate yourself or try to decelerate the competitor. Accelerating yourself shortens timelines and creates additional risk. The logic kind of forces you into: Well, you need to do capability degradation for the adversary. MAIM did a great job of framing that and pointing that out.
One of the issues with it, when we talk to some of the intelligence folks we're connected with, is that you can't actually put a bar of security where nation-states can go above this bar and hold your stuff at risk, but non-nation-states can't. So you can defend against one, but not the other. As we've seen, those sets of capabilities are actually not as distinct as many people think.
This idea that you can fine-tune security levels and do this kind of scalpel stuff is not viewed very encouragingly in those spaces. It doesn't mean it's not a really important part of the truth—it is—but there are some aspects missing that I think it would be really productive to discuss with Dan Hendrycks more deeply, even on your podcast, at some point in the future.
Nathan Labenz
I love it. We'll see if we can make that happen. Any other closing thoughts or assignments you want to give to the audience?
Jeremie Harris
None that comes to mind. We think of it as a living document, even though it's a point in time. It's not going to be perfect, and things change. We're really eager to get anybody's views and thoughts on it. We want to aggregate as much as possible.
We're maintaining a list of data-center security techniques that we think would make the difference, especially one-way doors. To the extent that you have thoughts on that side of things, we're always eager to hear from anybody in the intelligence community, special-operations community, and all that. Just any relevant community, right? AI security, AI folks, secure, governable chips, AI alignment people. Cool. Well, then they'll know where to find you. For now, Jeremy and Edward Harris, founders of Gladstone AI, thank you both for being part of the Cognitive Revolution. It is both energizing and enlightening to hear why people listen and learn what they value about the show. So, please don't hesitate to reach out via email at tcrturpentine.co or you can DM me on the social media platform of your choice. [Music]