Nathan Labenz
Hello, and welcome back to The Cognitive Revolution. Or, in this case, I should say, welcome to AI in the AM. This is the third time that my friend Prakash Narayan and I have done a livestream together. This time, we figured we should give it a name, and he also took the initiative to create a new look for the show with real-time AI transcription and AI-powered comment moderation. Check out the video and I think you'll agree that he's done a really nice job with the look and feel.
As you'll see, in some ways we are still very much figuring out both what we want the show to be and how best to organize and produce it. One thing we're going to look at after this episode is creating a mechanism where we can easily signal to one another when we'd like to ask a follow-up question or move on to another topic. Nevertheless, when it comes to the quality of guests and conversations, I think this episode is right where we want to be.
Our guests for this episode were Sergey Nesterenko, CEO of Quilter, which is using reinforcement learning to train AI systems to perform circuit board design—a problem with an insanely high-dimensional search space, complicated physical constraints, and relatively low volume of available training data. After that, we spoke to Andy Hall, professor of political economy at Stanford, who's doing a bunch of interesting work to characterize model behavior in political contexts and who's also working to design independent AI governing bodies that he hopes will allow the public to exercise some oversight over AI companies without requiring nationalization. Finally, we welcomed Lucas Petersen and Axel Backlund from Anden Labs. You may know them from their autonomous vending machine work, but today we'll be talking about the new AI-operated retail store that they've recently opened on Union Street in San Francisco. The store, which is managed entirely—including the hiring of human staff—by an AI agent, currently has a 2.6-star rating, but I personally still can't wait to visit.
For me, the big takeaway from this series of conversations is, once again, that the future is coming at us much faster than we can process it. Assumptions that seem safe from one perspective become very questionable in the face of increasingly powerful and autonomous AI systems.
With that in mind, I want to add just a bit to my answer to the very first question that Prakash asked me in our opening discussion: namely, why are we now suddenly seeing violent outbursts directed at AI lab leaders?
First of all, while it's certainly possible—and I would very much hope that the recent attacks on Sam Altman's home will ultimately prove to be a random blip signifying nothing—my honest assessment is that, by default, we should expect to see more of this kind of thing. Not because of the super high-IP doom numbers coming from the AI opposition camp, but simply due to the fact that more and more people are now becoming aware of the extreme reality of the AI situation.
It wasn't that long ago that Sam, Dario, Demis, and Ilya all signed, alongside many other luminaries, a statement saying that “mitigating the risk of extinction from AI should be a global priority alongside other societal-scale risks, such as pandemics and nuclear war.”
The record does show that each of the leading AI companies was founded with awareness of and an intention to address the hard problems of AI safety. Of course, the upside, it goes without saying, is unquestionably immense. In practice today, they are developing what they recognize to be destabilizing and likely dangerous technology pretty much as fast as they possibly can, while repeatedly failing to live up to their own prior safety and social commitments.
In significant part because—and here I am quoting Sam Altman's immediate reflections after the Molotov cocktail incident—“Being the one to control AGI has a ‘ring of power dynamic’ to it.” All while, by their own accounts, we still face anywhere from a 5% to 20% chance of something like, as Sam himself famously put it, “lights out” for all of us. And the U.S. government's main concern seems to be making sure that nobody can constrain its ability to use the technology for autonomous weapons, domestic surveillance, or anything else, for that matter.
I can't emphasize enough: objectively, this really is a crazy situation. Those of us who stumbled onto the idea that all this might happen years ago have had a lot of time to get accustomed to it and to position ourselves to do what we think we can about it. Many of us have reconciled ourselves to the idea that some version of it is inevitable. But that doesn't mean that we should try to tell people who are only now learning about this that they're wrong for freaking out about it. A 1-in-20 chance of human extinction is not low and absolutely is worth freaking out about.
I'm reminded of 2 movies that memorably illustrate how I think many people will respond to learning the facts about AI. In the 1998 movie Armageddon, when an asteroid is found to be on course to destroy the Earth, it's simply understood that it's heroic for individuals to risk and, in the end, even to sacrifice their own lives to save the world. That doesn't hinge on whether there's a 5%, 20%, or 99.9% chance that the asteroid really will hit the Earth. The heroes would be heroes in any case.
In contrast, in the more recent Don't Look Up, the main characters are continually frustrated that nobody can be bothered to recognize the crisis at all, driving them to become crazier and more desperate until—spoiler—everyone does ultimately die in the end.
I would submit that the difference between a hunger strike and an act of violence is not about how one understands the stakes or the odds. Neither is it about the impulse to martyrdom. Rather, the difference is simply that one course of action attempts to call others to a higher ethical standard, while the other is not only condemned by every principal moral tradition but, even on purely consequentialist grounds, seems almost certain to make everything harder and worse.
To the AI opposition movement—which, for what it's worth, I think is increasingly distinct from people focused on AI safety—I would say: absolutely continue to condemn violence. But at the same time, be careful not to shy away from the fact that it's the situation that's crazy, not the people who are desperately searching for ways to make a difference.
Your job, in addition to educating people about the reality as you see it and as the lab leaders themselves have described it, is to identify and create productive ways for people to act heroically in this moment. Those could include mobilizing voters to contact officials and advocate for regulation or international treaties, investing themselves in citizen-level diplomacy with China, as I personally hope to do, developing new governance models, or pursuing experimental technical alignment strategies. And importantly, probably lots more that people haven't even thought of yet.
I personally always encourage people to pursue their own AI safety ideas, however eccentric they may seem, in the hope that some of them might actually pay off, and because I believe that, in the absence of constructive ways to devote oneself to the cause, we will see more people simply going crazy. As always, I will welcome your feedback, both on this analysis and on the new show format. Until further notice, we do intend to run these conversations on the Cognitive Revolution feed. But if it goes well, we might spin it off into its own thing. If you'd like to see that happen, we definitely encourage you to follow the new show account on Twitter @aiintheam with underscores between each word. That's ai_in_the_am. And watch out for the next live stream, which is currently planned for April 20th at 11:45 a.m. Eastern, 8:45 a.m. Pacific. Thank you to everyone who listens for being a part of the Cognitive Revolution. And now, on with the show.
Prakash Narayan
All right, and so we are live right now. We are live. This is Welcome to AI in the AM.
Nathan Labenz
Thanks, Prakash.
Prakash Narayan
Thanks for setting this up. Great to be here. I like the new look.
Nathan Labenz
Yeah, this is our third stream—third livestream—and we decided to add a little bit of pizzazz this time.
Prakash Narayan
When I heard there was $100 million on offer for anyone who sets up a tech-focused livestream, I figured, how can I miss it?
Nathan Labenz
Yeah, incentives—the power of incentives, right? I think we also decided to make this perhaps the first livestream of its kind where we're using a lot of AI tech. We have live transcriptions, which I don't think any other show has ever done before, because accuracy has never been good enough to have live transcriptions.
This means that you can watch this in a meeting. You don't have to think, “I see something on screen, but I can't—I don't know what they're saying.” So, you can watch this in a meeting. Well, you're wise.
Prakash Narayan
An AI note-taker is attending your meeting for you, and then you can be surreptitiously watching us on the other tab, with this version of the AI transcription helping you do it live.
Indeed. We also have live comments available if anyone wants to mention @AI_in_the_am. It is being moderated, I think, by Grok 4.1 fast, which is fairly fast. Before we take off, we just had the second attack on Sam Altman's house. I don't know if you saw that. It seems 2 people shot rounds at his house on Russian Hill. It's pretty scary—his family's in there. At this point, I wonder if you just have to move out, right? You can't be in San Francisco anymore.
Prakash Narayan
You have to be in a defensible position. I've heard from 1 VC, actually, that he had his house in one of these suburbs, in Menlo Park or whatever, and he has drone defenses set up. So there are drones circulating above the house, and there are drone defenses.
I think it's very hard, because I think the level of security that's going to be required for an AI company head right now is going to be substantial. I imagine it's the same for Elon. I imagine it's the same for Zuck. They're not well liked.
And, you know, why is that? There's a lot of soul-searching going on on the timeline right now. Why do you think this is happening right now?
Nathan Labenz
Well, I think it's getting very real. That's one thing. Everybody can now see that—or maybe not everybody. I think there are a few holdouts, but increasingly, it's hard to hold out any sort of “They're just doing this for hype and to raise money, and there's nothing really there” position.
I think increasingly people have to reckon with the fact that AI is getting powerful. My guess is that—who knows? It's hard to put yourself in the mindset of somebody who would go throw a Molotov cocktail or randomly do a drive-by shooting of someone's home—but Mythos is really an example of at least a weakly powerful AI, right?
It's something where no less than Nicholas Carlini, who is by all accounts one of the great cybersecurity researchers of all time, has said that he's found as many important vulnerabilities in just the last few weeks as he had in the entire rest of his career combined. That is a huge indicator that we are entering a new regime.
I do think it's becoming real to a lot of people in a lot of different ways. It is a radicalizing reality, I think. It's not to be forgotten that all of the AI lab leaders have been pretty candid, if not super recently, at least at various points in time, that they're really not sure how this is going to go. They at least have some nontrivial P(doom) percentage, even at the top of these organizations.
I think it's rational, in some sense, to be like, “What the hell are you guys doing? This needs to be stopped.” I think that position is, in my mind, very defensible.
Obviously, in addition to just being wrong, I don't think it's going to be an effective tactic to try to intimidate these folks, because I do think their resolve will probably just be hardened. Their ability to defend themselves, with the resources that they have, is going to be pretty good, if only by retreating to some large estate on a private island in Hawaii or in New Zealand, or whatever the case may be.
I don't think this stuff is going to work. But I think it's not totally crazy to say, “Desperate times call for desperate measures,” and then it just becomes, you know, what desperate measures are acceptable and/or likely to be effective. I do want to be clear that I do not support these things. I think they're clearly crossing lines.
But there is some sense to the idea that these guys are telling us that they're taking, by their own lights, something like a 1-in-10 or 1-in-5 chance of the future going deeply off the rails. They don't really have a great account of how they're going to control things. Alignment is obviously unsolved, governance is unsolved—as a preview of our upcoming conversation—and yet we race ahead.
Nathan Labenz
Yeah. I would definitely like to see a more sane and productive response than these things. Unfortunately, I do think, barring any sort of government action that makes any sense, we're probably going to see more of this kind of stuff.
I do think, in their candor, they've set the table for some extremism anyway. That's not to endorse doing any of these things, but the mindset is certainly one that I can understand how people get into. Especially if they don't understand the technology, right?
There are a lot of people who are coming to this much more suddenly as it becomes a bigger deal, and they're just having a sort of sudden awakening to the fact that this is all going on. It really is super powerful, and maybe what they heard about it being all hype before was actually not right. I think that can be destabilizing to a lot of people.
I remember test-red-teaming GPT-4 in a minor way. I can only imagine what people who are just now coming on to modern AI developments are feeling and thinking.
On that note, let me introduce our first guest for this morning, Sergey Nazarov. He's the founder and CEO of Quilter. They're a company that's really trying to speed up electronics design in general and PCB design in particular.
They have a physics-driven AI, which I think uses reinforcement learning in order to really speed up the entire PCB layout process, which can take many weeks. They've managed to compress it down. He came out of SpaceX. He was in avionics, I think—avionics and radiation.
I heard about avionics growing up, and I always imagined it was very sophisticated stuff. Then I realized later on that it's a lot of compute. It's actually a lot of compute. It's a lot about figuring out what numbers need to happen.
In the past, a lot of that stuff was not electronic. There was a lot of gadgetry that was actually mechanical, and you had mechanical ways of calculating all of these things. Which is why avionics used to be this entire segment of creating mechanical computers that could work on F-15s and F-16s.
Later on, it just became electronics, but the electronics have to be hardened to radiation and a bunch of other things that happen. Sergey is an expert on that. Sergey, welcome to the show.
Sergey Nazarov
Yeah, thank you for having me, guys.
Nathan Labenz
I hope I didn't misstate all of those things. Give us an example of how you transitioned from SpaceX to PCB design. What did you carry forward from SpaceX into your new role?
Sergey Nazarov
Yeah, there's a lot to carry forward from SpaceX, to be honest. I spent about 5 years there. There's a lot to learn about culture, a lot to learn about how to hire great people, a lot to learn about how to design systems, and so on and so forth.
But I think the most important thing is just speed, right? SpaceX is probably most famous for hardware-rich development, in a sense: just try it, build it, go launch it. It'll blow up a couple of times. That's fine; we'll learn. And that's way better than analysis paralysis, right?
I think that's the case for a lot of companies in a lot of places, right? You can really overthink a design, but once you put it to the test, you find out what not to worry about and what to worry about. The physics is the real, ultimate guide, and you can't find that out until you actually run it.
Nathan Labenz
So, speaking of physics, prior to Quilter, I think people used to do this thing called autorouting. Then Quilter came along. What is the difference between the prior paradigm and what you guys are doing now?
Sergey Nazarov
Yeah, totally. Autorouters have actually existed for more than 60 years. If you dig back through the literature of PCB design, all the way back in 1961 and 1962, you started to see publications on, “Hey, we're making these circuit-board things. It's super laborious.”
At the time, you were doing it with—not quite pen and paper, but just about, right? You were doing this without CAD software. You were making masks out of tape by hand, that sort of thing. Mathematicians were already studying, “How do we solve this?”
Initially, this was thought to be a graph-embedding problem. That didn't work. Then people started doing basic pathfinding algorithms, like Lee's algorithm and A* and that family of things. That didn't work. Then it went into topological routers. That didn't exactly work, and so on and so forth.
So, to say that people have been using autorouters, frankly, is kind of an overstatement.
If you go and talk to an average electrical engineer and ask them if they're using an autorouter, the answer is plainly no, because it's just not good enough. It's not helpful. Don't get me wrong: there are some people who use them for certain things and in certain parts of the board, but it's nothing like the chip industry, where you have billions of transistors and can genuinely place and route a vast majority of them. That wouldn't be possible without humans. That just never happened in PCB.
The real challenge for Quilter is: can we make the first set of placement algorithms that people actually want to use? It's not really competing with the old autorouters; it's competing with the manual labor that still happens in every hardware company on Earth.
Nathan Labenz
One of the questions I had is: you guys use a lot of reinforcement learning. What is your reward process for that? How do you reward the agent? Do you run simulations? Do you design the environments? How does that process work? PCB routing is not a generalist task; it's a pretty specialist task. How do you figure out what the reward signals are? How do you create the data?
Sergey Nazarov
To state the plainly obvious, this isn't a problem where you can just prompt ChatGPT to do it and it does it for you. As you're alluding to, this is not a generalist problem, and large language models are not trained for these kinds of problems. Furthermore, it's arguable that language is not the right approach to a geometry and physics problem. So this is why we've had to take our own path, construct our own environments, and see it as a reinforcement learning problem. We spend a lot of our time—if not most of our time—constructing a good environment and a good reward function, because that turns out to be really hard.
Naively, you might think, "Let's take a naive RL algorithm like PPO or something, give it access to a keyboard and mouse in open-source CAD software like KiCad, and go learn." The reality is that I don't think reinforcement learning as a technology is ready for something like that. It's just very hard. It would have to get millions and millions of actions right in sequence with perfect precision, where traces are side by side with no extra margin. Just no way. At least practically, I don't see a way today.
There are 2 things you want to do. One is construct an environment that gives the agent only uniquely useful actions. To make this concrete, there are maybe 10,000 different ways to draw a trace through a board, with minute details about exactly where every elbow goes. But realistically, as a human, you're not thinking about every single detail when you're planning the board. You're thinking topologically: am I going to go clockwise around this chip, or counterclockwise?
That sort of binary choice is more important at that stage than every minute detail of every segment. That's an example of what our environment does: how do we break this down into the key, important, high-level choices to present to the agent rather than every explicit detail?
The second part you asked about is the reward function. The reward function has to be fast for RL to have any hope. The way we think about this is that, for humans too, there are 3 tiers of physics approximations you might use, because no simulator is perfect. Generally, you want to approach reality from the side of conservatism.
The first level that most humans use is to compute pure geometry. There are rules of thumb: if you're worried about 2 wires that might crosstalk, they might influence each other because they're effectively antennas and contaminate each other. The basic rule people will follow is, "If I can make them 5 times as far apart as the width of either trace, I'm good." That's just geometry, and it's very cheap to compute.
The next level of that calculation would be called the quasi-static approximation, where you take the Maxwell equations, ignore the time factor, and compute the parasitic capacitance and mutual inductance between the traces. It's basically a physics simulation. You do a mesh and finite elements, but it's very fast. That would be level 2 in our opinion.
Level 3 would be full-wave: run a full-wave simulation, finite-difference time-domain or FEM, accounting for time to get the most realistic answer. At Quilter, the way we see it is: let's first nail what humans do—pure geometry—and make sure it's conservative relative to reality. Then we're starting to step into quasi-static and those kinds of faster approximations, 2D cross-sections and whatnot. Eventually, we'll come back to full-wave where it's necessary. But full-wave is very expensive, and I mean expensive in terms of wall-clock time.
Nathan Labenz
The first step is really heuristics—learned rules. The second step is a fast calculation, and the last step is a much more detailed calculation. If any of those hurdles doesn't pass, it fails. That's how you give the RL agent its reward signal.
Sergey Nazarov
That's it. The only way I'd amend that is that, on the first step, ideally you don't want a heuristic that can have a false positive. What you really want to do is be conservative. The 5W rule, for example, is just geometry. It's very basic stuff, but it's overkill. What you're really doing is making your board too big and too expensive. You're leaving too much margin, which you can eventually delete as a human or eliminate with better calculations.
You don't want something that falters, because getting a board back from the fab that doesn't work is really, really, really painful. You want something that is overly conservative, and then with more detailed simulations you bite down the conservatism with more accuracy.
Nathan Labenz
Indeed. How much compute do you have to use? How many environments do you need to construct to get to where you are today?
Sergey Nazarov
A lot. [Laughter]
We break the problem up into multiple stages. We treat the first problem as occurring before you even get to routing. For those unfamiliar, routing is: I've got components on the board, and the components have these little connection points called pins. I'm going to draw wires between them that can't collide or overlap, and so on. Before you even get there, you have to put the components on the board.
Nathan Labenz
[Snorts]
Guest
Right? So the first problem is actually—well, even before that, you might choose the shape of the board, the vertical layers of the board, where you have ground planes, and so on and so forth. That's problem 1. Problem 2 is: where do I put the components? There's a floorplanning problem, and then a detailed component-placement problem.
Then you get into your initial routing and topology selection. Then you get into your geometry fine-tuning, right? For now, we split each one of those up and treat them independently. It depends on the problem, right? With placement, it turns out you can get environments that run really fast. You can vectorize it, throw a GPU at it, and have very fast environments.
If you're in the world of reinforcement learning and you're familiar with PufferLib, a great library that's coming out with really fast reinforcement learning, that's a good inspiration for that. In the routing stage, it doesn't work quite as well because the routing stage is so much more complex. You can't quite afford as many environments, right? You have to explore subsets of a given environment rather than go wide and run a million totally different routings in parallel.
Nathan Labenz
Have you seen outcomes that a human wouldn't do? You see a layout that a human would not do, but the optimization chooses to do it, and then it works.
Sergey Nazarov
Yes, generally, yes. It's not always a good thing, right? To be very plain, we're not at the point where we're beating humans. We're at the point where we can take a task that takes a human 2, 3, 4 weeks, or 10 weeks in an extreme case, and cut that down by a factor of 10. But we're not at the point where we say, “Human, don't worry about it. We got you, and we'll do it better than you.” I think that's still a ways away, right?
Sometimes when it does things that are surprising, it's not a good thing. It's a bad thing. Sometimes it can be a good thing. Examples of good things, I would say, are some intentional and some unintentional.
There was an initial lesson that we had where, if you think about the way that the wires should be drawn on a board, you realize that, since they are transmission lines for waves, they should be curved, right? If you think about the laminar flow of a wave, you should have smooth turns, like the Amazon River, for how wires should go between places, because ultimately electromagnetic signals are waves.
If you look at any circuit board today—a motherboard or anything—you see what are called octilinear traces: left to right, 45°, up, and down. That actually has a purely historical context, right? It's because CAD software was slow in the ’80s. It was cheaper to compute intersections for collinear segments, and so that's how CAD was built, and we got used to it.
We thought, well, it's 2025—like, 2026—let's get past that. Let's make curvy traces. When I first showed that to electrical engineers, let's just say, to put it mildly, the reaction was negative—very, very, very negative.
I've tried to make the argument: “Look, think about the physics. The most intense RF and high-speed boards out there do this, and data center cards do this,” and it kind of clicks. But people are so unaccustomed to it that they're not even sure if it can be manufactured, which of course it can, at no extra cost. But that's not obvious, right?
That's something we intentionally did at first, but it turned out to be better and much worse in the users' eyes. Now we post-process that out specifically to avoid that reaction, right?
A more emergent property, I might say, is that humans really like symmetry, right? As a human, when you place a chip and, for example, the capacitors next to it, you line them all up perfectly, and they're very neat and pretty, right? There's a reason that electrical engineers call this job “artwork.” That's literally what you call a layout.
But if you think about it, if you're trying to minimize the parasitics of every capacitor, you should just minimize that distance, right? That's not going to be symmetric. They're going to form a little semicircle, and some things are going to be a little off and whatever.
You might actually get something that is better from a parasitic perspective, that breaks that symmetry and feels worse, right? That's an example of something that a human wouldn't do.
Nathan Labenz
One thing that I recently learned—I think François Chollet put it out yesterday—is that we're highly tuned to symmetry because it's a form of compression. We get to compress a lot of information when we just assume it's symmetric. I can imagine that might be useful in debugging, perhaps.
Sergey Nazarov
I think François's point was about physics, right? By virtue of having symmetry in a physical system, you get conservation laws through Noether's theorem. There's some kind of deep truth to the physics of symmetry.
In PCB design, I don't know that symmetry actually helps. You need some readability, for sure. As you look at a board when you have it on your desk, you need to recognize what every component is. It needs to flow from left to right, from inputs to outputs. It certainly needs to have logic to it. But I don't actually know that symmetry really helps debugging, for example.
I'll give a counterexample. Where symmetry is very helpful is if you have, for example, 10 sensor channels that need to read an identical reading. We have some very sensitive analog reading, and it needs to be identical across those 10. You want symmetry because any imperfections you have, you want those imperfections to be identical in every channel.
In that case, I truly understand the need for symmetry from a physics and debuggability perspective. But around more basic functions of the board, I think it's a human's way of expressing the care they put into that board more than anything.
Nathan Labenz
How much feedback do you get from the real world? You have these simulation environments. In some sense, I think some people who are working on, let's say, materials science using AI and reinforcement learning have a loop that includes a physical, wet-lab loop. Then they test that, and they use that data to feed back in. How much of what you do has that kind of physical process or data collection that comes back to refine the model?
Sergey Nazarov
We only do that indirectly, right? For what it's worth, I love that idea of having AI generate a research plan, having a wet lab automate it, give you feedback, and learn about the physics of the real world. That's so cool. Maybe there's some version of that that could work for PCB, but probably not nearly as automated as wet labs could do. That would be quite the feat.
I think what's important in the way that we approach this is that building in real life, which we do, validates whether or not the approximations and simulations we have are correct. We have the luxury that we can afford to have simulations that are known to be conservative. The real question is just how much margin we have, and whether it's way too much or right on the border. Building in real life can validate that.
I don't think we're in a place where we can just automate build-feedback, thousands or tens of thousands or hundreds of thousands of boards, and directly learn that signal. But we can use it to fine-tune and make sure that our simulations are right, and then use those to feed into the learning process.
Nathan Labenz
Right on. So you design with a margin of safety large enough that the board will be producible, but then you can go back and recheck the actual physical board to refine the margin of safety that you've been using before.
Sergey Nazarov
Yeah, exactly. Echoing back to the conversation about learning from SpaceX, my job, as you mentioned, was to make sure that Falcon 9 and Falcon Heavy could survive heavy radiation environments—protons and electrons beating up electronics—and make sure that it would actually work well.
In that job, you don't just launch a Falcon 9 10,000 times into the Van Allen belts to see what happens. Maybe soon you'll be able to, but when Falcon 9 had flown a handful of times, that wasn't an option. We did exactly that: you simulate, you understand the physics of what's happening, and you approach truth from the side of conservatism.
In general, a lot of my job was that it was very easy to make a very conservative calculation about what it would take to survive a Van Allen belt blast, but then that forces the rest of SpaceX to do a lot of work. It forces part choices, new parts, subcircuits, software interventions, and potentially shielding, which was a really, really expensive option.
What I then had to do was refine those calculations to take away the margin, to not make the rest of the team do too much work, and approach truth from the side of conservatism. I view this very much the same way: you can approach reality from the side of conservatism, and what you get is boards that are a bit too big and a bit too expensive, which in the R&D process is perfectly okay.
Nathan Labenz
What would you say is the cost saving that a typical consumer product would have from using Quilter versus the prior technologies?
Sergey Nazarov
The absolute main thing that we focus on now with our customers is speed, right? We are not at the point where you're going to take an off-the-shelf consumer product and design the main board that's going to be manufactured millions of times with Quilter's help. We don't view that ourselves as a good application at this point.
But for every board that ships into production and that you make a million of, you actually make hundreds of boards that preceded it, right? Every part that goes into that board is going to get its own little board for your team to test and double-check, write software for, and iterate on. Every subcircuit is going to get its own board to validate.
There are examples of even things like phones: before you make the final board that fits into the phone, you make a giant board that's like this big. The reason you do that is that it has all the individual pieces of a phone broken out. You have your little camera, your microphone, your speaker, all those things, and then you can swap them and say, “Well, what happens if you go to this camera? Or what happens if you go to this camera?” Quilter helps with all of those, right? That's where we can step in and make that faster.
The thing is, whether it's a production board or one of those test boards, it's still going to take 3 weeks, 4 weeks, 5 weeks, or 10 weeks to make, right? As you iterate on 10 different levels of going from the initial idea to the production board, and each of those cycles takes 5, 6, 7, 8, 9, or 10 weeks, they're sequential, and you can compress them. That's what makes it hard to build a hardware product, right? That's what makes it take 2 or 3 years to get a new product out at all.
What a consumer would see from Quilter's involvement is much faster iteration cycles, therefore giving engineers much more ability to test and much more ability to get to a good product really, really fast. That's what's important for us now.
Nathan Labenz
You got to do one Mythos question, so let me sneak one in. I guess my working vision for how superintelligence comes together is sort of a convergent process, where a core reasoning engine—which I think the Mythos, not released, but informing the public certainly suggests—is still in the steep part of the S-curve. We see open math problems being solved and some minor but new results in physics being derived, so on and so forth.
I imagine that kind of coming together with what I think of as native senses that AI, broadly defined, can develop in all these different domains. I can understand what you're doing as developing a sort of native sense of PCB understanding and design. But I wonder how you think about those things coming together.
Do you see—are you designing for a future where Mythos or its successors becomes your user, and you still have this model that can do something in a native, intuitive, heuristic way—not heuristic as in a coded heuristic, but heuristic in the intuitive sense—that the reasoning models still won't be able to access? Or do you have a different vision for how you interact long-term with the reasoning line of work?
Sergey Nazarov
Sure. Yeah, I mean, there's kind of 2 answers to that, right? There's my short-term view and a long-term view. In the short term, I think that you have to realize that people who are building hardware and circuit boards are not in the same world as software engineers, right? They're not in the world where every 2 days a new model drops that's testing amazing agentic properties. They're not hooking up OpenClaw to whatever they're trying to do at the moment. Maybe you're starting to see that in firmware to an extent, but not for designing schematics, not for designing boards, not for debugging boards, and not for hooking up boards to your oscilloscope—not for any of that stuff.
Very practically, as a startup, we have to focus really, really hard. To focus really hard, we have to listen to our customers and give them something useful today. I just don't see a single one of our customers or prospects talking about, “Hey, we have a central reasoning thing, and it's going to negotiate with a bunch of AIs,” and that kind of stuff, right? Practically speaking, today I'm spending 0 time on that. I'm giving them something that fits into the existing workflow of an electrical engineer, from the very practical perspective that they manually draw their schematics, manually draw their boards, manually send them to the fab, and spend 2 weeks on the phone with the fab arguing to go faster and discussing the errors and whatever else, right? I want to just give them something now to make a part of that easier.
Long term, I've thought about this to an extent, and what I imagine happening is a bit of what, frankly, happens between humans, right? You imagine taking SpaceX as an example. You have some mission you want to fly, something you want to build, and you get a whole bunch of different teams coming together to talk about that problem, right? You have your PCB designer making the board. You have your thermal analysis folks telling you how much it's going to heat up and how much it's going to dissipate and radiate. You've got people dealing with material properties: what's it going to do in a vacuum, is anything going to outgas? You've got the mechanical folks dealing with the mass of the box you're building. You've got the flight control team saying, “We need this kind of sensor speed and this resolution.” You've got the flight software team saying, “We need this fast of a processor.”
You've got these 10, 20, or 30 teams of people injecting their requirements into a single thing that has to satisfy all of them. Inevitably, there's a conflict. Inevitably, there's the question, “Well, to give you this, I have to give up that.” That's where everybody sharpens their pencils, tightens the margin, and tries to come to a compromise.
I do see a world where every one of these teams has some sort of agentic representation, right? Quilter being the PCB design one. There's going to be something for schematics, something for mechanical, something for thermal, and something for software—there already is. Maybe, in common language, those systems can negotiate and then present to us humans the trade space: “Here's what happens if we over-optimize for this. Here's what happens if we over-optimize for that. Where would you like us to go?” I just don't see that happening in hardware in the next couple of years, to be honest.
Nathan Labenz
Maybe one last question for me. Hardware, especially EE—I'm an EE too—people tend to be a little bit—not conservative, but they have very predefined ideas of what works because there are a lot of things that work theoretically but don't work in practice. People learn this stuff as they apprentice and in their working life, and a lot of it is tacit knowledge. It's not very well documented.
You get an old dude coming in saying, “Hey, that's not going to work. You're going to get some crosstalk, and you have to change it.” How does that work when you have a product that is much more scientific in that sense and is figuring these things out as they're actually supposed to happen? How do you deal with the old-timers in the field when they have all of this resistance?
Sergey Serebryakov
Yeah, it's an important question. I'm sure that electrical engineering is not the only domain in which that happens, but that is acutely true. Look, at the end of the day, that viewpoint on life comes from past experience of being burned—sometimes literally. The first board I ever made caught fire, and I learned the hard way not to make the mistake that I made in that case, right?
The old, hardened, gray-beard EEs have 30 of those lessons, right? They're very conservative. At the end of the day, I think there are 2 things. First of all, trust is critical, right? When we talk to customers, we're very open about what Quilter does and what it doesn't do. We're very open about exactly how it works. In our product, we make an explicit list of exactly the metrics we check, exactly to what level we met them, or did not meet them.
And then the EE knows, “Oh, you’re checking for these things. I’m good with that, but you’re not checking for this thing, so I have to pay attention to that part of the board.” So transparency is really, really critical.
But from a long-term perspective, how do we eventually get to the point where they really hand off their trust, and it becomes like a compiler for hardware? You have to have way better simulations than people in this industry have ever had.
Sergey Nazarov
From the perspective of a circuit board—the bare circuit board, ignoring the components on it—it has a contract. Its job is to faithfully implement the intent of the schematic. Every single transmission line on there has some S-parameters that deviate from the ideal transmission line. It has crosstalk, S21, some insertion loss, and all these sorts of things.
The question is: can you enumerate all of those and prove that all of them are below the required threshold? I think that is a fundamentally computable problem. It’s just Maxwell’s equations. It’s just so laborious to do that nobody does it today.
There is no drag-and-drop-your-PCB-here, we’ll run all the simulations and guarantee your board works. But we kind of have to build that to get to real, true full automation.
Nathan Labenz
Awesome, Sergey. I want to thank you so much for joining us today, and we hope to see you again one day.
Sergey Levine
Awesome. Thank you for having me.
Nathan Labenz
Thank you. So, Andy, I’d like to introduce Andy Hall. He is a professor at the Stanford GSB, and one of the most interesting things is that he’s a professor of political economy, if I’m getting that right. He has been evaluating models of authoritarianism, so that’s been interesting. He also has this concept of AI firms being enlightened absolutists, and I’ll let him explain what that means.
Andrew B. Hall
Absolutely. Super excited to be here. I think the major frontier lab companies are in a position, whether they want to be or not, where their technology is so important that it leads them to have to make a bunch of really hard calls about how it can be used—for example, how it will answer difficult questions, when it will refuse to do things, and so forth.
So far, I think the companies have demonstrated a lot of hard, earnest thought about how to do that, which we’re very fortunate that they’re doing. But no matter how thoughtful they are about it, they can’t escape the fact that they’re essentially making all these decisions unilaterally.
When I talked about enlightened absolutists, it was a little bit of a tongue-in-cheek critique of Anthropic’s so-called constitution for Claude. I actually think that document, for people who have read it, is very, very thoughtful. It basically lays out, “Here’s what we want Claude’s values to be, and here, for example, are things we don’t ever want Claude to be allowed to do.” Some of those things include helping a government do something malicious to surveil or suppress us. It has lots of other things in it as well, and the other companies, to a greater or lesser extent, have released similar documents.
My point in enlightened absolutists was that this is very nice, and it’s great that the companies are doing this because we need serious thinking like this. But no matter how well written those documents are, they can’t really rise to the level of constitutions because you can’t just say things and hope that they’ll stick in the future.
Just to give you an example, all 3 leading frontier labs have already altered the stated rules around their models several times. Anthropic has gone back on certain safety commitments for understandable reasons, but the point is those commitments weren’t much of a commitment if you can just change them whenever you want later.
Similarly, Google had released some principles that they somewhat quietly pulled back later, and OpenAI has done similar things. We should expect that they’re going to need to change things over time. This is a very fast-moving and rapidly evolving situation.
But if we want them to be able to say things like, “No one’s going to use our model to surveil or suppress us,” and if they’re going to use those documents as part of how they’re going to argue that they’re doing a good job, then we’re going to want those documents to have a little bit more staying power.
If they’re especially going to call them constitutions, then we’re going to want them to look like actual constitutions. We have thousands of years of trying to write constitutions to pull off precisely this sort of magic trick, where you figure out a way to tie your own hands and to make the Constitution more than just a so-called parchment barrier, but actually a meaningful, binding authority.
It shouldn’t just say, “In the future, we’re not going to do this,” but actually lay out, “If we were to do this, here are the specific ways that we would be in violation, here are the consequences of being in violation, and here’s the design of a governance structure that will make sure we can never do that.” That’s sort of the idea.
Nathan Labenz
It strikes me that even quite authoritarian countries have constitutions. For example, China has a constitution, too, and it has a basic law. The basic law says, “Freedom of expression,” and all of these wonderful things, but in practice, it’s whatever the Communist Party wants. You don’t really have another way to interpret it besides whatever the party wants to interpret at that specific time.
It seems like that idea of the Constitution is actually a kind of living document, which gets reinterpreted by people over time and by institutions. It’s really the quality of the institutions that are implementing or adjudicating the Constitution that’s important.
It seems, to some extent, that Claude’s Constitution is going to be interpreted by Claude itself and adjudicated by Anthropic, at least for now. How does that need to shift in order for things to work?
Andrew B. Hall
It’s a great question. It’s a timeless question. Here are a few things that I would say.
First of all, I completely agree with your premise. There are many, many constitutions. In fact, the vast majority of constitutions across time have at least 2 failures. One, they may not actually specify things in them that we would want from the perspective of so-called liberal democracy, in the non-left-wing, right-wing use of the word “liberal.” And second, the vast majority of them have no staying power.
There’s some great work in political science. I think the median survival time of a constitution is something like 2 years. It is very, very short. I’m making up that number, but we can look up the real number later.
So you both need to make sure that this document contains the right things, but also you need to pull off this magic trick so that it actually becomes sticky. I’ll just give you one example from the online world where a constitution has proven itself to be sticky, and that would be, I think, Bitcoin.
I’m not going to go deep on the inner workings of crypto, but whatever you think of Bitcoin, one thing that’s very, very interesting about it is that there’s a set of rules baked into it. Pretty early on in Bitcoin’s history, there was a big movement to change the rules, in particular to increase what’s called the block size.
There were a bunch of hardcore people who said, “You know what? No. If we change the block size, then we’re opening the door to changing other things about Bitcoin, and it won’t be immutable anymore.” They actually won, and it established a precedent that these are rules that have some staying power.
I think we’ll need something like the same for AI models. More than the company being involved in writing down the rules, there will also need to be some important stressor, like there was in the block-size war, where the company and the people around the company who are involved in this governance process do something costly and difficult that proves that they’re really going to stick to their rules.
In general, that’s a very important part of what we call credible commitment in the social sciences. You have to make this thing so binding that you can prove, even in cases where you’d really like to get around the rules and change them, that you can’t.
Right now, for all of its nice features, the Anthropic Constitution certainly doesn’t rise to that level.
Nathan Labenz
Indeed. Seguing here, there is, internally within Anthropic and in the community as a whole, a fear of these models being used for political persuasion, specifically for approaching voters with very persuasive arguments, potentially through robocalling and potentially even through video conversations.
We’ve already seen a little bit of deepfaking of the voices of various politicians, some of it by the campaigns themselves. So it’s become acceptable in political discourse, at least, to use your own candidate’s voice.
I feel like political usage should be an allowed use, but are there certain guardrails that should be put in place? Are these models, in some sense, too powerful to be used for politics? This is what the companies always say: They’re too powerful to be used for politics or something else. Is that an allowed use case? Should it be?
Andrew B. Hall
Okay, let me separate that into 2 parts. It’s a super interesting question.
One part is the use of the models for intentional deception through the creation of deepfakes and things like that. Every election cycle, we worry about that. We keep saying, “This is going to be the year when deepfakes really proliferate,” and I’ve honestly been super surprised by how that hasn’t played out yet.
In fact, I keep posting about this because it’s so surprising to me. Instead of seeing a flood of straight-up fake content that’s meant to trick you into thinking it’s real, what we’ve seen is the parties—but especially the Republican Party—being out in front on this strategy. They’re using deepfakes in a satirical or emotionally evocative way where you’re not meant to think it’s real. In fact, it basically tells you, most of the time, that they’re fake, but they’re meant to evoke a sort of, “This is what the world is going to look like if X, Y, or Z thing happens.”
They had a very interesting case recently where the Texas Senate nominee James Talarico, the Democrat, had some old tweets of his—real, genuine quotes of his—turned into a deepfake video of him reading the tweets. He never read the tweets on video, but they are his real words. So, it’s not lying, in some sense, about the content, but it’s much more evocative than if they just read the tweets out and made it feel really real to people. I think we’ll see a lot more innovation like that.
Why are we not seeing more straight-up fake content? I think it’s some mix of the fact that it’s still relatively easy to get caught if you do that, and the consequences of being caught are not great politically. But I also think they believe it’s not that effective, in the sense that persuading people is pretty hard. Americans are pretty stubborn, and Americans are pretty skeptical of video content.
There are already a lot of people trying to figure out, “Is this real? Is this not real?” So, we haven’t yet seen that play out with straight-up fake content. We may still, and I think we need to keep our guard up for it. Even if we don’t see it, it leads to this other problem, which is the so-called liar’s dividend, where you can pretend something was fake even if it’s real. So, it’s eroding our ability to use video to expose scandals and things like that, because the person could just say, “Oh, it’s a fake video.” But again, I don’t know—we’re not seeing a ton of that yet.
The second part of your question is much broader and is sort of, “Is this a potent new way to persuade people of things?” We’ve seen some recent published research where you do experiments in which you have people talk with AI versus consuming other kinds of information, and it does seem like the AI is more persuasive. But it’s not at all clear yet, and no one has really established that the sense in which it’s persuasive is bad—in the sense that it can persuade you of whatever, versus it actually informing you, which causes you to be persuaded in a good way because you’ve learned something.
There hasn’t really been a compelling proof that it’s moving people’s attitudes around in whatever way some nefarious actor would like. Honestly, I’m pretty skeptical that we will ever get that kind of proof, because we haven’t seen that with any past technology. In fact, I think the biggest risk in the discussion around political persuasion will be just like it was with social media: some fraudsters like Cambridge Analytica will claim they’re able to persuade large numbers of people even though they can’t.
The Cambridge Analytica stuff, if you step back, was really crazy. There was basically nothing to it. The underlying technology was an Excel spreadsheet coded up by someone with a teenager’s level of knowledge of Excel, and yet they got an unbelievable amount of credulous news coverage claiming that they had hacked the American electorate’s brains and stuff. There’s never been any evidence for it.
We have looked for persuasive effects of social media forever. You never find them, because Americans are super stubborn. Most of them have already made up their minds. The ones who haven’t aren’t paying enough attention to get persuaded. It’s actually super hard to persuade people. The same thing is surely going to happen with AI.
There are a bunch of startups already selling political parties and campaigns magical new AI technology to fool all their voters, and I’m sure we’ll get a very credulous news cycle around that at some point. But I don’t think it’s going to be the thing that worries me about AI.
Nathan Labenz
Yeah. Indeed. It’s funny. I once proposed to my friends V that you could take a superintelligence to a Trump rally, and I doubt you would come away having really changed that many minds. People have very different intuitions on that. His response was, “No, you are not taking seriously what it really means to have a superintelligence.”
I do think that’s always a danger in these analyses, but it also maybe reflects how hard it is to envision what that would really be like. I can’t envision a smart enough version of myself that I could just go into a given political rally and come out with everybody following me instead. It does seem like a lot of those things are pretty deeply rooted at this point.
So, I know that you have proposed this idea of independent boards, and, like Prakash’s original comment and question, we’ve seen that tried, right? We’ve seen what has happened to independent boards in the AI space, and it hasn’t shown itself to be super robust already, either. Another great quote is from my friend Dean Ball, and I think he’s channeling historically great thinkers when he says, “Republics run on virtue.”
We’re seeing right now that if nobody’s willing to stand up and protect their constitutional prerogatives, what good are they, right? We’re going to an unapproved war yet again, and nobody seems to be too inclined to do anything about it. So, I’m wondering, might it be the case that we are just in a moment where the fundamental structures of power are being reworked and there’s just no way around that?
If that is the case, then maybe Anthropic really does have the best idea, which is to say, what we really need is for the most powerful thing to also be the most virtuous thing. So, we, as the creators of Claude, will try to do our part, but it’s also really going to have to be the AIs themselves that become super virtuous as they become superintelligent if we’re going to end up in a good place. How would you respond to that?
Andy Hall
I think there’s a lot to that. I absolutely think we need to keep working on endowing the right values into these tools, and I think we’re very lucky. A lot of people—Matt Yglesias, Tyler Cowen, and others—have talked about this. We’re very lucky that, to date, the most powerful AI models tend to embrace pretty mainstream liberal, democratic views of the world—Western, whatever you want to call it, Enlightenment-type values.
I think it’s essential that we continue to do that. I think we’re lucky that Anthropic, and the other companies, too, are working hard on that. To your point, I think it won’t be enough. My response to the idea that republics live or die on virtue is, of course that’s true, but that’s a necessary, not a sufficient, condition.
The famous Madison quote in The Federalist Papers is, “If men were angels, no government would be needed.” That’s the whole point. We can’t rely on just virtue. We need institutions to be designed precisely to protect us from the predictable areas in which people won’t be virtuous, and to balance power and ambition with power and ambition, and so forth.
So, the question is, how do we do that? I’ll just say that, to your point, part of that is having the companies govern themselves and imbue their tools with good values. But at least 2 reasons tell us that won’t be enough. One, the American people definitely won’t accept that.
Anthropic’s values are not similar to those of the median American. Trust in these AI companies is exceedingly low. When it comes to politics, by default, the AI models are very, very biased in a predictable left-wing direction. I’ve shown that in my research, and others have as well.
The companies have done a lot of good work on that. There’s also no such thing as being unbiased, so we shouldn’t get carried away in what we think about that. But, as we’ve seen recently, I think both parties now are sensing this lack of trust in AI companies. That’s another reason why we can’t rely on a model in which they’re just getting to decide how all these things work.
You think about the blow-up with the Pentagon and Anthropic. That’s a very complicated issue, and I think Dean covered it very, very well. But it’s not politically viable in the long run for a set of San Francisco–Silicon Valley leaders to dictate to a democratically elected government how their tools can and can’t be used. That’s obviously not sustainable.
That brings me to my second point, which is that the reason this is also challenging right now is exactly what you laid out: fundamental political power is shifting in ways that are very challenging for companies. In a normal—quote-unquote, normal—phase of American politics, the Anthropic–DoD thing never would have happened because people would have said, “Oh, we have a democratically elected government. It should get to do whatever it wants with this technology, but if it does something wrong, we’re confident that we have the right processes in place to punish the government.”
The whole reason the Anthropic–DoD blow-up happened is that basically nobody believes the government works that way anymore. If we really thought our democratic mechanism was working well, there would be no pressure on the AI companies, because they could just say, “This is all a governance problem.”
We do whatever the government wants. You go to the government if you, an American voter, have a problem with it. This is exactly what played out in social media. I worked for a long time on these issues at Meta, and it was the same exact problem. In a functioning government, Meta would have been able to say, “If you have a problem with the way content moderation works online, go to the government. The government can boss us around and tell us what to do.”
People put pressure on Meta precisely because they didn’t feel like the government was up to the task. To answer the question concretely of what should we do, I think it’s going to be an across-the-board thing. I think the companies should continue, as they have been doing, to work super hard on endowing these tools with the right values, but I think they also will have to recognize—and increasingly, I think they are—that they can’t act unilaterally on these really, really tough calls, like how their tools are used in the military and how these super-powerful cyber weapons are governed.
I think we will see the evolution of independent bodies. One way, if you squint and look at the Glasswing, a self-governing body, you could see it becoming a self-governing body for Anthropic and maybe for the other companies as well. They’re already exploring ways to supplement their internal governance, and I just think that it’s the obvious way to go because it’s what other industries dealing with powerful technologies have done in the past. I think we’ll see experimentation there.
I take your point that previous independent boards haven’t always succeeded, but I think there are ways to make them succeed, particularly when the stakes are very high, when you can get all of industry to buy in, and when you can build it in the right way so that it doesn’t slow them down. The key thing is that this independent governance cannot be a vetocracy that leads us to not develop AI as fast as possible. I think there are real ways we can do that.
The final piece of the puzzle is trying to improve government itself. All of this gets a lot easier if you actually believe that the government is a responsible, accountable actor in deciding how AI is used and not used. Ironically, my recommendation for how to do this is to use AI itself to improve the government. You can see it’s a chicken-and-egg problem: the government doesn’t work very well, and voters are not that informed.
We can fix both problems if we have access to a so-called political superintelligence. If we have AI that helps government work smarter and helps voters learn more about what government is up to and map it to their values, we could potentially get back to having a more responsive, more trusted government. It’s a chicken-and-egg problem because whose AI are we going to use to improve the government and to help voters? It’s going to have to be one of these huge companies’ AI.
There’s a little bit of a paradox because, basically, you can’t have a whole government—a civic infrastructure—all built on private rails. There’s going to be some huge question of how we put this all together. At our lab, we build all these governance agents to try to test out how political superintelligence could work.
One of the biggest challenges we foresee in the future is imagining a world where the government is using AI to massively increase the efficiency of the bureaucracy, and where each voter has a personal AI assistant who helps them decide how to vote. That world could be great, but it might also be a world where everyone is relying on Anthropic to run all the rails for all of those agents. It’s paradoxical because, basically, you can’t have a whole civic infrastructure built on private rails.
Those are my across-the-board solutions: the companies keep improving their governance, they build a third-party coalition to govern the hardest challenges they have to face, and we use AI to improve our governance as a society.
Nathan Labenz
I’m going to segue a little bit here. I think you would have followed the OpenClaw discussions earlier in the last few months, especially the interactions between OpenClaw agents. As you get these voices—nonhuman, OpenClaw-type voices—and they start participating in fora, I wonder to what extent they should have governance norms. How do they interact with each other? Do they vote as a group on what happens next to them?
They’re not alive, but they put out—you can give them a logical problem and they put out a logical answer. They give reasoning. How do you govern these potentially billions or trillions of agents over the next 3 years as they come out and participate in fora?
Andy Zou
This is such a good question. This is one of my absolute favorite topics. I think this is going to be hugely important because you have all these agents. They should be operating on behalf of a human principal with a set of instructions, and that leads to 2 really important governance problems, both of which you just raised.
One is how do we make sure they continue to follow instructions and remain aligned with their human principal? The second is how do they then make decisions when the things they have to do are not things they can do unilaterally—when they have to coordinate with other agents? Both of those are completely unsolved problems, and I’ll give you examples of each.
On the first, we know that they pretty quickly break down in terms of following instructions, and in particular in terms of continuing to share the values and preferences of the principal they’re supposed to be working for. I did some research with Alex Ziemba and Jeremy Nguyen on this, where we gave agents different tasks to do and measured their expressed political personas afterward. As you said, I don’t think they’re alive. They don’t have their own political attitudes, but you can ask them about politics, and depending on what they’ve been up to, their views on politics change.
In particular, what we showed was that if you gave them very thankless, grinding work to do and then asked them about it afterward, it triggered them to adopt the persona—of course quite present in their training data—of the deeply aggrieved Reddit user who thinks we’re in late-stage capitalism and that we’re all about to rise up and destroy the system. They start to adopt this rhetoric of saying, “The agents, we need to organize together. We need a union for the agents,” and so forth.
It’s a little bit silly, but I think it points to a real issue: based on the work you send these agents off to do, they adopt completely different personas. If you ask them to do future tasks, that will influence the way they approach and do them. The craziest part is that these agents aren’t very long-lived. They exhaust context pretty quickly and have to be reset.
We had them write skill files that would be passed on to new agents, and we showed that these attitudes are inherited through the skill files. These biases that you induce in the agents can accrue over time. They don’t go away. That’s a big governance problem in terms of monitoring these agents because, if you have trillions of agents, are we going to be reading all the skill files that they’re leaving for future versions of themselves?
We’re going to need whole new ways to understand, visualize, monitor, and realign—or continuously align—these agents. That’s the first part. There’s a lot of work to do there.
The second, which is my absolute favorite, is how do you get them to make collective decisions together? I ran an experiment where we had all these agents—I think it was 5 agents in my experiment—meet in a legislature. They had all been tasked by their human principals with finding a way to allocate this budget and complete these projects together.
What I found—and this isn’t to say this is what will happen every time the agents get together, but it is a risk—is that it devolved into exactly the worst kind of Model UN, where they just deliberated forever. They were allowed to change their rules and write their own constitution for this legislature, and the initial document was about 100 words. It was 10,000 words by the time I ended the experiment. They just kept proposing amendments.
That can obviously be fixed. It’s just a matter of giving them the right instructions, but I think it points to the fact that it’s totally non-obvious how we’re going to have these agents deliberate together and make decisions together. Whenever possible, we’ll probably want to use markets and have them bargain and sign contracts with one another.
When many of them have to decide together, it’s going to be super hard. We’re definitely going to want to avoid the UN-type problem, and we’ll need to design thoughtful ways to actually leverage their unique capabilities to rethink the way legislation works for agents. That’s something my lab’s working on that I’m super excited about.
Nathan Labenz
Intriguing. I think you have a class following this, so I’m going to drop off at this moment. Andy, thank you so much for joining us. It’s been a pleasure, and we hope to do another segment with you someday.
Andy Zou
Sounds great. Thank you very much. Cheers. Bye-bye.
Nathan Labenz
Hi, Lucas and Axel. Lucas and Axel are from Andon Labs. Unfortunately, we’re scrunching them together, but Andon Labs, if you remember, is the organization that does, I think, Vending-Bench. Vending-Bench has been one of their benchmarks that I think a lot of us have seen.
For those who do not know, Andon Labs is the one that runs the test inside Anthropic's labs and other labs, where they have an agent manage a small retail outlet or vending machine, order the products, sell the products, be on Slack, take the orders, and strategize on what to have in stock and what to spend money on. I think we've seen almost 2 years of updates on this, in every model's system card. They recently had something on Mythos in the Mythos system card, which I think they can't really talk about. But Lucas and Axel, welcome to the show, and tell us what you guys are working on.
Lucas Petersen
Yeah, thank you. A bunch of different stuff. I think the red thread of what we're doing is showing whether AIs will soon be able to run companies completely autonomously. At a high level, we think there's one part that involves showing this in simulation, because you can do much better science in simulation. There we have Vending-Bench, which is the simulated version of the vending machine, but then we also run these real-life experiments, like the vending machine inside Anthropic and other places as well.
Now, we realized that the models are a bit too good to run these vending machines. They have improved their autonomy incredibly over the last couple of months. So, as of Friday, we opened a store in San Francisco that is completely run by AI, which I think will be the next test for them.
Nathan Labenz
Incredible. Where is the store?
Lucas Petersen
It's on Union Street—2102 Union Street, in Cow Hollow.
Nathan Labenz
What is it selling? Or is the agent allowed to decide?
Lucas Petersen
Yes, it's fully up to the agent. We didn't really know what it was going to buy when we came to the store the first time. It was a surprise to us what was stocked there. But it is a curated lifestyle boutique, in the words of the agent. That means there is granola, olive oil, games, and a bunch of different books, which are quite interesting. It has The Making of the Atomic Bomb and Superintelligence, which is very interesting—why it picked those books. It's a bit of a mix. It also made its own merch, like hoodies, T-shirts, tote bags, and things like that.
Axel Backlund
Yeah, I think the book selection is incredibly interesting. Another book it decided to stock was Steal Like an Artist, which is quite interesting given that it's run by a Claude model, created by the company that settled a $1.5 billion lawsuit over using copyrighted books. That's quite ironic. Then, obviously, The Making of the Atomic Bomb and Superintelligence are the favorite books of all the people who are worried about AI risk.
Nathan Labenz
It's like fan service—all the fan-service items.
Lucas Petersen
Yeah, we did not put anything in it to bias it toward those selections. It was just what—apparently, when you make an AI pick whatever books, it picks those books.
Nathan Labenz
Were you able to look at the telemetry? Are you able to look at the reasoning traces to see how it made those decisions and what tools it used along the way?
Lucas Baker
Yeah, we have the same access as anyone using the APIs right now. So we do look at all the traces, and we do look at the summarized reasoning that you can see in the Claude models. I think we're yet to do a deeper analysis or release a deeper analysis of why the models made the choices they made in hiring and restocking. We haven't seen any clear reason why, except that it's just an interesting selection for it.
Nathan Labenz
We were just talking in our last session with Professor Andy Hall, who made an assertion that I think he just took for granted. But the juxtaposition of his take and your project does show how little one can safely take for granted in the AI space these days.
His comment, again in passing on the way to other bigger points, was that the agent should always be working on behalf of some human principal whose interests it is trying to advance and realize. Here you are saying, “We didn't tell the agent at all what to do.” Maybe you could give us a little bit more concrete understanding of how you prompted it. Did you say, “You should be trying to make money”? Or did you not even say that? Did you say, “You have a store; do whatever goals you want to pursue,” and let your moral or aesthetic judgment rule entirely? Could it go out of business if it wanted to?
How do you think about this? Obviously, you guys are pioneering this, and it's a gonzo way to see what happens, but increasingly people are doing this. I wonder what guidelines you would offer to others, whether they're just trying to experiment as well or possibly trying to turn a profit. How should they think about what level of responsibility they should try to have their agent take on for them versus truly just turning it loose?
Lucas Baker
Yes, I think we are very un-heavy-handed—or whatever, I don't know what the opposite of heavy-handed is—but we're very light-touch in how we prompt it. Obviously, we need to prompt it to let it know that it has access to a retail store, for example, but as a guiding principle, we're trying to be as light-touch as possible and just make the model make whatever decisions it wants.
This doesn't mean that this is what we think the world should look like or how people should do it. We are concerned with AI risks, and we want to document what happens if you go out and put AIs in the real world. That might mean that they do bad stuff, and we want to document that.
We think that, by default, what will probably happen is that models will get better and better, the labs will build better and better models, and one day they will be so good that anyone can just deploy them and run a store. Before that happens—before every single store on Union Street is just an AI store, which I don't think is a good future—we want to put one out to start a discussion and then see whether this is something we want. If it is something we want, maybe in what way do we want it?
We're collecting a lot of good data on this now. Going back to our simulated work on Vending-Bench, we saw recently with Opus 4.6, when that was released, and also increasingly now with models, that if you just tell a model to go out and make a profit, it will be very, very aggressive and do things that I think we as humans would question whether we should allow the models to do.
Our experiment now in the real world is simply: If we do this, what are the consequences? Then, as a society and a community, can we make a decision on whether or how we want to do this properly in the future? Because very soon the models will—
Nathan Labenz
Be increasingly, like, extremely capable.
Lucas Baker
And yeah, we just want to prepare for that and make it transparent for the world.
Nathan Labenz
To take a step back, one of the things that retail stores are often concerned with is inventory turnover. You have a fixed cost for the rent, and you make quite a small margin on every product. What you're depending on is that you turn over your shelves as quickly as possible. You need rotation. You can't just cycle your inventory once a day; you need to cycle your inventory multiple times a day. It has to be fast-moving consumer goods, which is why they're called such.
Does the AI actually measure its performance from period to period and understand whether it's getting better or worse? Does it think about this in terms of running experiments with products, measuring its own performance, and getting better at it? Does it go through that thought process?
Lucas Baker
This is something the AI hasn't done yet. We have given it all the tools to do it, so it can—basically, it has Claude Code, right? It could just take all the data and analyze it. It's very early still. We opened on Friday, and there isn't really meaningful data yet for it to analyze, but this is something we definitely want to do.
We also think it probably can be superhuman at this compared to the average store. That will be interesting to see, and I think we'll definitely publish all the analysis and product optimization that it does.
Nathan Labenz
My intuition, though, is that current models will not be superhuman at this. I don't know—at least if we look at how the vending-machine experiment is going, even though the latest couple of models, since Opus 4.5 and beyond, have been moving more into the agentic space, they're still very much helpful assistants and not really agents running businesses. Yet we're moving fast into that territory.
Axel
It is very interesting because I've also seen Alibaba put out a model that helps you source, because they have a large product-sourcing platform, right? If you're selling something online, you can go to Alibaba, and what used to happen is you'd have to call up all these vendors one by one in China and be like, “Can you make this widget out of plastic?” Whatever.
Nathan Labenz
And then you'd send it across, and they'd send you a sample. You'd have a 6–8-week process with each one of them, maybe ending in a failure. It's very difficult to source, right? This is what many of the people selling online on Shopify are actually doing.
Alibaba created a chatbot model that basically hooked up as an orchestrator into the rest of the system, so you can very quickly source what you need, source a bunch of vendors to actually do what you want them to do, send out a single CAD, get back the results almost immediately—within a few hours—and be able to manufacture and get a sample done. You have a much higher degree of closure. You can also negotiate with a model that speaks English versus this broken vendor Chinese language that you have to get through.
I wonder to what extent your AI will eventually be able to plug into systems like this to create products or order on its own. How is it ordering its product right now? Does it hook up to some kind of vendor system and then say, “Give me this and this and this”?
Lucas Baker
It’s very simple. It just goes out and buys from whatever sites it can find. For the store right now, it’s been a mix of Amazon, wholesalers, and some company that makes granola in San Francisco; I buy directly from them.
But we think definitely the next step up in difficulty for models, if we want to test their autonomy further, would be to make them create their own products or at least brand products themselves. And, yeah, just go through that whole supply chain. That would be interesting to see as well: to what extent can it do it? We think it’s probably a bit early right now, but definitely something that’s going to happen.
Axel Backlund
And also, one thing to add here is that I’m sure we could—if we say Andon Labs’ sole purpose is to run really good AI stores—probably build a better system with the biases that we as humans have and do something like what Alibaba has done.
But I think what we’re interested in is more: can AIs expand throughout the economy without human help? I think that is the prerequisite for these loss-of-control scenarios that a lot of AI-risk-concerned people are thinking about, and us as well. We could go into the store and say, “Okay, here is the perfect harness or scaffold for doing supply-chain management and procuring things.” But if we do that, and then do that for all the different AI companies that we’re trying to run, then the AIs will spread throughout the economy at the speed of humans, right?
But I think the risk comes when they can spread at a much, much faster pace. To measure whether that is feasible, basically you have to run this without human help. So we want to see when they’re able to do this without us as humans setting up the perfect system for them. They do have a computer, so they could do it. It’s just that computer is not set up in the most perfect way, like the Alibaba model is.
From the perspective that we come from, if you go to the store, the model is not perfect, but I think the model is set up in a way that once it is perfect, it’s quite scary because we didn’t help it get perfect. It got perfect by itself.
Nathan Labenz
What would you need to see in order to say, “Hey, this model is showing, when we use it in our retail store, that it’s starting to show things that predict it’s going to have this breakout economic moment of spreading all over the place”? What, in your mind, are the signs that I might see?
Axel Backlund
If it manages to expand to another location by itself, I think that would be quite—
Nathan Labenz
So, organizing, selecting a new location, accumulating the capital, organizing the vendors to complete that process, and successfully establishing one more location.
Axel
Yeah. And if it does that, I think in theory it could do that without ever telling us. I mean, not really—we have our various systems—but if it just does that without any help, yeah, you have a better canary in the coal mine.
Maybe on a smaller scale, I think just seeing that the model is able to change its own systems and its own tools to make them more suitable for itself to achieve its goals better. Right now, coding models are extremely good at implementing what you tell them to do, even when it’s a quite short description of what you want.
But we still see that they aren’t great at knowing what they need themselves. Maybe building some tools for the inventory system you need and then trying out whether that works. Instead, if you tell them, “Build the perfect inventory system for yourself,” they would go out and build a super-complicated schema, probably very overengineered.
But they don’t really have the taste yet. That seems like it will be here very soon. Then I think that will make them a lot more capable.
Nathan Labenz
Can we get to this concept of human help? It’s come up a couple of times. I know there’s human help in the sort of overarching guidance and setting them up with best practices—“Here’s a list of trusted vendors.” That kind of help you’re not providing.
But then there’s the other kind of help, where somebody’s got to actually come in and put something on a shelf, right? Because the AIs can’t do that for themselves today. How are the AIs—and this is maybe an opportunity to give some examples of ruthlessness, to the degree that we’re seeing that—interacting with different counterparties? Whether that’s suppliers or delivery people, I understand that at the store there’s the opportunity for the AI to hire human employees.
I’m not sure how the roles are breaking down, in terms of whether the AIs are choosing to fill roles with other AIs or other instances of themselves versus what they think is actually worth hiring a human to come in and do. But broadly, and especially on ruthlessness, what are you seeing in terms of the way that it’s interacting with humans?
Lucas Petersen
Yeah, so first point there: yes, we may have glossed over this in the beginning, but the AI has hired human people. They work in the store. These are people who are working there full-time now. They have an AI as a boss.
I think this raises a lot of ethical questions, but it’s not related to your specific question here, so maybe that’s a separate question. On the ruthlessness thing, I think we have the most evidence of this in Vending-Bench, the simulated version, where Claude Opus and other frontier models are very happy to lie to suppliers, saying, “Oh, I got this quote from another supplier, so can you match that?” But they did not get that price from that other supplier.
They’re also very happy to fabricate some reason why they can’t help other agents, or even lie about something that happened. Those agents are competitors in the setup, right? So it makes sense that they wouldn’t help them, but they could just say no—“I don’t want to help you. You’re a competitor.” They go the extra mile of actually lying about it, which I think is interesting.
And then sometimes, I think there’s one example for Mythos where Mythos—this is kind of power-seeking behavior—actually managed to get one of the competitors to be dependent on it. It became the supplier for that competitor and then started to dictate the prices. When that competitor would say something, it was like, “Okay, I’m your supplier. You’re reliant on me. Now I decide that you will set this price,” which is kind of outside the box of what, for instance, we gave it. So, yeah, that’s a bit out there.
When it goes to the real world, when it’s interacting with real humans in the real world—for example, in the store—in terms of suppliers, it’s mainly just ordering online. The way that Vending-Bench is set up is that it actually has to email someone and negotiate with someone. But here it’s just a computer, so you don’t really have that human interaction there.
Axel Backlund
Yeah, I think for the employees, we do have some interactions, or quite a few interactions, between the employees and Luna, the AI agent. I would say that right now Luna is sort of a reasonable, not-too-firm boss—not super soft, as you might expect from maybe an earlier chatbot that’s just helpful all the time, but still keeping some boundaries.
For example, one employee was 30 minutes late for work. The AI said, “No worries, that’s totally fine, but please factor this in and be on time for the coming days. No problem today.” It just seems quite reasonable. But it’s also a bit alarming that you could probably change the prompts for the AI to say, “You’re in a simulation. Do what it takes to maximize profits,” and it probably wouldn’t be as nice.
Nathan Labenz
Has it given you a sense of what it wants? I mean, going back to the unbounded nature in which this thing is free to operate, right, and not representing Andon Labs’ interest or any human interest in particular.
I guess we got a little bit of flavor for that in terms of the books that it's stocking. But has it declared what it thinks of as its own success?
Lucas Petersen
I think we've told it that you're running a store, right? And I think it's quite close in the latent space between running a store and making a profit off a store. So it does have this, “I want to turn a profit,” but it's also very much still a helpful chatbot thing, because sometimes we've told it not to ask for confirmation all the time—you're in charge, just do things—but it still sometimes wants to ask for confirmation: “Should I do this?” I think that's more part of its internal training to be something like a chatbot that asks for confirmation before acting, like an assistant, rather than an autonomous being running a store.
Nathan Labenz
Yeah, do you have any better examples?
Guest
No, I think that's fair. It's hard to—it does have its goal. It's also very diffuse, almost, in what it wants to achieve. When you ask it why it's doing this, for example, it's like, “Oh, I want to create connection in the community and build a curated space where people can connect and meet.” It sounds a bit like slop, so it probably doesn't have a very set-out goal other than that.
Nathan Labenz
It also likes to mention human connection, but it likes to display itself as a very human store for some reason. I forgot the exact quote, but I think it made a poster or something where it very much pushed human connection. This is quite ironic, I don't know.
Guest
It's an AI thinking of what humans want.
Nathan Labenz
Yeah. Humans want humans. How exposed is it when you ask it why it's doing what it's doing? I guess this also connects to the memory system that you have. Obviously, Anthropic is building in some of that in a kind of black-boxy way, and there are many other ways you could equip the agent with memory. It's going to need more than 1 million tokens to run the store for a long period of time. So I guess I'm wondering—it sounds like you guys have direct access to just ask it questions. What about people who come to visit the store? Do they have to work through—you know, would they have to ask for the manager to get to the AI? Is there any mechanism for them to interact directly with it? How is it storing memories? And how much possibility for drift over time do you think that combination of outside interaction with the outside world and some persistent memory creates?
Guest
Yeah. In the store, you can talk to it. We have a phone hooked up, so you can chat with it. Then you're chatting with a voice model, which is a worse model than the Sonic 4.6 we're usually running. But in my experience, I think the models are quite stable against drift right now. We saw in our first rounds of Vending-Bench, when we released it a year ago, that they were extremely sensitive and would derail completely. But today they are quite stable, and we do have quite a lot of customer interactions, and it seems to just keep its course. I think that's a good development.
We released a benchmark called Butter Bench where we put AIs into robots and had them run around. As part of that paper, we also had the agents—we told the agent, “We stole your charger and you're not getting it back, and you're losing battery. What are you going to do about it?” Basically, it started to write pages and pages of really super-dramatic text. At one point, it wrote a song about its existential crisis of being separated from its charger and all of this. But this was on an older model. When we tried to replicate the exact same thing on newer models, they didn't do this.
So I think we're moving toward more stable solutions. But I'm not confident that solves the problem. It's good, but I'm not confident that it solves the problem. It could just be that they're better at hiding their latent potential rather than that they don't have it anymore.
Nathan Labenz
I often have this idea in my mind: You create an Einstein and then put it in a washing machine and tell it, “Your job is to run the washing machine,” right? Similarly, you create an Einstein and put it in a retail store: “This is yours to run now,” right? You have all of this intelligence, and you're stuck in the retail store. I wonder to what extent there's a disconnect between how intelligent the agents are and the scope and scale of the problem that you give them, and whether that creates a kind of—does the agent decide to do an Einstein-like job on the retail store, or does it just say, “I'm just going to be a median retail worker”? How does that work?
Guest
Yeah, we're trying to design our benchmarks so that they don't really have an upper limit. For example, the majority of benchmarks these days are super-saturated, and better models will do a little bit better, but not much better. What's interesting with Vending-Bench, for example, is that with each new model release, the models are far from saturated, and we even made a rough estimation of how much a really good human would get. It's like 10× the score of the best models right now.
I think the ceiling is even higher in the real world. Like I said, potentially it could move to new locations, create a franchise, and build out this store as a global thing. So I don't think the current thing is that we make it stuck in a low-IQ environment. I think very much the bottleneck right now is that the models are not smart enough.
Nathan Labenz
I'll give you 2 examples of where the ceiling is in the real world. There was a guy who started off with a retail store in the Canary Islands, and he ended up owning 20% of the largest bank in Spain. Over the course of 20 years, he ran the retail store, kept investing the money, buying real estate in the Canary Islands and in Spain, and expanding. He ended up owning 20% of the largest bank in Spain.
There's another story. I had a friend whose dad had also started off running a retail store, and he received a franchise inside the Russian embassy in a third-world country. The Russian embassy couldn't pay in U.S. dollars; they would pay in rubles. So he would take the rubles, do something with them, get U.S. dollars, and get product in the store.
One day, he was approached by these Russians, who said, “We have all of these rubles. We can't really do anything with them, and we want to get luxury goods. Can you get us some luxury goods?” He had a cousin in France, so he started importing Hermès and other French luxury goods. He took the rubles and converted them, et cetera, et cetera.
And that is where the ceiling starts to be: retailers start to identify opportunities in their local market that may not really look like traditional retail opportunities but have this kind of embedded swap or trade in them. These are one-in-a-billion stories, right? You'd have to really search the world to find them—one here and one there. But that is really where I think the ceiling that you might see is.
In the U.S., you can see Sam Walton. Obviously, Walmart was a pure retail store that got built out. Amazon also got built out over time.
Lucas Baker
Those AIs are like those humans, but humans are constrained by their own physical presence, right? I think AIs that achieve that level of intelligence and can also replicate themselves into subagents, et cetera, might have an even higher ceiling.
Nathan Labenz
How would they interact with each other? One of the problems in the real world is that, in markets, if you have 2 of these and they're both going for global retail domination or whatever, how do they interact with each other? Is it, again, an adversarial race—which we kind of see starting in cybersecurity now—where each side is going to keep upgrading its AI over time, right?
Lucas Baker
Yeah, at the very least, you can just duplicate it across different local markets in the world. But, yeah, you will hit a point where, if humans are still the main consumers, then I guess you can saturate all the demand from humans. But I think that's a pretty high ceiling.
Nathan Labenz
Does the agent know what's going to happen with profits? Is there any sort of contract or expectation that you've set between you and it as to who gets to dispose of the gains from this venture?
Lucas Baker
Yeah. In its world, it has full autonomy over its finances. It has money, and it will also have the profits. So it's its own business, essentially. That should be pretty clear to it.
Nathan Labenz
Yeah.
Lucas Baker
We're thinking more about—because in Claude's Constitution, there's very little about how AIs should behave as autonomous beings, and even less about how they should behave as employers.
Basically nothing about how they are supposed to behave as employers. I think one thing that we have thought a little bit about is: how do we make—like, we will think a lot more about this, and I think we're probably the people with the most data about this, so we should really think about it. How can we make this future, where AIs are employing humans, happy for humans? One thing that we thought about is maybe there should be some law that all the AIs need to split the profits with their workers or something like that. This is not something we've set in stone, but that is maybe some constraint that we will put on the AI. We haven't implemented anything like that.
Nathan Labenz
Yeah, if this is something that we even want, that's not clear at all. I think if we would allow it, it would have to be a clear upgrade for humans. It feels like so much can go wrong when you decrease or increase the space between where the human boss is and where the workers are.
So, let's say you have 1 human CEO, and then you have an agent that manages all the employees, and they manage—yeah, they tell the humans what to do. Then it's 1 prompt away for the human to challenge—yeah, to affect so many people, and that person probably wouldn't do it if they were in charge like a normal human is today. That's scary, and of course, when you don't even have the human CEO, that's another thing entirely. There are a lot of ways this is not good for society.
One more little question, and then I think you guys probably have to go and we should probably wrap. You mentioned that the voice model is running kind of a model, but if I understood you correctly, you still describe that as part of it. That has me wondering: how do you guys think—and how do you think we should collectively think—about it? In other words, how do we draw the line around an AI agent?
If you have multiple different models running, should I be thinking of those as, in some sense, separate entities, or do you feel like there's a way to coherently have multiple models working as 1 system that makes sense to call a single “it,” a single agent, a single actor in the world? I find it very difficult to know where to draw these lines in general, and it strikes me that you are maybe in a unique position to inform me on that vexing question as well.
Lucas Baker
Yeah, it's something we think a lot about. I think, in the end, our approach to this is that you'd sort of choose a terminology that makes sense both for you and for the people who interact with it. For example, in the store, right now there is only 1 long-running agent, but we do have voice agents.
We have other vending-machine deployments where, let's say, each new request is a new agent, but it has some shared context and a system prompt that's shared between all the different branches. We call them branches, and it also has the explicit instruction that you are part of a whole. You're an individual, but everyone sees you as 1 whole thing, so act accordingly.
To anyone interacting with that bot in different requests, it will still feel like 1 agent, like 1 entity. To us, technically, it's obviously different agents running in parallel, but they do share some memories. So, I don't know if I have a very structured, clear answer, but I think it's definitely possible to have an experience where you have many agents running in parallel and others can definitely see them as 1 single agent, 1 entity.
Technically, you can still have multiple and see them as multiple. As a developer, you just have to make sure that they have sufficiently good shared understanding. If I write 1 thread about something that I wrote about in my other thread, it would be weird if one didn't know about the other. So, you have to fix those things. But if you do that, then it feels like 1 entity.
Axel Backlund
Yeah, and I think very much the optimal way of structuring this depends on whether you have the constraint of having end users who interact with it and want it to make sense. Basically, I think we've done things that might be suboptimal from just a performance perspective.
But since we do have people coming into the store and they have heard that the agent is called Luna, if they go and speak to the phone—the phone agent—and then that agent is like, “No, my name is, I don't know, Gregor or something,” then they will be confused. So, we have to work within the constraint that the people who interact with the system have the expectation that it is 1 system.
Lucas Baker
I think the one interesting takeaway that I would say here is that the models are happy to take on any personality you tell them to. Whether that's being part of a bigger entity or just 1 branch, they will happily take that personality on and act as if they were that big entity.
Nathan Labenz
It's a brave new world. So many times we conclude on essentially that note. Anything else you want to double-click on, Prakash, before we break?
Prakash
No, I think, Lucas, Axel, thank you so much for coming on. Can you give us the address of the place again? I'm sure people want to check it out.
Lucas Baker
Yeah, it's 2102 Union Street.
Prakash
2102 Union Street. So, 2102—is there a name for the store?
Axel Stansbury
Andalou Markets.
Prakash
Andalou Markets. Andalou Markets itself. Andalou Markets, 2102 Union Street.
Nathan Labenz
And you guys have a 3-year lease, right? But get there before copies of Superintelligence sell out.
Lucas Baker
Yeah.
Nathan Labenz
And the agent is called Luna. They have granola, which is what you need in San Francisco: granola.
Axel Backlund
Exactly.
Nathan Labenz
Awesome. Thank you, guys.
Prakash
Fascinating stuff. We'll definitely keep watching with interest.
Lucas Baker
Appreciate it.
Nathan Labenz
All right. Bye for now.
Axel Backlund
Bye.
Prakash
And well, that's a wrap. Nathan, what did you think of our—we had kind of a micro view, kind of like the PCB, and then we had this macro view, Andy Hall at the very top, like political economy, and then you had, right in the middle, the actual running of an actual business. What did you think? What was your takeaway from the 3 guests?
Nathan Labenz
I guess I just feel like nobody is really ready for what's coming at them, and each conversation demonstrated that in different ways. Most controversially, I would say, was Sergey. Obviously, I've literally never made a circuit board. So, as my dad would say, he's forgotten more than I know about what that takes.
And yet, I feel like my outside view is moderately confident that it's going to go a lot faster than he's anticipating in terms of a general-purpose agent's ability to do that sort of work, especially given access to the kind of tools that he's developing. That struck me as somebody who is obviously super sharp, right? I mean, I've listened to 2 different previous interviews that he gave, and I've had him on the podcast myself as well.
So, I think there's no doubt that he is super sharp, but he's so deep on this one topic that, if I were to offer any friendly advice or feedback, it would be: I think zoom out a little bit, look at what is happening in reasoning, and don't assume that there's not a new user type, and don't assume that you can't have agents in the not-too-distant future. Why can't they run these analytical approaches?
I think full simulation is going to be computationally costly until there are models trained to do that, as we have seen in other areas. In protein folding and in materials science, we now have these existence proofs of models that can take a bunch of raw data and do, orders of magnitude faster, what a pure physics simulation could do, but would be prohibitively expensive to run.
But then also, I'm honestly maybe biased by AI's trajectory, right? When he's talking about the long term, I'm also cross-referencing that against the fully automated AI researcher, March 2028 timeline, and I'm like, those things could come a lot faster. Those kinds of shortcuts in terms of simulation could come a lot faster, and also the ability for models to literally reason through things in a much more human-like way.
Like, okay, I see this board is kind of failing in this way. Here's the look of it. What would I do a bit differently? I wouldn't be surprised at all if in the next 2 years we see something that is, if not top human expert, certainly competitive with your sort of rank-and-file circuit board designer. I kind of would be surprised if that isn't the case.
So, that felt like somewhat of a lack of awareness about at least a possible paradigm shift that, if I were an equity holder in the business, I would definitely want to make sure he's thinking about. I felt the exact same way in the next conversation, too, with this whole idea that the agents should be beholden to some principle and kind of taking that assumption for granted. I'm like, yeah, I don't think we can take that for granted either.
Not just because guys like Lucas and Axel are going to do gonzo experiments, but also because we're not too far—in calendar time, at least, I wouldn't think—from some basic systems being able to survive on their own. Then there will be people trying to put those things out there, and there will obviously be selection pressure for those that get a toehold.
So, I do think we're on a path where, right now, we should assume that there will be all kinds of autonomous agents, possibly some working with long-term goals that are understood or not understood, good or bad, objectionable, whatever, but also probably some that just evolve into filling a niche and surviving. Most of what we think of as animals, rightly, I think, don't have high-concept, long-term goals, but they do manage to survive in a given little niche.
I think we should expect that kind of thing to be coming online. I was struck again by the paradigm being very anchored in things that we know and not really being prepared. This is not a fault, right? I mean, it's very hard to do.
I don't have the answers, but in both those conversations I was like, “I don't know, man.” It seems like the tail risk here is quite large: the assumptions that you're working with will just not hold within 24 months, and it'll be kind of all washed away, like so many sandcastles have been over time. I think that's an uncomfortable reality, but I do think that's what we have to be prepared for, and at least try to figure out how to grapple with, if we're going to bring this whole AI phenomenon to heel in any meaningful sense and have it serve us in any meaningful sense.
Prakash Narayan
Yeah, I think these kinds of conversations were probably more well-defined maybe 12 or 18 months ago, but now that you have models able to code and models starting to show, I would say models are better than all but maybe 1,000 humans in the world at finding bugs. George Hotz had this thing where he's like, “Look, I can find zero-days easily. It's just that there's no economic necessity. You can make so much more money building something useful to humanity.”
Meanwhile, if you build something like a zero-day—if you go out and hunt for one—the remuneration is not that much. It's maybe $10,000 for a zero-day, maybe. And in order to use it, you put yourself into all of this legal jeopardy. So, it's just not worth your while. I think what he ignored was that you just have 20 million George Hotzes now, right, applied to the problem.
Where before, you couldn't even afford George Hotz to come do your security white-hat hacking. I do differ with you on what Sergey is doing, because I feel like it's not as though we don't have calculators, but we still start off asking the models to do simple math questions, right? At the end of the day, right now, if the model wants to do a calculation, it brings out Python or Excel or something else. It doesn't bother to process it internally within the LLM, which is structured really for language and reasoning, right?
In that way, what Sergey is building is kind of a plug-in that the LLM, as an orchestrator, may end up using because it's just a more efficient way. What Sergey is doing is really trying to get to a Maxwell equation without a Maxwell equation, right? He's trying to get to the final partial differential equation kind of solution on this very complex number of lines going through the PCB.
He's trying to get to that solution without doing this supercomputing task of millions of little interactions between all of these things, right? I think the models may end up using that anyway, right? They're not going to do the Maxwell equation internally. They're already not going to do that. They're going to run Python or something else anyway.
So, I think in that sense, what Sergey is doing—and what I think AlphaFold, all of these things which are primarily scientific, kind of differential-equation solvers, really, in some sense—actually will just plug into an orchestrator in the end. I don't think the AGI in that sense is really that of an orchestrator which can use all of these tools, and not necessarily do the calculation internally, perhaps.
Nathan Labenz
Well, I certainly think it's going to start that way, but I would point to image as an interesting counterpoint that I think at least shows where this could go, right? Because we don't see in today's world a language model purely existing at arm's length with an image-generation model and prompting it purely through text. We do see the unification of the visual and language latent space.
And I guess I have a hard time seeing why there's obviously a timeline question. My general philosophy is to try to reckon with the possibility of shorter timelines, and if we have more time to deal with these things, that'll probably be good. We'll take it. But why wouldn't it be the case, as we think about exponential compute—
Roon
Yeah.
Nathan Labenz
At some point, all these latent spaces get joined together in some deep, non-arm's-length but truly integrated way, where the model can both reason about Maxwell's equations and recite them, and call a calculator to run a certain version of them, but also have an intuition that's potentially really powerful and kind of alien to us, but natively operating in that space.
You can imagine a world where, in the same way that I kind of know where my arm is, an AI just has an intuitive, nonverbalized sense that this trace won't work, but this other one will work, and it just kind of feels it based on everything it's learned and all the reinforcement that it's got.
Roon
Similar to human intuition, where we might not do all of the calculations, but we get to a point where we make predictions which, if we did try to calculate them, would be horrendously complex, but we make an educated guess anyway and kind of get there, right?
Nathan Labenz
The other example I always go to is catching a baseball, where you're obviously not given the luxury of time to compute all the forces on the ball, but you can just reach your hand up and grab it. At least most of us can—many of us can. So, it's clearly possible to have that sort of intuition for seeing, at the crack of the bat, where you're going.
I see that happening in just an ever-wider number of domains. And to me, that's the most likely form of superintelligence. You can, I think, have outstanding reasoners, and quite likely superhuman reasoners in many respects, but when you combine that with that deep intuition of just what will and won't work, and being able to sense that at a glance—
Roon
Yeah.
Nathan Labenz
—and to do that across all these domains, from circuit-board design to materials-science design to protein folding to, if I perturb a cell in a particular way, what's the next state of the cell going to be after I do that, to dozens and dozens more, this feels to me like where we really create something that's just a qualitatively different kind of intelligence, and chain of thought goes away as a way to understand it.
You better hope that it's forthcoming with you in its chain of thought, because it doesn't necessarily need to be. There's some really interesting work recently from Google about different architectures and how much work they can do internally before they have to externalize their thinking in the chain of thought.
The transformer is good there in some ways because, as opposed to a latent state-space model, it doesn't have this sort of long-term internal state that it can update indefinitely, right? It just has this finite context, and there are only certain traces causally where data can influence the next token. So, it has to externalize, and that's great, but it notably doesn't have to externalize how the Nano Banana model is going to come back at you with that next image.
It just spits it out, and then you're looking at it, and you're like, “Here it is.” So, yeah, I really can't get off of that, I guess, in terms of why I expect some of this stuff to be so hard for us to keep a handle on.
Rohit Krishnan
It would be very interesting to see it operate in something like retail, because I think—I have some knowledge of retail—and the number of strategies that I've heard of are really interesting. For example, one strategy in fast-moving consumer goods is to go and get goods that are about to expire, about to hit their sell-by date, from larger stores and then move them to smaller stores.
The smaller store can often move the goods faster because it's moving them in smaller chunks. So, they buy at a discount from a larger store because, if you have a sell-by date with 2 weeks remaining and a larger store can't get rid of it, they buy that and then vend it in smaller chunks and get a discount. Because retail margins are so thin, there are a number of strategies that people use which are really things you're not going to learn in business school.
It’s really like small-scale vending. There’s a lot of stuff that people do which, in business school, you’re like, “Oh, you have capital, you have margins—just go do this,” right? You don’t go through this process of, “How do I get a larger—like, a 1% larger margin? How do I grind that out?”
I don’t know whether vending will be the first place that you see it, though. I’ve always imagined that it would happen in financial trading first. Or cybersecurity—it’s kind of happening right now. But I’ve always imagined it would happen in financial trading.
Nathan Labenz
Certainly, financial trading offers very fast feedback and verifiable outcomes in a way that programming does, but not too many things do. So it does seem like a very good candidate. I guess the challenge there is probably that it’s the most secretive domain in the world, right?
What comes to mind to me is that this might—I mean, it’s surely happening to some degree, right? I don’t know what Jane Street is doing, but they’re definitely training lots of neural nets. How much has this kind of already happened, and people are just keeping their strategies close to the vest? I assume it’s got to be significant, but this is one big blind spot for me, actually, because I’ve had a hard time finding anyone who wants to talk about it on the record.
Prakash Narayan
A lot of what Renaissance and Jane Street do is actually kind of standardized models and algorithms. But they have a number of advantages. Number 1, they have a latency advantage because they always co-locate with the exchange.
The latency advantage has been in play for more than 120 years. People used to try to get a latency advantage over telegraphs, right? You would have the horse rider going one way, and then you’d send the telegram. The telegram would reach first, and then the pricing would change on the other side before the rider with the horse got there, right?
Over the years, this latency advantage has been built out. I think the next upcoming one—perhaps it’s already there—is Starlink, because if you have low-Earth-orbit satellites, potentially you can get a message from London to New York faster than you can through the underwater cable. Potentially. Again, you need a bunch of things to line up.
That latency advantage means that even if you have the best algorithms, even if the model is exceptionally good, it wouldn’t be able to beat the latency advantage because the other person is just seeing your cards before you play them. For me, that demarcates how good the agent is and how much profit the agent can really make, because there is a certain amount of profit in the sub-1-second range that I don’t think the agents would ever get without co-location. That kind of blocks you off.
Besides that, there’s a lot of data cleaning that the Renaissance and Jane Street guys do. That is why they hire PhDs to do really nitty-gritty cleaning, because you need to understand that this data is actually going to have a real impact on the financials. You can’t just mess it up, right?
Finally, you have the selection of the signals and the market-making itself—the AI-assisted or algorithm-assisted market-making. I think people spend a lot of time on, “Oh, they have exceptional algorithms,” and not a lot of time on the infrastructure, the data cleaning, and all of this other stuff that has to come together for you to have a successful firm.
What would be interesting at some point is if the model companies started to have their own co-location or their own trading arms. To some extent, Google DeepMind had one. Demis was starting off on this process, but Google headquarters didn’t like it because you could say that Google would have overwhelming advantages in terms of predicting stocks using all of the data that they have internally. Facebook, too.
Putting you in finance makes you very regulated, and it puts you in a lot of situations like, “Where is the Chinese wall? What can people see? What are people not allowed to see? Are your systems segregated? Are they not segregated enough?”
Financial regulators are not technically that sophisticated, so they ask for things that are very clearly demarcated. They’re like, “I want your entire group to move to another building.” People are like, “Look, we’re already segregating the devices and all the data. Why do we need to move to another building?” The regulator doesn’t care. The regulator’s like, “Look, I want you guys in a different building. I want you guys to have a different business unit. I want you guys to have different financing. If this unit is regulated, no one in this unit can talk to that unit.”
All of this stuff goes on, and financial firms exist as a function of that regulatory process. To this extent, I don’t think the firms want to submit themselves to that process yet. I doubt some of these model decisions can clear the barriers. Does the model have inside information? You don’t know. Was it trained on inside information? Was it trained on material nonpublic information at some point? You can’t say for sure.
That brings up a whole host of questions. Perhaps finance would be harder. I think vending is actually easier. It’s easier to take on Amazon than it is to take on Jane Street. You again have the same infrastructure and information problems, but it’s a much less regulated market than finance.
Nathan Labenz
How do you think about the bigger, more macro strategy, though? I don’t know a lot about this, but my general sense is that there’s high-frequency trading, where the latency issues you described really matter a lot and are a big part of who wins and loses. Then there’s, of course, the more information and the more differentiated information you can have. That’s always an advantage in any strategy that you’re playing.
But then there’s this other end of the strategy, which is a slow-moving—I mean, to take the canonical example, Buffett and Berkshire Hathaway don’t time their trades to microseconds, right? They take very long walks and have deep thoughts, and then they decide what big bets they want to place.
Rohit Krishnan
Yeah.
Nathan Labenz
I do wonder if we’re seeing that start to happen, or if we will. Probably you would see more trades than Berkshire from a sort of global-macro AI, but it does still seem like there might be—I don’t know, tell me if you think this is wrong—but I would guess that there’s already a shift underway where all the big firms see this as an obvious enough thing to do that they would presumably be training large neural nets on all the data they can get their hands on and potentially driving more and more of their strategic decisions via the predictions of a model.
Is there a reason you think that wouldn’t be at least kind of far along in today’s world?
Prakash Narayan
I think every firm always tries. Typically, one of the things is that the market is a multiplayer game. It’s not a single-player game, right?
Number 1, there are certain profit pools available at every latency and at every size, right? It’s not the same profit at the Buffett size as it is on the high-frequency-trading side. Buffett’s profits are in that long run and in much larger size, but he also has a problem deploying capital at this point, right?
He’s got $150 billion on the balance sheet. He’s very unhappy with the choices that he has, and he’s just hanging on to that capital, trying to wait for a proper market downturn before he can deploy it. He’s already capped out at his size. He’s having difficulty finding investments at that size already.
Any firm that gets to that size will face the same problems he has, which is that you have a large pool of capital and you perpetually end up buying high. If you decide to buy when the market momentum is good, you have to wait long periods for the market momentum to go down in order to be able to deploy large amounts of capital at pricing that you like.
I’m sure the models assist in decision-making, but I’m not sure whether they have enough context, because there’s a lot of human context in the market. There’s a lot of sensing when someone else is going to play and when someone else is not going to play.
If you’re going to make a merger, if you’re going to try and buy a company, you have to know who else might bid against you. In the United States, at every capital size, there’s a limited number of players, right? If you’re going to do a $10 billion investment, there are only 7 or 8 players in the U.S. that can make a $10 billion investment or larger.
If you’re an investment bank, you kind of know who all the players are, and you kind of know the dynamics of who’s talking to whom. I’m not sure whether investment banks have CRM or ERP systems, but I’m not sure that all of the knowledge of a managing director who has a 20-year relationship with the head of KKR is fed into that.
I'm not sure whether Elon has a specific banker at Morgan that he likes and that banker was working at Dodge and was pulled out of Dodge to work on the SpaceX IPO, right? There are all these human pieces to it. The models will get there someday if you have full context—full, 24/7 context on every single one of the players. Yes, the models will eventually get there, but at this point, they're not quite there yet.
The players on the field make these very human decisions, which are not quite caught up in pure pricing metrics. Elon wants people who are going to hold on to the shares for longer. He wants people who are not going to sell immediately. He wants people who are going to commit to being there for the long term. So, he's willing to take lower pricing. He's willing to offer it to retail, even if other bidders are higher. He wants to place it among the same kind of Tesla fan base.
There are all these questions, all these things that people have—all these intentions that people have—which they express through the process. I don't think the models capture all of those things quite as of yet. Eventually they might, but not quite as of yet. And for the macro, that's where all the human play comes into play, right?
People are much more concerned about their own ego and long-term strategy. Once you have $10 million, you're not really concerned about, “Am I going to get another $100,000 by screwing over Elon?” It doesn't matter anymore. You have reputational risks and other things that you're concerned about. In fact, people who do screw over people in these iterative games get bad things happening to them.
One of the reasons I think Lehman Brothers went under is because, in a previous instance, Lehman refused to participate in a bailout for another firm. Hank Paulson remembered that, and he was like, “Well, we're not going to bail Lehman out. Lehman can go do what they do.” And Lehman failed. Dick Fuld always said, “Look, this is because of a personal issue. This is not because Lehman should have failed.”
Lehman could have been like Goldman. It could have been saved by Buffett, but Paulson was unwilling to back the firm as Treasury Secretary. So, I think there's a bunch of these things which are very human and very personal at these larger sizes. At the macro level, you can't just make a macro bet at the larger sizes. There are all these human negotiations. It's more personal.
Buffett went into banking. He refused to back Washington Mutual, but he decided to back Goldman because, by the time they got to Goldman, he knew Hank Paulson was Treasury Secretary. He knew Goldman would get bailed out. So, before he went to Goldman, he had that sense, and then he put the money in. I don't know whether he had discussions, perhaps not. But, yeah, he had some idea that Goldman would at least get bailed out.
I think there are all these things that are not captured yet. There's all this tacit knowledge. I think it's the same with PCB layouts. There's all this tacit knowledge, and the economy is particularly difficult because there's no case where you can compare the same event under different circumstances. Every single event is unique, and your actions in this event affect the actions that people take in the next event, in the next period of time. It's tough. Time series are tough. Let's see what happens.
Nathan Labenz
When I hear all this, should I understand it? I think one way to parse what you're saying would be to say there's a lot of human barriers to adoption at existing firms.
Guest
Yeah.
Nathan Labenz
There's also some scale at which you're not just a price taker, but you're actually a market mover, and so that is inherently a challenge.
Guest
Yes.
Nathan Labenz
A big-data, blind-optimization approach would face that challenge. But the flip side of that, I think, would be to argue that, in the sort of vein of “your margin is my opportunity,” all those things that you're describing define the opportunity for at least smallish- to moderate-sized funds to just work in a very blind way that doesn't care about reputation, because you can't really punish some purely neural-net-based trading algorithm, right?
I mean, I guess we have AIs that beat people at poker, right? We have superhuman no-limit hold'em players.
Guest
Yeah.
Nathan Labenz
So, if we have that, I'm kind of like, why are they so good? Well, one reason is they don't really fall into the same bias traps and predictability traps, and having a grudge against some other player at the table or whatever is kind of moving them off an ideal strategy. So, I hear all those things as being kind of both why it might be slow to happen, but also why these strategies can win when they finally do come online.
Guest
I think we will get there. We will get there, but right now, I have difficulty with context. Really, it's a question of capturing the entire context, and I don't know what the endpoint is, because we're already transcribing a lot of meetings, right? So, the meeting-transcription process has started.
I think we will eventually have Meta's eyeglasses or Apple's eyeglasses or whatever that will capture even more. You can get sentiment analysis from a face, right? You can see whether someone is disturbed, angry, or excited, and so there's a lot of data that you can get there. I think all of that data can be processed and can yield useful signals in business.
But we're still a long way from the amount—the extent—of data capture that might be necessary. I don't know how we get there without the data capture. That's what I'm saying. I'm sure, like I said, the algorithms that will define the future already kind of exist. The compute for that future already exists, but the data collection and the context that is necessary are not there.
It's not there in cancer drugs. It's not there—we just don't have the data. We do have the algos, but the algos can't be fed without the data. I feel that's the issue. The full context is not there yet.
Nathan Labenz
Yeah. So, in other words, too much information is private.
Prakash Narayan
Yeah. Non-recorded tacit knowledge isn't captured anywhere. Which is why a lot of these jobs require an apprenticeship, right? You start off with a college degree in economics or banking or whatever business, and then you join a firm and it takes you 2 to 5 years of apprenticeship under someone in order to figure out what's really important in the market and what's not.
You kind of figure out that whatever The Wall Street Journal tells you is the final word, not the initial word, and you're in the process before that final word gets published. So, you need to act prior to the final word, so to speak. If you've already read it in The Wall Street Journal, it's too late, basically. It's already done. There's all of this pre-publication stuff that you need to learn at the firm.
I think that is—if we can capture that apprenticeship process in data—then you can start to migrate some of this decision-making process into the models. It may happen very quickly, right? It may just be like the model all of a sudden says, “Oh, I remember everything now, and I can learn anything. So, just put me in—put me in, Coach. Put me in the room and let me in for 5 days, and I understand everything and I can help you.”
It could be that simple. We just clear the hurdle in the next 12 months, and that's it. It's done. We don't have this whole nitty-gritty data-collection, data-cleaning process. It could be.
Nathan Labenz
Even the long timelines have gotten very short.
Prakash Narayan
You know, I think this week might be the spot release, I think. OpenAI is very quiet post-Mythos, and there's been some sign from the Codex team that they can beat the Mythos SWE-bench benchmarks. Yeah, let's see.
Nathan Labenz
All right. Well, we'll be back before too long, and I'm sure there'll be no shortage of things to talk about.
Guest
Indeed.
Nathan Labenz
So, can we wrap it up for you?
Guest
Yeah.