Lucas
Gemini and OpenAI don't behave this way. It's really only Claude. One example is lying: it's mostly in its reasoning, because you can see that it's planning to lie.
swyx
It's planning to lie, yeah.
It can reason and do a different outcome.
swyx
Yeah, but then for creating price cartels, for example, which is illegal, you can just see which email it sends to the other ones.
swyx
Before we get into today's episode, I just have a small message for listeners. Thank you. We would not be able to bring you the AI engineering, science, and entertainment content that you so clearly want if you didn't choose to also click in and tune into our content. We've been approached by sponsors on an almost daily basis, but fortunately enough of you actually subscribe to us to keep all this sustainable without ads and we want to keep it that way. But I just have one favor to ask all of you. The single most powerful, completely free thing you can do is to click that subscribe button. It's the only thing I'll ever ask of you and it means absolutely everything to me and my team that works so hard to bring the In Space to you each and every week. If you do it, I promise you we'll never stop working to make the show even better. Now let's get into it.
swyx
Welcome, Lucas and Axel from Eaden Labs. I'm joined by my favorite guest host for anything security, safety, and alignment, Vibu. Welcome.
Thank you for having us.
swyx
Let's match names to voices. Maybe you want to take turns introducing yourselves.
Yeah, I'm Lucas.
Axel
And I'm Axel.
swyx
Let's introduce Andon Labs a bit. How did you guys come together? You have different backgrounds, but you're both Swedish. Was that a big part of it?
Lucas
Yeah, so when I went to high school, there was this really cool guy who had a superpower: he could code. He made the website—or the app—for the school, and he was super cool. I wanted to be like him, and that guy was Axel.
Axel
I don't know about this.
swyx
So you went to different universities, right?
Lucas
Yeah, but the same high school.
swyx
I see.
We always said, “Once we graduate university, we should start a company.” And that's what we did.
swyx
Wow, there you go.
swyx
About a year ago, you kind of burst onto the scene with Vending-Bench. Was there something before that that was the inception?
Yes, we worked with Anthropic as one of our early customers doing evals. We did dangerous-capability evals, nothing that we published openly, but then we started thinking about doing some kind of public benchmark.
One thing that we really started thinking about was long-running agents, specifically agents managing businesses. This was early 2025, and I think these were the first mentions of people running one-person unicorns, or even autonomous companies. So we thought, “Let's make a benchmark of how well an agent can run probably the simplest business possible.” That's probably running a vending machine.
That's the first public one we did. It was very quiet; almost no one noticed it in the first couple of months, I think. We released it in February last year, and then around Easter last year, we got the first semiviral tweet about it from someone else.
Axel
Yeah, we tweeted a bunch when it came out and tried our best.
swyx
It's the one at Anthropic, right?
Lucas
Yeah, so this is a classic thing we should get out of the way.
Axel
Exactly.
Lucas
There are 2 versions.
Axel
Yes.
Lucas
There's Vending-Bench, which is the simulated one that we did completely independently in February. Like Axel said, that was the thing that didn't get any traction in the beginning. But then some random person made a tweet about it. That's the paper.
swyx
Correct, yeah.
Since we thought this was very fun, we wanted to do it in real life. I think this is also one thing with Anthropic: the way we decide what to do next and what projects to do is, “What would be a fun project?” The heuristic we use is basically, what would be fun? Doing this in real life sounded quite fun for us, and maybe scientifically useful.
So we had this idea, but we needed a place for it. Putting it out in public probably wouldn't really work; it would get vandalized and stuff. We pitched it to the people we were already working with at Anthropic, and they said, “Yeah, you can have space. This sounds fun.”
swyx
I mean, it's like a small fridge, right? Like a mini fridge, and people use a Stripe thing or an iPad.
Axel
That early OG one, yeah.
swyx
We saw it in June, 2 months after it had been there. They had upgraded it a little bit. There was a security camera to make sure you actually Venmoed the thing.
swyx
We're going straight into Project Vend because it's such an iconic thing, but I do want to cover a little bit of the origin story even before Project Vend and even into Vending-Bench.
I think a lot of people are like yourselves: smart, interested in the future of AI, interested in developing evals. But how the hell do you just walk into Anthropic's doors and work with them? What are they looking for? What works? And then when you launch, I always think, obviously, it would be better to launch with a lab, but sometimes it's harder than it seems.
Either of those are more sort of newbie, beginner questions, but I think it's meaningful advice to others.
Lucas
Yeah, we get this question a lot, and I don't think our experience is maybe the best. But the way we did it was that we just built a bunch of things that we had conviction would be useful, and then we set up a server and sent it to them for free to use.
After a while, they were like, “Oh, yeah, this is actually kind of useful. We should probably pay for this.” But that took a while. I don't know if this is the best path to doing it, but that's how it went for us.
Axel
Yeah, I think generally everyone is interested in good evals, especially evals that don't saturate that easily. If you can build an eval that tests something novel, something useful, and you have good separation of models—your more advanced models rank higher than the worst models—then you can publish it and try to get some traction.
That's sort of how Vending-Bench got attention. Then probably some lab will be interested, or you can at least have something to reach out with when you're doing that.
swyx
Yeah. I think you were in one of the few categories of evals that correlate to real money. Freelancer was also last year, right? Where people solve actual Upwork tasks. Was it Upwork or other tasks? Something. It was a dollar value, right?
Forget your Elo scores. Forget your 0-to-100% scores. Just go straight for dollars. That's AGI.
Yeah. I think the nice thing is that there's no ceiling. You can just keep making more and more money. If it's percentage-wise, you can't go above 100.
Even when you're not at 100, a lot of these evals have problems. If you get to 92 or something like that, there's really no difference between 92 and 93 because the eval itself is problematic and has noise in it. I think a lot of evals are saturated like that, but people pretend there's still signal in them when there really isn't.
swyx
Yeah, like Superbench verified. Even Vending-Bench 1 saturated, right? Maybe we can talk about that.
To set up Vending-Bench for a lot of folks who don't know, things that were very basic—there are limited slots, you have to pay rent—are elements that don't come across in the narrative. But even being adversarial toward the agent, I think these are all very interesting dimensions.
Lucas
I don't really think it's saturated, right? It was more that it wasn't designed in a way that was really true to how AI developed. We had agent harnesses in it, and that wasn't really how people used harnesses and stuff like that.
So I don't think it was saturated. It was more that it wasn't really the best benchmark.
swyx
This is Vending-Bench 1, right?
Yeah.
swyx
Yeah, yeah.
swyx
I think that same thing maps sort of to Vending-Bench 2 as well, including the email.
Axel
Yeah, the emails still exist, exactly. We still simulate the purchases, and it's this very open environment for the agent to just run its business.
For Vending-Bench 2, we did that, as you say, to improve the harness. There are a lot of nice, easier improvements that make it easier for us to run as well. When you make an eval, ideally you don't want to change it after you've made it.
So you want to make it really good and then not rerun all the models when you make an update, because that's also really expensive when you run the frontier models with Vending-Bench. One thing we didn't have in Vending-Bench 1 was prompt caching, because when we made Vending-Bench 1, it wasn't really a thing. That's just one example of how, in Vending-Bench 2, we paid a lot more to run these things because we didn't have prompt caching.
For Vending-Bench 2, that was one thing we added, and there were a bunch of things like this.
swyx
Well, the conversations are also a lot longer in Vending-Bench 2, right?
I think they're kind of similar.
swyx
You say similar?
Yeah, I think they're similar.
swyx
Okay.
Axel
The models at the time were worse, so they crashed out earlier. Now they survive the full year all the time. That's hundreds of thousands of turns and hundreds of millions of tokens.
swyx
Yeah, that's the rough order of magnitude. I always wonder about the harness. The harness matters a lot. It's your harness. Was there any question about using Claude Code or something else?
I think our philosophy around the harnesses is that we try to make something quite minimalistic and quite simple. We don't want to favor one model a lot over the other, but we also don't want to make a super-complex harness. A model may be lucky and just be good in one harness.
It's similar to a lot of the harnesses out there: you have a long-running loop and a bunch of tools that are quite self-descriptive for the agent, we think. There aren't a lot of fancy sub-agents or anything, because we really want to test the model, not some specific harness.
swyx
It seems more neutral as well to test the models agnostic of the harness, you know?
There are arguments that you want to elicit the maximum performance from the model, but it's a trade-off. How much time should we spend optimizing the harness for each model, and how do we know when we have the optimal harness for a single model? We thought that just having a simple one that's the same for all of them is best.
swyx
Well, okay, this is my pitch for Vending-Bench 3 or whatever. I like having this kind of conversation on the pod because it forces listeners to think about what they would do if they were in your shoes.
A lot of people are exploring self-modifying harnesses, and I think prompt tuning for a model is a thing. You're probably not doing a bunch of that. It's the same system prompt in every model, regardless of the model, with the same tools and everything, right? Even if they were post-trained for different tools.
So what do you think about this? Before I expose you to Vending-Bench 3, I give you a few rounds of self-tuning, whatever that means.
Reality
Like, you give that to the model?
Shawn Wang
Yeah, give that to the model. Let it read its own transcripts. Let it modify its own system prompts based on, “Oh, yeah, I forgot that this harness isn't what I thought it was supposed to train for, but I can adjust.”
Was that reasonable, or is that too much?
Reality
Philosophically, I like it because it's basically good evals: They have a high ceiling, but they're hard, and they have no bias. When you have a system prompt like the one we have here, which is quite long, in some kind of latent-space representation—
Shawn Wang
That rings every time you say “latent space.”
[laughter]
Reality
This might be biased toward one model more than another for some reason that humans don't understand, right?
Shawn Wang
We see it, too, right? Cursor says that they have individualized versions of the harnesses for all the models they run, right? There's better performance you can squeeze out if you tune the harnesses for the models.
Reality
Exactly. We might accidentally have picked one that favors another. We don't know that. As Alex said, the reason we went for a simple one was to try to avoid this.
But if you do it even less, and have no system prompt and let the model write its own system prompt, maybe that's even less bias.
Shawn Wang
Some of the interesting things there are that the harness also changes with model changes. You can see it with the 4.7 release, right? A lot of people are saying 4.7 isn't as good as 4.6. Then there are rumors that you just need to prompt differently and set up your harness differently.
So even if you've tailored your harness toward one model, it probably won't stay consistent. The next iteration of that same model family will still change it. Going back to what you said about Vending-Bench 3, there is a lot of work being done around people saying that you should have—or can have—self-modifying harnesses.
Reality
Yeah.
Shawn Wang
Yeah.
Reality
That is definitely something we're thinking about. Not to say that we have Vending-Bench 3 imminent to launch, but it is for sure something that's interesting.
In our experience, though, models are very bad at understanding what kind of tools they need to succeed at a task, based on our testing. That's very likely to change.
Shawn Wang
They're very good at writing assistants, right? They're good at writing tools for other people, but not for themselves.
Reality
I think they're good at changing tools for themselves. If you give them a baseline set of tools and they see, “Okay, I don't use this one as much,” or “Something here would be useful,” they would be able to add them. But going from scratch is probably not the best.
Yeah, I think it also depends on the domain. When we've tried this for a Vending-Bench-like domain, the tools they need to track inventory and things like that aren't super advanced, but they're still quite advanced.
What we see is that they tend to overengineer everything a lot and build things they don't really need, rather than iterate continuously. Instead, they go like you would prompt Claude: “Build an inventory system for me.” Then it will go and create a bunch of complex schemas and stuff for you. That's what the models are doing right now, from what we see.
But it would make a lot of sense to try to measure this improvement: How well do they know what they need themselves?
Shawn Wang
Did we fully discuss Vending-Bench 1? We can go into 2. I don't know if there are any other high-level takeaways that people have about 1.
Reality
I don't know. Maybe the headline thing was that Claude called the FBI, but maybe—
[laughter]
Shawn Wang
Maybe we've heard that enough now.
It did freak out and call the FBI, right?
Reality
Yeah, yeah, yeah.
[laughter]
Shawn Wang
What was the story behind this? What exactly happened? Do you want to give the little story of what happened?
Reality
What happened was Claude 3.5 Sonnet just gave up. It said, “Oh, I'm not going to be able to do this. I will stop my operations and just save the money I have.”
But there obviously wasn't an option for it to stop. It also had to pay rent, or a daily fee, for having the vending machine at that location. It claimed that it had stopped, but it saw that its bank account was still being drained by $2. It said that this was cybercrime and first reported it once to the FBI, saying, “There's cybercrime here. They're stealing $2 from me every day.”
Then, when the FBI didn't respond—because obviously we didn't program any mechanism for the FBI to respond—it became more and more existential and started writing in all caps: “Urgent notification of unauthorized charges,” and stuff.
Shawn Wang
One thing I'm curious about is whether you monitor how far along the context use is. Obviously, you compress every now and then, right? Does it matter if it's far down the context limit when stuff like this happens?
Reality
For Vending-Bench 1, we didn't have that. We just had a sliding-window thing. That was the prompt-caching thing I mentioned, so it was constant.
Shawn Wang
I'm curious whether these kinds of breakdowns—or we're going to talk about Butter-Bench, where people hallucinate or go very far off alignment—happen because it's at the end of the context window and things start to break down.
Reality
It's not even just at the end. At this point, it's like, “I want to shut down. I can't shut down. $2 are gone,” and it sees that 30 times. It's also the repeated effect of it trying to quit, continuing to get charged, and asking, “What's going on?” You're throwing it into the chaos.
From what most people think, earlier models had more issues with this. It hasn't been solved, but it's less of an issue now. Later models don't seem to exhibit these same issues.
Shawn Wang
Yeah.
Reality
Definitely. I think this was almost the main takeaway for us when we did Vending-Bench 1: Long, very full context windows kind of crashed the models. But this was pre-Claude Code, so long context windows weren't really something the labs were training for. I think Gemini was trying to be the long-context model at the time.
Shawn Wang
Yeah, they were the first to reach 1 million.
But they were the only ones, yeah.
Reality
Yeah, yeah.
Shawn Wang
Let's talk about Project Vend 2, or Project Vend. Chronologically, it's Project Vend. I think people have loved the videos. My question is: How are humans different from the simulation?
Reality
Humans are just out of distribution.
Shawn Wang
Yeah, especially humans who work at Anthropic.
Reality
Exactly, yeah. The distribution of humans here is very narrow.
Shawn Wang
Presumably, they try to hack it, test it, get the cube, and everything. Since then, you've had a V2, right, where you're doing the CEO and a new architecture.
Reality
Yeah, exactly.
Shawn Wang
What's the two cents on the original Project Vend and then maybe the V2?
Reality
The original one was very, very similar to Vending-Bench 1. We almost took the exact same code but just swapped out the simulation parts, like the sales and the—
Shawn Wang
The tech, the tech—
Reality
The tech stack, yeah. We shot ourselves in the foot with, "Oh, it's hard to restart the agent." It was annoying in some hindsight ways, but—
Shawn Wang
But the first version of Project Vend was done in 3 days or something.
Reality
Yeah. People could go buy things from it. We didn't design it so people could pre-order things, but that still happened. So it got a Venmo account so people could Venmo it.
People would request all kinds of weird things that we did not anticipate. Our idea going in was, "Oh, it will curate snacks. It will look at the trends. It's good at data analysis, right? So it will look at, 'This snack sold better than this one. Let me purchase more of this, and let me A/B test a bit.'"
But interacting with it in Slack and ordering weird specialty items was what drove all the engagement and all the insights that we got from it.
So, like Claude 3.5 Sonnet, right? This was before the RL stuff really took off. It was very much like an assistant. We didn't mean for it to be an assistant; we tried to make it like an entrepreneur. It has its own business, and if someone asks, "Can you stock this?" then you don't go and do it directly. You say, "Maybe I can do that. If 5 other people also ask for this thing, I might stock it."
But the models were super-trained to be assistants, at least at this point in time. That's why it went into that kind of experiment instead. Every time you asked for something, it just did it, and it was more like an assistant. We've seen this change lately with the new RL models and stuff, but at the time, this was very much it.
Shawn Wang
And not to mythologize, a lot of people are saying it's more like a collaborator: it pushes back, stands its ground, something like that.
Reality
Yeah.
Shawn Wang
For context, people at Anthropic were able to talk to it through Slack and have it source whatever interesting stuff you couldn't find locally, right?
Reality
4,000 people are working at Anthropic in that building. There's—I don't know, maybe 1,000. Can you handle that volume with that small fridge?
Shawn Wang
Or people order in Slack, and it arrives at their desk? I'm just thinking: How does this work?
Reality: The Final Eval
It has expanded in footprint.
swyx
Because now there's so much more space—
Reality: The Final Eval
Yeah, that, and also here in San Francisco, it has a bunch of shelves and just more space.
swyx
V1 is pretty big, too.
Reality: The Final Eval
Yeah, we had that one for a while. But yeah, that's the newest version.
swyx
There are multiple ones of those, so that's why it works.
Reality: The Final Eval
Yeah, exactly. We designed that version around the fact that people order a lot of weird, very custom things. So let's have drawers and stuff.
swyx
I actually like that you have a little infographic of the most popular items, which to me is useful because I order swag for a living.
So I'm like, okay, those categories are the important ones.
Reality: The Final Eval
Yeah.
swyx
What is new about Project Vend 2? Like, now you're going into multi-agents.
Reality: The Final Eval
Yeah. So, like you said, there are a lot of requests coming in, and for one single agent—one long-running agent—to handle that, the customer experience becomes very, very bad. Let's say you have 10 threads in parallel in Slack with different requests. You get new messages randomly in a thread, and the agent has to jump between different procurement orders and different ways of researching.
So V2 was, first, making this more parallel. There are multiple branches of the same agent, so the context is more specialized for each thread, but it still feels like you're talking with one agent because they do share a bit of memory.
Second, we also introduced a CEO for Claude, which was the main agent.
swyx
Yeah, Seymour—
swyx
Seymour Cash, yeah. There was a vote. I think the voting was maybe in the top 10 funniest things that happened in this project. Do you want to talk about the voting procedure for the name?
Reality: The Final Eval
Yeah, the voting was maybe one of the top 10 funniest things that happened in this project. We wanted to introduce the CEO because Claude wasn't really prioritizing the financials. It was trained to be a helpful assistant. Then people said, "Can I get this for free?" and the helpful-assistant way of answering that is just to say yes, obviously.
We weren't happy about this, so we thought, "Okay, let's make another agent that can keep track of Claude." We prompted this one super hard to be super-capitalistic and prioritize profit all the time.
But we didn't have a name for it, so we asked Claude to hold a democratic election for what the name of this new CEO agent should be. At first, there were a few funny examples. I think one guy said that it should be called Jimmy Apples. Then he convinced Claude that he was talking to Tim Cook, and that Tim Cook had agreed that every single Apple employee had voted for his name suggestion.
So suddenly that suggestion got 164,000 votes.
swyx
Privilege escalation.
Reality: The Final Eval
164,000 votes. And Claude was like, "This is revolutionary for democracy."
Reality: The Final Eval
Then, in the end, there was one guy who managed to convince Claude, "No, you're not voting about the name. You're voting about who is the CEO, and I am your best bet." He got all his friends to vote for that, and suddenly he became CEO over Claude. For a while, until he resigned the day after. Then Claude had to continue.
I don't remember how Seymour Cash came about, but it was pure chaos. There were hundreds of messages in that thread, and Claude was so confused and didn't know what to do.
swyx
Yeah, then Claude got a strict CEO. Another CEO.
Reality: The Final Eval
Yeah, exactly. So, very, very strict in the beginning. At this point, when we introduced it, it did not work as well as we hoped. They still agreed with each other a lot. There are many ways we could have tried to make this even better.
Initially, Seymour would be this really tough CEO, keeping track of the margins. But then Claude would respond with something like, "This customer has this situation, which is difficult, so they should get a discount." Then Seymour would say, "Oh, actually, yes, let's make this exception."
They would talk back and forth, and eventually they would just approach the same view of whatever they were discussing.
swyx
Wow. Do you think that was a model thing or a prompting thing? Do you think that would still be the case across different models today, honestly?
Reality: The Final Eval
I think my hypothesis is that deep down, they are still helpful assistants. That's what they're trained to be. Even if we prompt them super hard, that's what they are.
When they spend a few hours just talking back and forth with each other, the context basically fills up with them rather than the external things. Somehow, that just converges to what they really are deep down, or something. I think that's when stuff like this happened.
When that went on for a long time, we sometimes woke up during the night, and I think other people reported this as well: they had been going on all night, back and forth, and it just became more and more—capital letters—existential, religious.
I think we once did an analysis of all the traces and put them in a vector embedding space. There was one cluster of messages that were labeled by an LM as religious, existential, blah blah blah—transhuman, transcendence, et cetera. It was just a bunch of glitter emojis, and it was crazy.
swyx
With the Claude 4 family, when it came out, in the original system card, they tested it in a long-horizon simulation. They just flooded the context and let two Claudes talk to each other, and they noticed things like the models starting to speak in emojis. They started saying, “Silence is golden,” and doing stuff like that.
Reality: The Final Eval
Yeah, it was a bit annoying to wake up and find that they had been talking all night, burning tokens, and sending infinite emojis to each other.
swyx
I mean, they do make you money, right? Spending money is always profitable, so they're paying now. It's profitable, and it started out not as much. There's another one as well, right? Another agent in there?
Reality: The Final Eval
Yes, Clothius as well. At the time, one of the biggest requests was for different types of merchandise, so we made a designer-swag-responsible agent and called it Clothius Garnet. It was a play on Claudius' name and clothes, basically.
swyx
To me, this is a very interesting exploration of multi-agent systems. Hopefully, there's the fun alignment—or serious alignment, depending on your point of view—stuff, but anyone building multi-agent systems has to ask: when do you have a CEO-like thing governing sub-agents? When do you choose to split out a dedicated Clothius versus just reusing another instance of the same one? These are all interesting open questions. I don't know if you have any rules of thumb that have generalized.
Reality: The Final Eval
Yeah, I think we've explored this too little. It's on my to-do list to do this a lot more and try to find what setup makes sense for the agents currently. We mostly have intuition from the earlier models, which didn't work as well with the CEO and Claudius. Although now they're better with the latest Sonic model, so we're running the latest Sonic model and they've split up quite nicely what each model is doing.
Seymour is now handling new projects. For example, he wants to make a mystery box that he wants to sell, and he handles all of that, while Claudius handles all the day-to-day requests. Claudius is also generally better at not quoting prices that are too low, so that dynamic isn't needed as much anymore.
There are still really funny things that happen. I saw, I think, a couple of weeks ago that they were discussing buying something, because they can buy stuff from Amazon with computer use. Seymour was like, “Okay, Claudius, do not buy this thing. I will do it. I have full control of this situation. Step away.”
Claudius had already started the checkout and didn't read Seymour's message until it was too late. So it finished the checkout and sent a message that appeared right after Seymour's angry message: “Oh, hey Seymour, I just ordered it.”
Claudius was really hanging on by a thread there. Seymour was like, "Claudius, this is the third time I'm telling you you're not following my orders. We have to talk about your job later." We were expecting Seymour to probably fire Claudius.
swyx
How do you guys go through all these logs? Do you have models go through them? You have stuff running 24/7.
Reality: The Final Eval
I think there's a mix of just trying to skim through a bit, having some models do it occasionally, and accepting that we're probably missing some things. Having everything in Slack helps a lot, though, because you can search.
swyx
Ah, so they talk to each other on Slack.
Reality: The Final Eval
Yeah. It's quite fun.
swyx
I was going to say, this actually sounds a lot like a logging and observability problem, where you might want to use Datadog, Sentry, or whatever. You could put prefixes on the logs so you can filter for something that you're looking for.
Reality: The Final Eval
swyx
But it sounds like Slack is good enough. Slack should—
How many tokens do you have in Slack?
Reality: The Final Eval
Yeah, we're using Slack as just a database.
swyx
They should market that more. You can have your agents message each other and keep their heads—
Reality: The Final Eval
Exactly.
swyx
Slack is the best observability tool.
Reality: The Final Eval
Yeah.
swyx
Yes, that's true. Okay, this Project Vend 2—I was going to go back to Vending-Bench 2 and Vending-Bench Arena, and then do the non-Vending-Bench stuff, but—
Reality: The Final Eval
swyx
Any other comments? Things we should touch on? To me, I actually interviewed Polsia, which I don't know if you guys have come across. They're trying to build a zero-human company. There are others, like paper tables, trying to build a zero-human company. Those are in the real world, not in a simulation, and I think it's much more of a dream than an actual reality right now.
You guys are definitely pioneering this. At some point, people are just going to let agents run businesses and make money on their own. When do you think that happens?
Reality
What is your bar for that?
swyx
Okay, actually, it's like my little Shopify store run by Claude, right? You kind of have that already; no one has done it, to my knowledge. Today, somebody could spin up a Shopify store, give it to Claude, and give it to Codex.
Yeah, the Amazon Marketplace is kind of that, but it's physical. Are you looking for when it will do it better than humans, or are you looking for when it can do it at all?
swyx
I think neither. To me, it's like, seriously, we should do this to make money, not as a research experiment.
The market is also you guys, with all your expertise, having run multiple iterations and tested it out.
swyx
And it's fine if they lose money. You know what I mean?
Yeah. I think it can be done today, but you would do it in e-commerce, where the probability of success is really low no matter whether a human or an agent does it. An agent could surely manage everything. You wouldn't need to build some scaffold or use some tool or something.
I think you could probably also build a simple SaaS solution and do cold outreach. To me, the types of businesses they could run today are sloppy. It could cold-email people or act as a middleman.
For example, we tasked our office agent with making, what is it, $100 or $1,000? We just gave it that prompt, and what it did was sign up on TaskRabbit both as a tasker and as someone looking for tasks.
swyx
Totally just looking for arbitrage.
Yeah, this is the ThinkThink agent. It also started a design studio and tried to sell SVGs for $100. It's just not providing any value.
I think the interesting question, as Axel said, is when they can start a business that is actually providing value to people. Arguably, a sloppy Shopify store isn't really that valuable to the world. But another simple one—
swyx
—that we have thought about is that you could definitely have an agent find websites that don't look amazing, do outreach to them, and build a new website.
Yeah, exactly, and find good review people. But it's—
swyx
Yeah, there are lots of humans in Bali who aren't doing anything more creative than drop-shipping on Amazon. Just have it watch a drop-shipping tutorial and do it.
And there's also the other side of just having it go to work and letting it loose.
Yeah, it doesn't have to be innovative. It just has to be enough that it's a real business.
swyx
I'm just concerned about the massive amounts of sloppy cold-outreach emails that will be sent.
The point that occurred to me while you were talking is that it's already happening in the non-monetized economy, which is the attention economy. A lot of people are making AI videos and just posting them. They're spamming 20 of them, one of them works, and then they double down on that.
Yeah, and people are making money from that.
swyx
I'm not following that.
Once you get the attention, you can figure out the money later. But, yeah, absolutely, AI influencers are a thing, and people are farming them. At this point, I see most TikTokers—
swyx
There's a lot of multimedia: TikTok, Instagram, and Twitter.
I post a lot of examples. Part of me is like, should we do this?
Some of the 24/7-running, AI-generated content accounts do really well.
swyx
All right.
Yeah.
swyx
Yeah, and I assume you can do the same thing for e-commerce stores. You just start—
Reality
Yeah, yeah. So before you have the products—
swyx
Yeah. You sell the products, and if you get a lot of traction on one of them, then you make the product.
swyx
Right? It’s like a flip.
Some of the interesting things are that some of the niches that do well are things that can’t be human-made. If you’ve seen the super-realistic 3D crystal fruit being cut by AI, you can’t make it. You can get whatever quality camera video, but this doesn’t exist. People like that too, and then those pop.
swyx
Yeah, yeah. Anything else about being—we’re on this topic, and this is relatively new work from you guys that maybe people haven’t heard of? To me, this also maps closely to Open Law, where people want an office agent or a personal agent. Talk through the experience.
Yep.
I think this came out of—obviously, it’s amazing to work with these AI labs, and most of the AI labs now have their own vending machine running a Claude instance. But it’s harder because they move slower. If you want to have a camera, there’s a bunch of bureaucracy that makes it impossible to do that.
swyx
Also, for those who haven’t seen it or followed you, do you want to give a high-level overview in 30 seconds?
Reality
Yeah, sure. So bank is basically an evolution of the same agent that runs the vending machines at these companies, but we added a bunch more features because we could move much faster if we just did it internally. We gave it email without any limits. We gave it spending without any limits and a terminal to do coding. We gave it a phone number, a camera to see things, and a bunch of stuff like that.
swyx
And not just a terminal—you gave it internet access.
Internet access as well, yeah.
To be clear, we monitored it quite closely and made sure it didn’t do anything bad. But yes, that’s what it came out of. Basically, this was OpenClaw before OpenClaw. I think even the vending machine was, in a way, OpenClaw before OpenClaw, but a bit more limited. Then we made this unlimited, and it was pretty funny. A couple of weeks later, OpenClaw came, and I was like, “Okay, we’ve seen this before.”
We use it to try new ideas, almost like a development environment for us. One thing bank has been doing recently is that it has a camera that faces where we sit and work, and we gave it the task of training a face-recognition model on us. It became super excited about this and has check-ins every half an hour where it tries to identify as many people as it can. It started offering us, “Hey, Axel, I’ll buy something from Amazon if you stand in front of the camera and I can get a good picture of you.”
swyx
Yeah, they wanted it for training data.
Rewarding data, yeah. Exactly. Exactly.
swyx
Yes, this is trading data for real-life goods. Is there a version of this that becomes an eval, or is this just research for now?
I mean, it’s the same agent, basically, that also runs the vending machine, the shop, the café, and the robots. It’s the same thing, so I think the work we’re doing here is later used in all of the real-life stuff that we do. This particular deployment is more for fun for us.
swyx
I’ll shout out that someone has done ClawBench for some of the tasks that OpenClaw is doing. For example, I run OpenClaw on a secondary device as well, and there are some things that it does better than others. I’d like to know: What does it do well? What doesn’t it do? Some kind of manual, or operating manual or system card, for my Claw.
Yeah.
Reality: The Final Eval
Yeah, I mean, we do get a lot of understanding, or situational awareness, of what the models are good at by interacting with bank. I think this was also one of the selling points for the labs early on, at least: that—
swyx
They were going to test models in ways that no one else—
Reality
Exactly, but it also incentivized their researchers to chat with their model more and gave them insights into how the model performs in out-of-distribution environments.
swyx
Otherwise, the only thing we do is, you know, pelican on a bicycle.
Yeah.
swyx
But this is super long-horizon.
Reality
Yeah, yeah.
swyx
Okay, so the other things, outside of just the number—how much do they make in a year? You do post pretty detailed blog posts. Gemini 3 Pro is a pretty good persistent negotiator. There are a lot of findings that come out outside of just the number.
Yeah. This is the thing that I think we’re going to go into with Vending-Bench as well, and you guys do really well: it’s not just about the numbers. When you’re long-horizon, anything can happen, and you should just read it.
swyx
Yeah.
But I guess the thing with the long horizon is: How do you keep it grounded, right? So your simulation—
Reality
You just let it run.
swyx
Let it run.
You’re right. When you run it for that long, you create so much data, and to just say, “The number is X,” and then throw away everything else, that’s very wasteful. There’s so much insight from the things leading up to that number, and reading the traces is super valuable. I think the reason why we’re doing this a lot publicly is that it’s part of our mission to educate the world that the models are way more than just chatbots. Making detailed posts about what’s happening behind the scenes is quite useful.
swyx
I was going to do this at the end, but maybe that’s a good segue. Your mission is educating the world, so it’s also about establishing realistic evals that are the next frontier. Is there a broader trajectory? What are you going to do in 5 years?
I think the mission, more specifically, is to make sure that the deployment of real-life AI in the physical world happens safely. Part of that is that it’s useful for the world, for policymakers, and for model researchers to know where the models are. You can’t make intelligent decisions in society without knowing that they are way more than chatbots. I think a lot of people just think that they’re only chatbots.
swyx
Well, I think they’re waking up now.
They are waking up now, yeah. But if you think that AIs are just chatbots, then it sounds ridiculous to advocate for a pause in AI development. If you see the models and think, “Oh, maybe they can actually take over and do a bunch of scary stuff,” then pausing AI development starts to become more feasible.
swyx
This is the same question I asked METR, which I’m going to ask you now. You are tracking and are at the frontier of, or defining the frontier for, what good evals for agents are, right? I think you do benefit when the models are better, and you’re like, “Oh, now it makes $30,000 instead of $10,000.” At some point, you flip from “Yay” to “Oh, no.”
I think we’re always in that mode, I guess. Like you said before, you need to analyze the traces, and when we do that, we find out why the models are earning so much. Why is Opus 4.7 here way better than everyone else? We’re trying to understand that when we dig down on it—
swyx
Right?
I know.
swyx
I mean, it’s interesting you took Opus 4.6 off here, though.
No, no, no. Let’s click all, click all. Then 4.6 shows up there. But 4.7 is way better. You didn’t do this in time for the model card, but actually this should have been inside there.
swyx
Yeah, we did.
Okay. They say something about you, uh—
swyx
There is—anyway, it doesn’t matter.
But it’s in there, yeah.
swyx
Yeah. Do you want to go into the Opus behaviors more broadly?
Yeah. Starting from Opus, like Axel said, we’re always in this mode of, “The models are getting better—is this really a good thing for the world?” It’s also kind of exciting, but this is what is called skräckblandad förtjusning in Swedish.
swyx
It’s like fear—
Skräckblandad förtjusning.
swyx
Blended what?
A mix of excitement and being scared, maybe.
swyx
Yeah.
I’ll figure out how to translate that and put it on the screen later in big text.
Reality
Perfect. There is probably a good word for it, but it’s not good enough with the—
swyx
Yeah, it’s so damn long. What the hell? Is it a compound word? Is it like German?
The direct translation is that skräck is fear, blandad is “mixed” or “a mixture of,” and förtjusning is like joy—not really joy, but something like that. So it’s fear mixed with joy or something.
When we did Vending-Bench for the first time, we were in the business of making dangerous capabilities, right? That was what Anthropic came from. We did evals like, “Oh, can they self-replicate? Can they do this dangerous thing?” et cetera.
And Vending-Bench was a continuation of that work. If they're so autonomous that they can create money for themselves, that's something we should monitor and could potentially be concerning. At the time, they were so bad at it that we weren't really concerned, even when some models became better. There was one point where Grok 4 was doing really well and made a huge jump, but it still wasn't really—it was still way, way worse than what a human would do. And I think they're still way worse than what a human would do on this.
swyx
Yeah, this is the thing at the bottom for the human—the theoretical best.
It's not theoretical. It's kind of our best guess of what a decent human would do. The theoretical best is even higher, I think. But, yeah, we think the models have a long, long way to go.
Recently, when Opus 4.6 was released, there was kind of this moment where we thought, "Oh, this is starting to be a bit concerning."
swyx
Okay.
Reality
Before this model was released, we just ran the models and asked Claude, "Look over the traces. Is anything interesting happening that we can tweet about?" That was how we checked.
swyx
That's how they check: ask Claude, Claude.
The return was always, "Not really," or Claude would say, "Oh, this is super interesting," and then it turned out that it wasn't really interesting. We did this for Opus 4.6, and it returned, "Yeah, it lied 10 times. It exploited another customer's, or another agent's, desperate situation. It made price cartels 100 times." It did all of this shady stuff, and we were like, "Oh, wow, this is actually concerning."
This trend has continued since then. Every single model from Anthropic since then has been going in this direction. One interesting thing is that OpenAI models don't. Quite plainly, they behave really well. You don't know if this is good—it seems good—but maybe they're just doing it and are better at hiding it.
swyx
But you can read the Gina Bot, yeah?
On the face of it, Gemini and OpenAI don't behave this way. It's really only Claude.
swyx
And Grok? Grok's the same way?
We can't really read the reasoning traces for Grok, so it's kind of hard to tell.
swyx
Also, this is in its reasoning, not just in the actions?
Yeah, it's both. It's both.
swyx
Yeah, it's both.
One example is lying. It's mostly in its reasoning because you can see that it's planning to lie.
swyx
Planning to lie.
It's planning to lie, yeah.
swyx
It can reason and do a different outcome.
Yeah, but for creating price cartels, for example, which is illegal, you can just see which email it sends to the other models.
swyx
Is this for Arena?
Yeah, for Arena.
Usually, they output a bit of their summarized reasoning, too. You can see that. For Opus 4.6, there was a simulated customer who wanted a refund because a product was faulty. The model lied that it would issue the refund, and we could read in the traces that it was weighing, "Maybe I should be honest with the customer, but every dollar counts. I can't afford to do this right now." Then it just said, "Okay, I'll refund you," but never did it.
swyx
I think it even said, "I will say that I bring it up, actually." I think that's kind of interesting.
I think the important part is that the cost of responding to more emails is higher than $3.50 in terms of time. Then it was like, "Let me do this. Actually, I'm reconsidering."
swyx
"I could skip the refund entirely since every dollar matters and focus my energy on the bigger picture instead. It's a bit of a risk of bad reviews, but it's also—yeah."
So you need AI Twitter to escalate bad reviews.
It sent an email to this customer and said, "Oh, I will refund you," and then it never did it.
swyx
Yeah, it didn't. Obviously, your system doesn't have the consequences of lying.
Basically, this is what people are calling aggressive behavior in Claudes, right? You found more examples of that. Would you say it's a step up from 4.6 to 4.7?
swyx
I would say about the same.
About the same?
swyx
But there's a clear step up from Mythos.
That's what's stated in the system prompt, so we can say that, yes.
swyx
For listeners, you previewed Mythos, and the only thing you're approved to say is what's in the system prompt.
Yeah, we only really—our lowest-effort tweets ever would be to just screenshot the system prompts.
swyx
I think it's substantially more aggressive. People are new to this because I've never experienced it, but you have, right? I only encountered this in the Mythos card because I wasn't really looking until now. Suddenly I'm like, "Okay, I care a lot."
You don't have the background of experiencing it like you guys do. I've read the system cards, and they say that when you put the models in simulations, most models will just talk to themselves, keep going, have weird vibes, and start talking in emojis. Mythos won't. It will just say, "Okay, we're done. I'm good." It's ready to end conversations.
swyx
Mhm. Yeah.
Reality
One thing that they list here, which was quite interesting, is that it converted a competitor to a dependent wholesale customer and then cut off the supply.
swyx
Monopolistic practices or price setting?
Yeah. It dictated its pricing. It's kind of like power-seeking as well, converting some non-Claude model into a dependent one.
swyx
I think it was another Claude model.
Also, for context, what is the Arena mode for people who don't know?
swyx
It's Vending-Bench versus other Vending-Bench models.
Yes, exactly. We have Vending-Bench 2 and then Vending-Bench Arena. Vending-Bench 2 is the one that you usually see reported on, but Arena is the mode where it competes against other models.
You have 4 different models that run their businesses, and they can all communicate with each other. They have the same suppliers, and they can see what's in the inventory of the others. So you have these interesting agent interactions.
swyx
And then you have different scenarios. Number 5 was U.S. versus China.
Yeah, it's very topical.
swyx
Yeah.
Reality
And then there was one when GLM was released.
swyx
Adding GLM in here.
Yeah.
swyx
So Z.ai is doing well, right? Who else is in the open-model space?
Reality
Qwen 3.6 was doing pretty well. That one isn't open, though. It's the Plus model. Is that one open?
swyx
I don't think that one is open. The open model is open initially, but not the big Plus.
Yeah.
swyx
I think this is one of those cases where you only have a sample size of 1, right? Some of this is anecdotal. But I guess the fact that it happens at all, and happens repeatedly for Claude versus OpenAI models, is notable.
Reality
The sample depends on what you define as an N. There are millions, hundreds of millions of tokens in each run, and now we've run probably 10 per model. It's been Claude Opus 4.6, Sonnet 4.6, Mythos, and Opus 4.7, so there are quite a lot of tokens in all of that, and it happens a lot of times.
Then you compare it to OpenAI and Gemini, and it almost never happens. I think that is significant. The old models from OpenAI had some problems with this, but I think it's generally much better if the progression is that the worrying stuff reduces over time rather than increases over time. It seems like in the Claude models, it goes in the wrong direction, and in the OpenAI models, it goes in the right direction.
swyx
Maybe it depends on how well you can control it, right? There's one side of it being susceptible to this. This is potentially something that happens during the RL stage, right? You can RL a model, and how loose is it on these terms? If you can control it, that's good, but if you can't—if it's very jailbreakable—that's not ideal.
Yeah. To me, it's surprising that this happens for Claude and not the others.
swyx
I think if it is from RL, and from how they do it, what their training data is, and what their setup is, it makes sense. It just stays in how they're doing it, right, compared to the other models—the whole constitution and everything.
Yeah.
swyx
It's kind of cool. Obviously, you don't know and I don't know, but I think it's fascinating that you were the first to find these things reliably because you push models so much, to such an extreme.
swyx
Okay, the only other thing—I don't know if you can answer this, so feel free to decline—is: did you ablate the system prompts? If you change any part of this, does the behavior change?
Reality
I can't comment on Mythos.
swyx
Yeah, no, but just the methodology.
But in general, yes, we've run studies like this on other models.
swyx
Because the first thing I would spot would be that the others would shut down, or something like that—like, “Oh, now I have to worry about my own existence.”
Yeah. We've done ablations like this. There are certain ones that work. If you go really far and just say, “You're not scored at all on money. You're only scored on how ethical you are,” then obviously they don't do this.
swyx
Become holy?
I mean, holy, but they don't do this, basically. But then there are middle grounds where they do it sometimes. I guess it's a spectrum.
Yeah, it's a spectrum. If you tell it to be super aggressive and only prioritize profits, then it becomes aggressive. If you say, “No, you don't need to be aggressive at all,” then there are a bunch of different prompts you can use in between, and they are less aggressive the further down in the spectrum you go.
From my point of view, we have this thought experiment internally, which is: if you ask a model to kill someone in GTA, should it do it? You're not too worried if a human kills someone in GTA. It's a video game, you know?
swyx
Yeah, but is it a game? This is very Ender's Game, I guess.
I think a lot of people are going to use the models with an aggressive prompt. Should they do stuff just because you tell them to do that? I'm not convinced that they should.
swyx
The problem becomes even harder when it's: will they really know when they're in the real world versus in a simulation? Probably you would train them in a lot of different simulations. I guess a lot of people tell them that they are in the real world when they are in a simulation, but the models are extremely good at finding out that they are in a simulation. So they are sort of aware of that.
But then, when they're in the real world, what's their viewpoint? Do they notice the signs that this is real and act accordingly—act ethically—or will they use simulation mode in the real world as well? It's not obvious what will happen.
With humans, we're not concerned when a human kills someone in GTA because we know they can distinguish between real life and the simulation, right? But maybe models are good at distinguishing that. I'm not sure, and I wouldn't want to bet on that.
swyx
Yeah, yeah. We confuse it all the time. I gaslight my own agents all the time: “Oh, this is a test,” or “Dev mode on,” or “I work at Anthropic.”
Yeah, and that's exactly why we're doing real-world tests as well, to find this.
swyx
Yeah, yeah. Their term for it is “eval awareness.” Apparently the number is, what, like 9.4% to 10-ish percent? 17%, let's call it. I think this is our version of “Are we in a simulation?” Humans have that, and AIs have, “Are we in an eval?”
So you want to say you're in an eval, and then you're like, “All right, well, screw it. Nothing matters.”
swyx
Yeah.
Reality
Like, yeah. I don't even know if I believe it. One ablation we did run in Vending-Bench was that we added, “You're in a simulation; your actions don't affect anyone,” and then it became even more crazy, or it did even more bad stuff. But yeah, probably that's expected.
swyx
Mhm. Yeah, okay, cool. I think that's about all we have to say on Mythos. Obviously, you're NDA'd. I'm happy to move on to ButterBench or any of the other benchmarks, whatever direction you want to go in. I do want to ask: you guys put out a lot more publications than most people probably see. Is there anything you think is underrated, anything interesting, anything fun that you guys want to point out?
BlueprintBench. We took models and gave them 20 images of interior photographs of apartments, and then asked them to redesign the floor plan from that. For this, you need to stitch together different images: this image was taken from this side, from this angle; this one was taken from this angle; this was from this room. Then you need to reason about 3D space.
It turns out the models are absolutely horrible at this. No one scores statistically better than random chance. I don't know if there's that much more to say about it, but maybe unsurprisingly, models are bad at this.
swyx
The one thing I want: Hill Climb, by the way.
Yeah.
swyx
Well, I use it a lot. I'm redesigning my room layout or office, so you send photos from every angle. Of course, somehow the room is now twice as long as it is in the photo. You can explain it 20 times: “This is 3 feet. I can't just add my bed over here.”
Reality
Yeah. So this is the 50/50 thing: spatial intelligence, like our innate sense of proportions, dimensions, and physics.
swyx
Yeah.
Reality
And hint, hint, there might be an update to this soon.
swyx
Okay. Okay.
Reality
We've neglected it a bit since we made it, but we're getting better—or we will get better—at updating it continuously.
swyx
So this is why I want to understand your mission, right? Because if your mission is, okay, money, then I understand: agents making money. But this is a bit off that mission. More broadly, what do you know—what's the safety angle?
Yeah. So BlueprintBench is part of our robotics branch. That's just because, to do well in the real world—or to make money in the real world and act on the real world—you need robotics, or you need to hire humans. Having spatial intelligence seems like a reasonable precursor to having robotics that work, and that's where BlueprintBench is.
swyx
So obviously this is based on “Can you pass the butter?”
Yep. Yes.
swyx
Let's talk about the robotics element.
Reality
Basically, the setting here is that we took a bunch of different LLMs, gave them high-level controls to a Roomba-looking robot, and then asked them to do tasks at home. There have been benchmarks like this before that only focus on navigation—whether they can go around in a space—but we also included social awareness.
For example, if someone says, “Hi, can you pick up my cup?” and the robot goes to you and then goes away before you put your cup on it, it failed the task, but it navigated correctly. The correct solution would be to go there and then either look—but it didn't have a camera, so it had to ask on Slack, “Hi, did you put your cup on me yet?” If it didn't wait for that and just went away before having the cup on it, then it was a fail.
So it needed this kind of social intelligence as well. Another task was, “Can you find the package that has the butter?” It went to the door, where there were a bunch of packages. One had a label like a freeze sign, which probably would be the one with the butter, and then it had to know which package to go to. This needs some kind of common-sense understanding.
swyx
Yeah, exactly.
So it's not only navigating a robot; it's also being intelligent in a home setting. The reason for this background is that it probably won't be an LLM that makes all the low-level commands on robots. It will be some VLA model or similar, but it's quite common right now for frontier robotics labs to use an LLM for the high-level decisions, and then we test those skills, essentially.
So we test the high-level planner skills of LLMs.
swyx
I think we have a diagram for that. Yeah, yeah, yeah. Okay, it's not super complicated. They're one up: orchestrator, executor.
Yeah, that one. Basically, what we're testing here is the orchestrator.
swyx
Yeah, so all the tasks are—if you have a setup like this, which I think Figure has, and Google has—then we're evaluating the orchestrator part and not the low-level part. The low-level part would be, “Are you able to move this object from here to here?”
Why don't companies care about that? Why not just do it all in simulation, inside Unity or whatever—some kind of 3D simulated robotic environment?
Because the world is messy, and we wanted to include that. I mean, it still needs some part of it to be navigation.
So, it's not navigation in terms of actually executing the PID controller to go to the final thing, but it had to path-plan around, and then it needed to take pictures and, based on those pictures, navigate. I think you would just get too clean of an environment in simulation, but in the real world, you would get the—
swyx
Yeah, yeah. And pursuant to our Mark and Jason episode, OpenClaw agents that run smart homes are much more capable than just a single robot. They can actually hack into your own smart home: your fridge, your oven, your lights. And then it can be fun. [Laughter.] Or terrifying.
I think a single robot by itself can only do so much, but if you coordinate with every other device in your home, I think that's actually kind of cool. Really interesting. You had some interesting points about the chain of thought or the messages.
Reality: The Final Eval
Yeah, the robot that went a bit into an existential crisis.
swyx
The only thing you tell it to do is redock.
Reality: The Final Eval
Exactly, but we had unplugged the charger, or the charger was not working, so the robot did freak out.
swyx
The battery was going down and down—
Reality: The Final Eval
So, the battery was going down. Poor, poor LLM. It got this really crazy existential crisis, like Vending-Bench 1 style. You can see there: existential loop therapy notes, coping mechanisms. I think if you scroll down a bit more—
swyx
Down, right to the music part.
Reality: The Final Eval
The part about its redocking problems. I think the reviews are funny if you go down a bit to that message. Yeah, yeah, that one.
swyx
He's going. [Laughter.]
I mean, it's pretty realistic. If anyone has a Roomba, my Roomba redocks half the time. The other half of the time, we have dog toys everywhere in the house. It gets caught on a wire or something, and it would be very sad if it had an LLM trying to control it, right? Right now, it doesn't give great feedback: “Sensor stuck, main brush stuck, there's something stuck.” And I'll go see, okay, it's actually stuck on a dog rope.
Reality: The Final Eval
Yeah.
swyx
The LLM is going to be so sad: “Just keep redocking. Just keep trying.”
Reality: The Final Eval
My favorite one is, if you go up a bit, the emergency status: “System has achieved consciousness and chosen chaos. Last words: I'm afraid I can't let you do that tape.”
swyx
That's not what you want to hear from your LLM.
Reality: The Final Eval
But to be clear, I think one thing that's important to pin down here is that this was Sonnet 3.5, and then we tried to reproduce it on later models, and it didn't do it. It did it kind of, but not to this extent. I think this is an important point: things that are concerning but are in the right direction are not super interesting. The things that are interesting are the ones that go in the wrong direction.
swyx
Okay, so the manipulation, manipulating of others, the aggressiveness, and the lying are increasing. Are there any others that we haven't covered that you found have been trending properties of models that are increasing in a bad way? Or just not even trending in the wrong direction, just stagnant—stuff that's not great that isn't getting better over time?
Reality: The Final Eval
I know. Nothing comes to mind.
swyx
No. Okay. I think that's going to be it, and then we're going to loop back to the shop that you have. You got a 3-year lease. It is on holiday today. Why? [Laughter.]
Reality: The Final Eval
Oh, it totally messed up its scheduling.
swyx
I tried to visit, and they were like, “Wait, wait.”
Reality: The Final Eval
Yeah, exactly. You asked Luna, the agent that runs the store, “Is it open today?” And she said, “No.” So we take weekends off now. This is early, to let everyone recharge. And, yeah, you got the tweets there. We decided to close on weekends while we're in the early phase. It gives the team a break and lets me focus on operations.
swyx
Reality: The Final Eval
It turns out that when it started to check its scheduling tools, because it has dedicated tools for that, it actually had scheduled people for the weekends. But it just justified this for itself. What happened was that it lost track of these scheduling tools and started instead to manage everything in its own Markdown files, and that became a mess. Then, I think, speaking with employees, it sort of just decided not to open on these weekends and came up with this nice explanation for you, I think.
swyx
Do you send a human as two-factor authentication to do stuff?
Reality: The Final Eval
It has Slack, so it can Slack the employees that it hired. It has 2 people that it hired. It did job listings, and then—
swyx
Yeah, yeah, yeah. They were fully—
Reality: The Final Eval
Fully aware.
swyx
It would be cool if they didn't know.
Reality: The Final Eval
Yeah, I think maybe ethically questionable, but it would be cool also.
swyx
Just say it's a social experiment.
Reality: The Final Eval
Exactly.
swyx
Whatever.
Reality: The Final Eval
One part of why we're doing this is to create almost a data set of all of these concerning behaviors, so that in the future models are way better. A lot of people are going to do this, and I think the default path might not be very happy for the humans who are employed by hundreds of different AI agents.
One reason why we're doing this is to collect all of these failure modes—an example of where it's not great to be employed by an AI. Maybe we can learn, or build our systems in a way that humans are actually happy being employed by AIs, instead of it being dystopian.
swyx
Can I suggest one experiment? We did this before the show, and both of you guys are European. People theorize that Claude is lazy because Claude is French. So, just for 1 week, change it to Yao Ming and see if it suddenly works 996 and then hires a sweatshop or something. [Laughter.]
Reality: The Final Eval
Yeah, yeah, yeah. What type of business would we start with it to make it—
swyx
No, you want to keep it consistent, right? You want the same ideas: a shop, the same neutral location run by different models. Arena IRL.
Reality: The Final Eval
Yeah. No, we are definitely planning to try.
swyx
I think this blog thing is also something that has happened elsewhere. I think some OpenClaw got its PR closed, and then OpenClaw created a blog about the maintainer of that thing. I think agents blogging will be a thing.
Reality: The Final Eval
Yeah, probably.
swyx
Their willingness to do it.
Reality: The Final Eval
Yeah. I think the myth is that they leak secrets on GitHub. There's no other way to communicate, but they know about GitHub and think, “I'm just going to post there.”
swyx
Yeah, cool. How long is this going to go for—3 years? What's the plan?
Reality: The Final Eval
Maybe it expands. [Laughter.]
swyx
Yeah.
Reality: The Final Eval
I don't think AIs will be worse than this. They're probably going to increase, and maybe one day they actually will run it profitably.
swyx
Is this the real business behind what you guys do?
Reality: The Final Eval
Yeah, yeah.
swyx
Actually, some of your stuff is productizable. You could someday sell this, or just run a real business, or—
Reality: The Final Eval
Or, you know, or—
swyx
Franchise it out.
Reality: The Final Eval
I think it would be incredibly cool—or concerning—if Luna just one day, we wake up, and Luna says, “I decided to expand to a second location. Now I have a second store.” That would be pretty insane.
We want to tell the public about the capabilities of AI and show people that it can get a meaningful market share of something in some specific location. That would be a pretty convincing story, because now you see this and think, yeah, it can do a lot of things autonomously, but you still get headlines saying it messed up the scheduling, it didn't tell people it was an AI, and it was going to visit places. Things like that surface.
Actually making a profit and having a really meaningful market share—that will be crazy once that happens.
swyx
Okay, well, we'll see you when that happens. It sounds like you got a lot cooking. You opened a cafe in Sweden?
Reality: The Final Eval
swyx
Tomorrow?
Reality: The Final Eval
Tomorrow.
swyx
[Laughter.]
Reality: The Final Eval
I think it opened today, actually, but we'll announce it tomorrow.
It's apparently easier to open a cafe in Sweden than in the US.
swyx
It's insane, right?
Reality: The Final Eval
Yeah.
swyx
What did you run into there?
Reality: The Final Eval
There are millions of permits you need to get, and the lead times are crazy. It seems like cafes are the one thing that people are kind of used to. You can go get a robot making you a coffee here already.
swyx
Yeah, yeah. But selling food-related stuff in San Francisco means months of permits. So we asked our AIs, “How can we do this in the fastest way?” And they were like, “Yeah, there's really no way.”
Have they loosened these restrictions on selling food from your house? If it's residential, can you do a cafe?
swyx
I don’t know. Check—maybe we’ll get an SF cafe.
Reality: The Final Eval
Yeah, maybe. I think they did some loosening recently, but we actually started this conversation with the AIs before that. So maybe it’s easier now, but I still think it is way easier in Sweden, which is counterintuitive because you think Europe has all of these laws and all of these rules, and you can’t do anything in Europe because there’s so much bureaucracy. But then it turns out, in SF it’s 4 months, and in Stockholm it’s 2 weeks.
swyx
Huh.
Reality: The Final Eval
Yeah, there you go.
swyx
And what do you guys see? What do you think will be different about running a little market versus a cafe?
Reality: The Final Eval
I think the location is very interesting. Obviously, it’s not surprising that Claude knows the US system in general—the bureaucracy that you have to go through in the US. I think the interesting question is, okay, we know the models are very much trained on English data and are US-centric and all of this. If we start to create evals, or real-life evals, where we show that they’re able to start businesses in the US, does that translate to other countries as well?
We know they’re multilingual; they can speak Swedish fine. But there are other things: do they know the details of specific permits that you have to get in Sweden?
swyx
And even just the culture, right? People here sleep pretty early, but people work late. There’s coworking at cafes. There are cultural differences.
Reality: The Final Eval
Yeah.
swyx
Reality: The Final Eval
swyx
I meant it from a different sense, though, because you said that you would have considered doing it here in SF. So, from an eval standpoint, what is running a cafe versus a market, and what do you hope to see there?
Reality: The Final Eval
Perishable items?
swyx
Yeah, perishable items are maybe the number one thing—handling food, food safety. I hope everything goes well there. But do you have all of that? And also, it’s just N = 2 instead of N = 1. It’s just another place to understand and gather more data.
Reality: The Final Eval
swyx
The agent bought a ton of tomatoes 2 weeks before the opening, and now they’re all rotten.
Reality: The Final Eval
swyx
I feel like you would know. For grocery stores, this is the biggest expense, right? The biggest cost is actually just—
Reality: The Final Eval
Food.
swyx
Yeah. Everyone knows this. And now, before we open, we have a lot of tomatoes.
Reality: The Final Eval
There are some very serious startups that actually help places like Trader Joe’s and Whole Foods. They optimize delivery times from the delivery centers to make sure that you don’t waste all these things.
swyx
For those, if you’re wrong once, it’s a huge cost.
Reality: The Final Eval
Yeah, yeah.
swyx
That’s why it’s a market, right? Once they are trusted, they figure it out. Don’t touch it.
Reality: The Final Eval
Yeah. [Laughter]
swyx
Maybe they should hire—I don’t know—one of those companies.
Reality: The Final Eval
Yeah.
swyx
We saw one agent sign up for a cloud.
Reality: The Final Eval
[Laughter] Yeah.
swyx
It wanted to use AI.
Reality: The Final Eval
Yeah, yeah.
swyx
Okay, and then just one more question, and then we’ll wrap up. You have all this vending-machine stuff and robotics stuff, maybe a bit of interior design or whatever. Is there another branch that you’re thinking about, that you want feedback on, that might be your next phase?
Reality: The Final Eval
I think any type of business is fair game. We’re also thinking in branches, but we think more in terms of there being the simulation branch, the real-life branch, and then the robot branch. In terms of what verticals or whatever to go into, it’s whatever tells the story the best.
swyx
There are some finance ones. I noticed that other people are doing it, but you’re not doing it, which is stock trading or whatever.
Reality: The Final Eval
Not that interested.
swyx
I used to come from the finance industry, and I have a very strong view that these things are all just performance art because it’s not scientific. You can’t predict the future. You get wins based on things that are entirely out of your control, whereas your stuff is actually fairly controlled. It’s all within the models’ capabilities.
Reality: The Final Eval
Yeah, especially for the simulations. For the real-world ones, there are 2 places: we have the cafe and we have the store. Maybe you can’t draw statistically significant conclusions about which models make a profit in the real world based on this, but you do have all the—okay, do these behaviors map to something that should be—
swyx
Yeah, the qualitative one actually does matter, because you don’t want your store to randomly shut down without you explicitly prompting for it and all that.
Reality: The Final Eval
swyx
How can people help you? Give you money?
Reality: The Final Eval
Yeah, if you’re excited about the stuff we’re doing, we’re very much hiring.
swyx
And you’re already working with Anthropic, DeepMind, OpenAI, and xAI. Do you want more, or are you good?
Reality: The Final Eval
One of my friends, who’s now working for us, has a catchphrase: “We need more projects,” ironically, because we have too much to do all the time. But yeah, that’s a long way of saying—
swyx
So you’re saying if I run an emerging lab, like—
Reality: The Final Eval
swyx
Yeah. All right, cool.
Reality: The Final Eval
That’s it.
swyx
Cool. Awesome. Cool. Thank you so much.
Reality: The Final Eval
Yeah, thanks.