Alessio Fanelli
You don't write code. You talk to an agent, and it goes and does it for you. At best, you review it. That's probably largely not even what you're doing.
What's happening is that we're changing our work to make the agents effective in that model. The agent didn't really adapt to how we work; we adapted to how the agent works. All of the economy has to go through that exact same evolution. Right now, it's a huge asset and an advantage for the teams that do it early and are wired into doing this, because you'll see compounding returns. But it's going to take a while for most companies to actually get this deployed.
We're back in the Chroma studio with Chroma CEO Jeff Huber. Welcome. Returning guest, but now guest host.
swyx
It's a pleasure.
Alessio Fanelli
Wow. How did you get upgraded to that? He's the perfect guy to be guest host for you.
swyx
That makes sense, actually. You love context. We both really love Context.
Alessio Fanelli
We really do. And we're here with Aaron Levie. Welcome.
Every Agent Needs a Box
Thank you. Good to be here.
Alessio Fanelli
We've all met offline and chatted a little bit, but it's always nice to have these conversations in person. You just started off with so much energy. You're super excited about agents.
Every Agent Needs a Box
I love agents.
Alessio Fanelli
Yeah. OpenClaw just got bought by OpenAI. Well, not bought, but you know what I mean—some sort of acquihire.
Speaker 3
Executive hire.
Alessio Fanelli
Executive hire. Okay.
Executive hire. Hey, that's my term. (Laughter.) What are you pounding the table on with agents? You have so many insightful tweets.
Speaker 3
The thing that we get super excited about, which I think should be relatively obvious, is that we've built a platform to help enterprises manage their files—their corporate files, the permissions governing who has access to those files, and the sharing and collaboration around them.
All those files contain really important information for the enterprise. They might have your contracts, research materials, marketing information, or memos. All that data has predominantly been used by humans, but there's been one really interesting problem: humans only really work with their files during an active engagement with them. Then they go away, and you don't really see them for a long time.
All of a sudden, with the power of AI and AI agents, all that data becomes extremely relevant as an ongoing source of answers to new questions and data that can transform into something else that produces value in your organization. It contains the answer for the new employee who's onboarding and needs to ramp up on a project. It contains the answer to the right thing to sell a customer when you're having a conversation with them. It contains the roadmap information that's going to produce the next feature.
All that data that we previously just stored and occasionally forgot about because we were only working on the new, active stuff becomes valuable to the enterprise. It's going to become extremely valuable to end users because now they can have agents go find what they're looking for and produce new value and new data from that information.
It's also going to become incredibly valuable to agents because agents can roam around and do a bunch of work, and they're going to need access to that data as well. Sometimes that will be an agent working on behalf of you—effectively as you—accessing all the same information you have access to and operating as you in the system.
Other times, there are going to be agents that are effectively autonomous and run on their own. You'll collaborate and work with them kind of like you would with another person. OpenClaw is the most recent, and maybe the first real, version of what that could look like that has updated everybody's view of this landscape: I have an agent on its own system, on its own computer, with access to its own tools. I probably don't give it access to my entire life. I communicate with it like I would with an assistant or a colleague, and it has this sandbox environment.
All of that has massive implications for a platform that manages enterprise data. We think it's going to transform how we work with all of the enterprise content we work with, and we just have to make sure we're building the right platform to support that.
Speaker 2
The shorthand I put it is: as people build agents, everybody's just realizing that every agent needs a box.
Speaker 3
Yes.
Speaker 2
It's nice to be called Box and just give everyone a box. (Laughter.) If we can make that go viral, I think that terminology—
Alessio Fanelli
The tag: “Every agent needs a box.”
Speaker 2
Every agent needs a box. If we can make that the headline of this, I'm fine with it.
Alessio Fanelli
That's the billboard.
Speaker 2
Exactly. Every agent needs a box. I like it. Can we ship this?
Alessio Fanelli
My work here is done. (Laughter.) I got the value I needed out of this podcast.
Speaker 3
The thing that we think about is that, whether you think the number is 10× or 100× or whatever it is, we're going to have some order of magnitude more agents than people. That's inevitable. It has to happen.
The question is: What infrastructure is needed to make all those agents effective in the enterprise? How do you make sure they're well governed, that they're only doing safe things with your information, and that they're not being exposed to data they shouldn't have access to?
There are going to be spectacularly crazy security incidents involving agents because you'll prompt-inject an agent and find your way through the CRM system to pull out data that you shouldn't have access to. It's just going to happen all over the place.
So how do you make sure you have the right security, permissions, access controls, and data governance? We don't yet know exactly how we're going to regulate some of these agents. If you think about an agent in financial services, does it have the exact same financial requirements as a human, or is the risk fully on the human who was interacting with or created the agent? Those are all open questions.
No matter what, there is going to need to be a layer that manages the data they have access to, the workflows they're involved in, and pulling up data from multiple systems. This is the new infrastructure opportunity in the era of agents.
Speaker 2
You have a piece on agent identities, which I think was today. A lot of the security people are talking about it right now. I always think of it as: You need the human you, and then you need the agent you. I don't know if it's that simple, but is Box going to have an opinion on that, or are you just going to be the storage layer and let Okta or Cerbos handle that?
Speaker 3
I think we're going to have an opinion, and we'll work with wherever the contours of the market end up. The reason we're going to have an opinion, probably more than on other topics, is because one of the biggest use cases for why your agent might need its own identity is file-system access. Thus, we have to think about this pretty deeply.
Unless you're in our world, thinking about this particular problem all day long, you might wonder, “Why is this such a big deal?” Sometimes people say, “Just give the agent an account on the system and treat it like every other type of user on the system.”
The problem is that I, as Aaron, don't really have any responsibility over anybody else's Box account in our organization. I can't see the Box account of any other employee I work with. I'm not liable for anything they do, and they have strict privacy requirements around everything they're able to work on.
Agents don't have those properties. The person who creates the agent is probably going to take on a lot of the liability for what that agent does, at least for the foreseeable future. The agent doesn't deserve any privacy because it can't be fully autonomously operated, and it doesn't have any legal responsibility.
Thus, you can't just say, “I'll create a bunch of accounts, work with those agents, and talk to them occasionally.” You need oversight of that. The question is, how do you have a world where you sometimes have oversight of an agent, but what if that agent goes and works with other people? If someone else is collaborating with the agent on something, you shouldn't have access to what they're doing.
We have all these new boundaries that we're going to have to figure out. So far, we've been in easy mode. We've hit the easy button with AI: The agent is just you. When you're in Claude Code, Cursor, and Codex, you're the agent. You're authenticating into your services, and it can do everything you can do. That's the easy mode.
Every Agent Needs a Box
The hard mode is agents kind of running on their own. People check in with them occasionally. They’re doing things autonomously. How do you give them access to resources in the enterprise without dramatically increasing the security risk and the risk that you might expose the wrong thing to somebody? These are all the new problems that we have to solve.
I like the identity layer and identity vendors as being a solution to that. But we’ll need some opinions as well, because so many of the use cases are these collaborative file system use cases. How do I give an agent a subset of my data and give it its own workspace as well? It’s going to need to store its own information that would be relevant for it, and how do I have the right oversight into that?
swyx
One thing which I think is kind of what you think about is that you know how humans work, right? I may not just give you access to the whole file. I might sit next to you and scroll to one part of the file [laughter] and just show you that one part.
Every Agent Needs a Box
Partial file access.
swyx
Well, I’m just saying, I think RBAC does seem to be dead, right? If you want to say something is dead, probably RBAC is dead. The OAuth story to me seems incredibly unsolved and unaddressed by the existing state of AI vendors.
Every Agent Needs a Box
Yeah, I think we’re taking this to an extreme that we probably need to solve for. We built an access control system that was kind of its own little world for a long time. The idea was this: It’s a many-to-many collaboration system where I can give you any part of the file system, and it’s a waterfall model. If I give you something higher up in the system, you get everything below it.
That created immense flexibility, because I can point you to any layer in the tree, but then you’re going to get access to everything below it. That mostly is working in this world. But you do have to manage this issue: How do I create an agent that has access to some of my stuff and somebody else’s stuff as well, and which parts do I get to look at as the creator of the agent?
These are just brand-new problems. When there was a human there, that was really easy to do. If the 3 of us were all sharing, there’d be a Venn diagram where we’d have an overlapping set of things we’d shared, but then we’d have our own ways that we shared with each other. But in an agent world, somebody needs to take responsibility for what that agent has access to and what it’s working on.
These are some of the most boring problems for 98% of people on the internet, but they will be the problems that make the difference between whether you can actually have autonomous agents in an enterprise context that are not leaking your data constantly.
swyx
No, I mean, I run a very, very small company for my conference, and we already have data-sensitivity issues. Some of my team members cannot see the others, and I can’t imagine what it’s like to run a Fortune 500 and have to worry about this. I’m just kind of curious: You talk to a lot of companies. Are 70% or 80% of the Fortune 500 your customers?
Every Agent Needs a Box
Yep. 67%.
swyx
Just something I’m rounding up.
Every Agent Needs a Box
Yes, I appreciate the rounding. We’re projecting to the end of the year.
swyx
Thank you. There you go. You do make it sound like—we’ve got to be honest—we’re taking way too long to get to 80%.
Every Agent Needs a Box
Well, no. I mean, this is actually the stark reality that unfortunately pours a little water on the party. We all in Silicon Valley have the absolute best conditions possible for AI ever. I think we all saw the Dwarkesh podcast with Dario and this idea of AI coding. Why has that taken off, and why are we not yet fully seeing it everywhere else?
If you just enumerated the list of properties that AI coding has and compared it to other knowledge work, let’s go through a few of them. Generally speaking, when you bring on a new engineer, they have access to a large swath of the codebase. There’s very little friction: A new engineer comes on, and they can find the stuff they need to work with. It’s a fully text-in, text-out medium. It’s just going to be text at the end of the day, so it’s really great in terms of what the agent can work with. Obviously, the models are super-trained on that data set.
The labs themselves have a really strong, self-reinforcing positive flywheel around why they need to do agentic coding deeply. So then you get better tooling and better services. The actual developers of the AI are daily users of the thing they’re working in, versus there are probably only 7 Claude Cowork legal plug-in users at Anthropic on any given day, while there are a couple thousand Claude Code users every single day. Think about which one they’re getting more feedback on all day long.
You just go through this list. Everybody who’s a developer is, by definition, technical. They can go install the latest thing. We’re all generally online—or at least the weird ones are—and we’re all talking to each other and sharing best practices. That’s already 8 differences versus the rest of the economy. Every other part of the economy has 6 to 7 headwinds relative to that list.
You go into a company and you’re a banker in financial services. You have access to a tiny little subset of the total data that’s relevant to doing your job, and you have to start talking to a bunch of people to get the right data. Sally didn’t add you to that deal-room folder, and the information is actually in a completely different organization that you now have to go and sort of run into. You have this endless list of access controls and security, as you talked about.
You have a medium that is not just text, right? You have a Zoom call where you’re getting all of the requirements from the customer. You have a lot of in-person conversations, and you’re doing in-person sales. How do you ever digitize all of that information?
I think a lot of people got upset with this idea that the codebase has all the context. I don’t know if you followed some of that conversation that went viral: It’s not that simple; the codebase doesn’t have all the knowledge. But you’re a lot better off than you are with other areas of knowledge work. We have documentation practices and write specifications. Those things don’t exist for 80% of the work that happens in the enterprise.
That’s the divide that we have. AI coding has fully reached escape velocity in terms of how powerful this stuff is, and then we’re going to have to find a way to bring that same energy and momentum to all these other areas of knowledge work, where the tools aren’t there, the data isn’t set up to be there, and the access controls don’t make it that easy.
Context engineering is an incredibly hard problem because, again, you have access-control challenges. You have different data formats. You have end users who are going to need to be trained through this, as opposed to adopting these tools in their free time. That’s where the Fortune 500 is. We have to be prepared as an industry for a multi-year march to bring agents to the enterprise for these workflows.
And I think probably the thing we’ve learned most in coding, which the rest of the world isn’t yet ready for—I mean, they’ll have to be ready for it because it’s inevitably going to happen—is this: If you think about the practice of coding today versus 2 years ago, it’s probably the most changed workflow in maybe the history of time, in terms of how quickly it has changed.
Has any workflow in the entire economy changed that quickly in terms of the amount of change? At least in any knowledge-worker workflow, there’s very rarely been an event where 1 piece of technology and 1 work practice has so fundamentally changed what you do. You don’t write code; you talk to an agent, and it goes and does it for you. At best, you review it, and even that is probably largely not what you’re doing. What’s happening is we are changing our work to make the agents effective in that model. The agent didn’t really adapt to how we work. We basically adapted to how the agent works.
swyx
All of the economy has to go through that exact same evolution. The rest of the economy is going to have to update its workflows to make agents effective and to give agents the context they need, to figure out what kind of prompting works, and to figure out how to ensure that the agent has the right access to information to execute on its work.
This is not the panacea people were hoping for, where the agent drops in and automates your life. You have to basically re-engineer your workflow to get the most out of agents, and that’s just going to take multiple years across the economy.
Every Agent Needs a Box
Right now, it's a huge asset and an advantage for the teams that do it early and are kind of wired into doing this because you'll see compounding returns. But that's just going to take a while for most companies to actually get deployed.
swyx
I love pushing back. I think that a lot of technology consultants love to hear this sort of thing, right? To embrace AI and get to the promised land, you must pay me so much money to adopt the prescribed way of conforming to the agents. I worry that you will be eclipsed by someone else who says, “No, come as you are, and we'll meet you where you are.”
Alessio Fanelli
And what was the thing that went viral a week ago? OpenAI is probably hiring FDEs to go into the enterprise, and Anthropic is embedded at Goldman Sachs.
swyx
So if the labs are having to do this, if the labs have decided that they need to hire FTEs and professional services, then I think that's a pretty clear indication that there's no easy mode for workflow transformation. To your point, I think this is actually a market opportunity for new professional services and consulting firms that are like agent-build firms. They go into organizations, figure out how to re-engineer your workflows to make them more agent-ready, get your data into the right format, and reconstruct your business process.
You're not doing most of the work. You're telling agents how to do the work, and then you're reviewing it. But I haven't seen the thing that can just drop in and let you avoid those changes. I don't know how that kind of sales pitch goes over. You're saying things like, “Well, in my nice, beautiful walled garden, here's this beautiful Box account that has everything.” And I'm like, “Well, most real life is extremely messy, poorly named, and outdated.”
Every Agent Needs a Box
100%. I mean, we agree that getting to the beautiful garden is going to be tough.
There's also the other end of the spectrum, where I just think it's a technical impossibility to solve. The agent truly cannot get enough context to make the right decision in an incredibly messy environment. There's no AGI that will solve that.
So we're going to have to land somewhere in between, where we all collectively get better at documentation practices, having authoritative, relatively up-to-date information, and putting it in the right place. Agents will certainly cause us to be much better organized around how we work with our information, simply because the severity of an agent pulling the wrong data will be too high. The productivity gain you'll miss out on by not doing this will be too high as well. Your competition will just do it, and they'll have higher velocity.
We see this a lot firsthand. We build a series of agents internally that can have access to your full Box account. You give it a task, and it can go find whatever information you're looking for and work with it.
Thank God for the progress in models, but if you gave that task to an agent 9 months ago, you'd just get lots of bogus answers. It would say, “Here are 5 documents that all kind of smell like the right thing.” But you're putting it on the clock because its system prompt says, “Be pretty smart, but also try to respond to the user,” and it's going to respond. Then you're like, “Ah, it got the wrong document.”
Alessio Fanelli
Yeah. Doesn't work.
Every Agent Needs a Box
It doesn't work. Opus 4.6, Gemini 3.1 Pro, and whatever the latest GPT-5.3 will be are getting better and better. They're using better judgment.
All of these updates to the agentic tool-use and search systems are showing very real progress. The agent can almost smell that something is a little fishy when it's getting an answer. We have this process where we have it fan out, do a bunch of searches, pull up a bunch of data, and then it has to do its own ranking of the right documents that it should be working with.
The intelligence level of a model 6 months ago would just be throwing a dart: “I'm going to grab these 7 files, and I hope that's the right answer.” Something like Claude Opus 4.5, and now Claude Opus 4.6, is like, “No, that one doesn't seem right relative to this question because I'm seeing some signal that's contradicting the document, where it would normally be, and who should have access.” It's doing all that kind of work for you.
But it still doesn't work if you just have a total wasteland of data. It's just not possible, partly because a human wouldn't even be able to do it. Basically, if a really, really smart human could not do that task in 5 or 10 minutes for a search-retrieval-type task, your agent is not going to be able to do it any better.
Alessio Fanelli
You see this all day long.
swyx
This touches on a thing that I'm passionate about, which is context engineering. I'm just going to let you ramble or riff on context engineering, if there's anything. You did really good work on context rot, which has really taken over as the term that people use and reference.
Every Agent Needs a Box
100%. We all think about the context-rot problem.
Alessio Fanelli
Yeah, there's certainly a lot of ranking considerations. Agentic search, I think, is incredibly promising. I was trying to generate a question, though. I think I have a question right now.
Every Agent Needs a Box
I think there was this moment 2 years ago, before we knew where the gotchas were going to be in AI, when someone was like, “Infinite context windows will just solve all of these problems, because you'll just give the context window all the data.” It's like, okay, maybe in 2035 this is a viable solution.
First of all, it would simply cost too much. We just can't give the model the 5,000 documents that might be relevant and have it read them all. I've seen enough to start believing in crazy stuff, so I'm willing to say, sure, 10 years from now we'll have infinite context windows at a thousandth of the price of today. Let's believe that's possible. But we're in reality today.
Today we have a context-engineering problem. I've got 200,000 tokens that I can work with—or I don't even know what the latest graph is before massive degradation. Okay, I have 60,000 tokens that I get to work with where I'm going to get accurate information.
That's not a lot of tokens for a corpus of 10 million documents that a knowledge worker might have across all the teams, projects, and people they work with. I have 10 million documents, which maybe is 5 pages per document or something like that. I'm at 50 million pages of information, and I have 60,000 tokens.
How do I bridge the 50 million pages of information with the couple hundred that I get to work with in that token window? This is such an interesting problem, and that's why so much work is actually just search systems and databases. That layer has to get so locked in.
Models are getting better, and importantly, they're getting better at knowing when they've done a search and found the wrong thing. They go back, check their work, and find a way to balance appeasing the user versus double-checking.
We have this one test case where we ask the agent to go find 10 pieces of information.
Alessio Fanelli
Is this a complex-work eval?
Every Agent Needs a Box
This is actually not an eval. This is just a set of internal benchmark scenarios we have every time we update our agent. We have one where I ask it to find all of our office addresses, and I give it the list of 10 offices that we have.
There's not 1 document that has this. Maybe there should be. That would be a great example of the kind of thing that, over time, companies start to have: these canonical key areas of knowledge that we need to have. We don't seem to have this 1 document that says, “Here are all of our offices.” We have a bunch of documents that have, like, “Here's the New York office,” and whatever.
Alessio Fanelli
So you ask this agent, and you say, “I need the addresses for these 10 offices.” Okay?
Every Agent Needs a Box
By the way, if you do this on any public chat model, the same outcome is going to happen, but for a different kind of query. You say, “I need these 10 addresses.” How many times should the agent go and do its search before it decides whether there's just no answer to this question?
Often, especially with the lower-tier models, it'll come back and give you 6 of the 10 addresses and just say, “I couldn't find the other 4.”
Alessio Fanelli
It doesn't know what it doesn't know.
Should it just keep reading every single file in your entire Box account until it exhausts every single piece of information?
swyx
Expensive.
Alessio Fanelli
These are the new problems that we have. So, let's say a new Opus model is like, “Okay, I’m going to try these types of queries. I didn’t get exactly what I wanted. I’m going to try again.” At some point, it’s going to stop searching because it’s determined that no amount of searching is going to solve this problem. It’s just not able to do it.
That judgment is a really new thing that the model needs to be able to have. When should it give up on a task? You just don’t—it can’t find the thing. That’s the real world of knowledge-work problems.
This is the stuff that coding agents don’t have to deal with, because you’re not usually asking them about existing information; you’re always creating net-new information coming right out of the model, for the most part. Obviously, they have to know about your codebase, your specs, and your documentation, but when you deploy an agent on all of your data, you now have all of these new problems that you’re dealing with.
swyx
Our follow-up research to Context Rot is actually on search. We’ve stress-tested frontier models and their ability to search, and they’re not actually that good at searching, right?
Alessio Fanelli
So, you’re highlighting this explore-exploit trade-off: not everything works. [laughter]
swyx
Well, somebody has to. Can I throw out one more thing that is different from coding and the rest of knowledge work that I failed to mention? One other key point is that—
Every Agent Needs a Box
At the end of the day, whether you believe we’re in a slop apocalypse or whatever, if you’ve built a working solution, that is ultimately what the customer is paying for. Whether I have a lot of slop, a little slop, or whatever, I’m sure there are lots of codebases we could go into in enterprise software companies where it’s just crazy slop that humans created over a 20-year period. But the end customer just gets this little interface. They can type into it, and it does its thing.
Knowledge work doesn’t have that property. If I have an AI model generate a contract, and I generate a contract 20 times, and all 20 times it’s just 3% different, that kind of slop introduces all new kinds of risk for my organization that the code version of that slop didn’t introduce.
How do you constrain these models to just the part that you want them to work on and just do the thing that you want them to do? In engineering, you can’t be disbarred as an engineer, but you could be disbarred as a lawyer. You can do the wrong medical thing in healthcare. There’s no equivalent to that in engineering.
swyx
You want there to be, because I’ve considered software—
Alessio Fanelli
Oh, is that—
swyx
Civil engineering. There is, right?
Alessio Fanelli
Civil engineering. Sure. Oh, yeah, for sure. But in any of our companies, you’ll be forgiven if you took down the site. We’ll do a rollback, and you’ll be in a meeting, but you have not been disbarred as an engineer. We don’t change your computer science degree.
swyx
Yeah, exactly. So, now maybe we collectively, as an industry, need to figure out what you’re liable for—not legally, but in a management sense—with these agents. All sorts of interesting problems have to come out.
In knowledge work, those are the real hostile environments that we’re operating in. I do think a lot of last year’s 2025 story was the rise of coding agents, and I think the 2026 story is definitely knowledge work. OpenClaw and Claude Cowork are just the beginning. The next one is going to be absolute craziness.
Alessio Fanelli
It is, and it’s going to be—again, this is going to be a wave where we try to bring as many of the practices from coding as possible, because that will clearly be the forefront: tell an agent to go do something, give it access to a set of resources, and be responsible for reviewing it at the end of the process.
That, to me, is the template that goes across knowledge work. Claude Cowork is a great example. OpenClaw is a great example. You can sort of see what Codex could become over time. These are some really interesting platforms that are emerging.
swyx
Okay. We touched on evals a little bit. You had the report that you were going to bring up, and then I was going to go into Box’s evals, but go ahead and talk about your agentic-search thing.
Alessio Fanelli
Yeah. Mostly, I think the insight is that one frontier model is not good at search. Humans have this natural explore-exploit trade-off where we understand when to stop doing something.
Humans are also pretty good at forgetting, actually, and pruning their own context, whereas agents are not. In an agent’s context history, if it knew something was bad—and even if you can see in the reasoning trace that it probably wasn’t a good idea—if it’s still in the trace, it’s still in the context, and it’ll still do it again.
I think pruning is going to be a really big thing. It’s already becoming a thing, right? Letting models self-prune their context windows.
swyx
So, don’t leave the mistake in there.
Alessio Fanelli
Cut out the mistake, but tell it that it made a mistake in the past so it doesn’t repeat it.
swyx
Yeah, but cut it so it doesn’t get distracted by it again, because it will repeat its mistake just because it’s been in the context so much, even if it knows. It’s like, “Oh, this is a great thing to go try,” even if it knows it didn’t work. Yeah, exactly. So, there’s a bunch of stuff there.
Alessio Fanelli
Groundhog Day inside these models.
swyx
I’m going to keep doing the same wrong thing every time. You’re trying to fit a manifold in latent space, which is kind of what we’re doing, right? Program synthesis is one thing we’re doing. Certain facts might be overly pinning it to certain sectors of latent space.
And so—[music]—plug Latent Space. Our editor adds a bell every time you say that. You have to remove those links to give it the freedom to do what it needs to do. But, yeah, release more soon.
Alessio Fanelli
That’s awesome. Yeah, that’ll be cool.
swyx
We’re a cerebral podcast. People listen to us and think really deeply, so we try to keep it subtle.
Alessio Fanelli
Okay, fine.
swyx
You guys have talked about your Box thing, but you’ve also been promoting APEX agents and complex work. Wherever you want to take this, just how do you—
Alessio Fanelli
APEX is obviously in our corpus of agent evals. We supported that by opening up some data for them around how we see these data workspaces in the regular economy. How do lawyers have a workspace? How do investment bankers have a workspace? What kind of data goes into those? We partnered with them on their APEX eval.
Our own eval is actually relatively straightforward. We have a set of documents in a range of industries. We previously did this as a one-shot test of just the model, and then we realized that, based on where everything’s going, it’s got to be more agentic.
Now it’s more of a test of both our harness and the model. We have a rubric of a set of things that it has to get right, and we score it. You’re seeing these incredible jumps in almost every single model in its own family—Opus, Sonnet 4.6 versus Sonnet 4.5.
swyx
Yeah, we have this up on screen.
Alessio Fanelli
Okay, cool. It was like a 15-point jump, I think, on the overall score.
swyx
Yes.
Alessio Fanelli
Anthropic doesn’t know anything about it; it’s completely held out from Anthropic. This isn’t in any public data, which has its benefits. This is just a private eval that we do, and then we happen to show it to the world. You can’t train against it, and I think it’s just as representative of its reasoning capabilities, what it’s doing at test-time compute, thinking levels, and all the context-rot issues—so many interesting capabilities that are now improving.
One sector that you have that’s interesting—people are roughly familiar with healthcare and legal, but you have the public sector in there.
swyx
What’s that? What is that?
Alessio Fanelli
Yeah, and we actually test against maybe 10 industries. We usually end up just highlighting a few that we think have interesting gains. Public sector is one: a lot of government-type documents.
swyx
What are government-type documents? Government filings, tax returns—
Alessio Fanelli
Probably not tax returns. [laughter] It would be more what the government would be using as data. Think about research, that type of dataset. And then we have financial services for things like data rooms and what would be in an investment prospectus.
swyx
That one you can dog-food.
Every Agent Needs a Box
Yeah. Exactly. Exactly. Yes. [laughter] So we run the models now in more of an agent mode, but still with limited capacity, and just try to see, on a like-for-like basis, what the improvements are. Again, we just continue to be blown away by how good these models are getting.
swyx
Yeah. I think every serious AI company needs something like that: “This is the work we do; here’s our company eval.” If you don’t have it, you’re not a serious AI company.
Every Agent Needs a Box
There are 2 dimensions, right? There’s how the models are improving—which model you should recommend a customer use and which one you should adopt—but then every single day we’re making changes to our agents, and you need to know—
swyx
If you regress—yeah, I’ve been fully convinced that the whole agent observability and eval space is going to be a massive space. I’m super excited for what Braintrust is doing, excited for LangSmith, all the things.
I think what you’re going to see—I mean, this is literally every enterprise. Right now, the AI companies are the customers of these tools. Every enterprise will have this. You’ll just have to have an eval of all of your work. You’ll have an eval of your RFP generation, an eval of your sales-material creation, and an eval of your invoice processing.
As you buy or use new agentic systems, you’re going to need to know: What’s the quality of your pipeline? Yeah.
Every Agent Needs a Box
Huge, huge market with agent eval.
swyx
Yeah. I’m going to shout out your team a bit. Your CTO, Ben, did a great talk with us last year, and he’s going to come back again for World’s Fair. Just talk about your team—brag a little bit. I think people take these eval numbers and pretty charts for granted, but there are lots of really smart people at work doing all this.
Every Agent Needs a Box
Yeah.
swyx
Biggest shout-out?
Every Agent Needs a Box
The biggest shout-out is that we have a couple of folks, Ditya and Siddharth, who kind of run this. They’re a tag-team duo on our evals. Ben, our CTO, is heavily involved; Yash, head of AI; and a bunch of other folks.
Eval is one part of the story, and then the full AI agent team is core to this whole effort. There are probably a few dozen people who are the epicenter, and then you just have layers and layers of concentric circles. There’s a search team that supports them, and an infrastructure team that supports them, and it’s starting to ripple through the entire company. But there’s that core agent team that’s a pretty close-knit group.
swyx
The search team is separate from the infrastructure team?
Every Agent Needs a Box
I mean, we have every layer of the stack that we have to do, except for pure public cloud. We store—I don’t even know what our public numbers are—but you can just think about it as a lot of data being stored in Box.
We have every layer of the stack: How do you manage the data, the file system, the metadata system, the search system—all of those components? They all have to understand that now you’ve got this new customer, which is the agent.
They’ve been building for 2 types of customers in the past: They’ve been building for users, and they’ve been building for applications. Now you’ve got this new agent user, and it comes in with different sub-properties sometimes. Maybe sometimes we should do embeddings—an embedding-based search—versus your typical semantic search.
You have to build the capabilities to support all of this. We’re testing stuff, throwing things away. If something doesn’t work and isn’t relevant, we just move on. It’s total chaos, but all of those teams are supporting the agent team, which is coming up with its requirements: What do we need?
swyx
Yeah. We just came from a fireside chat where you talked about how you’re doing this. It’s kind of like an internal startup within the broader company. The broader company is about 3,000 people, but there’s this core team of—well, here’s the innovation center, and every company is kind of run this way.
Every Agent Needs a Box
I want to be sensitive. I don’t call it the innovation center only because I think everybody has to do innovation. There’s a part of the company that is sort of do-or-die for the agent wave.
swyx
Yeah.
Every Agent Needs a Box
It only happens to be more of my focus simply because it’s existential that we get it right.
swyx
Yeah.
Every Agent Needs a Box
All of the supporting systems are necessary. All of the surrounding, adjacent capabilities are necessary. The only reason we get to be a platform where you’d run an agent is because we have a security feature, a compliance feature, or a governance feature that some team is working on.
But that’s not going to be the make-or-break of whether we get agents right. That already exists, and we need to keep innovating there. I don’t know what the exact, precise number is, but it’s not 1,000 people and it’s not 10 people.
There’s a number of people who are the startup within the company—the make-or-break team for everything related to AI agents leveraging our platform and letting you work with your data. That’s where I spend a lot of my time. Ben, Yash, Diego, and Terry are just some of the people across the team who are working on this.
swyx
Yeah. Amazing. How do you think about read workflows over your Box data—generative search, questions, queries, and so on—but what about write or authoring workflows?
Every Agent Needs a Box
Yes, I’ve probably revealed too much, actually, now that I think about it. I guess I would just make it a little bit conceptual because I’ve already said things that aren’t even generally available, but we’ve kind of danced around them publicly.
Hopefully nobody watches this. [laughter] These are tidbits for the highly engaged to go figure out exactly what our line of thinking is.
swyx
Yeah.
Every Agent Needs a Box
I would say that, as a place where you have your enterprise content, there’s a use case where I want to have an agent read that data and answer questions for me. Then there’s a use case where I want the agent to create something, use the file system to create something, store data that it’s working on, or have various files that it’s writing to about the work it’s doing.
We do see it as a total read-write problem. The harder problem so far has been read-only, because you have that 10-million-to-1 ratio problem. Writes are a lot easier; that’s just going to come from the model, and we’ll put it in the file system and use it.
It’s a technically easier problem. The part that isn’t necessarily technically hard—it’s just not yet perfected in the ecosystem—is building a beautiful PowerPoint presentation. That’s still a hard problem for these models. These formats were just not built for—
swyx
They’re working on it.
Every Agent Needs a Box
They’re working on it. Everybody’s working on it.
swyx
Everybody launches like, “Well, we do PowerPoint now.”
Every Agent Needs a Box
We’re getting a lot better each time. But then you’ll do this thing where you ask it to update 1 slide, and all of a sudden the fonts will be just a little bit different on 2 of the slides, or it moved some shape over to the left a little bit.
These are the kinds of things that, in code, you might not really care about if you care about how beautiful the code is. The end user doesn’t notice all those problems. In file creation, the end user instantly sees it. You’re like, “Yeah, but paragraph 3—you literally just changed the font on me. It’s a totally different font midway through the document.”
Those are the kinds of things that you run into a lot on the content-creation side.
We are going to have native agents that do all of those things. They’ll be powered by the leading models and labs. But the thing that I think is probably going to be a much bigger idea over time is any agent on any system using Box as a file system for its work.
In that kind of scenario, we don’t necessarily care what it’s putting in the file system. It could put its memory files, its specification documents, whatever its Markdown files are, or it could generate PDFs. It’s just a workspace that’s sandboxed off for its work.
People can collaborate in it. It can share with other people. We’re thinking a lot about the right way to deliver that at scale.
swyx
I wanted to come into the AI transformation or AI operations side of things. One of the tweets that you wanted to talk about—this is just me going through your tweets, by the way.
Every Agent Needs a Box
Okay.
swyx
I mean, this is me reading them one by one. [laughter] You’re the easiest guest to prep for because you already have, like, “This is what I’m interested in.”
I'm like, okay, well—
Every Agent Needs a Box
Are we going to get to February, January, or something? Where are we in the timeline? How far back are we going?
swyx
Can you describe Box's set of skills? That's one of the extremes: if you just turn everything into a Markdown file, then your agent can run your company.
Every Agent Needs a Box
Like, you just have to write the right sequence of words to—
swyx
Yes.
Every Agent Needs a Box
—to do it.
swyx
Oh, sorry. Is that—
Every Agent Needs a Box
So I think the question is: What if we documented everything the way that you said? Let's get all the Fortune 500s prepared for agents, and everything's in golden, nicely filed away, and everything. What's missing? What's left? You've run your company for a decade, like—
swyx
I think the challenge is that that information changes a week later.
Every Agent Needs a Box
And because something happened in the market for that customer or us as a company, that now has to go get updated. These systems are living and breathing, and they have to experience reality and updates to reality, which right now is probably going to be humans giving them the updates. There is this piece, “Context Graphs,” that was kind of very viral. I thought it was super provocative. I agreed with many parts of it, and I disagreed with a few parts around—it’s not going to be as easy as if we just had the agent traces, then we can finally do that work, because there’s so much other stuff happening that we haven’t been able to capture and digitize.
I think they actually represented that in the piece, to be clear. But there’s a lot of work. You just can’t have only skill files for your company, because there’s going to be a lot of other stuff that happens—
swyx
—change over time.
Every Agent Needs a Box
Yeah.
swyx
Most companies are practically apprenticeships. Like every new employee who joins the team, you spend 1–3 months ramping them up. Yes.
Every Agent Needs a Box
All that tacit knowledge is not written down.
swyx
Yes.
Every Agent Needs a Box
But it would have to be if you wanted to give it to an agent, right? So that seems to me like—
swyx
One is, I think you’re going to see, again, a premium on companies that can document this much. There’ll be a huge premium on that, because can you shorten that 3-month ramp cycle to a 2-week ramp cycle? That’s an instant productivity gain.
Can you dramatically reduce rework in the organization because you’ve documented where all the stuff is and where the answers are? Can you make your average employee as good as your 90th-percentile employee because you’ve captured the knowledge that’s in the heads of those top employees and made that available?
So you can see some very clear productivity benefits if you had a company culture of making sure your information was captured, digitized, put in a format that was agent-ready, and then made available to agents to work with. Then you have this reality that, at a 10,000-person company, mapping that to the access structure of the company is just a hard problem. Not every piece of information that’s digitized can be shared with everybody, so now you have to organize that in a way that actually works.
Every Agent Needs a Box
There was a pretty good piece called “Your Company as a Filesystem.” I don’t know, did you see that one?
swyx
Nope.
Alessio Fanelli
Yes, you saw it. Yeah.
Every Agent Needs a Box
I’d actually be curious about your thoughts on it. It’s an interesting metaphor. We agree with it because that’s how we see the world, and we have it up. It’s all about how we’re already organized in this permission-structure way, and these are the natural ways that agents can now work with data.
It’s an interesting metaphor, but I do think companies will have to start thinking about how they digitize more of that data. What was your take?
swyx
The company is probably like an ACL-compliant file system, which I’m guessing Box is, right?
Every Agent Needs a Box
Yes.
swyx
Yeah, which you have a great piece on.
Every Agent Needs a Box
Well, I want to rewind a little bit to the graph word. You said that’s the magic trigger word for us. I always ask, “What’s your take on knowledge graphs?” Because every database person always wants to see what they think. There have been knowledge-graph cycles, and you’ve seen it all.
swyx
I actually am not the expert in knowledge graphs, so you might need to—
Every Agent Needs a Box
You don’t need to be an expert. I think it’s just like, how seriously do people take it? Is there a lot of potential in knowledge graphs?
swyx
Can I understand first whether it’s a loaded question, in the sense of whether you’re super pro, super con, super anti, or medium?
Every Agent Needs a Box
I see pros and cons, but I think your opinion should be independent of mine.
swyx
No, totally. I just want to see what I’m stepping into. I know it’s a huge trigger word for a lot of people in our audience, and they’re trying to figure out why this is such a hot topic now.
Every Agent Needs a Box
Because a lot of people get graph religion, and they’re like, “Everything’s a graph. Of course you have to represent it as a graph.” Or, “How do you solve your knowledge changing over time? Well, it’s a graph.”
I think there’s that line of work, and then there are a lot of people who are like, “You don’t need it.” Both are right.
swyx
Yeah, and what do the people who say you don’t need it argue for?
Every Agent Needs a Box
Markdown files. Simplicity.
swyx
Versus structure versus less structure, right?
Every Agent Needs a Box
I think the tricky thing is, again, when this gets met with real humans, they’re just going to their computer. They’re working with some people on Slack or Teams. They’re sharing some data through a collaborative file system and Google Docs or Box or whatever.
I certainly like the vision of most knowledge-graph, futuristic ways of thinking about it. It’s just that it’s 2026, and we haven’t seen it play out yet. I remember, like—I actually don’t even know how old you guys are, but to show my age, I remember 17 years ago everybody thought enterprises would just run on wikis.
Confluence actually took off for engineering, for sure, unquestionably, but this was like everything would be in the wiki. Based on our general internal style of what we were building, we were just like, “I don’t know. People just want a workspace. They’re going to collaborate with other people.”
swyx
Exactly. So you were anti-knowledge graph.
Every Agent Needs a Box
Not anti, not anti.
swyx
I’m not anti. I think your search system—I just think these are 2 systems that probably—
Every Agent Needs a Box
I’m not in any religious war. I don’t want to be in anybody’s YouTube comments on this. This is not a fight for me.
swyx
We love your comments. Get in the comments.
Every Agent Needs a Box
Okay, but it’s mostly just a virtue of what we built, and we just continued down that path. That was what we pursued. But this is not existential for you.
swyx
Great.
Every Agent Needs a Box
We’re happy to plug into somebody else’s graph. We’re happy to feed data into it. We’re happy for agents to talk to multiple systems. Not our fight.
swyx
Yeah.
Alessio Fanelli
But I need your answer. You know, graphs are nerd bait. It’s very effective, nerd.
Every Agent Needs a Box
See, this is one opinion, and then I’ve—
Alessio Fanelli
I think the actual graph structure is emergent in the mind of the agent, in the same way it is in the mind of the human, and that’s a more powerful graph because it actually evolves over time.
Every Agent Needs a Box
Tell me how to graph. I’ll figure it out myself.
swyx
Exactly. Okay, all right. And what’s yours?
Alessio Fanelli
I like the wiki approach. I’m actually—obviously, I spend some of my time at Cognition, which you know very well, and they’ve had a lot of success with DeepWiki. It powers a lot of Devin’s brain. It’s super powerful, and it’s useful for humans, but it’s useful for agents.
swyx
Yes. Tell me if you think I’m wrong on this, but it’s not much of an access-control-structure issue. You get the whole codebase, and everybody gets—
Alessio Fanelli
Before I speak too much, there may be some enterprise controls on the enterprise DeepWiki offering that I’m not familiar with, but I don’t have anything on the public side. I think almost every agent should have its own wiki that it’s updating, and that’s persistent memory, and that is a very weak knowledge graph.
swyx
And you could strengthen it if you want more structure, but you may not need it. Markdown files having links and wiki style, right? But very effective, right, Lindy?
Alessio Fanelli
Yep.
swyx
Okay, last couple of questions, but feel free to jump in, or if you want to rant. I see you as a very interesting and unusual founder. You’re of 2 worlds: you’re of Silicon Valley, but you’re also of the Fortune 500s. I feel like your founder mode is very different from Brian Chesky’s founder mode, and I’m curious if you have any reflections on how you operate as a founder.
Every Agent Needs a Box
What would his founder mode be?
swyx
Don't delegate.
Every Agent Needs a Box
Right. And how would you put me?
swyx
You do delegate.
Every Agent Needs a Box
Ah, okay. I see. I don't know that Brian and I would be that far removed from each other when you get to the specifics.
There's a whole bunch that I delegate. Ninety percent of the work that happens at Box is fully delegated. We've got great leaders running all that stuff. It's just too much for my brain to handle. And probably 70% to 80% of the work at Box—I’m going to make up all the numbers here—I only need to really look at about 5% of that for high-leverage decisions.
You know, what's the marketing message that we think is going to resonate with customers? That's a little bit of a high-leverage thing that we do in marketing, but most marketing activities I don't get involved in. What's our sales pitch? Maybe I'll be involved in that a little bit. What are roughly the investments or push we're going to do in certain verticals? That's about 5% of the total bandwidth of the key areas of sales or go-to-market.
So 70% to 80% of the company, I can just do about 5% of it, and then operationally we've got great leaders and they're going to execute on that. We collaborate on the 5%. It's not like I'm just making up a decision and saying to go and do it.
Then there's this part that is the existential part of the business, which is: if we don't do this right, we're out of business. By virtue of just being a founder, you get kind of sucked into that part of the work because you can feel it. You can just see how the AI tsunami could wipe you out if you make just 2, 3, 4, 5 wrong decisions in this space.
A couple of wrong architecture decisions, a couple of wrong AI feature decisions, a couple of wrong API platform decisions, and you might be out of the game a year from now. You just feel it in your bones. You know this. We feel this all day long in this space, given what's happening.
And so, in that area, you can't delegate in a classic sense. You still need to make sure you've got great leaders and strong hires and people that have high agency because they want to be able to own part of the strategy and the roadmap, or else you can't hire good people. But there's going to be a lot of little micro-forks in the road that will compound to determine whether you succeed or fail.
Your founder energy just automatically draws you into those because they are the determining decisions of your company's future. That's kind of where I spend my time. You have to do it in a collaborative way again, because if you are only dictatorial, you eventually won't be able to hire the best people because they won't want to work in that environment.
But you also can't abdicate all the responsibility, because the risks are simply too high. You have to somehow add some value. The value I add is that I've seen 20 years of this business, so I think I can piece together what I expect the value propositions are going to be and how customers will react to certain things. That's what I can bring to the table.
Then you have this kind of existential fear that, if I get it wrong, it's all on me anyway. I don't get to blame the engineer who was working on that project. It's all my fault, right? At the end of the day, it will be my fault if it doesn't work. By virtue of that liability and responsibility, you just get pulled into needing to make sure it's all going according to how you think it needs to end up.
swyx
I don't know how Brian would answer that, I guess.
Alessio Fanelli
Yeah, it's a long essay. It's an interesting essay. People should go and compare and contrast your answer versus his. I do think that systems have a way of letting entropy get to them. If you step away for too long, you need to have a way to check in and go, “Well, do I need to come back in, or are we good?”
Every Agent Needs a Box
People are going to tell you things are good, but they're not good.
Alessio Fanelli
Yes. 100%.
Every Agent Needs a Box
Yeah.
Alessio Fanelli
And I'm actually a fan of process for that 70% to 80%.
Every Agent Needs a Box
So that 70% to 80%, the process is: you're going to do a quarterly business review, you're going to have a brand check-in, and you're going to do those things. You're going to make sure that you're seeing all the right episodes of what's changing and how it's evolving, and make sure it's going in the right direction.
Then there are some areas where it's like, no, it's 24/7. I guarantee that after this podcast, at 11 p.m., I'll be doing a Zoom with Ben and probably some other people because we're going to be talking about agents and new platform features. That's—you’re just in the cauldron, grinding on that side.
swyx
Yeah. That's extremely realistic as to what it's like, and I just want to have people hear your perspective on what it—
Every Agent Needs a Box
And this is like—you read the post about everybody having agents running over the weekend, and it's like, you just—I mean, first of all, anybody crazy enough to come to Silicon Valley, we don't bring good news about the healthiness of our environment right now. You have to know what you're signing up for.
But there's a real issue: do I have enough agents running?
swyx
Yeah, I made a meme that was semi-viral for me about this.
You can't even enjoy a party these days because you're working with your tokens.
Every Agent Needs a Box
There's compute out there that you're not utilizing. What the hell?
swyx
I paid for the $200. I'm going to spend the $200.
Every Agent Needs a Box
Yeah.
swyx
I'm going to spend $6,000 out of $200.
Every Agent Needs a Box
Exactly. Exactly. Exactly.
swyx
We need to make Anthropic very unprofitable.
Every Agent Needs a Box
Okay. Yeah. We're not doing a good enough job.
Alessio Fanelli
Cool. I have a closing question, unless you—
swyx
I have a question. I've asked this question in private before, but I'm going to ask it again. It's a question that Tyler Cowen asks guests on his podcast: what is the Aaron Levie production function?
Every Agent Needs a Box
And—
swyx
Oh, I love that. I love this question because there are so few people who I think are good at both executing and distilling and putting good ideas into the ether. You put a lot of good ideas into the ether. What is the Aaron Levie production function that allows you to do that versus others? How do I get that information?
Every Agent Needs a Box
I can give you a variant: what goes into Aaron Levie—
swyx
And what goes out, and how does it turn inside?
Every Agent Needs a Box
I'm just trying to think of it because there are some very— I just read a lot of Twitter as well, and so I—
swyx
And you spend a lot of effort, too. You don't see great mini-essays from Brian Chesky every day, but you do from you.
Every Agent Needs a Box
Oh, yeah. You're kind of weird in that way.
swyx
Maybe he's healthier than me, actually. We should just text him to see if he's got a—
Every Agent Needs a Box
I think he does work out, right? He has bigger muscles. Well, that's the thing. I work out less than him and I tweet more than him. That's how we're balancing things out.
Every Agent Needs a Box
I mostly the way I just think about it is uh is just um you know there's there's lots of work that's happening in the business. I'm getting to see the all the problems that we are running into constantly and I'm trying to uh be a little bit of a create a flywheel between what we're doing internally what what what then we talk about uh getting a feedback loop on that and seeing other people's you know experiences of what they're doing bring that back into the business and and so I just see like my job as as you know hopefully being able to kind of connect the dots of of what's going on in the world with what's going on in box and then I just happen to tweet about that along the way. Um because
swyx
It's all you, and there's no editor?
Every Agent Needs a Box
Yeah. Wow.
At the time, I tried to get an internship between my freshman and sophomore year at a film production company in New York. I got the internship, and then I emailed my liaison—the guy who sponsored me for the internship—and said, “Hey, I'd like to do a blog of my summer internship where I blog about being an intern at a production company in New York.”
About half a day later, they emailed me back saying they'd rescinded the internship.
swyx
No.
Every Agent Needs a Box
Yeah, because I showed a lack of judgment on professionalism or whatever. Even the idea that I would ask that question raised red flags: “Who the fuck is this guy?”
Anyway, I only say that to say that, to me, building in public is just a natural thing. I just go through the day, we deal with interesting problems, I tweet about them, and I get information back in the process. I see your work, I see a bunch of folks' work, and try to incorporate that back into Box.
My job is to try and connect all these things together and make it useful.
swyx
And you're, I mean, the number-one spokesperson, right? So you do have to be out there.
Every Agent Needs a Box
Yeah, but I would kind of be doing it whether or not. I don't really think of it as a job requirement as much as I just like social media.
swyx
You're so good at it.
Every Agent Needs a Box
Yeah.
swyx
It's so hard to believe. Okay, so do you get up at 5:00 a.m. with coffee? Is that your secret?
Alessio Fanelli
How do you work? Do you actually do it in the back of Waymos? Do you do it that way? How do you do this?
Every Agent Needs a Box
No, it's mostly that, though. I have a commute home each night. I try to see my kids most weekdays before I have to hop back online, so there's a 20-minute window where I can distill the information that's happened and ask, “Is there anything I learned today that would be interesting to throw out there, or anything that I saw?”
Then, probably somewhere between 7:30 and 9:00 p.m., I finally get a chance to look through the feed and see, “Did anything crazy happen in AI?” That will also catalyze something. That's the best I can kind of do.
Alessio Fanelli
Yeah. Okay, thanks. Now I know your cutoff is 8:00 p.m. I will try to get AI news out before 8:00 p.m. so I can help him do his thing. Basically, if I don't see it before 8:00 or 8:30, I'm not going to tweet it or something, because then I'm back on Zoom after that.
swyx
I wasn't planning on asking this, but you've mentioned the film stuff.
Every Agent Needs a Box
Yeah.
swyx
One of my favorite parts of researching you was that you got the idea for Box from the Paramount lot, pushing paper. Are you a film guy?
Every Agent Needs a Box
I would say I used to be more of a film guy.
swyx
What are your favorites, if you want to list off any?
Every Agent Needs a Box
Kind of the classic wannabe film-student classics.
swyx
Are we talking Scorsese, Pulp Fiction, Magnolia—
Every Agent Needs a Box
Requiem for a Dream? Basically, if there was an art-house film in the 1990s to the early 2000s, that was my genre. That got me into thinking, “Wouldn't it be cool to do film?” Then I thought maybe I could connect digital into it, but could you do film online? That just seemed too hard from a licensing standpoint. Then, obviously, Netflix existed, so I was never quite able to fully connect the dots on those things.
But the internship at Paramount was one catalyst for starting Box, because we were using just traditional enterprise software, and I was like, “Wow, it's really hard to share data—files going back and forth.” The same thing was happening in school as well, and so that all led to Box, basically.
swyx
A24 is kind of driving the resurgence of independent film, I guess, in the face of all the Marvel slop.
Every Agent Needs a Box
100%.
swyx
I was thinking about this the other day, and A24 is certainly the best example of this today. They just don't—you know, it's hard to make a film like No Country for Old Men or There Will Be Blood. What is that movie today?
Every Agent Needs a Box
What is a brand-new movie that you just watch and you're like, “What did I just watch?”
My 6-year-old's movie benchmark is Forrest Gump.
swyx
Which was iconic in its time.
Every Agent Needs a Box
Yep, 100%.
swyx
Never again.
Every Agent Needs a Box
Yeah. We did not make—we don't know how to make Forrest Gump anymore. We'll try to make the sequel, though, at some point, for sure.
swyx
I'm fine with—
Every Agent Needs a Box
No, Forrest Gump has a kid. Yeah, yeah, he's still around.
swyx
Exactly. I think Forrest Gump having a grandkid would be a good movie. What is the grandkid of Forrest Gump doing in 2026?
Every Agent Needs a Box
Goes tropical.
swyx
Yeah.
Every Agent Needs a Box
I definitely want to see more good movies out there. I'm a little bit conflicted on AI and film.
swyx
Oh, let's do that.
Every Agent Needs a Box
The world does not need more slop in AI entertainment, but I'm in a mode where I think AI is generally going to be a pure positive. If I were me 25 years ago in high school, I would definitely be making a full-production film that had explosions and car chases, but then there would be people who would show up there. I think that ability to just get to be Spielberg is completely amazing, and democratizing that is incredible.
I'm concerned about how you make sure that we still get P.T. Anderson along the way, and whether we can make sure that those guys continue to exist. Interestingly, I never saw it, but Darren Aronofsky has either put out or is going to put out an AI film. Even some of the best artists are starting to adopt this.
What I don't want to do is just be in this TikTok feed of films, where it's like, “Oh, there's a film about the car chase that does this thing.” We don't need that. This should be a form of entertainment and art. Let's use AI to accelerate the production process, do the really hard CG work that you had to spend way too much money on previously, and test out all kinds of new plot ideas.
Alessio Fanelli
Yeah, previs.
Every Agent Needs a Box
Yeah, background, and it's incredible. All those things are super incredible. I still like the idea—it's very nostalgic, but I still like the idea that there's a camera and a person, and a person who says, “Action.” Hopefully, we can surround AI around that. We'll see how that plays out.
swyx
Yeah. One of the things that Stability AI made an impression on me with was, “Well, at least now we can remix Game of Thrones season 8 and make it again like it was meant to be, not rushed.”
Every Agent Needs a Box
Yeah. I have a 6.5-year-old, and you see a lot of these kids' movies and you're like, “Yeah, that probably will be AI.” I don't totally know the job math, because I don't know how many animators there are today. But I actually think, weirdly, we could be producing more high-quality, maybe even slightly educational, kids' entertainment.
Maybe that's a positive: you could just have a Pixar for things where kids learn stuff. It used to be these very low-fi lesson things.
swyx
I mean, we had Teletubbies. That was so slow.
Every Agent Needs a Box
We could have way more of that. Maybe every animator who's making a Pixar film today is now responsible for more content, and they've got AI agents running. I think there are some optimistic scenarios on the entertainment side. There are a lot of great use cases for generative media.
swyx
Yeah. Edutainment as well.
I guess one question I have is kind of a self-serving one, almost like an advice question. One of the things I really enjoyed researching about you was that Michael Arrington had some influence on the Box journey because you went to his house party.
Every Agent Needs a Box
Yes.
swyx
And that's how you got funding.
Every Agent Needs a Box
Yes.
swyx
One of Michael Arrington's—that's a deep cut, right?
Every Agent Needs a Box
Yeah, very deep cut. That's a 2006 deep cut.
swyx
Do you want to tell that story? I don't know if you've told it.
Every Agent Needs a Box
It's not even much of a story.
swyx
That's like a random intro, right?
Every Agent Needs a Box
Well, he used to have house parties. TechCrunch had these house parties, and it was probably no different from somebody having a house party in San Francisco. You just go and meet the VCs and founders. I don't want to make up examples, but there would be Chad Hurley over there pitching YouTube to people. That's just how it worked.
It was this era where all these new companies were emerging, and I met our first investor in Silicon Valley at one of these house parties, Emily Melton, who then brought us into DFJ. That became our Series A. It was all because of Arrington's backyard party.
swyx
One of my aspirations for Latent Space is to be as helpful and influential, or whatever, as TechCrunch was in the day.
Every Agent Needs a Box
Yeah.
swyx
What would a new TechCrunch today look like? What should I do?
There used to be TechCrunch Disrupt. I could do that with my conference, but I haven't done it yet.
Every Agent Needs a Box
Well, I mean, I think useful. I don't know. Actually, interestingly, I would argue that Disrupt came after that deep-cut period. I think Disrupt ended up being catalyzing. I think Cloudflare launched at Disrupt—is that the story?
swyx
Okay, okay.
Every Agent Needs a Box
I think anytime you can be a launchpad, that's great, because it draws in people who are in that creative moment. Whether it needs to be a contest or just everybody gets 5 minutes and you're fundraising—
Alessio Fanelli
Who knows? But, for what it's worth, I don't have that much advice because I think you're already doing it effectively. I just watch the YouTube videos late at night from the events. I haven't been to one of your events, but from the camera angles, it looks like everybody's there.
What's great is that people are going to be in the audience as 2 random people, and they'll be like, “The next big AI company will come from people coming to a meetup because they were like, ‘I came in from Chicago, and I'm from Poland. Let's go do a startup.’” That's the magic of the Valley. X43 [?] found his co-founder at AI Engineer, and I know of at least one marriage that's—
swyx
Wow, you have marriages already.
Alessio Fanelli
I never heard that about—
swyx
That's my favorite KPI.
Alessio Fanelli
Wow, we have AI marriages at the AI Engineer conferences.
swyx
These are about humans, to—
Alessio Fanelli
Clear. That's a very good clarification.
swyx
I like that you have to check.
Alessio Fanelli
Yes, that's a very good clarification.
swyx
No, but I think you're an insightful business leader with a lot of thoughts on media. I just figured I would—
Alessio Fanelli
Media is such an interesting space right now because with the go-direct model, every company is going to have to be a media company.
swyx
You are the OG go-direct.
Alessio Fanelli
Yeah. But we're still—I think what you guys are doing, and I don't even know all the overlapping relationships, but I watched your videos of your events, and it's clearly the new format, right? Companies have to become channels to communicate with audiences.
I think the resurgence—maybe “resurgence” is a bad word because it implies it declined—but DevRel is hot. It's the hottest thing of all time right now.
swyx
I'd like it if you could produce a freaking factory of DevRel people. There's just unlimited jobs right now on the other end of that.
Alessio Fanelli
Because everybody needs their services and APIs to be used by agents, and so we have to all find a way to be like, “Hey, look at me. Agent, please come over here, agent.” That's going to be a content game. How do you get the agents to see your stuff?
swyx
And know your APIs? This is a new world that we are in, and it's going to be a completely digital-marketing kind of world that we're in.
Alessio Fanelli
Yeah. For what it's worth, I'm trying to help by doing little writing boot camps and basically turning them into DevRel boot camps. It's a demand-and-supply problem: there's huge demand and no supply.
swyx
Why is there no supply?
Alessio Fanelli
The really good ones work for themselves.
swyx
Uh-huh.
Alessio Fanelli
The creator economy screwed you over.
swyx
So I see.
Alessio Fanelli
The most talented guys are making millions and just working for themselves while they work for you.
swyx
Good. We don't want them to make that much money. [laughter]
Alessio Fanelli
We need to be able to hire people.
swyx
Do what some companies are doing—not saying it's my situation exactly—but give them equity. It should probably be worth more, just sort of helping them out.
Alessio Fanelli
They are getting—
swyx
Oh, sorry. As full-time employees or not?
Alessio Fanelli
Part-time.
swyx
You need full-time.
Alessio Fanelli
I'm part-time.
swyx
Yeah, but you're an N of 1. We also need people who are full-time.
Alessio Fanelli
My classic joke, or observation, was when HubSpot bought The Hustle, the newsletter business, and then they bought My First Million, the podcast—you must know Sam.
swyx
He's obsessed with this guy.
Alessio Fanelli
So my conclusion was that every company must either build or buy a media company, right? Until you realize that you have to take it that seriously—that you are running a media business in your company—you will never be good at it.
swyx
Yes, 100%.
Alessio Fanelli
Yeah.
swyx
No, we're very much taking that seriously.
Alessio Fanelli
No, we are all engineers here.
swyx
No, that's the headline.
Alessio Fanelli
Okay, yeah.
swyx
DevRel is the future job. We're all just going to be doing DevRel in some form.
Alessio Fanelli
I mean, what is DevRel?
swyx
Developers are ruling the earth. What is DevRel? I don't know.
Alessio Fanelli
No, it's DevRel.
swyx
Yeah, okay.
Alessio Fanelli
Isn't it just glorified consulting? That's the downside.
swyx
Sure. I mean, I guess nobody can actually fully define this, but I think it's micro-DevRel. You're in the company, you're helping them with the services, and you're doing a little extra implementation.
But, yeah, I think we're all—the thing that's going to happen on the leverage of software is that we're going to produce far more output of code, and thus features, per dollar. On the other end of this, we're going to end up spending probably just as much on how you get all of that stuff to the customer.
That's going to create a new set of roles that we are all doing, partly because there's so much choice now that you have to fight for attention, or because the stuff is changing so quickly that you have to technically help your customers along the journey.
I just laugh when people say you don't need to be an engineer or that you shouldn't do computer science. I actually think that's still one of the most protected job categories, because things are only getting more technical and harder. Anybody in a technical position is in the best position to get agents deployed, get them built, get them adopted, and build the custom-code software for the IT system—all of that.
Alessio Fanelli
My classic founding story of why I picked “AI Engineer” as a title and as a theme for this podcast and my conference was that, back in early 2023, someone came to me and said, “I'm all in on AI. What should I do?” I just looked at her and was like, “God damn it, there's nothing you can do. Engineers are about to get so much more powerful than you. You don't even understand.”
swyx
Tell me that's a good idea. Should she go and then learn how—
Alessio Fanelli
No, I didn't say any of that to her. I'm not that honest.
swyx
I hope somewhere out there she did go to some online academy and learned. But there's a lot of people who believe AI too much, and then they're like, “Well, you don't need to learn to code, so I won't learn to code,” and then there's—
Alessio Fanelli
There's a bunch of us who are just in that sweet spot where we can code and wield AI a thousand times more effectively than you can. Yeah.
swyx
And, like, who's going to win here?
Alessio Fanelli
I think this was another tweet, but it was the observation that software engineering for the past 30 years was the primary career track for technical, high-agency people who wanted to have a large, outsized impact on the world.
Software was a means to do that effectively. So, with AI, is it that AI could eat software engineering, or, say, software engineering could eat all these other domains and disciplines?
swyx
Those same principles then get applied to every other field, right?
Alessio Fanelli
Yeah, exactly. Yeah. I mean, GTM engineering is that. And, well, this is the thing: anybody who believes that an enterprise is going to build its own software for all of its problems must be the most long on computer science as a discipline of all time.
Most of the economy does not have enough engineers to maintain all those systems, update all those systems, figure out the relationship between the business problem and what the code needs to do, and actually manage that. That's a very pro-engineering-job argument for what the future is going to look like.
I'm still back and forth on whether you're really going to build all these things versus using prepackaged software, but no matter what, there's going to be 10 to 100 times more code. I think you can be very long engineering right now, purely on the dimension that software is going to become increasingly more important once agents are turning everything into software.
swyx
All right. 3 software guys say software. [laughter]
Alessio Fanelli
Not biased at all.
swyx
Okay.
Alessio Fanelli
But you're an inspiration. Such a pleasure.
swyx
All right. Good to be here.